<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <title>Harshith Kethavath — Blog</title>
  <subtitle>Writing by Harshith Kethavath on ML research, infrastructure, and things worth thinking about.</subtitle>
  <link href="https://kethavath.com/feed.xml" rel="self" />
  <link href="https://kethavath.com/" />
  <updated>2026-08-19T00:00:00Z</updated>
  <id>https://kethavath.com/</id>
  <author>
    <name>Harshith Kethavath</name>
  </author>
  <entry>
    <title>You Are Already Devoted</title>
    <link href="https://kethavath.com/blog/devotion.html" />
    <updated>2026-08-19T00:00:00Z</updated>
    <id>https://kethavath.com/blog/devotion.html</id>
    <content type="html">&lt;p&gt;A woman at 6:47 in the morning, as soon as she wakes up reaches for her phone. Her thumb knows to unlock, swipe and scroll. Across the town, a man ties his shoe lace in the dark. He has been doing this every morning for six years. He ran through rain, through grief, through a broken toe that he told no one about. And at five in the evening, a regular takes his stool at a bar and orders the same thing. The bartender pours it before he even asks. It is the most important appointment to him in his life, which he has not missed in years.&lt;/p&gt;
&lt;p&gt;Humans are worshipping creatures. This is not a claim about religion, it is about the thing that organizes your attention, structures your day, and benefits from your sacrifices, call it god. Whatever gets your first waking hour and last thought before you sleep. Whatever you would defend a little too quickly, if someone suggested it is a problem. Finding yours requires no Sherlock Holmes, look at where time and money gets invested in.&lt;/p&gt;
&lt;p&gt;David Foster Wallace once said &amp;quot;In the day to day trenches of adult life there is actually no such thing as not worshipping. Everybody worships. The only choice we get is what to worship.&amp;quot; I think most of us never make the choice at all.&lt;/p&gt;
&lt;p&gt;Devotion is not some tap you can just turn off, it is a river. If you do not give it a channel, it will make its own, and our attention like water will follow the path of least resistance. Your environment picks it for you. Your friends pick it for you. Your wounds pick it for you. The algorithm, which knows your weaknesses, will pick it for you.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;There is this tweet I read: &amp;quot;Addiction is proof that you are capable of intense devotion. You just have a false god.&amp;quot;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Think of what addiction actually requires. The addict organizes an entire life around a single thing. Rises for it, lies for it. Sacrifices sleep, money, health, and love for it. Returns to it every day, against all punishments, past every consequence, with a consistency most of us can never bring to anything. Remove the object and look only at the behavior, and you will find a discipline most Olympians could not match: total commitment, perfect attendance, a life built entirely of ritual.&lt;/p&gt;
&lt;p&gt;We usually describe addiction as weakness, as a failure of will, as a lack of discipline. The tweet says the opposite, and I think it&#39;s right. The addict does not lack commitment, they have plenty of it, just pointed at the wrong thing. You cannot shame a river to make it flow backwards, you can only offer a better channel.&lt;/p&gt;
&lt;p&gt;So, if devotion is the problem, it can be the exit as well. You do not need to become a less intense person. You cannot quit devotion, you can convert it. Conversion sounds dramatic: lightning, tears, a before and after. I think in practice it is very boring.&lt;/p&gt;
&lt;p&gt;Your subconscious consumes everything you feed it. It absorbs the songs you play, the things you say to yourself over and over, the company you keep. It cannot distinguish between what is truly happening and what is being imagined in your mind. Which means, whatever you do daily is, quite literally, talking to the deepest part of you about what reality is.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;As the Bhagavad Gita teaches: “You are what you believe in. You become that which you believe you can become”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is why the way back to a true god is not a decision, it is a practice. bell hooks wrote that &amp;quot;love is not a lightning strike but a practice, a discipline of deliberate choices, repeated&amp;quot;. The same is true for any kind of devotion worth keeping. Take the run when you don&#39;t want to. Write the page, even badly rather than not writing at all. Pray. Meditate. Sleep early. None of these feel like transformation when you look at them on a scale of a day. That is the point. Rituals are never impressive right there in the moment; they are only impressive in aggregate.&lt;/p&gt;
&lt;p&gt;None of this means willpower conquers everything. Anyone who has been through a real addiction knows this, and there is no shame in needing other people to get out. The most durable recovery programs do not ask their members to become less devoted. They rebuild the whole architecture of devotion around a new god: a higher power, a daily meeting, a sponsor&#39;s phone number. The genius was never the restraint, it was the conversion.&lt;/p&gt;
&lt;p&gt;And once the conversion happens, compounding does the rest. A ritual which is kept for a week is a nice try, for a year is a personality, and for a decade is a life. The life you end up with is not decided in the dramatic moments, it is the sum of thousands of unwitnessed mornings, each with a small vote for one god or another.&lt;/p&gt;
&lt;p&gt;This is also where manifestation stops being woo woo. People roll their eyes when they hear about vision boards, the affirmations, writing the same sentence a hundred times. But look at the mechanism underneath. Manifestation is just a ritual that feeds your subconscious a reality that does not exist yet. Repeat the picture often enough and the deepest part of you starts treating it as already true. And once you believe something is true, you act like it. The belief shapes your actions, and your actions shape your outcomes. Nothing supernatural has to happen. You are simply choosing, on purpose, what to feed the part of you that believes everything.&lt;/p&gt;
&lt;p&gt;The same is true of the small things. Choose not to wallow. Speak to yourself the way you would speak to someone you love. Keep company with people who make you want more from your life. On their own these look like minor changes. But each one is a message to your subconscious, repeated daily: this is a person worth taking care of. Say it often enough and you will start to believe it. And a person who believes they are worth taking care of makes very different choices than one who does not. Self care is not a treat. Done right, it is devotion to yourself, living with intention instead of by accident.&lt;/p&gt;
&lt;p&gt;And after all of it, there is one last step: letting go. You do the rituals. You feed your mind carefully. You show up every morning. And then, as &lt;a href=&quot;https://jessyin.world/writing/to-have-the-life-of-your-dreams-you-need-to-live-with-intention&quot;&gt;jess yin puts it&lt;/a&gt;, you trust that the universe will take what you have given it and run its course. Not because the universe is guaranteed to be kind, but because gripping every outcome was never devotion. It was control.&lt;/p&gt;
&lt;p&gt;Go back to the three people from the beginning. I let you judge them, and you probably did, and you were probably at least one-third wrong. The scroll might be love: a daughter checking the weather in her mother&#39;s city. The run might be the only rope holding a man above his own grief. And the stool might be the loneliest form of belonging there is, a room where somebody knows your name, at six dollars a glass.&lt;/p&gt;
&lt;p&gt;Same hunger in all three. Same magnificent, dangerous capacity. The only variable, the only one there has ever been, is the object, and what the object gives back.&lt;/p&gt;
&lt;p&gt;You will worship today. You worshipped yesterday. You will find it again tomorrow morning before your eyes are fully open. The question was never whether you are capable of devotion. Addiction proves you are. So does love. So does every habit you have ever kept in secret.&lt;/p&gt;
&lt;p&gt;You are already devoted. The only question left is to what.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>What Cloud Segmentation Taught Us About Vision-Language Models</title>
    <link href="https://kethavath.com/blog/cloudprompts.html" />
    <updated>2026-04-01T00:00:00Z</updated>
    <id>https://kethavath.com/blog/cloudprompts.html</id>
    <content type="html">&lt;p&gt;Most production AI systems today reach for pretrained models and prompt them.
A recent &lt;a href=&quot;https://arxiv.org/abs/2512.04123&quot;&gt;study&lt;/a&gt; found 70% of production AI systems rely on prompting rather than weight tuning.
The logic is simple: prompting is cheap, labeling is expensive, and if the pretrained model is good enough,
language can steer it the rest of the way.&lt;/p&gt;
&lt;p&gt;For natural images, this mostly works.
For satellite imagery, we found it doesn&#39;t work at all, and the cheap alternative to prompting turned out to be cheaper than we expected.&lt;/p&gt;
&lt;p&gt;This post summarizes our paper,
&lt;a href=&quot;https://arxiv.org/abs/2604.08956&quot;&gt;&lt;em&gt;Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift&lt;/em&gt;&lt;/a&gt;,
accepted at EarthVision 2026 (CVPR Workshop).&lt;/p&gt;
&lt;h2&gt;The setup&lt;/h2&gt;
&lt;p&gt;We picked a task where the domain shift from natural images is severe on two axes at once: cloud segmentation in Sentinel-2 satellite imagery.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Visually:&lt;/strong&gt; overhead perspectives, multispectral sensors, atmospheric phenomena that blend into haze and shadows that lack hard edges. Nothing like the object-centric photos CLIP was trained on.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Linguistically:&lt;/strong&gt; vocabulary like &amp;quot;optically thin cirrus&amp;quot; or &amp;quot;cloud shadow&amp;quot; barely exists in CLIP&#39;s image-caption training corpus.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We used &lt;a href=&quot;https://huggingface.co/CIDAS/clipseg-rd64-refined&quot;&gt;CLIPSeg&lt;/a&gt;, a promptable segmentation model built on a frozen CLIP backbone, and evaluated it on &lt;a href=&quot;https://huggingface.co/datasets/isp-uv-es/CloudSEN12Plus&quot;&gt;CloudSEN12+&lt;/a&gt;, the largest expert-labeled cloud segmentation dataset available. The question was simple: under this kind of compound shift, can prompt engineering alone compensate, or is supervised adaptation necessary?&lt;/p&gt;
&lt;h2&gt;Finding 1: Every prompt we tried made things worse&lt;/h2&gt;
&lt;figure style=&quot;max-width: 500px; margin-left: auto; margin-right: auto;&quot;&gt;
&lt;img src=&quot;https://kethavath.com/assets/cloudprompts/prompts.png&quot; alt=&quot;Segmentation results across prompt variants&quot;&gt;
&lt;figcaption&gt;mIoU scores across all 60 prompt variants. Every engineered prompt falls below the zero-shot baseline.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;We designed 60 prompt variants (15 per class) spanning four strategies: minimal single-word labels, domain-specific meteorological terms, appearance descriptors, and contextual phrases. &amp;quot;White cloud,&amp;quot; &amp;quot;wispy cloud,&amp;quot; &amp;quot;bright white opaque cloud,&amp;quot; &amp;quot;shadow beneath cloud,&amp;quot; and so on.&lt;/p&gt;
&lt;p&gt;Every single one underperformed the simple class-label baseline of 0.255 mIoU. The worst engineered prompt scored 0.07 mIoU, a 73% relative degradation. The best engineered variants clustered around 0.20 to 0.22, still clearly below baseline.&lt;/p&gt;
&lt;p&gt;The most interesting failure came from negation. Prompts like &amp;quot;not cloud&amp;quot; or &amp;quot;not haze&amp;quot; produced the worst results of all. This isn&#39;t a quirk; it&#39;s architectural. CLIP&#39;s contrastive training does use negative examples, but those teach the model that specific images and captions don&#39;t match, not that the word &amp;quot;not&amp;quot; is a semantic operator. The text encoder never saw &amp;quot;not cloud&amp;quot; paired with cloud-free images during training. The embedding for &amp;quot;not cloud&amp;quot; remains dominated by &amp;quot;cloud.&amp;quot; The word &amp;quot;not&amp;quot; carries no learned visual meaning.&lt;/p&gt;
&lt;p&gt;This matters for a broader point: learnable prompt methods like CoOp and CoCoOp optimize within the same embedding space. If the space itself is misaligned with the target domain, prompt optimization runs into the same ceiling. The bottleneck isn&#39;t the prompt strategy. It&#39;s the visual encoder.&lt;/p&gt;
&lt;h2&gt;Finding 2: The crossover point is shockingly low&lt;/h2&gt;
&lt;figure style=&quot;max-width: 500px; margin-left: auto; margin-right: auto;&quot;&gt;
&lt;img src=&quot;https://kethavath.com/assets/cloudprompts/low%20data.png&quot; alt=&quot;mIoU across data percentages&quot;&gt;
&lt;figcaption&gt;mIoU as a function of training data percentage for LoRA and FFT.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Having established that prompting can&#39;t bridge the gap, we asked the complementary question: how much labeled data does it take to beat the zero-shot baseline?&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Eight images.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;At 0.1% of the training set, approximately 8 labeled patches, both LoRA and full fine-tuning already surpass the zero-shot baseline on aggregate mIoU. At 5 to 10% of the data (roughly 425 to 850 images), we recover about 85% of the maximum achievable mIoU. Beyond 30%, additional labels yield diminishing returns that rarely justify the annotation cost.&lt;/p&gt;
&lt;p&gt;This fundamentally changes the deployment math. The usual argument for zero-shot is &amp;quot;labeling is expensive, so we prompt.&amp;quot; But if eight images is the crossover, that argument doesn&#39;t hold, even for a small research team or a practitioner working alone. Labeled data isn&#39;t the expensive alternative to prompting; for domains with real distribution shift, it&#39;s the worthwhile path.&lt;/p&gt;
&lt;h2&gt;Finding 3: LoRA vs. FFT is a task-structure decision, not a compute tradeoff&lt;/h2&gt;
&lt;figure style=&quot;max-width: 700px; margin-left: auto; margin-right: auto;&quot;&gt;
&lt;img src=&quot;https://kethavath.com/assets/cloudprompts/cm.png&quot; alt=&quot;Confusion Matrices across adaptation strategies&quot;&gt;
&lt;figcaption&gt;Confusion matrices for zero-shot, LoRA, and FFT at 100% training data. Diagonal entries represent per-class correct classification rates.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The conventional framing of LoRA vs. full fine-tuning is about resources: LoRA trades a little accuracy for a lot of parameter efficiency. Our results complicate that framing.&lt;/p&gt;
&lt;p&gt;Across the full data sweep, FFT outperformed LoRA by a consistent 0.03 to 0.09 mIoU. But the aggregate number hides where the gap actually lives.&lt;/p&gt;
&lt;p&gt;For spectrally distinct classes, clear sky and thick cloud, LoRA and FFT perform nearly identically. Both raise clear-sky classification from 0.59 (zero-shot) to ~0.92. Both push thick cloud from 0.32 to 0.88. When the classification boundary is visually unambiguous, the low-rank subspace is expressive enough.&lt;/p&gt;
&lt;p&gt;For spectrally ambiguous classes, the picture changes sharply. On thin cloud, FFT reaches 0.61 classification accuracy against LoRA&#39;s 0.49, a 12-point gap. On cloud shadow, it&#39;s 0.69 vs. 0.56, a 13-point gap. These classes require fine-grained reshaping of representations to capture subtle spectral overlap: semi-transparent cloud layers against varied surface reflectance, shadow regions against dark terrain. Those distinctions are multi-dimensional in embedding space and exceed what a constrained low-rank subspace can express.&lt;/p&gt;
&lt;p&gt;The practical rule that falls out: if your task is well-defined boundaries (separating cloud from no-cloud, say), LoRA is the right default. If your task involves spectrally ambiguous classes that downstream pipelines actually depend on, and in atmospheric correction, thin cloud and cloud shadow absolutely do, FFT&#39;s extra parameter cost is worth it.&lt;/p&gt;
&lt;h2&gt;The supervision dip: a warning about low-data fine-tuning&lt;/h2&gt;
&lt;figure style=&quot;max-width: 700px; margin-left: auto; margin-right: auto;&quot;&gt;
&lt;img src=&quot;https://kethavath.com/assets/cloudprompts/per%20class%20low%20data.png&quot; alt=&quot;Per class IoU across data percentages&quot;&gt;
&lt;figcaption&gt;Per-class IoU as a function of training data percentage for LoRA and FFT.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One finding surprised us and deserves its own mention. For thin cloud and cloud shadow, supervised adaptation initially &lt;em&gt;degrades&lt;/em&gt; performance below the zero-shot baseline before recovering.&lt;/p&gt;
&lt;p&gt;Both classes dip at 0.5 to 1% labeled data and recover at 2.5 to 5%. The mechanism: too few representative examples aren&#39;t enough to reshape embeddings toward the target distribution, but are enough to disrupt whatever coherent structure existed in the zero-shot embedding. You get the worst of both regimes, you&#39;ve broken the zero-shot representation without building a supervised one in its place.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The aggregate mIoU curve doesn&#39;t show this, because gains on clear sky and thick cloud mask it. If you&#39;re fine-tuning with under 1% labeled data, monitor per-class performance, not just the mean.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Stratified sampling strategies that guarantee minority-class representation in low-data subsets are a straightforward mitigation.&lt;/p&gt;
&lt;h2&gt;What this means for practitioners&lt;/h2&gt;
&lt;p&gt;If you&#39;re deploying vision-language models to a domain that diverges significantly from natural image pretraining such as satellite imagery, medical imaging, microscopy, industrial inspection, the defaults from the natural image literature may not apply. A few concrete takeaways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Don&#39;t assume prompting will bridge a real domain gap.&lt;/strong&gt; Test a zero-shot baseline, but budget for supervision. The cost of a few hundred labels is almost certainly lower than the cost of deploying a model that fails on your minority classes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Small label budgets go further than you&#39;d expect.&lt;/strong&gt; Targeting 5 to 10% of a reasonable dataset often recovers most of the achievable performance. Full labeling is rarely the right target.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Choose adaptation based on where your hard classes live.&lt;/strong&gt; If your difficult classes are spectrally or visually ambiguous, prefer full fine-tuning. If they&#39;re well-separated, LoRA is fine.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Watch per-class metrics in low-data regimes.&lt;/strong&gt; Aggregate mIoU can hide real harm to the classes you probably care about most.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The broader claim is this: for specialized imagery, the choice to prompt rather than fine-tune has mostly been inherited from a different deployment context. It deserves to be re-examined.&lt;/p&gt;
&lt;figure style=&quot;max-width: 500px; margin-left: auto; margin-right: auto;&quot;&gt;
&lt;img src=&quot;https://kethavath.com/assets/cloudprompts/samples.png&quot; alt=&quot;Qualitative segmentation samples&quot;&gt;
&lt;figcaption&gt;Qualitative segmentation results across three test samples.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The code, and models are available on &lt;a href=&quot;https://github.com/uga-gaim/2026_CVPRW_CloudPrompts&quot;&gt;GitHub&lt;/a&gt; and &lt;a href=&quot;https://huggingface.co/collections/uga-gaim/2026-cvprw-cloudprompts&quot;&gt;HuggingFace&lt;/a&gt;. This work was done at the University of Georgia&#39;s &lt;a href=&quot;https://uga-gaim.github.io/&quot;&gt;Geoinformatics and AI Modeling Lab&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
</feed>