<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>👨🏽‍💻 Blog on Amir Kiani</title><link>https://amirkiani.xyz/posts/</link><description>Amir Kiani (👨🏽‍💻 Blog)</description><language>en-us</language><lastBuildDate>Sun, 03 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://amirkiani.xyz/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>We are all chefs now</title><link>https://amirkiani.xyz/posts/we-are-all-chefs-now/</link><pubDate>Sun, 03 May 2026 00:00:00 +0000</pubDate><guid>https://amirkiani.xyz/posts/we-are-all-chefs-now/</guid><description>&lt;p>One of my not-so-secret fantasies has been to become a professional chef. I&amp;rsquo;m sure that I&amp;rsquo;m not alone :)&lt;/p>
&lt;p>Cooking beautifully ties science, art, and culture together. I learned to cook from my mom, who is an incredible chef. Having a lovingly slow-cooked meal with great ingredients can easily become among a person&amp;rsquo;s most cherished life experiences.&lt;/p>
&lt;p>Lately I&amp;rsquo;ve been thinking that making software has become much closer to making meals. The similarities are hard to miss:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Anyone can cook.&lt;/strong> Not everyone will open a restaurant, but most people can put a nutritious meal on the table. Claude Code has done something similar for software. The barrier to producing something that works is now astonishingly low.&lt;/li>
&lt;li>&lt;strong>Ingredients matter more than technique.&lt;/strong> A great chef with bad fish makes a bad dinner. In software, the model you choose, the problem you pick, and whether anyone actually wants what you&amp;rsquo;re making will determine the outcome far more than how clever your implementation is.&lt;/li>
&lt;li>&lt;strong>It&amp;rsquo;s now about the craft.&lt;/strong> For years, software engineering was mostly about learning frameworks and writing code. Most of our time went to &lt;em>how&lt;/em> to build, not &lt;em>what&lt;/em> to build. That has flipped. You can spend months with a big team building the wrong thing, and no amount of AGI will save you. I&amp;rsquo;ve watched it happen. You can also have a small, beautiful idea, use AI, and ship something that changes lives in a week.&lt;/li>
&lt;/ul>
&lt;p>So yes, in many ways, I&amp;rsquo;m now a chef. The interesting question isn&amp;rsquo;t whether I can cook. It&amp;rsquo;s what I want to cook.&lt;/p>
&lt;p>A lovingly cooked gourmet meal, or McDonald&amp;rsquo;s?&lt;/p></description></item><item><title>What I learned trying to build a VC-backed startup</title><link>https://amirkiani.xyz/posts/startup-lessons/</link><pubDate>Mon, 22 Sep 2025 00:00:00 +0000</pubDate><guid>https://amirkiani.xyz/posts/startup-lessons/</guid><description>&lt;p>&lt;img src="https://amirkiani.xyz/images/posts/startup-lessons/amir-marta-minujin.jpg" alt="alt_text" title="Marta Minujin installation">&lt;/p>
&lt;figcaption>Rana took this photo of me jumping around inside Marta Minujin's mattress-padded room &lt;a href="https://copenhagencontemporary.org/en/marta-minujin/" target="_new">installation&lt;/a> at Copenhagen Contemporary. It felt like an apt description for this post.&lt;/figcaption>
&lt;p>Life is a series of experiments.&lt;/p>
&lt;p>After 14 years of working for others, building my own company was a long-awaited experiment that had always piqued my interest.&lt;/p>
&lt;p>This is a brief summary of what I&amp;rsquo;ve learned trying this idea.&lt;/p>
&lt;p>It’s mostly intended as a capsule for my future self to read. But it could also serve as a source of information for others who might also be thinking about starting a company.&lt;/p>
&lt;p>It is of course a sample of one and “your mileage may vary”.&lt;/p>
&lt;h2 id="some-important-context" >Some important context
&lt;span>
&lt;a href="#some-important-context">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>There are a million ways to do a startup.&lt;/p>
&lt;p>Different variables, such as whether you raise funding or not, are a solo founder or team up with others, or work on a consumer product or enterprise, would have a significant impact on how things might pan out. The experiment I ran had the following specific characteristics:&lt;/p>
&lt;ul>
&lt;li>I built a &lt;em>venture-backed&lt;/em> startup in Silicon Valley in 2024&lt;/li>
&lt;li>I chose a rather complex domain: the Health AI industry&lt;/li>
&lt;li>I tried to team up with multiple co-founders, but for a wide range of reasons -- usually timing or misaligned values/expectations -- I ended up being a solo founder&lt;/li>
&lt;li>Before the startup, I was in big tech (Google)&lt;/li>
&lt;/ul>
&lt;h2 id="why-i-started-a-company" >Why I started a company
&lt;span>
&lt;a href="#why-i-started-a-company">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>The most important lesson that I discovered through building the company was how much of an exercise in self-discovery and reflection it was.&lt;/p>
&lt;p>Working as an employee, one inherits a well-marked path to walk on. But as a CEO and especially a solo founder, one needs to intentionally and frequently choose/define the foundational aspects of the business.&lt;/p>
&lt;p>It starts with high-level and shiny things like the company&amp;rsquo;s mission &amp;amp; vision and goes all the way to the decision-making process, choice of investors, commercial strategy, and even boring details such as which service to use for payroll, email, or the marketing website.&lt;/p>
&lt;p>Every decision has a major impact on your work life. And you, &lt;em>no longer your employer&lt;/em>, are responsible for the outcome of these decisions.&lt;/p>
&lt;p>Over and over, one is forced to define who one actually is.&lt;/p>
&lt;p>For me, some of the reasons for starting the company were:&lt;/p>
&lt;ul>
&lt;li>I wanted to create my own culture: one that was aligned with my values.&lt;/li>
&lt;li>I wanted to work on something that genuinely made the world a better place.&lt;sup id="fnref:1">&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref">1&lt;/a>&lt;/sup>&lt;/li>
&lt;li>I wanted to leverage my generalist nature (engineering, research, product), which was rather under-utilized when I was working in a specialized role at a big tech company.&lt;/li>
&lt;li>I wanted to experiment with building an entire company with the help of AI. There were a lot of &lt;a href="https://www.forbes.com/sites/markminevich/2025/08/20/the-billion-dollar-company-of-one-is-coming-faster-than-you-think/">wild claims&lt;/a> about this being a possibility. I had cautiously assumed that everything was &lt;em>a lot&lt;/em> easier with AI and that there was a non-zero chance that AI would &lt;em>exponentially&lt;/em> improve through the course of my experiment. If that was true, then having a blank slate to draw on was going to be a massive advantage.&lt;/li>
&lt;li>I was curious to experience being a founder. If you talk to people who know me well, they would very likely tell you that trying a bit of everything -- and being a master of none -- is very much a part of my character 🙂&lt;/li>
&lt;li>It felt like an opportunity loss &lt;em>not to&lt;/em> try doing a VC-backed startup in San Francisco, the mecca of startups, while I was relatively young and able.&lt;/li>
&lt;/ul>
&lt;h2 id="what-i-learned-about-myself" >What I learned about myself
&lt;span>
&lt;a href="#what-i-learned-about-myself">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>I don&amp;rsquo;t think I&amp;rsquo;ve gained more insight about myself in any other experiment of similar length. And the majority of my experience was not by any means the exhilarating and fun ride that is portrayed by most. It was painful and confusing at times. But it was also incredibly satisfying at other times.&lt;/p>
&lt;p>It was worth it.&lt;/p>
&lt;p>My learning journey started right before I even embarked on the path. Given the weight of this decision to start a company, I decided that I needed to clear my head before committing to such a major decision. To do this, I did a 10-day Vipassana. A year later, I still feel the magnitude of that experience. 10 days of watching my &amp;ldquo;monkey brain&amp;rdquo; going from one thought to another. From one emotion to another. There are a lot of lectures about Buddhism in Vipassana sessions. Over and over you are reminded that cravings and aversions are the source of human suffering. And the point of life is to escape this cycle of suffering.&lt;/p>
&lt;p>I came out of the Vipassana a different person. I was in fact seriously considering that I &lt;strong>should not&lt;/strong> start a venture-backed company 🙂 The whole exercise felt like an ego-centric pursuit that I would easily cling to and cause immense suffering. It felt like a distraction from my -- unexpectedly started -- spiritual journey.&lt;/p>
&lt;p>For better or worse, my enlightenment episode wasn’t able to stop me from pursuing this journey.&lt;/p>
&lt;p>In retrospect, I was pretty disoriented after this intense meditation course and should have given myself much more time to reintegrate into society. I think Vipassana was a very potent, though incredibly helpful, shock to my system.&lt;sup id="fnref:2">&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref">2&lt;/a>&lt;/sup>&lt;/p>
&lt;h2 id="at-times-ai-gave-me-an-illusion-of-progress" >At times, AI gave me an illusion of progress
&lt;span>
&lt;a href="#at-times-ai-gave-me-an-illusion-of-progress">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>It&amp;rsquo;s hard to make something out of thin air. Especially when you are alone in thinking about it.&lt;/p>
&lt;p>At first, AI was a major band-aid for me in these situations. Every time I&amp;rsquo;d feel stuck on something or was daunted by a major decision, I&amp;rsquo;d turn to my three trusted friends: Claude, ChatGPT, and Gemini 🤖 These interactions were honestly helpful on many occasions. But I do think there was a lack of durability in some of what came out of this experience. At first glance, the slides, wireframes, apps, and business ideas seemed very promising and professional. But like the actual taste of a McDonald’s sandwich that looks incredible on a billboard, the reality did not match my illusions.&lt;/p>
&lt;p>In addition, the fact that it was so easy to spend days, weeks, months mindlessly &amp;ldquo;building&amp;rdquo; prevented me from doing some important reality checks. There is something really addictive about the AI coding process.&lt;sup id="fnref:3">&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref">3&lt;/a>&lt;/sup> Others have &lt;a href="https://ideia.me/programming-is-a-drug">written&lt;/a> about this.&lt;/p>
&lt;p>Towards the end, I intentionally switched to doing a lot of thinking by writing on paper and talking to real users. But the low barrier of using AI as a crutch was always hard to avoid.&lt;/p>
&lt;h2 id="i-didnt-find-something-that-would-work-for-me" >I didn&amp;rsquo;t find something that would work for me
&lt;span>
&lt;a href="#i-didnt-find-something-that-would-work-for-me">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>The harsh reality is that I failed at my startup (it’s totally ok though! 🙂).&lt;/p>
&lt;p>Aside from the well-known challenges of building complex and high-stakes solutions with GenAI, I couldn&amp;rsquo;t find a path that simultaneously:&lt;/p>
&lt;ul>
&lt;li>was a venture-scale business&lt;/li>
&lt;li>was something that I was deeply interested to work on for a long time&lt;/li>
&lt;li>had high enough traction to justify continuing with&lt;/li>
&lt;/ul>
&lt;p>I &lt;a href="https://amirkiani.xyz/posts/yari-care-nav">wrote&lt;/a> about some of the things I tried in the past. I tried a few other options after the pivot as well. But at some point I decided that I had enough data to make the decision to stop the experiment.&lt;/p>
&lt;p>In case it is useful for you, I&amp;rsquo;ve also &lt;a href="https://github.com/akiani/timeline-oss">open sourced&lt;/a> the iOS app that I ended up shipping which got modest traction.&lt;/p>
&lt;h2 id="i-miss-working-with-other-smart-people" >I miss working with other smart people
&lt;span>
&lt;a href="#i-miss-working-with-other-smart-people">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>If you&amp;rsquo;re a strong technical person in the Bay Area working in big tech, it&amp;rsquo;s pretty hard to justify quitting your job for a startup whose entire bank account is less than your yearly income. One might say that I did this, so there must be others. But I was not able to convince others to quit their job to join me. Some of this was probably a timing issue. Another reason was, however, that I wasn&amp;rsquo;t able to prove to these folks that they would have remotely similar expected value of gains by working in the domain that I had picked.&lt;/p>
&lt;p>Unsurprisingly, I couldn&amp;rsquo;t prove this to myself either.&lt;/p>
&lt;p>What I do miss now is deep collaboration with some of the folks that I worked with in the past. Not just engineers but also researchers, designers, clinicians, and product managers.&lt;/p>
&lt;p>I of course do not miss the many levels of middle management, hours of unnecessary meetings, and the strange politics of big tech 😉&lt;/p>
&lt;h2 id="i-found-out-about-my-hidden-values" >I found out about my &lt;em>hidden&lt;/em> values
&lt;span>
&lt;a href="#i-found-out-about-my-hidden-values">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>When I was contemplating getting off the VC-backed startup train, I benefited from the help of a wonderful friend and mentor.&lt;/p>
&lt;p>One of the exercises that I went through was to re-read one of my favorite books, &lt;a href="https://designingyour.life/books-designing-life-original-book/">Designing Your Life&lt;/a> and go through the exercise of &lt;a href="https://designingyour.life/insights/the-magic-of-odysseys-prototyping-your-future-with-designing-your-life/">Odyssey Planning&lt;/a>. I tried to do this without thinking about what happens to the company. What I realized going through the exercise was that no future came up that pointed to an aspiration to be a CEO, a founder, or a &lt;em>sole&lt;/em> decision maker for others. Instead what did come up were aspirations for balance, authenticity, health, presence, and to have a rich life. Nothing came up that required massive amounts of money or control.&lt;/p>
&lt;p>This exercise was a major catalyst for making my mind up about ending the startup experiment. &lt;sup id="fnref:4">&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref">4&lt;/a>&lt;/sup>&lt;/p>
&lt;h2 id="i-feel-so-incredibly-lucky-and-spoiled" >I feel so incredibly lucky (and spoiled)
&lt;span>
&lt;a href="#i-feel-so-incredibly-lucky-and-spoiled">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>On one hand, I am -- yet another -- failed startup founder 🙃&lt;/p>
&lt;p>I didn’t go &lt;em>all in&lt;/em> to make my company successful. I didn’t make a dent in the universe. I earned a fraction of what I would have earned had I simply stayed in my past job. I didn’t even spend most of the funding that I raised…&lt;/p>
&lt;p>On the other hand, I feel &lt;em>incredibly&lt;/em> lucky to have spent the last year:&lt;/p>
&lt;ul>
&lt;li>Intentionally building an entire company from scratch&lt;/li>
&lt;li>Testing many assumptions&lt;/li>
&lt;li>Getting experience in fundraising and interacting with investors -- from micro-VCs to the largest investors in Silicon Valley&lt;/li>
&lt;li>Learning a lot more about hiring and people management&lt;/li>
&lt;li>Learning more about the healthcare and AI space&lt;/li>
&lt;li>Getting back to my roots as a researcher and software engineer by doing months of focused development work&lt;/li>
&lt;li>Reading tens of books&lt;/li>
&lt;li>Meditating for hundreds of hours&lt;/li>
&lt;li>Running tons of miles&lt;/li>
&lt;li>Eating healthy homemade food for most days and being close to my wife and dog&lt;/li>
&lt;/ul>
&lt;p>And I got paid doing all that.&lt;/p>
&lt;p>I genuinely don’t regret this failure and have a pretty good idea of the next life experiments that I want to run!&lt;/p>
&lt;p>&lt;br/>&lt;br/>Lastly, and most importantly, I am so grateful for fellow friends, family, employees, advisors, and investors who helped me spend this year in the startup land.&lt;br>
&lt;br/>
&lt;br/>&lt;/p>
&lt;div class="footnotes" role="doc-endnotes">
&lt;hr>
&lt;ol>
&lt;li id="fn:1">
&lt;p>Turns out this is really hard to do if you are building a VC-backed company. I strongly recommend bootstrapping and not raising funds if this is truly your intention. It is also extremely hard to compete in a market where others don’t play by similar rules.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;li id="fn:2">
&lt;p>In case you are thinking about doing a Vipassana and want to hear about my experience, I have a lot more to share!&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;li id="fn:3">
&lt;p>During a few months of coding, I used Cursor so much that I ended up in the top 1% users in San Francisco and got invited to their office for a meetup 😂&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;li id="fn:4">
&lt;p>It might be worth noting that I ended up not spending most of the raised funds and returned the remaining amount to our VC.&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;/ol>
&lt;/div></description></item><item><title>The challenges of building LLM-powered clinical care navigation apps</title><link>https://amirkiani.xyz/posts/yari-care-nav/</link><pubDate>Fri, 06 Jun 2025 12:00:00 -0700</pubDate><guid>https://amirkiani.xyz/posts/yari-care-nav/</guid><description>&lt;p>&lt;strong>TL;DR&lt;/strong>: I spent 9 months building LLM-powered apps to help cancer patients navigate clinical care—chatbots with EHR integrations, appointment prep tools, and iOS apps (that Apple kept rejecting). I found real use cases and built working prototypes, but I discovered that LLMs aren&amp;rsquo;t ready for safe clinical use by patients. The core structural issues—context building, grounding, and usefulness—won&amp;rsquo;t be solved by the next model release. This is my personal experience; LLMs remain amazing for other use cases like prototyping the very applications illustrated in this post.&lt;/p>
&lt;br/>
&lt;hr/>
&lt;br/>
&lt;p>About 9 months ago, I set out to test a hypothesis:&lt;/p>
&lt;blockquote>
&lt;p>Given that people &lt;a href="https://www.usertesting.com/resources/reports/consumer-perceptions-ai-healthcare">already use&lt;/a> general-purpose chatbots for health concerns, one could theoretically build a business around a specialized, LLM-powered app for clinical care navigation.&lt;/p>&lt;/blockquote>
&lt;p>And presumably, this specialized approach could lead to higher quality, accuracy and usefulness compared to the off-the-shelf general solutions.&lt;/p>
&lt;p>I chose an area that was very close to my heart: the problem of navigating the tough journey of cancer care.&lt;/p>
&lt;p>This writeup captures—in perhaps an excessive amount of detail—my learnings from this nine month journey and the challenges I discovered.&lt;/p>
&lt;h1 id="the-use-cases" >The Use Cases
&lt;span>
&lt;a href="#the-use-cases">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h1>&lt;p>I spent a few months researching the most common user needs in this space to find a problem-solution fit. I talked to physicians, nurses, specialists, patient advocates, and of course, many patients.&lt;/p>
&lt;p>My research led me to several key use cases that seemed feasible, desirable, and viable as product directions:&lt;/p>
&lt;ol>
&lt;li>Understanding one&amp;rsquo;s diagnosis and medical records&lt;/li>
&lt;li>Figuring out what questions to ask providers and preparing for visits&lt;/li>
&lt;li>Dealing with treatment side effects&lt;/li>
&lt;li>Researching alternative treatment paths&lt;/li>
&lt;li>Patient self-advocacy and empowering caregivers&lt;/li>
&lt;li>Capturing and debriefing interactions with providers&lt;/li>
&lt;/ol>
&lt;h1 id="the-prototypes" >The Prototypes
&lt;span>
&lt;a href="#the-prototypes">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h1>&lt;p>One of the most fun parts of this journey has been the ability to go back to coding after years of &amp;ldquo;product managing.&amp;rdquo; The tools I used (mostly Cursor + Claude) kept getting better almost monthly. And I grew with them. I&amp;rsquo;ve developed a pretty sophisticated symbiosis with AI-coding toolkits that almost warrants its own article—though there&amp;rsquo;s already a lot being written about this, so I&amp;rsquo;m not sure I have unique insights to add. Being able to turn ideas into real working prototypes in a matter of hours/days was very transformative in my explorations.&lt;sup id="fnref:1">&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref">1&lt;/a>&lt;/sup>&lt;/p>
&lt;p>Here&amp;rsquo;s what we built:&lt;sup id="fnref:2">&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref">2&lt;/a>&lt;/sup>&lt;/p>
&lt;h2 id="1-a-series-of-mini-apps" >1. A Series of Mini Apps
&lt;span>
&lt;a href="#1-a-series-of-mini-apps">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;ul>
&lt;li>One that facilitated talking to a chatbot specially designed to pull data from authoritative sources (e.g. cancer.org) and annotating complex medical terms in its responses to meet the average user&amp;rsquo;s medical literacy level&lt;/li>
&lt;li>One that facilitated talking to a PDF export (thousands of pages) of a patient&amp;rsquo;s MyChart data while linking to the specific parts of the PDF file to validate citations in the model&amp;rsquo;s responses&lt;/li>
&lt;li>One that helped people get a &amp;ldquo;second opinion&amp;rdquo; on a question using Deep Research features from different platforms while passing their entire health records through a MyChart PDF export&lt;/li>
&lt;/ul>
&lt;div style="text-align: center; margin: 2rem 0;">
&lt;img src="https://amirkiani.xyz/images/posts/yari-care-nav/small-app.png" alt="Mini App Screenshot" style="width: 80%; border-radius: 12px; box-shadow: 0 4px 8px rgba(0,0,0,0.1); max-width: 500px;">
&lt;p style="margin-top: 0.5rem; font-style: italic; color: #666; font-size: 0.9rem;">One of the mini apps (chatbot) that annotated medical terms and cited authoritative sources.&lt;/p>
&lt;/div>
&lt;hr>
&lt;h2 id="2-the-first-major-web-app" >2. The first major Web App
&lt;span>
&lt;a href="#2-the-first-major-web-app">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>After building the series of mini apps, I ended up building a full web app that allowed people to connect to their provider&amp;rsquo;s EHR instance through a FHIR interface so they could:&lt;/p>
&lt;ul>
&lt;li>Generate an overview of their care and summarize their health records with an LLM&lt;/li>
&lt;li>Generate insights about their journey that they could use as conversation starters with an LLM&lt;/li>
&lt;li>Ask questions about their history&lt;/li>
&lt;li>Create a list of notes and insights to discuss with their providers&lt;/li>
&lt;/ul>
&lt;div class="carousel-container" id="basic-webapp-carousel">
&lt;div class="carousel-wrapper">
&lt;div class="carousel-track" data-carousel="basic-webapp-carousel">&lt;div class="carousel-slide active">
&lt;img src="https://amirkiani.xyz/images/posts/yari-care-nav/basic-web-app/2.png" alt="Screenshot 1" loading="lazy">&lt;/div>&lt;div class="carousel-slide">
&lt;img src="https://amirkiani.xyz/images/posts/yari-care-nav/basic-web-app/1.png" alt="Screenshot 2" loading="lazy">&lt;/div>&lt;div class="carousel-slide">
&lt;img src="https://amirkiani.xyz/images/posts/yari-care-nav/basic-web-app/3.png" alt="Screenshot 3" loading="lazy">&lt;/div>&lt;div class="carousel-slide">
&lt;img src="https://amirkiani.xyz/images/posts/yari-care-nav/basic-web-app/4.png" alt="Screenshot 4" loading="lazy">&lt;/div>&lt;/div>&lt;div class="carousel-dots">&lt;button class="carousel-dot active" onclick="currentSlide('basic-webapp-carousel', 1 );">&lt;/button>&lt;button class="carousel-dot" onclick="currentSlide('basic-webapp-carousel', 2 );">&lt;/button>&lt;button class="carousel-dot" onclick="currentSlide('basic-webapp-carousel', 3 );">&lt;/button>&lt;button class="carousel-dot" onclick="currentSlide('basic-webapp-carousel', 4 );">&lt;/button>&lt;/div>&lt;/div>
&lt;/div>
&lt;script>
if (typeof window.carouselInitialized === 'undefined') {
window.carouselSlides = {};
function preloadCarouselImages(carouselId) {
const container = document.getElementById(carouselId);
const images = container.querySelectorAll('.carousel-slide img');
images.forEach(img => {
if (!img.complete) {
const preloadImg = new Image();
preloadImg.src = img.src;
}
});
}
function moveCarousel(carouselId, direction) {
const container = document.getElementById(carouselId);
const slides = container.querySelectorAll('.carousel-slide');
const dots = container.querySelectorAll('.carousel-dot');
if (!window.carouselSlides[carouselId]) {
window.carouselSlides[carouselId] = 0;
}
slides[window.carouselSlides[carouselId]].classList.remove('active');
if (dots[window.carouselSlides[carouselId]]) {
dots[window.carouselSlides[carouselId]].classList.remove('active');
}
window.carouselSlides[carouselId] += direction;
if (window.carouselSlides[carouselId] >= slides.length) {
window.carouselSlides[carouselId] = 0;
} else if (window.carouselSlides[carouselId] &lt; 0) {
window.carouselSlides[carouselId] = slides.length - 1;
}
slides[window.carouselSlides[carouselId]].classList.add('active');
if (dots[window.carouselSlides[carouselId]]) {
dots[window.carouselSlides[carouselId]].classList.add('active');
}
}
function currentSlide(carouselId, slideNumber) {
const container = document.getElementById(carouselId);
const slides = container.querySelectorAll('.carousel-slide');
const dots = container.querySelectorAll('.carousel-dot');
slides.forEach(slide => slide.classList.remove('active'));
dots.forEach(dot => dot.classList.remove('active'));
window.carouselSlides[carouselId] = slideNumber - 1;
slides[window.carouselSlides[carouselId]].classList.add('active');
dots[window.carouselSlides[carouselId]].classList.add('active');
}
function initializeTouch() {
document.querySelectorAll('.carousel-track').forEach(track => {
const carouselId = track.dataset.carousel;
let startX = 0;
let startY = 0;
let distX = 0;
let distY = 0;
track.addEventListener('touchstart', (e) => {
const touch = e.touches[0];
startX = touch.clientX;
startY = touch.clientY;
});
track.addEventListener('touchmove', (e) => {
e.preventDefault();
});
track.addEventListener('touchend', (e) => {
const touch = e.changedTouches[0];
distX = touch.clientX - startX;
distY = touch.clientY - startY;
if (Math.abs(distX) > Math.abs(distY) &amp;&amp; Math.abs(distX) > 50) {
if (distX > 0) {
moveCarousel(carouselId, -1);
} else {
moveCarousel(carouselId, 1);
}
}
});
});
}
function initializeCarousels() {
initializeTouch();
document.querySelectorAll('.carousel-container').forEach(container => {
preloadCarouselImages(container.id);
});
}
if (document.readyState === 'loading') {
document.addEventListener('DOMContentLoaded', initializeCarousels);
} else {
initializeCarousels();
}
window.carouselInitialized = true;
}
&lt;/script>
&lt;p style="margin-top: 0.5rem; font-style: italic; color: #666; font-size: 0.9rem; text-align: center;">Screenshots from the first web app that allowed users to connect with their EHR system and chat with the results&lt;/p>
&lt;hr>
&lt;h2 id="3-an-iphone-app-the-apple-saga" >3. An iPhone App (The Apple Saga)
&lt;span>
&lt;a href="#3-an-iphone-app-the-apple-saga">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;div style="text-align: center; margin: 2rem 0;">
&lt;video width="60%" controls autoplay muted loop playsinline style="max-width: 600px; border-radius: 12px;">
&lt;source src="https://amirkiani.xyz/images/posts/yari-care-nav/yari-ios-demo.mp4" type="video/mp4">
&lt;p>Your browser doesn't support video playback. &lt;a href="https://amirkiani.xyz/images/posts/yari-care-nav/yari-ios-demo.mp4">Download the video&lt;/a> to view it.&lt;/p>
&lt;/video>
&lt;p style="margin-top: 0.5rem; font-style: italic; color: #666; font-size: 0.9rem;">One of the native iOS apps that I built to ground an LLM on the user's HealthKit medical records&lt;/p>
&lt;/div>
&lt;p>Most of my time was spent building a native iOS app (in Swift) that did everything above but also let users connect the app to their medical records through &lt;a href="https://developer.apple.com/documentation/healthkit/accessing-health-records">Apple&amp;rsquo;s HealthKit&lt;/a>. Why go the iOS route? I discovered that fetching user data through a third-party FHIR data acquisition service was prohibitively expensive. It also required users to trust not only us, but also our third-party service with their sensitive health information. Meanwhile, Apple had built a solution with the most reach in the US, Canada, and UK for fetching patient records—and it was FREE. So I figured it was worth building a snappy native iOS app on top of this functionality. I added features to record appointments, transcribe them, and discuss them with an LLM grounded in both the appointment and the user&amp;rsquo;s health history.&lt;/p>
&lt;p>Apple never approved the app. I never got a straight answer as to why, but I suspect it was because they did not want to allow off-device processing of users&amp;rsquo; health data, even with clear explanations and user consent.&lt;/p>
&lt;p>I even built a much simpler version that did a single LLM call to generate the user&amp;rsquo;s health journey (&lt;a href="https://amirkiani.xyz/images/posts/yari-care-nav/yari-timeline-ios-narrated-demo.mp4">demo&lt;/a>). I labeled it with all sorts of disclaimers and drafted a &lt;a href="https://amirkiani.xyz/data/posts/yari-care-nav/policy.md">very privacy-centered Privacy Policy&lt;/a>. The policy indicated that data would only be processed on a HIPAA-compliant backend and never stored beyond processing time on Google&amp;rsquo;s Vertex AI, which we had a HIPAA BAA (Business Associate Agreement) with. My apps are still in review and have been rejected 10+ times. It&amp;rsquo;s starting to feel like Apple is simply stalling me. I still don&amp;rsquo;t know whether I can run a remote HIPAA-compliant LLM on HealthKit medical records data even without storing any patient data. Given what I will explain later in this article, while this issue was a catalyst, it was not the root cause of my decision to pause this direction.&lt;/p>
&lt;hr>
&lt;h2 id="4-another-web-app" >4. Another Web App
&lt;span>
&lt;a href="#4-another-web-app">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;div style="text-align: center; margin: 2rem 0;">
&lt;video width="100%" controls autoplay muted loop playsinline style="border-radius: 12px;">
&lt;source src="https://amirkiani.xyz/images/posts/yari-care-nav/yari-web-demo.mp4" type="video/mp4">
&lt;p>Your browser doesn't support video playback. &lt;a href="https://amirkiani.xyz/images/posts/yari-care-nav/yari-web-demo.mp4">Download the video&lt;/a> to view it.&lt;/p>
&lt;/video>
&lt;p style="margin-top: 0.5rem; font-style: italic; color: #666; font-size: 0.9rem;">The final web app built as the final attempt to test the original hypothesis&lt;/p>
&lt;/div>
&lt;p>In response to Apple&amp;rsquo;s App Store rejections, I recreated the iOS app as a web app. By this time, Claude 4 had come out and I knew what I wanted to make, so this iteration went by at least 5 times faster than the iOS app. I also added a way to detect health &amp;ldquo;memories&amp;rdquo; from chatbot conversations (using function calls) to reduce the barrier of asking users to trust us with all their medical records. I later added the ability to bring in MyChart data to this app too.&lt;/p>
&lt;p>I thought about open-sourcing all these apps. But their reliance on ever-changing web services makes them quickly obsolete.&lt;/p>
&lt;h1 id="reflecting-on-the-limitations" >Reflecting on the Limitations
&lt;span>
&lt;a href="#reflecting-on-the-limitations">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h1>&lt;p>While any attempt to summarize the limitations of LLMs is doomed to become obsolete within weeks due to their seemingly exponential growth, I&amp;rsquo;d like to think that I gave this idea an honest try.&lt;/p>
&lt;p>If I had to sum up my learnings in one sentence: &lt;em>LLMs aren&amp;rsquo;t safe for direct clinical use by patients—not yet.&lt;/em>&lt;/p>
&lt;p>While I had different product iterations ready for public testing, my conscience didn&amp;rsquo;t allow me to put them up for people to freely try. I would literally dream of someone using them, getting wrong results, and feeling responsible for the suffering they would experience. Maybe this mentality makes me unqualified for working in healthcare.&lt;/p>
&lt;h2 id="the-root-causes" >The Root Causes
&lt;span>
&lt;a href="#the-root-causes">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>As usual, there are a million reasons why a startup idea might fail. It&amp;rsquo;s expected for startups to pivot a few times until they find something that works. Much of it is timing, founder problems, and systemic external issues (like the mess that is the US Healthcare system). But I have enough data to say, with 90%+ confidence, that I cannot make this idea work within a reasonable timeframe. This would be true even if every other factor was held constant. Sure, it might work if LLMs became hallucination-free and perfectly steerable tomorrow. But knowing a bit about the technology behind them, I&amp;rsquo;m not sure this will happen anytime soon (though I&amp;rsquo;ve been working on and waiting for that day for two years and counting).&lt;/p>
&lt;p>The way I break down the key failure reasons in my head:&lt;/p>
&lt;ol>
&lt;li>Impossible context building&lt;/li>
&lt;li>Impossible grounding&lt;/li>
&lt;li>Limited usefulness&lt;/li>
&lt;/ol>
&lt;p>The rest of this writeup explains these three core challenges.&lt;/p>
&lt;h3 id="1-impossible-context-building" >1. Impossible Context Building
&lt;span>
&lt;a href="#1-impossible-context-building">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h3>&lt;p>The promise that patients should own and have access to their personal records dates back many years. There are countless trials and startups trying to make this dream a reality. Some have failed. Some have been acquired by big-tech companies (e.g., &lt;a href="https://techcrunch.com/2016/08/22/apple-acquired-gliimpse-a-personal-health-data-startup/">Gliimpse&lt;/a> and &lt;a href="https://ir.invitae.com/news-and-events/press-releases/press-release-details/2021/Invitae-to-Acquire-Ciitizen-to-Strengthen-its-Patient-Consented-Health-Data-Platform-to-Improve-Personal-Outcomes-and-Global-Research/default.aspx">Ciitizen&lt;/a>, both built by the same founder who is now creating &lt;a href="https://selfiie.com">Selfiie&lt;/a> 🤨).&lt;/p>
&lt;p>But what&amp;rsquo;s clear is that no single person I know has access to all their data. Yes, you can log into MyChart and get an &lt;em>intentionally convoluted&lt;/em> dump of your records—a mix of PDFs, structured data, and clinical notes that would challenge even the most sophisticated parser. But that&amp;rsquo;s just one health system. What about the urgent care visit from five years ago? The specialist you saw in another state? The dental records that might be relevant to your cancer treatment?&lt;/p>
&lt;p>The fragmentation is staggering. Each provider uses different systems, different formats, different standards. Even within the same hospital network, data might be siloed across departments. The promise of interoperability remains largely that—a promise. And without comprehensive context, any LLM system is working with a dangerously incomplete picture.&lt;sup id="fnref:3">&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref">3&lt;/a>&lt;/sup>&lt;/p>
&lt;h3 id="2-impossible-grounding" >2. Impossible Grounding
&lt;span>
&lt;a href="#2-impossible-grounding">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h3>&lt;p>I tried almost every tool in the toolbox: RAG (Retrieval-Augmented Generation), Function Calls, MCPs (Model Context Protocols), passing everything to the models with large context windows. There&amp;rsquo;s always some part of the system that fails. And the errors aren&amp;rsquo;t recoverable or transparent, which makes them really hard to reliably avoid.&lt;br>
It&amp;rsquo;s not like traditional systems where you actually get an error. The model just makes a best &amp;ldquo;guess&amp;rdquo; and &lt;em>confidently&lt;/em> gives an inaccurate response. Or it gives a response that exposes the prompt and inner workings of the system. I see this frequently with ANY system that has an open chatbot interface.&lt;br>
While this issue is forgivable in low-stakes use cases, when users ask a system for information that could lead to life-or-death outcomes, this behavior breaks their trust. Imagine talking to your nurse and at some point in a deep conversation, they just &lt;em>confidently&lt;/em> stop making sense. Your relationship with that provider is over.&lt;/p>
&lt;p>And yes, I know the theory that as long as your solution is better than the alternative human option, it&amp;rsquo;s alright. But in a clinical setting, that idea is &lt;em>dangerous&lt;/em>. Yes, ChatGPT and Claude already operate on that premise. But at least they&amp;rsquo;re not marketing a health product. If they did, I bet they&amp;rsquo;d be in massive legal trouble. I, however, was setting out to make a health product and don&amp;rsquo;t have the legal team that big tech companies have.&lt;/p>
&lt;h3 id="3-limited-usefulness" >3. Limited Usefulness
&lt;span>
&lt;a href="#3-limited-usefulness">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h3>&lt;p>Even when everything worked in terms of context assembly and grounding, the likelihood of users actually knowing what to do with the model&amp;rsquo;s output was low.&lt;/p>
&lt;p>Let&amp;rsquo;s take the case of hard decisions about &lt;em>surgery versus chemo&lt;/em> for a cancer patient.&lt;/p>
&lt;p>Assume the model knows everything about the patient&amp;rsquo;s &lt;em>medical&lt;/em> history. It&amp;rsquo;s asked whether the patient should do chemo or surgery. First, this question is medical advice. So from a legal perspective, we should punt.&lt;/p>
&lt;p>If we punt, the patient assumes the system is useless because they could just use ChatGPT or Claude and get an answer.&lt;/p>
&lt;p>So we shouldn&amp;rsquo;t punt? The best we can do is what ChatGPT does: give a best-effort answer sandwiched in disclaimers. My argument is that for hard questions, any answer is tricky.&lt;/p>
&lt;p>The trickiness arises from the fact that &lt;strong>there is no easy right answer&lt;/strong>.&lt;/p>
&lt;p>I got to this conclusion when researching my late aunt&amp;rsquo;s case. She was diagnosed with Stage 4 pancreatic cancer. Her family had to decide if she should get chemo or join an immunotherapy clinical trial. The answer to this question is so multifaceted. The LLM&amp;rsquo;s eagerness to please makes it very likely to start spewing information to the user, giving a gigantic list of options, and listing their pros and cons. Sure, that&amp;rsquo;s better than no information, but the truth is Stage 4 pancreatic cancer is &lt;em>barely survivable&lt;/em>. My aunt passed away about a month after her diagnosis even though she was in her sixties.&lt;/p>
&lt;p>In that situation, the patient and caregivers are looking for &lt;em>clarity and less information&lt;/em>. Not more options and content to overwhelm them. The right &amp;ldquo;answer&amp;rdquo; for them was really to spend more quality time with each other rather than endure more pain and suffering through treatment.&lt;/p>
&lt;p>This taught me that the most critical healthcare decisions aren&amp;rsquo;t just clinical—they&amp;rsquo;re deeply human, contextual, and often about quality of life rather than quantity. It&amp;rsquo;s impossible to give a good answer if you just focus on clinical aspects of someone&amp;rsquo;s life. There are moral, religious, and personal aspects that need to be considered. And this takes us back to the first problem: it&amp;rsquo;s impossible to meaningfully know that context at large scale.&lt;/p>
&lt;h1 id="the-business-model-reality-check" >The Business Model Reality Check
&lt;span>
&lt;a href="#the-business-model-reality-check">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h1>&lt;p>Aside from these structural challenges we still haven&amp;rsquo;t even considered answering the fundamental question: how might one make money from this idea?&lt;/p>
&lt;p>The usual business models for direct-to-consumer health tech companies are:&lt;/p>
&lt;ol>
&lt;li>Selling the data&lt;/li>
&lt;li>Showing ads&lt;/li>
&lt;li>Charging for subscriptions&lt;/li>
&lt;li>Partnerships (i.e. B2B2C) with providers, payers or pharma&lt;/li>
&lt;/ol>
&lt;p>Options 1 and 2 are out of the question for me. Health data is different. It&amp;rsquo;s &lt;em>a part of you&lt;/em>.&lt;/p>
&lt;p>Option 3 is the most common path, but faces two major challenges: (a) patients notoriously have low willingness to pay for health tools, and (b) the product needs to be strictly better than alternatives. When we frequently punt on hard questions while ChatGPT and Claude do not, it&amp;rsquo;s tough to justify the subscription cost.&lt;/p>
&lt;p>Option 4 is an area that I also explored, but frankly, due to the lack of feasibility and usefulness I described above, I didn&amp;rsquo;t see a major path for success there either.&lt;/p>
&lt;h1 id="so-now-what" >So, Now What?
&lt;span>
&lt;a href="#so-now-what">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h1>&lt;p>Are there other clinical health problems that LLMs could help with? Absolutely. Does the idea of building a &lt;em>patient-facing clinical cancer care navigation app&lt;/em> make sense for an early-stage startup? Not yet. And my reasoning is based on the &lt;em>structural issues&lt;/em> I&amp;rsquo;ve outlined which I don&amp;rsquo;t think will disappear with the next LLM model release.&lt;/p>
&lt;p>A question that led me to clarity is: &lt;em>How would I now deal with a personal cancer case? How would I use technology for it?&lt;/em>&lt;/p>
&lt;p>My answer, after thinking about this use case for so long, is that I would actually just use off-the-shelf tools. I would upload all the important parts of my data (personalized context assembly) to a Claude Project and run Deep Research agents on my questions. I would record my appointments on my iPhone&amp;rsquo;s Voice Memos and upload them to that Claude project to reference them. I would also use Claude to think about the best questions to ask for my upcoming appointments.&lt;/p>
&lt;p>Yes, it&amp;rsquo;s not HIPAA compliant, but so what? They &lt;a href="https://privacy.anthropic.com/en/articles/10023555-how-do-you-use-personal-data-in-model-training#:~:text=We%20will%20not,in%20to%20training.">say&lt;/a> they&amp;rsquo;re not training on my data, and their claims on respecting data rights are as good as any other company. Their security measures are better than a small startup, even if the startup has gone through a HIPAA checklist and has the documentation for it. And if I&amp;rsquo;m too worried, I can only give them parts of my data that aren&amp;rsquo;t identifiable.&lt;/p>
&lt;p>Is this something that only highly capable users can do? Yes. But my argument is that only those highly capable people can actually process these answers. Just because chainsaws might now cost zero doesn&amp;rsquo;t mean we should give one to everyone.&lt;/p>
&lt;h1 id="whats-next" >What&amp;rsquo;s Next?
&lt;span>
&lt;a href="#whats-next">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h1>&lt;ol>
&lt;li>I wrote this long writeup for myself to make sense of my past 9 months. I would be delighted if it helps someone else or at least helps others start where I left off. There is a major problem-solution fit for using LLMs for navigating complex diseases. The big question for me wasn&amp;rsquo;t whether people should use them (assuming they know what they&amp;rsquo;re doing). The big question was whether someone should build a new company around this concept &lt;em>today&lt;/em>.&lt;/li>
&lt;li>I want to teach the useful things I&amp;rsquo;ve learned to caregivers and patients. If they&amp;rsquo;re at the edge of knowing how to use these tools, I want to make them more productive. An educational module could achieve that with relatively low cost. I might also do some talks at cancer support groups.&lt;sup id="fnref:4">&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref">4&lt;/a>&lt;/sup>&lt;/li>
&lt;li>I&amp;rsquo;ll probably try something else in the oncology space or explore the non-clinical aspects of care navigation (e.g. financial navigation). But I need to do more market research first.&lt;/li>
&lt;/ol>
&lt;p>Would I do everything the same if I could go back 9 months? Probably. How would I do it differently? I would be nicer to myself. I had picked a very complex problem. I should have had more grace for my struggles in navigating it.&lt;/p>
&lt;p>I hope you found this writeup useful.&lt;/p>
&lt;div class="footnotes" role="doc-endnotes">
&lt;hr>
&lt;ol>
&lt;li id="fn:1">
&lt;p>As a former PM, I think we can now replace much of our traditional PRDs with AI-coded prototypes paired with clear evaluation criteria and initial test sets.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;li id="fn:2">
&lt;p>I would not have been able to do the work I mention in this article without the generous and heartfelt help from a group of incredible collaborators. While this list is incomplete, I would like to thank Daniel Matiaudes (my design partner), Joan Venticinque, Dr. Jeanne Shen, Dr. Matthew Lugren, Srecko Dimitrijevic, Jenny Rizk, Jake Knapp, Jeanette Mellinger, John Zeratsky, Eli Blee-Goldman, Dr. Howard Kleckner, Rob Tufel, Thushan Amarasiriwardena, Henry Schurkus, Anna Brezhneva, Naxin Wang, Cody Sam, and of course my wife Rana, who was my very patient sounding board over the last 9 months. I also want to cherish the memory of my beautiful and loving aunt who passed away while I was working on this idea.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;li id="fn:3">
&lt;p>You might rightly counter that even doctors do not have the full context for a patient&amp;rsquo;s case. My argument is that doctors—due to their training, specialization and contextual understanding of the local systems they operate in—have a much better ability to figure out missing pieces of information. The first thing that a doctor usually does is ask for the full details of your relevant tests and symptoms. The first thing an LLM does is … well, it eagerly tries to shower you with tokens.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;li id="fn:4">
&lt;p>If you&amp;rsquo;re interested in having me speak at your cancer support group or organization, feel free to reach out! Contact details are on my &lt;a href="https://amirkiani.xyz/about/#contact">about page&lt;/a>.&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;/ol>
&lt;/div></description></item><item><title>On mortality, music and perfection anxiety</title><link>https://amirkiani.xyz/posts/mortality/</link><pubDate>Tue, 14 Jan 2025 00:00:00 +0000</pubDate><guid>https://amirkiani.xyz/posts/mortality/</guid><description>&lt;p>&lt;small>&lt;b>Note:&lt;/b> This post is going to be &lt;em>i&lt;/em>m&lt;em>p&lt;/em>er&lt;em>f&lt;/em>ect in so many ways. And that’s probably ok.&lt;/small>&lt;/p>
&lt;p>&lt;img src="https://amirkiani.xyz/images/posts/mortality/jackson-hole.jpg" alt="alt_text" title="Grand Teton National Park">&lt;/p>
&lt;figcaption>A picture I took during our recent trip to the truly beautiful Grand Teton National Park&lt;/figcaption>
&lt;p>About 2 years ago, after months of dealing with chronic stomach ache and digestive problems, I finally managed to get my gastroenterologist (GI) specialist at UCSF to prescribe me an endoscopy and colonoscopy. Before then, my GI specialist was (understandably) suggesting that I should just take more fiber in my diet and should instead work on reducing my stress levels, which were likely the cause of my psychosomatic symptoms. Given my young age and lack of severe symptoms, there was no reason to be worried, but having tried almost everything on my end, I just felt like something was fundamentally off with my insides.&lt;sup id="fnref:1">&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref">1&lt;/a>&lt;/sup>&lt;/p>
&lt;p>I woke up from my (first) colonoscopy on August, 8, 2022 feeling weirdly calm. The country was still recovering from the COVID pandemic and the hospitals still didn’t allow anyone other than the patients inside. The first face I recognized while laying in the hospital bed was my GI surgeon. In a soft and calm tone she said: “Hi Amir. I am really glad you came in for this. We found &lt;em>a lot of&lt;/em> polyps in your stomach and colon. We removed most of them but you have to come again in six months for a second surgery. You should do regular screening going forward.”&lt;/p>
&lt;p>I felt a mild combination of shock and relief. I was however a bit too high on drugs at that moment for my usual problem solving mode to kick in. My loving wife, who had been waiting nervously for me to get discharged at the hospital parking lot, then picked me up and we rode back home. We now had to wait for a pathologist to analyze the extracted tissue from my body and tell us if it was cancer or not.&lt;/p>
&lt;p>I waited for almost a week with no news. I then decided to call UCSF to inquire about my results. After a brief hold on the call, the nurse told me that my pathology diagnosis was ready &lt;em>2 days&lt;/em> after my operation but was held up to be released by my specialist whose first available appointment was &lt;em>another month away&lt;/em>!&lt;/p>
&lt;p>Completely freaked out by my predicament, I asked for a copy of my pathology report and started Googling my way into understanding it. At the time, ChatGPT was not yet launched – though coincidentally I was part of a small team working on an internal yet less capable version of Gemini at Google (and we had no idea it was such a big deal 🙂)
&lt;br/>&lt;br/>&lt;/p>
&lt;p>&lt;img src="https://amirkiani.xyz/images/posts/mortality/path-report.png" alt="alt_text" title="image_tooltip">&lt;/p>
&lt;figcaption>An excerpt from my first pathology results&lt;/figcaption>
&lt;p>The key word was “&lt;a href="https://www.cancer.org/cancer/diagnosis-staging/tests/biopsy-and-cytology-tests/understanding-your-pathology-report/colon-pathology/colon-polyps-sessile-or-traditional-serrated-adenomas.html">sessile serrated adenoma&lt;/a>”: a &lt;em>precancerous&lt;/em> polyp in the colon that can develop into &lt;em>colorectal cancer&lt;/em>. After my second colonoscopy and a total of 13 polyps – some up to 2 cm in diameter – removed from my colon, I was diagnosed with &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC7429648/">Serrated Polyposis Syndrome&lt;/a> which &lt;a href="https://pubmed.ncbi.nlm.nih.gov/34089849/">recent studies&lt;/a> indicate has a ~20% chance of colorectal cancer (CRC). CRC is the &lt;a href="https://seer.cancer.gov/statfacts/html/common.html">second most fatal&lt;/a> cancer but has a much &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6791134/#:~:text=The%20United%20States%20(US)%20is,%25%20and%2065%25%2C%20respectively.">better survival rate&lt;/a> if caught early.&lt;/p>
&lt;p>So to summarize, my &lt;em>obsessive health anxiety&lt;/em> (and access to great care at UCSF) might have led to me preventing a likely fatal disease in my thirties.&lt;/p>
&lt;p>Having had my first serious brush with my mortality, I started to really reevaluate my life. I read &lt;a href="https://www.amazon.com/Good-Enough-Job-Reclaiming-Life/dp/059353896X">The Good Enough Job&lt;/a> (now a &lt;a href="https://www.youtube.com/watch?v=cfKFbh8LPvU">TED talk&lt;/a>) and started thinking more deeply about how I had been spending my life at work. In my next ~1.5 year at Google, which coincided with a 6% layoff at Google impacting many people that I personally knew, I started wondering if I was really spending my life on something that genuinely mattered to me and my values. I spent a few months thinking about what I should do next, but it felt like I did not have the capacity to work on one of the most stressful projects at the company while building an alternative career in parallel. I decided to quit with no concrete plans on what was next. It felt extremely scary but also very empowering.&lt;sup id="fnref:2">&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref">2&lt;/a>&lt;/sup>&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Going was dying, and staying was dying. When we get to junctures like that, we had better choose the dying that enlarges rather than the one that keeps us stuck.&amp;rdquo;
— JAMES HOLLIS&lt;/p>&lt;/blockquote>
&lt;p>After quitting my job, I ended up taking a short career break and did a few things that I had always dreamt of never got a chance to do. Two noteworthy items were:&lt;/p>
&lt;ol>
&lt;li>Making music for a few months and publishing four songs with professional musicians &lt;a href="https://open.spotify.com/artist/7DgBe0wieIJDgc2XJrLpMJ?si=c1WnswO0QnW1am5fjr-1mQ">on Spotify&lt;/a> (more on each song &lt;a href="https://instagram.com/amirkiani.music">on my Instagram&lt;/a>)&lt;/li>
&lt;li>A 10 day &lt;a href="https://www.dhamma.org/en/index">Vipassana&lt;/a> in the mountains of Idaho, getting attuned with my body’s sensations, and a crash course on &lt;a href="https://en.wikipedia.org/wiki/Impermanence_(Buddhism)">Anicca&lt;/a>.&lt;/li>
&lt;/ol>
&lt;p>Both of these activities were extremely imperfect. Music making made me feel incredibly vulnerable. I felt like the entire world would judge my work for its shortcomings (I still feel the same most times). But I had a goal of making at least a handful of professional recordings and keeping them for my future self. Looking back, this process was incredibly rewarding and productive. Almost each song was an improvement over the last one. I tried different genres and I realized what registers my voice sounds best on. I also learned how to work with artists – whom could be incredibly moody at times 😆&lt;/p>
&lt;p>&lt;img src="https://amirkiani.xyz/images/posts/mortality/spotify-songs.png" alt="alt_text" title="Spotify album covers for my four songs">&lt;/p>
&lt;figcaption>Album covers for my four songs done by my wonderful friend &lt;a href="https://www.instagram.com/alimation/" target="_blank">@alimation&lt;/a>&lt;/figcaption>
&lt;p>Vipassana was also a major mental challenge in accepting imperfection. Sitting for 10 hours for 10 days while simply focusing my attention on my body&amp;rsquo;s sensations, not reading, not writing, not speaking or even making eye contact with anyone, and skipping dinner everyday, was incredibly rough. I questioned every aspect of the course on an hourly basis and considered quitting multiple times. But when it was done, it felt like it made sense. I don&amp;rsquo;t know if I will ever have the luxury of doing this practice again, but I will never forget the serenity and genuinely &amp;ldquo;happy&amp;rdquo; feeling I had at the end of the 10th day. I will also not forget the sense that it felt like I had just found a way to look &lt;em>inward&lt;/em> after 36 years of my very outward focused presence in the world.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;That which seems like a false step is just the next step.&amp;rdquo;&lt;br>
— AGNES MARTIN&lt;/p>&lt;/blockquote>
&lt;p>For most of my time away from work, I was thinking about what to do next. With a nudge from &lt;a href="https://jakeknapp.com/">one of our awesome investors&lt;/a>, I ended up taking a major leap of faith and decided that what I should do next was gathering what I had learned in my years of working in Healthtech and AI to build something that helps cancer patients. Something that I would want if (or when) I would be diagnosed with cancer.&lt;/p>
&lt;p>I would be totally lying if I said that this endeavor has been a piece of cake. As our investor &lt;a href="https://www.youtube.com/watch?v=IpwQeS1ASlI&amp;amp;t=551s">beautifully put it&lt;/a> in a recent podcast episode, I suffer from a massive amount of – hereditary – &lt;em>Perfection Anxiety&lt;/em>: &lt;em>the deep discomfort of having something that you want to make amazing and yet realizing that you will probably screw it up somehow.&lt;/em> I am happy to confirm that I have found &lt;a href="https://www.thesprintbook.com">Design Sprints&lt;/a> to be a genuinely useful method for dealing with my perfection anxiety. My wonderful team and I just chip away at our goal, one Sprint at a time.&lt;/p>
&lt;h2 id="my-one-person-kayak" >My one-person kayak
&lt;span>
&lt;a href="#my-one-person-kayak">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>The motivation for this long, disorganized and imperfect story was me reading the wonderful book, &lt;a href="https://www.amazon.com/Meditations-Mortals-Embrace-Limitations-Counts/dp/0374611998">Meditation for Mortals&lt;/a>, this year. The gist of the book is that we are simply on a path to our death. And if we &lt;em>accept&lt;/em> this undeniable fact, then we can – counterintuitively – reduce a lot of our anxieties. One of the chapters is called &amp;ldquo;Kayaks and superyachts.&amp;rdquo; The author elegantly says:&lt;/p>
&lt;blockquote>
&lt;p>To be human … is to occupy a little one-person kayak [as opposed to a fancy superyacht], borne along on the river of time towards your inevitable yet unpredictable death. It&amp;rsquo;s a thrilling situation, but also an intensely vulnerable one: you&amp;rsquo;re at the mercy of the current, and all you can really do is to stay alert, steering as best you can, reacting as wisely and gracefully as possible to whatever arises from moment to moment … The challenge, then, is simple, though for many of us also excruciating: What&amp;rsquo;s one thing you could do today? … Because the irony, of course, is that just doing something once today, just steering your kayak over the next few inches of water, is the only way you&amp;rsquo;ll ever become the kind of person who does that sort of thing on a regular basis anyway.&lt;/p>&lt;/blockquote>
&lt;p>I&amp;rsquo;ve been thinking about a theme, a resolution, for this year. And I think I have finally found it: &lt;strong>accepting imperfection&lt;/strong>. I am inherently the most imperfect being. I am on my little kayak and on my path to an eventual, undeniable, death. We &lt;em>all&lt;/em> are. All I need to do is to do one thing a day that I find meaningful. It will be imperfect (like this very long post) and likely wrong. Even when I think I&amp;rsquo;ve taken the correct path, it is just a rationalization because we just don&amp;rsquo;t know what the world would have been had we taken an alternative path.&lt;/p>
&lt;p>Accepting my very real and liberating (if not always comfortable) lack of control gives me hope that I can somehow live with my perfection anxiety and make things that matter.
&lt;br/>&lt;br/>
Thanks for reading.&lt;/p>
&lt;!-- Footnotes themselves at the bottom. -->
&lt;h2 id="notes" >Notes
&lt;span>
&lt;a href="#notes">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;div class="footnotes" role="doc-endnotes">
&lt;hr>
&lt;ol>
&lt;li id="fn:1">
&lt;p>I learned later on that the key that led to my GI specialist succumbing to prescribing me a colonoscopy was my mention of my aunt passing away due to stomach cancer in young age, which I believe made it more likely to get insurance to pay for my procedure due to family history.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;li id="fn:2">
&lt;p>I have to give due credit to my wife who supported me during this really strange life period. I probably would have been lost without her.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink">&amp;#x21a9;&amp;#xfe0e;&lt;/a>&lt;/p>
&lt;/li>
&lt;/ol>
&lt;/div></description></item><item><title>LarryGPT – State of LLM Fine-tuning in August 2024</title><link>https://amirkiani.xyz/posts/larrygpt/</link><pubDate>Thu, 01 Aug 2024 00:00:00 +0000</pubDate><guid>https://amirkiani.xyz/posts/larrygpt/</guid><description>&lt;p>&lt;img src="https://amirkiani.xyz/images/posts/larrygpt/larry-standing.webp" alt="Photo of Larry David talking.">&lt;/p>
&lt;figcaption>“Curb Your Enthusiasm,” created by Larry David&lt;/figcaption>
&lt;br/>
&lt;p>&lt;strong>TL;DR:&lt;/strong> I explored the journey of fine-tuning two small LLMs: &lt;em>Llama 3.1-8B-Instruct&lt;/em> vs &lt;em>GPT-3.5-Turbo&lt;/em> on a fun dataset of GPT-4o generated conversations in the style of Larry David. All the &lt;a href="https://github.com/akiani/larry">code&lt;/a> for this project and &lt;a href="https://akiani.github.io/larrygpt/eval.html">evaluation results&lt;/a> are openly accessible. This approach could provide a major cost/latency/privacy advantage if one is interested in small models for &lt;em>specialized&lt;/em> use cases.
&lt;br/>
&lt;br/>&lt;/p>
&lt;h2 id="background" >Background
&lt;span>
&lt;a href="#background">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>While LLMs continue to get multi-modal, larger and more “intelligent”, there is an emerging realization that a single &lt;em>massive, centralized and generalist&lt;/em> LLM may not be the right answer for all use cases. AGI may not in fact be a singular &lt;a href="https://en.wikipedia.org/wiki/HAL_9000">HAL&lt;/a>-like super intelligence but a rather &lt;em>decentralized&lt;/em> collection of &lt;em>small and specialized&lt;/em> models thoughtfully integrated across our technological presence.&lt;/p>
&lt;p>Current state of the art LLMs are also &lt;em>expensive&lt;/em> to run – even the &lt;a href="https://artificialanalysis.ai/models/llama-3-1-instruct-405b">open weights ones&lt;/a> – and asking users to share their data with a centralized service is not always the most privacy-aware option. In addition to the cost and privacy limitations, not all expected model behaviors are easily “promptable” and it is sometimes easier to “show” a small model what we need rather than “tell” it to a large generalized model in multiple paragraphs of instructions. Enter model fine-tuning. Done in a &lt;a href="https://huggingface.co/blog/mlabonne/sft-llama3#%E2%9A%96%EF%B8%8F-sft-techniques">variety of ways&lt;/a>, this technique enables further training of LLMs with high quality example data that captures the expected behavior. Even though fine-tuning should &lt;a href="https://www.tidepool.so/blog/why-you-probably-dont-need-to-fine-tune-an-llm">only be considered as a last resort&lt;/a>, there are in fact specific &lt;a href="https://platform.openai.com/docs/guides/fine-tuning/when-to-use-fine-tuning">scenarios&lt;/a> for which fine-tuning might be a reliable and cost-efficient solution – with my favorite one being &lt;a href="https://machinelearning.apple.com/research/introducing-apple-foundation-models">Apple’s&lt;/a> use of a 3B parameter &lt;em>on-device&lt;/em> model for the new Siri.&lt;/p>
&lt;p>This long-form post captures the culmination of my multiple attempts at LLM fine-tuning over the past year by demonstrating two state of the art methods: fine-tuning &lt;strong>(1) GPT-3.5-Turbo&lt;/strong>, Open AI’s cheapest yet &lt;em>closed source&lt;/em> model vs &lt;strong>(2) Meta’s &lt;a href="https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct">Llama-3.1-8B-Instruct&lt;/a>&lt;/strong>, which was the top performing small/medium size &lt;em>open weight&lt;/em> model at the time of this writing.&lt;/p>
&lt;p>All the code and data for this project is available on &lt;a href="https://github.com/akiani/larrygpt">this repository&lt;/a> and in case you’d rather jump to the final results, you can view those &lt;a href="https://akiani.github.io/larrygpt/eval.html">here&lt;/a>.
&lt;br/>
&lt;br/>&lt;/p>
&lt;h2 id="larrygpt" >LarryGPT
&lt;span>
&lt;a href="#larrygpt">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>One of the use cases for which fine-tuning is indeed relevant is when we want to set the tone, style or other qualitative aspects of a model’s responses. Given my admiration, empathy – and perhaps unfortunate self-identification – with Larry David’s character in &lt;a href="https://en.wikipedia.org/wiki/Curb_Your_Enthusiasm">Curb Your Enthusiasm&lt;/a>, I figured it would be a fun exercise to fine-tune a small/medium LLM to impersonate him while responding in a &lt;em>structured format&lt;/em> and &lt;em>annotating&lt;/em> the range of emotions in his quintessential rants.&lt;/p>
&lt;p>Here’s an example output that we want such model to generate:&lt;/p>
&lt;p>&lt;strong>User:&lt;/strong> &amp;ldquo;Larry, what are your thoughts on tipping?&amp;rdquo;&lt;/p>
&lt;p>&lt;strong>Model:&lt;/strong>&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-json" data-lang="json">&lt;span style="display:flex;">&lt;span>[
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;Oh, tipping. &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;irritated&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;Don&amp;#39;t even get me started on tipping. &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;exasperated&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;Isn&amp;#39;t that why you have a job? &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;confused&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;To get paid by your employer? &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;confused&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;Why is it my responsibility? &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;irritated&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;I mean, I&amp;#39;m already paying for the food! &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;exasperated&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;Now I have to pay you for bringing it to me? &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;exasperated&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;It&amp;#39;s like paying extra for breathing air! &amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;sarcastic&amp;#34;&lt;/span>},
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> {&lt;span style="color:#f92672">&amp;#34;text&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;Ridiculous.&amp;#34;&lt;/span>, &lt;span style="color:#f92672">&amp;#34;emotion&amp;#34;&lt;/span>:&lt;span style="color:#e6db74">&amp;#34;frustrated&amp;#34;&lt;/span>}
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>]
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="dataset" >Dataset
&lt;span>
&lt;a href="#dataset">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>To generate a dataset for this project, I decided to go with the common approach of &lt;em>synthetic data generation&lt;/em>. In simple words, I used GPT-4o to generate a series of &lt;em>conversation starter&lt;/em> prompts and for those I generated a high quality set of responses by setting a System prompt that described the intended behavior to GPT-4o. I then used this dataset to train the two &lt;em>smaller&lt;/em> LLMs. You can read the detailed data generation code in &lt;a href="https://github.com/akiani/larry/blob/main/fine-tune-gpt3.5.ipynb">this notebook&lt;/a>.&lt;/p>
&lt;h2 id="training-gpt-35-turbo---closed-source-fine-tuning-as-a-service" >Training GPT-3.5-Turbo - Closed source fine-tuning as a service
&lt;span>
&lt;a href="#training-gpt-35-turbo---closed-source-fine-tuning-as-a-service">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>I have done model fine-tuning on multiple occasions using both proprietary and open-source tools but I have to say that the OpenAI’s Fine-tuning API is by far the simplest, fastest and most robust method I have come across. The detailed process is described &lt;a href="https://platform.openai.com/docs/guides/fine-tuning">here&lt;/a> on their developer guide which I strongly recommend reading but you can also follow my code &lt;a href="https://github.com/akiani/larry/blob/main/fine-tune-gpt3.5.ipynb">here&lt;/a>.&lt;/p>
&lt;p>In short, all I needed to do was to create a &lt;a href="https://jsonlines.org/">JSONL&lt;/a> format of my training and validation data, upload the files to via their API and simply kick-off training by running this one Python function:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>job &lt;span style="color:#f92672">=&lt;/span> client&lt;span style="color:#f92672">.&lt;/span>fine_tuning&lt;span style="color:#f92672">.&lt;/span>jobs&lt;span style="color:#f92672">.&lt;/span>create(
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> training_file&lt;span style="color:#f92672">=&lt;/span>train_file&lt;span style="color:#f92672">.&lt;/span>id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> validation_file&lt;span style="color:#f92672">=&lt;/span>valid_file&lt;span style="color:#f92672">.&lt;/span>id,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> model&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#34;gpt-3.5-turbo-1106&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> suffix&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#34;larry_david&amp;#34;&lt;/span>,
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> seed&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">42&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Once the training was done, I got a nice “Fine-tuning job … successfully completed” email and the model was ready to be used for inference both via the Chat Playground as well as the API. I then ran the validation set against the newly created model to generate the model responses to be compared with other paths.&lt;/p>
&lt;p>&lt;img src="https://amirkiani.xyz/images/posts/larrygpt/openai-playground.png" alt="Talking to fine-tuned model on OpenAI Playground">&lt;/p>
&lt;figcaption>OpenAI’s Console using GPT-3.5-turbo (left) vs Fine-tuned version (right)&lt;/figcaption>
&lt;h2 id="training-llama-31-8b-instruct-open-weights-model-trained-using-open-source-libraries" >Training Llama 3.1-8B-Instruct: Open weights model trained using open source libraries
&lt;span>
&lt;a href="#training-llama-31-8b-instruct-open-weights-model-trained-using-open-source-libraries">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>The open source ecosystem around LLM training, serving and evaluation is &lt;a href="https://github.com/underlines/awesome-ml/blob/master/llm-tools.md#fine-tuning--training">incredibly rich&lt;/a>. For this project, I used &lt;a href="https://github.com/unslothai/unsloth">Unsloth AI&lt;/a>&amp;rsquo;s approach which through a wide range of clever hacks enables LoRA training for 8B parameter models on a &lt;em>free-tier&lt;/em> Google Colab notebook (&lt;a href="https://github.com/akiani/larrygpt/blob/main/larry-llama.ipynb">code&lt;/a>). This means the cost of this training and inference for this model was practically &lt;em>zero dollars&lt;/em>. That is of course excluding the cost of 10s of hours of &lt;em>developer time&lt;/em> that goes into debugging an ever changing/unstable complex set of tools and randomly crashing free Google Colab kernels 😆&lt;/p>
&lt;p>The most tricky parts of fine-tuning Llama 3.1 were: getting the training data into the expected &lt;code>chatml&lt;/code> format, sorting out nuances of tokenization, dealing with memory issues due to apparent lack of memory garbage collection in the libraries, and figuring out the right values for a large number of &lt;a href="https://huggingface.co/docs/trl/v0.9.6/en/sft_trainer#trl.SFTTrainer">hyper parameters&lt;/a>. These were of course issues that could be likely abstracted away with simple wrappers on top of existing libraries.&lt;/p>
&lt;p>&lt;strong>Open Source LLM’s Developer Experience&lt;/strong>&lt;/p>
&lt;p>Though I was able to read the API docs and fine-tune the GPT-3.5 model in less than an hour, understanding the many steps and edge-cases for training Llama 3.1 8B took me about 1-2 days of focused work. In other words, what was gained due to lower cost of training using the open weights model comes at the price of the much worse developer experience. Having said that, I expect the open source toolchain to get better and I am almost certain that there are hundreds of LLM startups working on building a one-click Llama training service similar to Open AI 😬&lt;/p>
&lt;h2 id="results" >Results
&lt;span>
&lt;a href="#results">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>Even if we fully solve model fine-tuning, model &lt;em>evaluation&lt;/em> is perhaps the most complex and unsolved part of the LLM development process. How do we ensure that the information in the training data was in fact “learned” by the model? How should we properly compare alternative models when much of the evaluation criteria is so &lt;em>qualitative&lt;/em>? Having spent most of my time in my previous job on model evaluation, I think I will leave my thoughts on evaluation for a future post to reduce the chance of this post turning into an actual textbook. But for the sake of completeness, I built a &lt;a href="https://akiani.github.io/larrygpt/eval.html">simple evaluation UI&lt;/a> – with the help of my brilliant friend, Claude Sonnet 3.5 🙂 – which allows us to go through a set of validation data which I randomly held-out from the full generated dataset to be used for evaluation.
&lt;br/>
&lt;br/>&lt;/p>
&lt;p>&lt;img src="https://amirkiani.xyz/images/posts/larrygpt/eval-screenshot.png" alt="Talking to fine-tuned model on OpenAI Playground">&lt;/p>
&lt;figcaption>Simple Side-by-Side evaluation UI&lt;/figcaption>
&lt;h2 id="-and-the-winner-is" >🥁 And the winner is&amp;hellip;
&lt;span>
&lt;a href="#-and-the-winner-is">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>First, I would like to appreciate that you have read this far in this long post! As mentioned at the beginning of this writing, my goal in this project was to contrast the &lt;em>experience&lt;/em> of fine-tuning a closed source model with the open weights one in 2024. When it comes to results, I think it’s worth crowning the winner in a few different categories:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Developer Experience and Speed:&lt;/strong>
&lt;ol>
&lt;li>OpenAI’s API for model training and inference is definitely ahead of competition (though, as I mentioned, this is not so hard to replicate for open weights models)&lt;/li>
&lt;li>The actual fine-tuning job for OpenAI took ~16 minutes. Strangely however, the model inference took twice the time compared to the vanilla GPT-3.5-Turbo.&lt;/li>
&lt;li>The fine-tuning job for Llama 3.1-8B took ~9 minutes for the same number of epochs as GPT-3.5-Turbo but inference was &lt;em>much&lt;/em> slower (~10 seconds on Google Colab vs 3 seconds on OpenAI API). This is likely the outcome of the cheap GPUs used on my Google Colab. Grok has shown that one can run inference on Llama 3.1-8B with the incredible speed of &lt;a href="https://groq.com/">1000+ tokens/second&lt;/a>. In contrast, the best inference speed I got from the vanilla GPT-3.5-Turbo was around 150 t/s and half of that for the fine-tuned model.&lt;/li>
&lt;li>I do believe that the key reason that this was a relatively easy and managable process for the open weights model was the small model size. This process could be much harder and more expensive for larger models – though I am not sure that fine-tuning a very large model makes a lot of sense for most use cases.&lt;/li>
&lt;/ol>
&lt;/li>
&lt;li>&lt;strong>Cost:&lt;/strong> Llama 3.1-8B and the resources used in this specific demonstration were practically free. Hence Llama 3.1-8B wins this category. Through this exercise I realized that OpenAI charges &lt;a href="https://openai.com/api/pricing/">twice more&lt;/a> for inference on top of fine-tuned models. I did try to do this exercise with GPT-4o-mini first but I did not have access to fine-tuning because I wasn’t a high-tier customer.&lt;/li>
&lt;li>&lt;strong>Results Quality:&lt;/strong> As you can see for yourself (&lt;a href="https://akiani.github.io/larrygpt/eval.html?index=5">example&lt;/a>), we were able to successfully train both GPT-3.5 as well as Llama 3.1 to generate GPT-4o level quality responses in our desired tone while sticking to the expected output format. Llama 3.1 generated incorrect format in 30%+ of the eval cases before fine-tuning and this issue was &lt;em>fully fixed&lt;/em> after fine-tuning. GPT-3.5 had 2% incorrect formatting both before and after fine-tuning.&lt;/li>
&lt;/ol>
&lt;h2 id="conclusion" >Conclusion
&lt;span>
&lt;a href="#conclusion">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>The following are my key takeaways from this project:&lt;/p>
&lt;ol>
&lt;li>For specialized and narrow use cases, small models can indeed be fine-tuned with high quality data from larger models showing significant improvements in generation quality, cost and speed&lt;/li>
&lt;li>OpenAI’s fine-tuning API is simple and robust but a bit expensive and closed sourced. This makes the use cases of cost reduction and privacy constraints less relevant. It would perhaps be a major value propostion if OpenAI publicly shared a small open-weight model that could be fine-tuned using their high-quality data generation and model evalution capabilities on their Platform.&lt;/li>
&lt;li>It is very much possible to fine-tune state of the art small open weight models on a free-tier Google colab and use them in production, even if you are a rusty former software engineer turned PM like me. This is incredibly exciting! 🤩
&lt;br/>
&lt;br/>&lt;/li>
&lt;/ol>
&lt;p>Again, thanks for reading this long post and I hope that you have learned something. Feel free to reach out to me via email (&lt;img alt="Contact address" src="https://amirkiani.xyz/images/contact.png" class="contact-photo">) if you would like to chat more about this topic or &lt;a href="https://amirkiani.xyz/newsletter">subscribe to my newsletter&lt;/a> to get notified about the future posts!&lt;/p>
&lt;br/>
&lt;br/></description></item><item><title>Building AI Dialer</title><link>https://amirkiani.xyz/posts/ai-dialer/</link><pubDate>Fri, 12 Jul 2024 00:00:00 +0000</pubDate><guid>https://amirkiani.xyz/posts/ai-dialer/</guid><description>&lt;p>TL;DR: I built &lt;a href="https://github.com/akiani/aidialer">AI Dialer&lt;/a>, a full stack Python app for &lt;em>interruptible&lt;/em> and &lt;em>near-human quality&lt;/em> AI phone calls by stitching LLMs, speech understanding tools, text-to-speech models, and Twilio’s phone API. This post is a long-form companion piece to &lt;a href="https://github.com/akiani/aidialer">this repository on Github&lt;/a>.
&lt;br/>
&lt;br/>&lt;/p>
&lt;img alt="AI Dialer Screenshot" src="https://amirkiani.xyz/images/posts/ai-dialer/ai-dialer.png" style="border-radius: 20px"/>
&lt;figcaption>AI Dialer's web user interface&lt;/figcaption>
&lt;div style="text-align:center">
&lt;audio controls style="width: 70%">
&lt;source src="https://amirkiani.xyz/audio/ai-dialer.m4a" >
Your browser does not support the audio tag.
&lt;/audio>
&lt;figcaption>Example audio recording with AI Dialer&lt;/figcaption>
&lt;/div>
&lt;h2 id="background" >Background
&lt;span>
&lt;a href="#background">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>Phone calls are &lt;a href="https://www.bankmycell.com/blog/why-millennials-ignore-calls#editor">time-consuming and anxiety-inducing&lt;/a>. But their ubiquitous and synchronous nature has resulted in their continued existence for &lt;a href="https://en.wikipedia.org/wiki/Invention_of_the_telephone">more than a century&lt;/a>. Phone calls are also a significant portion of “back-office” tasks that currently amount to billions of dollars in spending across enterprise industries such as &lt;a href="https://www.grandviewresearch.com/industry-analysis/medical-automation-market">healthcare&lt;/a>.&lt;/p>
&lt;p>Built upon the widespread general availability of AI technologies for text and speech processing, numerous &lt;a href="https://www.bland.ai">startups&lt;/a> have set out to build the next generation of AI-powered callers. This project is my personal exploration of this use case which culminated in a full-stack Python application that facilitates interruptible and near-human quality AI phone calls and supports LLM function calling during a phone call.&lt;/p>
&lt;h2 id="how-it-works" >How it works
&lt;span>
&lt;a href="#how-it-works">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;img src="https://amirkiani.xyz/images/posts/ai-dialer/ai-dialer.svg"/>
&lt;figcaption>Simplified architecture of AI Dialer&lt;/figcaption>
&lt;p>With heavy inspiration from &lt;a href="https://github.com/twilio-labs/call-gpt/tree/main">this project&lt;/a> from Twilio Labs, this application wires up the following key components:&lt;/p>
&lt;ol>
&lt;li>Phone Service: makes and receives phone calls through a virtual phone number&lt;/li>
&lt;li>Speech-to-Text Service: converts the caller’s voice to text – so that it can be passed to LLMs – and understands speech patterns such as when the user is done speaking or interrupts the system&lt;/li>
&lt;li>Text-to-text LLM: understands the phone conversation, can make “function calls” and steers the conversation towards accomplishing specific tasks specified through a specified “system” message&lt;/li>
&lt;li>Text-to-Speech Service: converts the LLM response to high-quality speech&lt;/li>
&lt;li>Python Web Server: provides end-points for
&lt;ol>
&lt;li>Answering calls using Twilio’s Twilio Markup Language (Twilio ML)&lt;/li>
&lt;li>Enabling audio streaming from and to Twilio through a per-call WebSocket&lt;/li>
&lt;li>Interacting with the basic Streamlit web UI&lt;/li>
&lt;/ol>
&lt;/li>
&lt;li>Python Web UI: provides a way to initiate calls and specify system/initial messages for LLMs, follow the conversation in real-time, and listen to the call recordings&lt;/li>
&lt;/ol>
&lt;h2 id="why-was-this-complicated" >Why was this complicated
&lt;span>
&lt;a href="#why-was-this-complicated">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;h3 id="streaming" >Streaming
&lt;span>
&lt;a href="#streaming">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h3>&lt;p>The complexity of this project comes from the fact that each constituent service (Twilio, LLM, TTS, …) introduces a meaningful amount of latency to the overall process. The only way to minimize latency is to &lt;em>stream&lt;/em> content from one service to another as soon as a “chunk” of data is available to pass to the next service. But this chunking does need to be done with a bit of calculation. As an example, the way a sentence is pronounced is both a function of how it starts as well as how it ends. So we need to break the LLM output by sentence as they are generated. We also cannot send the user’s query to the LLM before they have stopped talking, therefore we cannot start generating the LLM response before the user has come to a &lt;a href="https://developers.deepgram.com/docs/endpointing">natural pause&lt;/a>.&lt;/p>
&lt;h3 id="parallelism" >Parallelism
&lt;span>
&lt;a href="#parallelism">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h3>&lt;p>The caller and callee can talk at the same time and all pieces of the system should be kept busy as data becomes available from their upstream service. This requires the system to be implemented in a fully parallel architecture. While the &lt;a href="https://github.com/twilio-labs/call-gpt/tree/main">original&lt;/a> Node-based implementation of this project heavily leverages Node’s &lt;a href="https://nodejs.org/en/learn/asynchronous-work/the-nodejs-event-emitter">native&lt;/a> Event-driven programming, I had to implement this pattern using Python’s &lt;a href="https://docs.python.org/3/library/asyncio.html">asyncio&lt;/a> library – which involved getting a good amount of help from my LLM programmer [friends](see Acknowledgement) 🙂&lt;/p>
&lt;h3 id="interruptions" >Interruptions
&lt;span>
&lt;a href="#interruptions">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h3>&lt;p>This project does not take advantage of OpenAI’s GPT-4o’s Audio-to-Audio &lt;a href="https://www.youtube.com/watch?v=1uM8jhcqDP0">feature&lt;/a> because it was not released at the time of its creation. Interruptions are handled by reactively breaking the flow of audio generation if speech is detected when a Twilio audio stream is in progress and resetting the underlying services. This process does work relatively well but could result in speech disruptions mid-word which could sound different from how humans deal with interruptions.&lt;/p>
&lt;h2 id="open-challenges-and-opportunities" >Open challenges and opportunities
&lt;span>
&lt;a href="#open-challenges-and-opportunities">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>While the first version of this project does demonstrate a rather impressive starting point and provided a fun learning opportunity, I did uncover a few open challenges and opportunities through the course of building this service including:&lt;/p>
&lt;h3 id="challenges" >Challenges
&lt;span>
&lt;a href="#challenges">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h3>&lt;ol>
&lt;li>Phone audio is very lossy. Phone land lines (and Twilio) are &lt;a href="https://en.wikipedia.org/wiki/G.711">encoded&lt;/a> with 8000 samples per second and 8-bit quantization using the G.711 codec (A-law or μ-law). The potential environmental noise, microphone distance to the speaker’s mouth and accents also add additional complexities to this audio loss. While advanced speech to text models do exist that can overcome these challenges, a real-time speech processing engine does have to prioritize speed over quality. This also loss sometimes results in awkward situations in calls when the system misunderstands the users.&lt;/li>
&lt;li>The initial delay to get to the first sentence for LLMs and first byte for TTS are simply impossible to overcome in this design
&lt;ol>
&lt;li>I wrote a basic &lt;a href="https://gist.github.com/akiani/84fb82a7bd62dab047eef7b4cce6d8ef">script&lt;/a> to measure the latency distribution for OpenAI (GPT-4o) and Anthropic (Claude 3.5 Sonnet) in generating one small &lt;em>sentence&lt;/em> over 10 trials. Based on my measurement, Claude Sonnet 3.5 took &lt;code>Avg: 763.72 ms, Std: 294.35 ms, Min: 512.37 ms, Max: 1467.52 ms&lt;/code> and GPT-4o took
&lt;code>Avg: 459.85 ms, Std: 123.68 ms, Min: 290.15 ms, Max: 635.97 ms&lt;/code>. This means on average 0.5-1 second of response latency is due to LLM generation&lt;/li>
&lt;li>On top of the LLM, the speech understanding and generation also adds some latency (another 1-2 seconds depending on what service is used with Eleven Labs taking significantly longer).&lt;/li>
&lt;li>Network time for round-trip requests between different services and the web-server does add up to another 1-2 seconds.&lt;/li>
&lt;/ol>
&lt;/li>
&lt;li>User’s tone of voice is lost during speech to text conversion. This can be really challenging to overcome especially when the user is angry or frustrated on the calls (which happens to be common)&lt;/li>
&lt;li>TTS speech speed is not controllable. Imagine the user asks the system to repeat a number slowly. The current design simply is not capable of this.&lt;/li>
&lt;li>While the support for two LLMs (GPT-4o and Claude Sonnet 3.5) and two TTS (Eleven Labs and Deepgram) is implemented, there is a wide range of differences between the services of the same category. For example, Claude requires a much more verbose and specific system prompt to stick to a brief and conversational tone whereas GPT-4o is a bit too eager to give information about the task even when the user has not yet asked for it. The two models are also different in their tendency to run (and even the expected format) for function calls. Deepgram’s TTS is much faster compared to Eleven Labs but it unfortunately comes with the cost of strange mispronunciations (e.g. when a phrase has both words and numbers).&lt;/li>
&lt;li>Enabling robot phone calls is a double-edged sword. In a world that misinformation and robocalls are rampant, creating yet another tool for these activities should be done with extra care and responsibility. My assumption is that this responsibility is enforced via Twilio’s own &lt;a href="https://www.twilio.com/en-us/legal/aup">Acceptable Use Policy&lt;/a> and &lt;a href="https://www.twilio.com/en-us/help/abuse">Report Abuse&lt;/a> system. I however do believe that projects such as this could help lower the barrier for irresponsible usage even though this is far from my intentions.&lt;/li>
&lt;/ol>
&lt;h2 id="opportunities" >Opportunities
&lt;span>
&lt;a href="#opportunities">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;ol>
&lt;li>One straightforward opportunity could be to try the future version of GPT-4o and remove the Speech-to-Text and Text-to-Speech modules. I am however not certain that this version of GPT-4o would be as “controlled” as the text version as the user could really bias the model’s output via significant control over the range of emotions and tone of speech. This “openness” might in fact be a liability, especially in an enterprise setting.&lt;/li>
&lt;li>The current implementation does not chunk audio streaming. This is a straightforward addition that could be added.&lt;/li>
&lt;li>The user interface is currently very bare and could be expanded to support addition of function calls via UI, saving prompts and stateful memory (via using a database for storing call logs).&lt;/li>
&lt;/ol>
&lt;h2 id="acknowledgement" >Acknowledgement
&lt;span>
&lt;a href="#acknowledgement">
&lt;svg viewBox="0 0 28 23" height="100%" width="19" xmlns="http://www.w3.org/2000/svg">&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" fill="none" stroke-linecap="round" stroke-miterlimit="10" stroke-width="2"/>&lt;/svg>
&lt;/a>
&lt;/span>
&lt;/h2>&lt;p>This project would have not happened without &lt;a href="https://github.com/twilio-labs/call-gpt">this great TypeScript example&lt;/a> from Twilio Labs. Claude Sonnet 3.5, GPT-4o, and &lt;a href="https://aider.chat/">Aider&lt;/a> also provided ample help in writing parts of this code base. Additional thanks to my friend Mona for trying the source code and proofreading this post before publication.&lt;/p></description></item><item><title>Hello World</title><link>https://amirkiani.xyz/posts/hello-world/</link><pubDate>Wed, 20 Dec 2023 19:54:42 -0800</pubDate><guid>https://amirkiani.xyz/posts/hello-world/</guid><description>&lt;br/>
&lt;p>I have been writing for most of my life. I recently realized that most of my professional writings are owned by and hidden behind the corporations where I have worked at through my career. And as for my personal writings, they have become buried in numerous social media accounts and mostly inaccessible given the sheer volume of noise in these systems.&lt;/p>
&lt;p>While this is not my first attempt at blogging – I used to write blogs since the early 2000s – it is my most conscious attempt to date.&lt;/p>
&lt;p>Having lived the age of social media from its beginning to today, I feel like a more distributed and self-published way to share information is the proper way, at least for me, to share my thoughts with the world.&lt;/p>
&lt;p>I would be also remiss if I didn&amp;rsquo;t mentioned that &lt;a href="https://www.theverge.com/2023/10/23/23928550/posse-posting-activitypub-standard-twitter-tumblr-mastodon">this Podcast episode&lt;/a> and it&amp;rsquo;s detailed covering of &lt;a href="https://indieweb.org/POSSE">Publish (on your) Own Site, Syndicate Elsewhere&lt;/a> did not catalyze the creation of this website.&lt;/p>
&lt;p>Anyways, thank you for visiting and I look forward to sharing my learnings and thoughts with you.&lt;/p>
&lt;p>I will be sharing links to the posts on this blog on my various social media, but if you are old school, please feel free to either subscribe to my blog&amp;rsquo;s &lt;a href="https://amirkiani.xyz/posts/index.xml">RSS Feed&lt;/a> or &lt;a href="https://amirkiani.xyz/newsletter">sign up for my newsletter&lt;/a>.&lt;/p></description></item></channel></rss>