<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Artificial Intelligence on Posit Open Source</title>
    <link>https://opensource.posit.co/topics/artificial-intelligence/</link>
    <description>Recent content in Artificial Intelligence on Posit Open Source</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en-us</language>
    <lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://opensource.posit.co/topics/artificial-intelligence/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Introducing shinyreact: React UI backed by a Shiny server</title>
      <link>https://opensource.posit.co/blog/2026-09-30_introducing-shinyreact/</link>
      <pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-30_introducing-shinyreact/</guid>
      <dc:creator>Barret Schloerke</dc:creator><description><![CDATA[<p>We&rsquo;re excited to introduce <a href="https://posit-dev.github.io/shinyreact/" target="_blank" rel="noopener">shinyreact</a>, a new package for R and Python. It lets you write the UI of a Shiny app in React, with any component library on npm, while reducing your Shiny server code to data-only logic.</p>
<p>shinyreact splits a Shiny app along a clean line. The Shiny server does reactive computation. The UI is a <a href="https://react.dev" target="_blank" rel="noopener">React</a> client that you own. shinyreact is the bridge between them, and it ships zero UI components of its own.</p>
<p>You can install it from CRAN or PyPI:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-1" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-1-1">R</a></li>
<li><a href="#tabset-1-2">Python</a></li>
</ul>
<div id="tabset-1-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">install.packages</span><span class="p">(</span><span class="s">&#34;shinyreact&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-1-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">pip install shinyreact</span></span></code></pre></div></div>
</div>
</div>
<p>shinyreact is new, and so is the way of building Shiny apps it proposes. We intend to keep the API small, but expect it to evolve as we learn from early adopters.</p>
<h2 id="shiny--react">Shiny + React?
</h2>
<p>For most Shiny apps, defining the UI in your app.py or app.R file allows you to construct a complete, production-ready app in just a few lines of code.</p>
<p>The trouble starts when the design asks for something Shiny and <a href="https://rstudio.github.io/bslib/" target="_blank" rel="noopener">bslib</a> don&rsquo;t have: a unique layout, richer interaction, or a component from a non-Bootstrap design system. At that point, Shiny hasn&rsquo;t had the right tool for the job.</p>
<p>React is that tool, for three reasons:</p>
<ul>
<li><strong>Ecosystem.</strong> React is the most widely used UI library on the web. Design systems, charts, tables, and maps are all one <code>npm install</code> away.</li>
<li><strong>The right model.</strong> React components are functions of state, which fits Shiny&rsquo;s reactive model naturally. When the server sends new data, the UI re-renders efficiently.</li>
<li><strong>AI assistance.</strong> LLMs have trained on an enormous amount of React code. Ask a frontier agent for a UI and it will produce better React than it will bespoke Shiny UI. This aligns with Shiny&rsquo;s goal that app developers should never be required to write low-level HTML or JavaScript themselves.</li>
</ul>
<p>Later in the post, we&rsquo;ll discuss a genomics app that renders a 584,000-cell UMAP on the GPU from a plain Shiny server. First, the basics.</p>
<h2 id="old-faithful-with-shinyreact">Old Faithful with shinyreact
</h2>
<p>Here is the classic Old Faithful histogram app as a shinyreact app:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-2" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-2-1">R</a></li>
<li><a href="#tabset-2-2">Python</a></li>
</ul>
<div id="tabset-2-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">shiny</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">shinyreact</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># Set up the page UI using shinyreact</span>
</span></span><span class="line"><span class="cl"><span class="n">ui</span> <span class="o">&lt;-</span> <span class="nf">page_react</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">server</span> <span class="o">&lt;-</span> <span class="kr">function</span><span class="p">(</span><span class="n">input</span><span class="p">,</span> <span class="n">output</span><span class="p">,</span> <span class="n">session</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">  <span class="n">x</span> <span class="o">&lt;-</span> <span class="n">faithful</span><span class="o">$</span><span class="n">waiting</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">  <span class="n">breaks</span> <span class="o">&lt;-</span> <span class="nf">reactive</span><span class="p">({</span>
</span></span><span class="line"><span class="cl">    <span class="nf">seq</span><span class="p">(</span><span class="nf">min</span><span class="p">(</span><span class="n">x</span><span class="p">),</span> <span class="nf">max</span><span class="p">(</span><span class="n">x</span><span class="p">),</span> <span class="n">length.out</span> <span class="o">=</span> <span class="n">input</span><span class="o">$</span><span class="n">bin_count</span> <span class="o">+</span> <span class="m">1</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="p">})</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">  <span class="c1"># Use `reactive_output()` to send data to the client</span>
</span></span><span class="line"><span class="cl">  <span class="n">output</span><span class="o">$</span><span class="n">dist_data</span> <span class="o">&lt;-</span> <span class="nf">reactive_output</span><span class="p">({</span>
</span></span><span class="line"><span class="cl">    <span class="n">bins</span> <span class="o">&lt;-</span> <span class="nf">hist</span><span class="p">(</span><span class="n">x</span><span class="p">,</span> <span class="n">breaks</span> <span class="o">=</span> <span class="nf">breaks</span><span class="p">(),</span> <span class="n">plot</span> <span class="o">=</span> <span class="kc">FALSE</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="nf">list</span><span class="p">(</span><span class="n">breaks</span> <span class="o">=</span> <span class="nf">I</span><span class="p">(</span><span class="n">bins</span><span class="o">$</span><span class="n">breaks</span><span class="p">),</span> <span class="n">counts</span> <span class="o">=</span> <span class="nf">I</span><span class="p">(</span><span class="n">bins</span><span class="o">$</span><span class="n">counts</span><span class="p">))</span>
</span></span><span class="line"><span class="cl">  <span class="p">})</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nf">shinyApp</span><span class="p">(</span><span class="n">ui</span><span class="p">,</span> <span class="n">server</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-2-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">import</span> <span class="nn">numpy</span> <span class="k">as</span> <span class="nn">np</span>
</span></span><span class="line"><span class="cl"><span class="kn">import</span> <span class="nn">pandas</span> <span class="k">as</span> <span class="nn">pd</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shiny.express</span> <span class="kn">import</span> <span class="nb">input</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shinyreact</span> <span class="kn">import</span> <span class="n">reactive_output</span><span class="p">,</span> <span class="n">set_react_page</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># Set up the page UI using shinyreact</span>
</span></span><span class="line"><span class="cl"><span class="n">set_react_page</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">x</span> <span class="o">=</span> <span class="n">pd</span><span class="o">.</span><span class="n">read_csv</span><span class="p">(</span><span class="s2">&#34;faithful.csv&#34;</span><span class="p">)[</span><span class="s2">&#34;waiting&#34;</span><span class="p">]</span><span class="o">.</span><span class="n">to_numpy</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nd">@reactive_output</span>
</span></span><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">dist_data</span><span class="p">():</span>
</span></span><span class="line"><span class="cl">    <span class="n">breaks</span> <span class="o">=</span> <span class="n">np</span><span class="o">.</span><span class="n">linspace</span><span class="p">(</span><span class="n">x</span><span class="o">.</span><span class="n">min</span><span class="p">(),</span> <span class="n">x</span><span class="o">.</span><span class="n">max</span><span class="p">(),</span> <span class="nb">input</span><span class="o">.</span><span class="n">bin_count</span><span class="p">()</span> <span class="o">+</span> <span class="mi">1</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="n">counts</span><span class="p">,</span> <span class="n">_</span> <span class="o">=</span> <span class="n">np</span><span class="o">.</span><span class="n">histogram</span><span class="p">(</span><span class="n">x</span><span class="p">,</span> <span class="n">bins</span><span class="o">=</span><span class="n">breaks</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="p">{</span><span class="s2">&#34;breaks&#34;</span><span class="p">:</span> <span class="n">breaks</span><span class="o">.</span><span class="n">tolist</span><span class="p">(),</span> <span class="s2">&#34;counts&#34;</span><span class="p">:</span> <span class="n">counts</span><span class="o">.</span><span class="n">tolist</span><span class="p">()}</span></span></span></code></pre></div></div>
</div>
</div>
<p>Two things are different from a traditional Shiny app.</p>
<ol>
<li><strong>The UI is one line.</strong> <code>page_react()</code> (or <code>set_react_page()</code> in Shiny Express) serves the React client that lives in your app&rsquo;s <code>www/</code> directory. There&rsquo;s no <code>sliderInput()</code> or <code>plotOutput()</code>. The <code>www/</code> directory contains the static assets for the React app: <code>ui.js</code> and <code>ui.css</code> (when available).</li>
<li><strong>The output is data.</strong> <code>reactive_output()</code> has no matching UI function. This is a new concept for the Shiny ecosystem! <code>renderPlot()</code> sends an image for <code>plotOutput()</code> to place. <code>reactive_output()</code> sends plain JSON, here the histogram&rsquo;s <code>breaks</code> and <code>counts</code>. Like any render function, <code>reactive_output()</code> re-executes each time its reactive dependencies change. The server sends facts and React decides how to present them.</li>
</ol>
<p>Now the client:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-tsx" data-lang="tsx"><span class="line"><span class="cl"><span class="c1">// src/ui.tsx (which compiles to www/ui.js)
</span></span></span><span class="line"><span class="cl"><span class="kd">function</span> <span class="nx">App() {</span>
</span></span><span class="line"><span class="cl">  <span class="kr">const</span> <span class="p">[</span><span class="nx">binCount</span><span class="p">,</span> <span class="nx">setBinCount</span><span class="p">]</span> <span class="o">=</span> <span class="nx">useShinyInput</span><span class="p">(</span><span class="s2">&#34;bin_count&#34;</span><span class="p">,</span> <span class="mi">30</span><span class="p">);</span>
</span></span><span class="line"><span class="cl">  <span class="kr">const</span> <span class="nx">bins</span> <span class="o">=</span> <span class="nx">useShinyOutputValue</span><span class="p">(</span><span class="s2">&#34;dist_data&#34;</span><span class="p">,</span> <span class="kc">null</span><span class="p">);</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">  <span class="k">return</span> <span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="p">&lt;</span><span class="nt">main</span> <span class="na">className</span><span class="o">=</span><span class="s">&#34;layout&#34;</span><span class="p">&gt;</span>
</span></span><span class="line"><span class="cl">      <span class="p">&lt;</span><span class="nt">label</span> <span class="na">htmlFor</span><span class="o">=</span><span class="s">&#34;bin_count&#34;</span><span class="p">&gt;</span><span class="nb">Number</span> <span class="k">of</span> <span class="nx">bins</span><span class="o">:</span><span class="p">&lt;/</span><span class="nt">label</span><span class="p">&gt;</span>
</span></span><span class="line"><span class="cl">      <span class="p">&lt;</span><span class="nt">input</span>
</span></span><span class="line"><span class="cl">        <span class="na">id</span><span class="o">=</span><span class="s">&#34;bin_count&#34;</span>
</span></span><span class="line"><span class="cl">        <span class="na">type</span><span class="o">=</span><span class="s">&#34;range&#34;</span>
</span></span><span class="line"><span class="cl">        <span class="na">min</span><span class="o">=</span><span class="p">{</span><span class="mi">1</span><span class="p">}</span>
</span></span><span class="line"><span class="cl">        <span class="na">max</span><span class="o">=</span><span class="p">{</span><span class="mi">50</span><span class="p">}</span>
</span></span><span class="line"><span class="cl">        <span class="na">value</span><span class="o">=</span><span class="p">{</span><span class="nx">binCount</span><span class="p">}</span>
</span></span><span class="line"><span class="cl">        <span class="na">onChange</span><span class="o">=</span><span class="p">{(</span><span class="nx">e</span><span class="p">)</span> <span class="o">=&gt;</span> <span class="nx">setBinCount</span><span class="p">(</span><span class="nb">Number</span><span class="p">(</span><span class="nx">e</span><span class="p">.</span><span class="nx">target</span><span class="p">.</span><span class="nx">value</span><span class="p">))}</span>
</span></span><span class="line"><span class="cl">      <span class="p">/&gt;</span>
</span></span><span class="line"><span class="cl">      <span class="p">&lt;</span><span class="nt">Histogram</span> <span class="na">bins</span><span class="o">=</span><span class="p">{</span><span class="nx">bins</span><span class="p">}</span> <span class="p">/&gt;</span>
</span></span><span class="line"><span class="cl">    <span class="p">&lt;/</span><span class="nt">main</span><span class="p">&gt;</span>
</span></span><span class="line"><span class="cl">  <span class="p">);</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span></span></span></code></pre></div></div>
<figure>
<img src="https://opensource.posit.co/blog/2026-09-30_introducing-shinyreact/hello-app.gif" data-fig-alt="A Shiny app with a range slider labeled Number of bins on the left and a histogram of Old Faithful waiting times on the right. As the slider moves from 30 to 8, 45, 20, and back to 30, the histogram redraws with that many bars and the caption updates to match." alt="The Old Faithful app: dragging the bin-count slider re-renders the histogram as the server sends new breaks and counts." />
<figcaption aria-hidden="true">The Old Faithful app: dragging the bin-count slider re-renders the histogram as the server sends new breaks and counts.</figcaption>
</figure>
<p>If you&rsquo;ve written React before, this is an ordinary component. <code>Histogram</code> is whatever you like: a hand-written SVG, a charting library, or a component from your design system. The only shinyreact-specific parts are two hooks:</p>
<ul>
<li><code>useShinyInput()</code> works like React&rsquo;s <code>useState()</code>, except that the value is also sent to the server as <code>input$bin_count</code> (or <code>input.bin_count()</code> in Python). Calling <code>setBinCount()</code> updates the UI and triggers the server&rsquo;s reactive graph.</li>
<li><code>useShinyOutputValue()</code> reads the value of a <code>reactive_output()</code>. When the server recomputes <code>dist_data</code>, the component re-renders with the new data.</li>
</ul>
<p>Those two hooks cover the vast majority of apps.</p>
<h2 id="ids-and-json-are-the-contract">IDs and JSON are the contract
</h2>
<p>The client and server share exactly two things: IDs and JSON values. If you write the client in TypeScript, each hook takes an optional type for its value, such as <code>useShinyInput&lt;number&gt;(&quot;bin_count&quot;, 30)</code>, and your editor will then flag a mismatch that <code>Shiny.setInputValue()</code> never could.</p>
<p>Here is the full round trip for the Old Faithful app:</p>
<ol>
<li>The client calls <code>useShinyInput(&quot;bin_count&quot;, 30)</code>, which sends <code>{&quot;bin_count&quot;: 30}</code> to the server.</li>
<li>The server reads <code>input$bin_count</code>, runs its reactive graph, and computes <code>dist_data</code> using <code>reactive_output()</code>.</li>
<li>The client receives the <code>dist_data</code> result as <code>{&quot;dist_data&quot;: {&quot;breaks&quot;: [...], &quot;counts&quot;: [...]}}</code>, and <code>useShinyOutputValue(&quot;dist_data&quot;)</code> hands it to React.</li>
</ol>
<p>That narrow boundary is what makes a shinyreact app easy to reason about. The server doesn&rsquo;t know or care how the histogram is drawn, and the client doesn&rsquo;t know how the bins are computed. Each side can be reviewed, tested, and rewritten independently, whether a person or an agent wrote it.</p>
<p>When you need more, a few other hooks are available. <code>useShinyOutputStatus()</code> tells you when an output is recalculating so you can show a skeleton, and <code>useShinyMessageHandler()</code> receives one-off messages pushed from the server with <code>send_message()</code>. See the <a href="https://posit-dev.github.io/shinyreact/js/" target="_blank" rel="noopener">JavaScript API reference</a> for the full list.</p>
<h2 id="you-dont-have-to-write-the-react-yourself">You don&rsquo;t have to write the React yourself
</h2>
<p>A fair reaction to all of this is, &ldquo;but I chose Shiny so I wouldn&rsquo;t have to write JavaScript!&rdquo;. That&rsquo;s still the goal. What&rsquo;s changed is that today&rsquo;s AI agents are very good at writing React, far better than they are at writing custom Shiny UI, because there is so much more React in the world for them to learn from.</p>
<p>With shinyreact, your job is to own the server, which is where your data and domain logic live, and to describe and review the UI. To make that concrete, both packages ship <a href="https://posit-dev.github.io/shinyreact/articles/agent-skills.html" target="_blank" rel="noopener">Agent Skills</a>:</p>
<ul>
<li><strong><code>shinyreact-build-app</code></strong> scaffolds a new shinyreact app from a description.</li>
<li><strong><code>shinyreact-convert-app</code></strong> opens an existing Shiny app in a browser, describes what it does in plain English, and then rewrites the UI in React against the same server.</li>
</ul>
<p>The day-to-day loop is familiar. You edit <code>src/ui.tsx</code> (or ask an agent to), the build step writes <code>www/ui.js</code>, and you reload the running Shiny app to see the change. The server side is unchanged: <code>runApp()</code> or <code>shiny run</code>, exactly as before.</p>
<p>Node.js isn&rsquo;t required, since a client can be a single <code>www/ui.js</code> file with no build step. However, we do strongly recommend it as it gives you a proper development environment: a build step gets you TypeScript, linting, and formatting.</p>
<h2 id="keep-what-you-already-have">Keep what you already have
</h2>
<p>shinyreact doesn&rsquo;t ask you to throw away the rest of the Shiny ecosystem.</p>
<ul>
<li><strong>Existing outputs.</strong> <code>renderPlotly()</code>, <code>render.data_frame</code>, and other render functions work as before. Drop a <code>&lt;ShinyOutput id=&quot;...&quot; /&gt;</code> into your React tree, and the output&rsquo;s JavaScript and CSS dependencies are delivered automatically.</li>
<li><strong>Modules.</strong> <code>ShinyModuleProvider</code> namespaces hook IDs to match a server-side module.</li>
<li><strong>Bookmarking.</strong> URL and server bookmarking seed the initial values of <code>useShinyInput()</code>.</li>
</ul>
<h2 id="testing-at-every-layer">Testing at every layer
</h2>
<p>Because the client and server only share IDs and JSON, each layer can be tested on its own. The server is the layer most Shiny developers care about, and it needs no browser at all. Use <code>shiny::testServer()</code> in R, or the new <code>local_server</code> pytest fixture in <a href="https://opensource.posit.co/blog/2026-09-22_shiny-python-1-8">Shiny for Python 1.8</a>: set inputs and assert on the JSON that comes out.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">test_histogram</span><span class="p">(</span><span class="n">local_server</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="n">local_server</span><span class="o">.</span><span class="n">set_inputs</span><span class="p">(</span><span class="n">bin_count</span><span class="o">=</span><span class="mi">10</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="n">data</span> <span class="o">=</span> <span class="n">local_server</span><span class="o">.</span><span class="n">get_output</span><span class="p">(</span><span class="s2">&#34;dist_data&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="k">assert</span> <span class="nb">len</span><span class="p">(</span><span class="n">data</span><span class="p">[</span><span class="s2">&#34;counts&#34;</span><span class="p">])</span> <span class="o">==</span> <span class="mi">10</span></span></span></code></pre></div></div>
<p>The other layers have their own tools:</p>
<ul>
<li><strong>The client.</strong> Ordinary JavaScript unit tests, with whichever test runner you (or your agent) prefer.</li>
<li><strong>The wire.</strong> <code>wire_tap()</code> for <a href="https://rstudio.github.io/shinytest2/" target="_blank" rel="noopener">shinytest2</a> and <code>WireTap</code> for Playwright record the JSON crossing the websocket during a browser test. That gives you an end-to-end assertion on what the client actually sent and what the server actually returned, without reaching into the rendered DOM.</li>
<li><strong>The behavior.</strong> Each example app ships a <code>FEATURES.md</code>: a nested list where every leaf is one checkable claim about the app, written in plain English. A person can read it as a spec, and an agent with a browser can walk it and turn each claim into a deterministic check against the running app.</li>
</ul>
<p>The <a href="https://posit-dev.github.io/shinyreact/articles/testing.html" target="_blank" rel="noopener">testing article</a> covers all of them.</p>
<h2 id="in-the-wild-plotomics-live">In the wild: Plotomics Live
</h2>
<p>Over the summer, Shiny intern <a href="https://www.samuelbharti.com" target="_blank" rel="noopener">Samuel Bharti</a> built a <a href="https://posit-shiny-showcase-bioinformatics.share.connect.posit.cloud/" target="_blank" rel="noopener">collection of bioinformatics Shiny apps</a>. Most of them are plain Shiny and bslib. The one that reached for shinyreact did so because the visualizations demanded it.</p>
<p><a href="https://posit-plotomics-live.share.connect.posit.cloud/" target="_blank" rel="noopener">Plotomics Live</a> (<a href="https://github.com/samuelbharti/plotomics-live" target="_blank" rel="noopener">source</a>, <a href="https://doi.org/10.5281/zenodo.21936926" target="_blank" rel="noopener">DOI</a>) is a 26-page gallery of GPU-accelerated genomics visualizations, from oncoplots to a one-million-point Xenium spatial view and an interactive 584,000-cell UMAP. Large data skips JSON entirely and moves as compact binary typed arrays straight to the GPU. Because React owns the component, a new selection updates the data in place without re-mounting the visualization or reallocating GPU buffers.</p>
<p><video controls autoplay loop muted playsinline src="https://opensource.posit.co/blog/2026-09-30_introducing-shinyreact/plotomics-live.mp4" class="w-full border rounded" title="Plotomics Live: one million Xenium detections rendered with WebGL, with hover tooltips and a legend of marker classes"></video></p>
<p>In Samuel&rsquo;s words:</p>
<blockquote>
<p>R stays the analysis engine, React becomes the visualization layer, and shinyreact removes the custom JavaScript bindings, manual message passing, and serialization code that used to sit between them.</p>
</blockquote>
<h2 id="whats-next">What&rsquo;s next
</h2>
<p>We&rsquo;re working on two directions next:</p>
<ul>
<li><strong>Embedding React components in existing apps</strong>, so you can adopt shinyreact one piece at a time without porting a whole app.</li>
<li><strong>Wrapping shinyreact in your own package</strong>, so you can build a component once and ship it the way bslib ships its components.</li>
</ul>
<h2 id="learn-more">Learn more
</h2>
<ul>
<li>Documentation: <a href="https://posit-dev.github.io/shinyreact/" target="_blank" rel="noopener">posit-dev.github.io/shinyreact</a></li>
<li>Source and example apps: <a href="https://github.com/posit-dev/shinyreact" target="_blank" rel="noopener">github.com/posit-dev/shinyreact</a></li>
<li>posit::conf(2026) talk slides: <a href="https://schloerke.com/presentation-2026-09-15-posit-conf-shinyreact/" target="_blank" rel="noopener">Beyond Bootstrap: Building Custom Shiny UI with React</a></li>
</ul>
<p>Give shinyreact a try, and please <a href="https://github.com/posit-dev/shinyreact/issues" target="_blank" rel="noopener">let us know</a> what you build and what breaks.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-30_introducing-shinyreact/feature.png" length="310650" type="image/png" />
    </item>
    <item>
      <title>Multiple tables, saved conversations, and take-home dashboards: querychat R 0.4.0 and Python 0.9.0</title>
      <link>https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/</link>
      <pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/</guid>
      <dc:creator>Carson Sievert</dc:creator>
      <dc:creator>Garrick Aden-Buie</dc:creator><description><![CDATA[<p>I&rsquo;m thrilled to share the latest <a href="https://posit-dev.github.io/querychat" target="_blank" rel="noopener">querychat</a> release for both R (v0.4.0) and Python (v0.9.0).
Grab the latest from CRAN or PyPI:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-1" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-1-1">R</a></li>
<li><a href="#tabset-1-2">Python</a></li>
</ul>
<div id="tabset-1-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">install.packages</span><span class="p">(</span><span class="s">&#34;querychat&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-1-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">pip install -U querychat</span></span></code></pre></div></div>
</div>
</div>
<p>This release adds several headline features, including support for multiple tables, <code>data-dict.yml</code>, a full-page chat layout, support for pins, and a new <code>/handoff</code> command.
It also builds on <a href="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1">shinychat&rsquo;s recent momentum</a>.
As a result, querychat gets chat features like history and file attachments basically for free.
<code>querychat_app()</code> provides a quick and useful way to start chatting with data and getting bespoke <a href="https://opensource.posit.co/blog/2026-06-17_querychat-ggsql">ggsql visualizations</a>, and it now uses shinychat&rsquo;s <code>page_chat()</code> for a full chat app experience.</p>
<p>See the <a href="https://github.com/posit-dev/querychat/blob/main/pkg-r/NEWS.md" target="_blank" rel="noopener">R release notes</a> and the <a href="https://github.com/posit-dev/querychat/blob/main/pkg-py/CHANGELOG.md" target="_blank" rel="noopener">Python changelog</a> for the complete list, including <a href="#a-few-changes-for-existing-apps">a few changes for existing apps</a> if you&rsquo;re upgrading.</p>
<h2 id="full-page-chat-layout">Full-page chat layout
</h2>
<p><code>querychat_app()</code> / <code>QueryChat.app()</code> now put the chat front and center (built on shinychat&rsquo;s <code>page_chat()</code>), leaving more breathing room for things you create within the chat.</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-2" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-2-1">R</a></li>
<li><a href="#tabset-2-2">Python</a></li>
</ul>
<div id="tabset-2-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">palmerpenguins</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nf">querychat_app</span><span class="p">(</span><span class="n">penguins</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-2-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">querychat</span> <span class="kn">import</span> <span class="n">QueryChat</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">palmerpenguins</span> <span class="kn">import</span> <span class="n">load_penguins</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">qc</span> <span class="o">=</span> <span class="n">QueryChat</span><span class="p">(</span><span class="n">load_penguins</span><span class="p">(),</span> <span class="s2">&#34;penguins&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">qc</span><span class="o">.</span><span class="n">app</span><span class="p">()</span></span></span></code></pre></div></div>
</div>
</div>
<script src="https://fast.wistia.com/player.js" async></script>
<script src="https://fast.wistia.com/embed/xe4aalo9yw.js" async type="module"></script>
<style>wistia-player[media-id='xe4aalo9yw']:not(:defined) { background: center / contain no-repeat url('https://fast.wistia.com/embed/medias/xe4aalo9yw/swatch'); display: block; filter: blur(5px); padding-top:75.21%; }</style>
<p><wistia-player media-id="xe4aalo9yw" aspect="1.3296296296296297"></wistia-player></p>
<p>A view of the actual data is always accessible via the data source drawer on the right-hand side.
In the case of <a href="#multiple-tables">multiple tables</a>, you&rsquo;ll see the active table<sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup>, as well as other available tables below it.</p>
<img src="https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/multi-table.png" alt="A view of QueryChat.app() with the data source drawer opened." class="shadow rounded" />
<p>The new <code>page()</code> method brings this same full-page chat layout to your own apps.
Your users get the chat front and center, and you can still add custom views on other <code>pages</code>, in the <code>drawer</code>, or in the <code>sidebar</code>.</p>
<p>Learn more about <a href="https://posit-dev.github.io/querychat/r/articles/build.html" target="_blank" rel="noopener">building custom apps in R</a> and <a href="https://posit-dev.github.io/querychat/py/build.html" target="_blank" rel="noopener">Python</a>.</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-3" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-3-1">R</a></li>
<li><a href="#tabset-3-2">Python</a></li>
</ul>
<div id="tabset-3-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">qc</span> <span class="o">&lt;-</span> <span class="n">QueryChat</span><span class="o">$</span><span class="nf">new</span><span class="p">(</span><span class="n">penguins</span><span class="p">,</span> <span class="s">&#34;penguins&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">ui</span> <span class="o">&lt;-</span> <span class="n">qc</span><span class="o">$</span><span class="nf">page</span><span class="p">(</span><span class="s">&#34;Penguins Explorer&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-3-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">querychat.express</span> <span class="kn">import</span> <span class="n">QueryChat</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">palmerpenguins</span> <span class="kn">import</span> <span class="n">load_penguins</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">qc</span> <span class="o">=</span> <span class="n">QueryChat</span><span class="p">(</span><span class="n">load_penguins</span><span class="p">(),</span> <span class="s2">&#34;penguins&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">qc</span><span class="o">.</span><span class="n">page</span><span class="p">(</span><span class="s2">&#34;Penguins Explorer&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
</div>
<h2 id="conversation-history">Conversation history
</h2>
<p>Another major improvement is persistent conversation history (<a href="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1#return-to-earlier-conversations">mostly thanks to shinychat</a>).
In addition to starting new chats and returning to previous ones, conversations now persist across page reloads and timeouts.
As a result, it is now much more difficult to lose your progress.</p>
<img src="https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/history.png" alt="A view of QueryChat.app() with the history sidebar opened." class="shadow rounded" />
<p>Also, now that shinychat supports <a href="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1#edit-a-message-and-compare-answers">editable messages</a>, <a href="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1#stream-responses-and-show-thinking">canceling responses</a>, <a href="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1#attach-files">file attachments</a>, and more, querychat does too.</p>
<h2 id="multiple-tables">Multiple tables
</h2>
<p>querychat now supports multiple tables in a single chat instance.
If those tables reside in a singular source, like a database, you can add them all in one fell swoop with the <code>add_tables()</code> method.</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-4" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-4-1">R</a></li>
<li><a href="#tabset-4-2">Python</a></li>
</ul>
<div id="tabset-4-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">querychat</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">qc</span> <span class="o">&lt;-</span> <span class="n">QueryChat</span><span class="o">$</span><span class="nf">new</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">qc</span><span class="o">$</span><span class="nf">add_tables</span><span class="p">(</span><span class="n">db</span><span class="p">,</span> <span class="nf">c</span><span class="p">(</span><span class="s">&#34;customers&#34;</span><span class="p">,</span> <span class="s">&#34;orders&#34;</span><span class="p">,</span> <span class="s">&#34;order_items&#34;</span><span class="p">))</span></span></span></code></pre></div></div>
</div>
<div id="tabset-4-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">querychat</span> <span class="kn">import</span> <span class="n">QueryChat</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">qc</span> <span class="o">=</span> <span class="n">QueryChat</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">qc</span><span class="o">.</span><span class="n">add_tables</span><span class="p">(</span><span class="n">db</span><span class="p">,</span> <span class="p">[</span><span class="s2">&#34;customers&#34;</span><span class="p">,</span> <span class="s2">&#34;orders&#34;</span><span class="p">,</span> <span class="s2">&#34;order_items&#34;</span><span class="p">])</span></span></span></code></pre></div></div>
</div>
</div>
<p>querychat&rsquo;s query and visualization tools handle joins across these tables, so a single question can span all of them.
To write a query like the one below, the LLM first needs to know what&rsquo;s in each table: column names, types, and value ranges.
So it starts by fetching the schema of each table it needs (&ldquo;Fetch schemas&rdquo;), then generates the query with that metadata in mind.</p>
<img src="https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/cross-join.png" alt="A chat asking for average order value by acquisition channel. The LLM fetches schemas for the customers, orders, and order_items tables, then runs a SQL query that joins all three." class="shadow rounded" />
<p>In a custom app, the new <code>table()</code> method gives your server code reactive access to any table, including whatever filters the LLM has applied to it.
That means you can keep building your own plots and views in Shiny, and your users can drive them just by chatting.</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-5" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-5-1">R</a></li>
<li><a href="#tabset-5-2">Python</a></li>
</ul>
<div id="tabset-5-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">output</span><span class="o">$</span><span class="n">order_price</span> <span class="o">&lt;-</span> <span class="nf">renderPlot</span><span class="p">({</span>
</span></span><span class="line"><span class="cl">  <span class="n">orders_tbl</span> <span class="o">&lt;-</span> <span class="n">qc</span><span class="o">$</span><span class="nf">table</span><span class="p">(</span><span class="s">&#34;orders&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="c1"># LLM can perform filter queries on $df()</span>
</span></span><span class="line"><span class="cl">  <span class="n">orders_df</span> <span class="o">&lt;-</span> <span class="n">orders_tbl</span><span class="o">$</span><span class="nf">df</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">  <span class="nf">hist</span><span class="p">(</span><span class="n">orders_df</span><span class="o">$</span><span class="n">price</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="p">})</span></span></span></code></pre></div></div>
</div>
<div id="tabset-5-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="nd">@render.plot</span>
</span></span><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">_</span><span class="p">():</span>
</span></span><span class="line"><span class="cl">  <span class="n">orders_tbl</span> <span class="o">=</span> <span class="n">qc</span><span class="o">.</span><span class="n">table</span><span class="p">(</span><span class="s2">&#34;orders&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="c1"># LLM can perform filter queries on .df()</span>
</span></span><span class="line"><span class="cl">  <span class="n">orders_df</span> <span class="o">=</span> <span class="n">orders_tbl</span><span class="o">.</span><span class="n">df</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">  <span class="n">plt</span><span class="o">.</span><span class="n">hist</span><span class="p">(</span><span class="n">orders_df</span><span class="p">[</span><span class="s2">&#34;price&#34;</span><span class="p">])</span></span></span></code></pre></div></div>
</div>
</div>
<h2 id="provide-context-data-dict">Provide context: <code>data-dict</code>
</h2>
<p>querychat does its best to gather context from the data itself.
When the LLM fetches a table&rsquo;s schema, it gets whatever metadata querychat can compute from the data.
That&rsquo;s a good start, but in practice it often isn&rsquo;t enough.
Column names can be cryptic, coded values need decoding, and nothing in the data says what &ldquo;active customer&rdquo; means to your business or how tables relate.
In the <a href="#multiple-tables">example above</a>, the LLM had to infer from column names alone that <code>orders.customer_id</code> points to <code>customers.id</code>.</p>
<p>A <strong>data dictionary</strong> is how you fill in what the data can&rsquo;t say about itself.
It&rsquo;s a YAML file that follows the <a href="https://data-dict.tidyverse.org/" target="_blank" rel="noopener">data-dict</a> spec.
Alongside plain-English descriptions, it has its own fields for column types, allowed values, keys, and the relationships between tables.
This is now the preferred way to describe your data:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">tables</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">customers</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">One row per customer.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">columns</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- <span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">acquisition_channel</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="l">enum</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">values</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="l">organic, paid_search, social, referral]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">How the customer first found us.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">orders</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">One row per order.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">columns</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- <span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">customer_id</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="l">number(id)</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">constraints</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="l">foreign_key]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">order_items</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">One row per item in an order.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">columns</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- <span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">price</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="l">number(quantity)</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">Item price in USD.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">relationships</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span>- <span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">Each order belongs to one customer.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">cardinality</span><span class="p">:</span><span class="w"> </span><span class="l">many-to-one</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">join</span><span class="p">:</span><span class="w"> </span><span class="l">orders.customer_id = customers.id</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span>- <span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">Each order has one or more items.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">cardinality</span><span class="p">:</span><span class="w"> </span><span class="l">many-to-one</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">join</span><span class="p">:</span><span class="w"> </span><span class="l">order_items.order_id = orders.id</span></span></span></code></pre></div></div>
<p>Pass it in as <code>data_dict = &quot;dictionary.yml&quot;</code> (R) / <code>data_dict=&quot;dictionary.yml&quot;</code> (Python).
When the LLM fetches a table&rsquo;s schema, any column you&rsquo;ve documented comes straight from your dictionary, with nothing left to infer.
querychat only computes metadata from the data for the columns your dictionary doesn&rsquo;t cover.</p>
<h2 id="extract-insights-handoff">Extract insights: <code>/handoff</code>
</h2>
<p>Over the course of a conversation, querychat tends to produce a pile of results, some more useful than others.
The useful ones deserve to live on in a reproducible artifact that doesn&rsquo;t depend on the chat app.</p>
<p>That&rsquo;s the idea behind the new <strong><code>/handoff</code></strong> slash command.
It&rsquo;s available in every querychat app, with no setup required.
When a user types <code>/handoff</code> into the chat input, a wizard opens where they select the results that matter, choose an output format (e.g., Quarto, marimo, Shiny, Jupyter), and add any presentation instructions for the LLM to follow when it generates the handoff document.</p>
<img src="https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/handoff-wizard.png" alt="The handoff wizard" class="shadow rounded" />
<p>When the user finishes the wizard, the handoff document&rsquo;s source code streams into a code editor, where they can revise it by hand or with AI assistance.
A download button then gives them a zip bundle with the handoff document, a README file, and the data sources (if they&rsquo;re small enough).</p>
<img src="https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/handoff-download.png" alt="The handoff editor" class="shadow rounded" />
<h2 id="chat-with-pinned-data">Chat with pinned data
</h2>
<p>querychat can now chat with data pinned to a <a href="https://pins.rstudio.com/" target="_blank" rel="noopener">pins</a> board.
Pass the board and the pin name, and querychat reads the pin (parquet, CSV, JSON, RDS, and more) and uses its title, description, and tags as the starting data description:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-6" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-6-1">R</a></li>
<li><a href="#tabset-6-2">Python</a></li>
</ul>
<div id="tabset-6-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">pins</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">querychat</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">board</span> <span class="o">&lt;-</span> <span class="nf">board_connect</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nf">querychat_app</span><span class="p">(</span><span class="n">board</span><span class="p">,</span> <span class="s">&#34;my_pin&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-6-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">pip install <span class="s2">&#34;querychat[pins]&#34;</span></span></span></code></pre></div></div>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">import</span> <span class="nn">pins</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">querychat</span> <span class="kn">import</span> <span class="n">QueryChat</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">board</span> <span class="o">=</span> <span class="n">pins</span><span class="o">.</span><span class="n">board_connect</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">qc</span> <span class="o">=</span> <span class="n">QueryChat</span><span class="p">(</span><span class="n">board</span><span class="p">,</span> <span class="s2">&#34;my_pin&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">qc</span><span class="o">.</span><span class="n">app</span><span class="p">()</span></span></span></code></pre></div></div>
</div>
</div>
<p>For more control, such as setting the table name used in SQL, use the new <code>PinSource</code> class directly.
Multiple pins, or pins mixed with ordinary data frames, also work together in one chat.
Everything is materialized into a shared DuckDB connection behind the scenes, so the LLM can join and filter across all of it.
See the <a href="https://posit-dev.github.io/querychat/r/articles/data-sources.html" target="_blank" rel="noopener">data sources guide for R</a> and <a href="https://posit-dev.github.io/querychat/py/data-sources.html" target="_blank" rel="noopener">Python</a> for details.</p>
<h2 id="a-few-changes-for-existing-apps">A few changes for existing apps
</h2>
<p>This release also includes a handful of breaking changes, mostly around how querychat manages connections and bookmarking now that history is built in.
If you&rsquo;re upgrading, skim the <a href="https://github.com/posit-dev/querychat/blob/main/pkg-r/NEWS.md" target="_blank" rel="noopener">R NEWS</a> or <a href="https://github.com/posit-dev/querychat/blob/main/pkg-py/CHANGELOG.md" target="_blank" rel="noopener">Python CHANGELOG</a> breaking-changes sections before you do.</p>
<h2 id="learn-more">Learn more
</h2>
<ul>
<li><a href="https://posit-dev.github.io/querychat/py/" target="_blank" rel="noopener">querychat documentation</a> (<a href="https://posit-dev.github.io/querychat/r/" target="_blank" rel="noopener">R</a>) &mdash; full guides on data sources, context, tools, and deployment</li>
<li><a href="https://data-dict.tidyverse.org/" target="_blank" rel="noopener">data-dict</a> &mdash; the data dictionary spec querychat now reads</li>
<li><a href="https://ggsql.org" target="_blank" rel="noopener">ggsql</a> &mdash; the grammar of graphics for SQL that powers querychat&rsquo;s visualizations</li>
<li><a href="https://posit-dev.github.io/shinychat/py/" target="_blank" rel="noopener">shinychat</a> (<a href="https://posit-dev.github.io/shinychat/r/" target="_blank" rel="noopener">R</a>) &mdash; the chat UI toolkit querychat builds on</li>
<li><a href="https://posit-dev.github.io/chatlas/" target="_blank" rel="noopener">chatlas</a> (<a href="https://ellmer.tidyverse.org" target="_blank" rel="noopener">ellmer</a>) &mdash; the underlying LLM tool-calling libraries</li>
<li><a href="https://github.com/posit-dev/querychat" target="_blank" rel="noopener">Source on GitHub</a> &mdash; issues, discussions, and contributions welcome</li>
</ul>
<h2 id="acknowledgements">Acknowledgements
</h2>
<p>We thank everyone who contributed to these releases, for opening issues, submitting pull requests, and providing feedback:
<a href="https://github.com/gadenbuie" target="_blank" rel="noopener">@gadenbuie</a>,
<a href="https://github.com/hadley" target="_blank" rel="noopener">@hadley</a>,
<a href="https://github.com/iainwallacebms" target="_blank" rel="noopener">@iainwallacebms</a>,
<a href="https://github.com/iamYannC" target="_blank" rel="noopener">@iamYannC</a>,
<a href="https://github.com/jnhyeon" target="_blank" rel="noopener">@jnhyeon</a>,
<a href="https://github.com/kolabearafk" target="_blank" rel="noopener">@kolabearafk</a>, and
<a href="https://github.com/thisisnic" target="_blank" rel="noopener">@thisisnic</a>.</p>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>By default, the active table is the first one supplied.
However, if the LLM is prompted to show a filtered/sorted view of a table, then that table becomes active (and the drawer will automatically open).&#160;<a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-29_querychat-tables-handoff/featured.png" length="227292" type="image/png" />
    </item>
    <item>
      <title>AI Newsletter: New releases from ellmer, shinychat, and commons</title>
      <link>https://opensource.posit.co/blog/2026-09-18_ai-newsletter/</link>
      <pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-18_ai-newsletter/</guid>
      <dc:creator>Sara Altman</dc:creator>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<div class="callout callout-tip" role="note" aria-label="Tip">
<div class="callout-header">
<span class="callout-title"><strong>Subscribe to the AI Newsletter!</strong></span>
</div>
<div class="callout-body">
<p>The AI newsletter is published as an RSS feed. Follow it in your favorite reader:</p>
<p><a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener noreferrer" class="btn-shortcode inline-flex mb-5 mr-5 items-center px-4 py-3 text-sm leading-5 gap-2 rounded-lg bg-blue-400 !text-white font-semibold align-middle hover:bg-blue-500 transition no-underline">Subscribe via RSS</a></p>
<p><strong>Want the newsletter as an email?</strong> Paste the feed URL, <a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener">https://opensource.posit.co/tags/ai-newsletter/index.xml</a>, into a free RSS-to-email service such as <a href="https://blogtrottr.com/" target="_blank" rel="noopener">Blogtrottr</a>, <a href="https://feedrabbit.com/" target="_blank" rel="noopener">Feedrabbit</a>, or <a href="https://follow.it/" target="_blank" rel="noopener">Follow.it</a>, and each new issue will arrive in your inbox.</p>
</div>
</div>
<p>Last week, three packages in Posit&rsquo;s open-source AI stack shipped significant releases. We introduced commons 0.1.0, and ellmer and shinychat both received substantial updates. These releases are part of a broader effort to make it as easy as possible to build modern chat applications in R and Python.</p>
<p>In this newsletter, we&rsquo;ll take a quick tour of what&rsquo;s new.</p>
<h2 id="introducing-commons">Introducing commons
</h2>
<p><strong>commons, a new framework for building trustworthy self-service data analysis agents in R and Python, is now on CRAN.</strong> Read the full announcement <a href="https://opensource.posit.co/blog/2026-09-15_commons-0-1-0">here</a>.</p>
<p>The Python package is currently in a pre-release beta stage, with more features arriving over the next few weeks.</p>
<p>If you&rsquo;re a data analyst, data scientist, statistical programmer, or other data practitioner, you likely have extensive domain knowledge and a collection of <em>trusted code</em> that you already use in analyses, apps, reports, and packages. The core idea behind commons is that we can leverage this trusted code to improve an agent&rsquo;s correctness.</p>
<p>A commons agent first searches for a trusted calculation. If it finds one that can answer the user&rsquo;s question, it can run that vetted code and the answer is deterministically marked as verified.</p>
<img src="https://opensource.posit.co/blog/2026-09-18_ai-newsletter/images/commons-01-traffic-trend.gif" title="A commons agent answering a question with a trusted calculation." data-fig-alt="A commons agent answers how site traffic is trending by finding and running a trusted calculation. The resulting chart and answer are marked as verified." />
<p>If it doesn&rsquo;t find a trusted calculation, the agent searches trusted context before writing custom R, Python, or SQL. The answer is either given a citation or marked as &ldquo;untrusted,&rdquo; depending on whether the agent provides a verified citation that supports its approach.</p>
<p>The model doesn&rsquo;t decide how trustworthy its answer is. commons assigns each label deterministically based on the analysis path taken by the agent.</p>
<img src="https://opensource.posit.co/blog/2026-09-18_ai-newsletter/images/trust-flow.svg" class="column-page" data-fig-alt="Flow diagram showing how commons routes questions. It first searches trusted calculations. If it finds one, it runs the calculation and returns a verified answer. Otherwise, it searches trusted context, writes custom code, and returns either a cited or lower-trust answer." />
<p>commons also ships with an agent skill to help you create a commons agent and functions for analyzing your users&rsquo; conversations.</p>
<h2 id="shinychat-v050-r-and-v071-python">shinychat v0.5.0 (R) and v0.7.1 (Python)
</h2>
<p><strong>shinychat v0.5.0 for R and v0.7.1 for Python bring together more of what you need to build a complete chat application.</strong> Several of these shinychat updates also made commons possible!</p>
<p>Read the full blog post <a href="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1">here</a>. There are many more updates worth checking out.</p>
<h3 id="page_chat"><code>page_chat()</code>
</h3>
<p>Use <code>page_chat()</code> instead of the bslib <code>page_*()</code> functions when you want the chat to be the center of your application. <code>page_chat()</code> creates a full-window, chatbot-oriented layout with support for navigation pages, conversation history, an <a href="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1#artifact-drawer">artifact drawer</a>, and more.</p>
<img src="https://opensource.posit.co/blog/2026-09-18_ai-newsletter/images/shinychat-page-chat.png" title="A shinychat app built `page_chat()`." class="column-page" data-fig-alt="A full-window Site traffic assistant built with page_chat, showing a traffic-trend question, a compact calculation activity row, the assistant&#39;s answer, and the chat input." />
<h3 id="conversation-history">Conversation history
</h3>
<p>Conversation history is enabled by default when you use <code>chat_server()</code> in R or <code>Chat(client=...)</code> in Python, allowing users to start a new conversation, switch between saved conversations, search them, rename them, and delete them.</p>
<img src="https://opensource.posit.co/blog/2026-09-18_ai-newsletter/images/shinychat-history.png" title="Previous conversations shown in the chat history sidebar." class="column-page" data-fig-alt="The Site traffic assistant with its history sidebar open, showing controls to search or start a conversation and three saved traffic-analysis conversations beside the active chat." />
<h3 id="readable-tool-calls-and-citations">Readable tool calls and citations
</h3>
<p>This shinychat release also includes several improvements for understanding how a model arrived at its response.</p>
<p>One such improvement is readable tool calls. By default, related tool calls are grouped into compact, single-line &ldquo;activity rows&rdquo;, keeping them from overwhelming the conversation. You can still inspect the individual tool calls by expanding a row.</p>
<p>shinychat also displays citations returned by providers&rsquo; built-in web-search and web-fetch tools.</p>
<img src="https://opensource.posit.co/blog/2026-09-18_ai-newsletter/images/shinychat-tools-citations.png" title="A grouped activity row and an open citation." data-fig-alt="A Shinychat answer with a compact grouped activity row for searching documentation and querying the warehouse, plus an open citation identifying sessions_daily as the canonical site-traffic source." />
<h2 id="ellmer-050">ellmer 0.5.0
</h2>
<p><strong><a href="https://opensource.posit.co/blog/2026-09-14_ellmer-0-5-0">ellmer 0.5.0</a> is now on CRAN.</strong> ellmer makes it easy to work with LLMs from R.</p>
<p>Read the full announcement <a href="https://opensource.posit.co/blog/2026-09-14_ellmer-0-5-0">here</a>. Many of the features made available in this release are also available in recent releases of <a href="https://github.com/posit-dev/chatlas/releases" target="_blank" rel="noopener">chatlas</a>, ellmer&rsquo;s sibling package in Python.</p>
<h3 id="citations">Citations
</h3>
<p>When a model uses a supported built-in web tool, ellmer now returns and displays the provider-supplied citations. This works with <code>claude_tool_web_search()</code>, <code>claude_tool_web_fetch()</code>, <code>google_tool_web_search()</code>, and <code>openai_tool_web_search()</code>, helping you identify the sources behind the model&rsquo;s answer.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">&lt;-</span> <span class="nf">chat_openai</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">register_tool</span><span class="p">(</span><span class="nf">openai_tool_web_search</span><span class="p">())</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">  <span class="s">&#34;What is the most recent version of ellmer on CRAN? Look it up.&#34;</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; The most recent CRAN release of **ellmer** is **version 0.5.0**, published</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; **September 4, 2026**.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ([cran.r-project.org](https://cran.r-project.org/package%3Dellmer))[1]</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Sources</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; [1] CRAN: Package ellmer: https://cran.r-project.org/package%3Dellmer</span></span></span></code></pre></div></div>
<h3 id="tokens-and-costs">Tokens and costs
</h3>
<p>Managing tokens and costs is an important part of working with LLMs.</p>
<p>Ever want to know how many tokens an input will take before sending it? For supported providers, you can now use <code>chat$token_count()</code> to estimate input token use. Instead of actually sending the request to the model, it sends the request to the provider&rsquo;s token-counting API.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">&lt;-</span> <span class="nf">chat_openai</span><span class="p">(</span><span class="n">model</span> <span class="o">=</span> <span class="s">&#34;gpt-5.6-luna&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">prompt</span> <span class="o">&lt;-</span> <span class="nf">content_pdf_file</span><span class="p">(</span><span class="s">&#34;example-document.pdf&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">token_count</span><span class="p">(</span><span class="n">prompt</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; [1] 252</span></span></span></code></pre></div></div>
<p>This estimates only the tokens used by the input and does not predict the number of output tokens, so it won&rsquo;t represent the total round-trip count.</p>
<p>Companies frequently release new models and change their prices. Use the new function <code>models_update_prices()</code> to download and cache the latest pricing data from the ellmer GitHub repo. Cost estimates reported by <code>token_usage()</code>, <code>Chat$get_cost()</code>, <code>Chat$get_tokens()</code>, and printed <code>Chat</code> objects use this data.</p>
<h3 id="send-files-to-the-model">Send files to the model
</h3>
<p>It&rsquo;s often useful to send files as part of a chat. You can now send CSV, Markdown, code, and other text-based files to a model with <code>content_document_file()</code> and <code>content_document_url()</code>. For large files or files reused across multiple turns, if you&rsquo;re using <code>chat_openai()</code>, <code>chat_anthropic()</code>, or <code>chat_google_gemini()</code>, use <code>chat$file_upload()</code> instead. It uploads the file once and returns a reference for <code>$chat()</code>. This avoids repeatedly sending the file and reduces token usage and cost.</p>
<h2 id="solid-improvements-for-custom-agents">Solid improvements for custom agents
</h2>
<p>Taken together, these releases make it easier to build more complete custom agents. ellmer manages model interactions in R, shinychat provides the user-facing chat interface, and commons adds a framework for data analysis agents that builds upon these two.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-18_ai-newsletter/images/hero.png" length="543199" type="image/png" />
    </item>
    <item>
      <title>Introducing commons</title>
      <link>https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/</link>
      <pubDate>Tue, 15 Sep 2026 13:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/</guid>
      <dc:creator>Simon Couch</dc:creator>
      <dc:creator>Sara Altman</dc:creator>
      <dc:creator>Josh Taillon</dc:creator><description><![CDATA[<p>We&rsquo;re hootin&rsquo; and hollerin&rsquo; to share <a href="https://posit-dev.github.io/commons/" target="_blank" rel="noopener">commons</a>, an R and Python package that helps data scientists build trustworthy data analysis agents.</p>
<video class="column-page" autoplay loop muted playsinline controls preload="metadata" aria-label="Screen recording of a commons agent answering 'How is traffic trending for our site?' by running a trusted calculation and displaying a chart showing daily site visits increased 22%.">
  <source src="https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/commons-01-traffic-trend.mp4" type="video/mp4">
  Your browser does not support embedded videos.
</video>
<p>The package is built on <a href="https://ellmer.tidyverse.org/" target="_blank" rel="noopener">ellmer</a>, <a href="https://posit-dev.github.io/chatlas/" target="_blank" rel="noopener">chatlas</a>, and <a href="https://github.com/posit-dev/shinychat/" target="_blank" rel="noopener">shinychat</a>, Posit&rsquo;s open source LLM stack. You can use whatever model you want from any of the providers supported by those packages with it.</p>
<p>To install the R package, run:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">install.packages</span><span class="p">(</span><span class="s">&#34;commons&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
<p>To install the Python package from PyPI, run:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-py" data-lang="py"><span class="line"><span class="cl"><span class="n">pip</span> <span class="n">install</span> <span class="n">commons</span></span></span></code></pre></div></div>
<p>The Python package is currently in a pre-release beta stage, but you can install and play around with it today, with more features arriving over the next few weeks.</p>
<h2 id="design-philosophy">Design philosophy
</h2>
<p>AI agents for data analysis can range from overly cautious and narrowly correct to wildly and confidently incorrect. commons provides a framework for you to design more trustworthy analysis agents by providing them with access to existing <strong>trusted code</strong>, while still allowing them enough flexibility to answer novel, realistic questions.</p>
<p>If you are a data analyst, data scientist, statistical programmer, or other data practitioner, you likely have a deep understanding of your problem domain and a collection of trusted code you depend on for your analyses and use to create apps, reports, and packages. The core idea behind commons is that we can improve an agent&rsquo;s correctness by giving it the right access and documentation to run this code that you have already vetted.</p>
<style>
.commons-inline-icon {
  display: inline-block;
  height: 1.25em;
  margin: 0 0.08em;
  vertical-align: -0.25em;
  width: 1.25em;
}
</style>
<p>When answering questions, commons agents first search through a pool of trusted code. If the agent finds an appropriate piece of trusted code, it can invoke it directly, and its response will be tagged with a green shield icon <img src="https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/trusted-icon.svg" class="commons-inline-icon" alt="">. If it doesn&rsquo;t, it will search through relevant context before writing its own SQL, R, or Python. If the agent can find trusted context that justifies its approach, it can provide a citation <img src="https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/citation-mark.svg" class="commons-inline-icon" alt=""> to it at the end of its answer, which will be deterministically checked by commons. Otherwise, the answer is marked with a small warning label <img src="https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/warning-icon.svg" class="commons-inline-icon" alt="">.</p>
<img src="https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/trust-flow.svg" class="column-page" alt="Flow diagram. A commons agent searches trusted calculations. If it finds a relevant calculation, it runs the trusted calculation and returns a verified answer. Otherwise, it searches context, writes SQL or R, and returns either a cited or untrusted answer.">
<p>Notably, the agent itself does not decide how to label a response. commons instead labels answers deterministically, based on the path the agent takes to get it to its answer.</p>
<h2 id="get-started">Get started
</h2>
<p>To get started with the R package, check out the <a href="https://posit-dev.github.io/commons/r/articles/commons.html" target="_blank" rel="noopener">introductory vignette</a>. The package ships with an <a href="https://posit-dev.github.io/commons/r/articles/commons.html#working-with-the-agent-skill" target="_blank" rel="noopener">agent skill</a> to help you hook your trusted code and context up to the agent.</p>
<p>The Python package (although still in beta) is based on the same ideas, and the resulting apps will look <em>very</em> similar regardless of whether you use R or Python—they literally share the same CSS! Check out the <a href="https://posit-dev.github.io/commons/py/" target="_blank" rel="noopener">Python package site</a> to learn more.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-15_commons-0-1-0/commons.png" length="1439234" type="image/png" />
    </item>
    <item>
      <title>Complete chat applications in shinychat: R 0.5.0 and Python 0.7.1</title>
      <link>https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/</link>
      <pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/</guid>
      <dc:creator>Garrick Aden-Buie</dc:creator>
      <dc:creator>Carson Sievert</dc:creator><description><![CDATA[<p>We&rsquo;re excited to announce <a href="https://posit-dev.github.io/shinychat/r/" target="_blank" rel="noopener">shinychat v0.5.0 for R</a> and <a href="https://posit-dev.github.io/shinychat/py/" target="_blank" rel="noopener">shinychat v0.7.1 for Python</a>.
This release brings the pieces of a complete chat application together around the conversation itself.</p>
<p>shinychat is a toolkit for building complete, conversation-centered chat applications with Shiny.
The R package pairs with <a href="https://ellmer.tidyverse.org/" target="_blank" rel="noopener">ellmer</a>, and the Python package pairs with <a href="https://posit-dev.github.io/chatlas/" target="_blank" rel="noopener">chatlas</a>.
Install the latest releases from CRAN or PyPI:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-1" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-1-1">R</a></li>
<li><a href="#tabset-1-2">Python</a></li>
</ul>
<div id="tabset-1-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">install.packages</span><span class="p">(</span><span class="s">&#34;shinychat&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-1-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">pip install -U shinychat</span></span></code></pre></div></div>
</div>
</div>
<p>We cover a lot in this post, and there&rsquo;s even more in the releases.
See the <a href="https://github.com/posit-dev/shinychat/blob/main/pkg-r/NEWS.md" target="_blank" rel="noopener">R release notes</a> and the <a href="https://github.com/posit-dev/shinychat/blob/main/pkg-py/CHANGELOG.md" target="_blank" rel="noopener">Python changelog</a> for the complete list of changes, including <a href="#a-few-changes-for-existing-apps">a few changes for existing apps</a> if you&rsquo;re upgrading.</p>
<h2 id="build-a-chat-application">Build a chat application
</h2>
<div class="w-full aspect-4/3">
      <video
        src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/complete-app.mp4"
        class="w-full h-full object-contain"
        title="A complete chat application: the history sidebar lists saved conversations, the assistant answers with a tool activity row and citations, and the artifact drawer opens beside the chat with a plot"
        controls></video>
    </div>
<p>A useful chat application needs more than a text box and a streaming response.
Your users need a way to return to an earlier conversation, start a new one, correct a question, compare answers, inspect sources, and see what the model is doing when it calls a tool.
They may also need to upload a file, open a preview, or move between the chat and the rest of the application.</p>
<p>shinychat gives you sensible starting points for building that experience.
Pair it with <a href="https://ellmer.tidyverse.org/" target="_blank" rel="noopener">ellmer</a> in R or <a href="https://posit-dev.github.io/chatlas/" target="_blank" rel="noopener">chatlas</a> in Python, and you can get a working chat app running with little setup.
The chat application model has three layers:</p>
<ol>
<li><code>page_chat()</code> gives you a full-window chat app with space for navigation, history, tools, and supporting content.</li>
<li><code>chat_ui()</code> lets you place chat wherever it fits best in your application.</li>
<li><code>chat_server()</code> for R or <code>Chat(client=...)</code> for Python connects your app to an <code>ellmer</code> or <code>chatlas</code> client and enables the integrated chat features.</li>
</ol>
<p>When you want a fully custom experience or need a model client other than ellmer or chatlas, the lower-level pieces are still available for you to assemble yourself.</p>
<h2 id="start-with-page_chat">Start with <code>page_chat()</code>
</h2>
<p>When chat is the center of your application, use <code>page_chat()</code>.
It gives your users a full-window experience with a <a href="#create-a-chat-app">chat home</a>, <a href="#complete-application">navigation pages</a>, <a href="#complete-application">sidebars</a>, <a href="#toolbars">toolbars</a>, <a href="#return-to-earlier-conversations">conversation history</a>, and an <a href="#artifact-drawer">artifact drawer</a>.
Users can move to a settings or sources page while their conversation keeps working and streaming.
The <a href="https://posit-dev.github.io/shinychat/r/articles/get-started.html" target="_blank" rel="noopener">Get started</a> guide for R and the <a href="https://posit-dev.github.io/shinychat/py/page-chat.html" target="_blank" rel="noopener">Page chat</a> guide for Python walk through the full layout.</p>
<h3 id="create-a-chat-app">Create a chat app
</h3>
<p>Build a shinychat application starts similarly in both languages:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-2" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-2-1">R</a></li>
<li><a href="#tabset-2-2">Python</a></li>
</ul>
<div id="tabset-2-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">shiny</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">shinychat</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">ui</span> <span class="o">&lt;-</span> <span class="nf">page_chat</span><span class="p">(</span><span class="n">title</span> <span class="o">=</span> <span class="s">&#34;Assistant&#34;</span><span class="p">,</span> <span class="n">id</span> <span class="o">=</span> <span class="s">&#34;chat&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">server</span> <span class="o">&lt;-</span> <span class="kr">function</span><span class="p">(</span><span class="n">input</span><span class="p">,</span> <span class="n">output</span><span class="p">,</span> <span class="n">session</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">  <span class="n">client</span> <span class="o">&lt;-</span> <span class="n">ellmer</span><span class="o">::</span><span class="nf">chat_openai</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">system_prompt</span> <span class="o">=</span> <span class="s">&#34;You are a helpful assistant.&#34;</span>
</span></span><span class="line"><span class="cl">  <span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">  <span class="nf">chat_server</span><span class="p">(</span><span class="s">&#34;chat&#34;</span><span class="p">,</span> <span class="n">client</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nf">shinyApp</span><span class="p">(</span><span class="n">ui</span><span class="p">,</span> <span class="n">server</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-2-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">chatlas</span> <span class="kn">import</span> <span class="n">ChatAnthropic</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shinychat.express</span> <span class="kn">import</span> <span class="n">Chat</span><span class="p">,</span> <span class="n">page_chat</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">client</span> <span class="o">=</span> <span class="n">ChatAnthropic</span><span class="p">(</span><span class="n">system_prompt</span><span class="o">=</span><span class="s2">&#34;You are a helpful assistant.&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">=</span> <span class="n">Chat</span><span class="p">(</span><span class="nb">id</span><span class="o">=</span><span class="s2">&#34;chat&#34;</span><span class="p">,</span> <span class="n">client</span><span class="o">=</span><span class="n">client</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">page_chat</span><span class="p">(</span><span class="n">title</span><span class="o">=</span><span class="s2">&#34;Assistant&#34;</span><span class="p">,</span> <span class="nb">id</span><span class="o">=</span><span class="s2">&#34;chat&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
</div>
<p>With just a few lines of code, you&rsquo;ll have a working chat app backed by a live LLM.
Passing a client to <code>chat_server()</code> in R, or to <code>Chat()</code> in Python, does all the hard work for you, fulling connecting your app to the model client and giving you a complete multi-user chat application<sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup>.</p>
<p>For a personal chat UI you can use while you develop locally, pass an ellmer client to <a href="https://posit-dev.github.io/shinychat/r/reference/chat_app.html" target="_blank" rel="noopener"><code>chat_app()</code></a> in R, or a chatlas client to <a href="https://posit-dev.github.io/shinychat/py/api/Chat.html" target="_blank" rel="noopener"><code>Chat(client=...)</code></a>, and then call <code>.app()</code> in Python.</p>
<h3 id="welcome-users">Welcome users
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/greeting-suggestions-greeting.png" data-fig-alt="A new chat with a short welcome message and a grid of three suggestion cards beneath it." />
<p>When you&rsquo;re app opens, don&rsquo;t leave your users hanging with an empty chat canvas, gree them with <code>chat_greeting()</code> (<a href="https://posit-dev.github.io/shinychat/r/reference/chat_greeting.html" target="_blank" rel="noopener">R</a>, <a href="https://posit-dev.github.io/shinychat/py/api/chat_greeting.html" target="_blank" rel="noopener">Python</a>)!</p>
<p>Greetings can be used to explain the application, set expectations, and give users a useful first step before they write their first message.
By default, they disappear when the user starts chatting, but you can set <code>persistent = TRUE</code> in R or <code>persistent=True</code> in Python to keep one at the top of the conversation history.</p>
<p>Greetings are written in markdown and can even provide actionable suggestions.
Users can click a suggestion to fill the input, ready to edit before sending, or send it immediately.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-markdown" data-lang="markdown"><span class="line"><span class="cl"><span class="gu">## Welcome!
</span></span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">What would you like to do?
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">*</span> &lt;span class=&#34;suggestion submit&#34;&gt;Summarize my data&lt;/span&gt;
</span></span><span class="line"><span class="cl"><span class="k">*</span> &lt;span class=&#34;suggestion&#34;&gt;Create a plot&lt;/span&gt;
</span></span><span class="line"><span class="cl">* &lt;span class=&#34;suggestion&#34;&gt;Explain this code&lt;/span&gt;</span></span></code></pre></div></div>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/greeting-suggestions-fill-input.png" data-fig-alt="Clicking a suggestion card fills the chat input with the suggested prompt, ready to edit before sending." />
<p>You don&rsquo;t have to greet your users with the same message every time, you can use LLMs to generate fresh custom greetings.
To learn more, we&rsquo;ll point you to the <code>chat_greeting()</code> documentation pages (<a href="https://posit-dev.github.io/shinychat/r/reference/chat_greeting.html" target="_blank" rel="noopener">R</a>, <a href="https://posit-dev.github.io/shinychat/py/api/chat_greeting.html" target="_blank" rel="noopener">Python</a>), but it&rsquo;s worth noting that dynamic greetings can stream into the chat like any other response.</p>
<div class="w-full aspect-4/3">
      <video
        src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/greeting-stream.mp4"
        class="w-full h-full object-contain"
        title="A generated greeting streams into the empty chat: the welcome message arrives word by word, then two suggestion cards appear"
        controls></video>
    </div>
<h2 id="return-to-earlier-conversations">Return to earlier conversations
</h2>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/history-list.png" data-fig-alt="The conversation history drawer open beside the chat, listing several named conversations under Today with a search field and a New conversation button." />
<p>One of the biggest features to arrive in this release is conversation history, giving your chat app the ability to save and return to previous conversations.
It will also persist the current conversation across page reloads and other disconnects, virtually eliminating the possibility of losing your progress.
As usual, when you connect shinychat with an ellmer or chatlas client, conversation history is wired up and enabled for you!</p>
<h3 id="save-conversations">Save conversations
</h3>
<p>The history drawer lets users:</p>
<ul>
<li>Start a new conversation.</li>
<li>Switch between saved conversations.</li>
<li>Search conversations.</li>
<li>Rename a conversation.</li>
<li>Delete a conversation.</li>
<li>Return to the conversation that was active when they last opened the app.</li>
</ul>
<p>shinychat generates a short title once the conversation has enough content.
Users can replace that title, and title generation never overwrites a manual rename.</p>
<div class="panel-tabset">
<ul id="tabset-3" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-3-1">Rename</a></li>
<li><a href="#tabset-3-2">Search</a></li>
</ul>
<div id="tabset-3-1">
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/history-actions-menu.png" data-fig-alt="The menu on a saved conversation with options to rename and delete it." />
</div>
<div id="tabset-3-2">
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/history-search.png" data-fig-alt="Typing in the history drawer search field narrows the conversation list to matching titles." />
</div>
</div>
<p>You can <code>history_options()</code> in R or <code>HistoryOptions</code> in Python to configure the conversations that shinychat saves.
The main options are:</p>
<ul>
<li><code>restore_mode</code>, which controls which conversation opens when a user returns to the app:
<ul>
<li><code>&quot;browser&quot;</code> is the default. It returns that browser to its most recent conversation without changing the URL.</li>
<li><code>&quot;url&quot;</code> puts the active conversation ID in the address bar, so users can bookmark or share a specific conversation.</li>
<li><code>&quot;bookmark&quot;</code> restores the conversation with the rest of the app state when your app uses Shiny server bookmarking.</li>
</ul>
</li>
<li><code>store</code> controls where shinychat saves conversations. Use <code>&quot;memory&quot;</code> for local development or tests, or <code>&quot;file&quot;</code> to save them on disk.</li>
<li><code>title</code> controls how the automated conversation titles are generated.</li>
</ul>
<p>For example, this configuration stores conversations on disk and puts the active conversation ID in the URL:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-4" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-4-1">R</a></li>
<li><a href="#tabset-4-2">Python</a></li>
</ul>
<div id="tabset-4-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">history</span> <span class="o">&lt;-</span> <span class="nf">history_options</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">  <span class="n">restore_mode</span> <span class="o">=</span> <span class="s">&#34;url&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">  <span class="n">store</span> <span class="o">=</span> <span class="s">&#34;file&#34;</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nf">chat_server</span><span class="p">(</span><span class="s">&#34;chat&#34;</span><span class="p">,</span> <span class="n">client</span><span class="p">,</span> <span class="n">history</span> <span class="o">=</span> <span class="n">history</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-4-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shinychat</span> <span class="kn">import</span> <span class="n">Chat</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shinychat.types</span> <span class="kn">import</span> <span class="n">HistoryOptions</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">history</span> <span class="o">=</span> <span class="n">HistoryOptions</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">restore_mode</span><span class="o">=</span><span class="s2">&#34;url&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">store</span><span class="o">=</span><span class="s2">&#34;file&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">=</span> <span class="n">Chat</span><span class="p">(</span><span class="s2">&#34;chat&#34;</span><span class="p">,</span> <span class="n">client</span><span class="o">=</span><span class="n">client</span><span class="p">,</span> <span class="n">history</span><span class="o">=</span><span class="n">history</span><span class="p">)</span></span></span></code></pre></div></div>
</div>
</div>
<p>On Posit Connect, conversation history is included with the platform and is enabled automatically when you provide a model client.
The default configuration uses Connect&rsquo;s <a href="https://docs.posit.co/connect/user/structuring-content/#persistent-storage-on-posit-connect" target="_blank" rel="noopener">persistent storage</a> and scopes conversations to the authenticated user.
That gives every user a private conversation history without an additional history service or per-user setup.</p>
<p>In every restore mode, shinychat keeps the transcript in its configured store instead of putting the full conversation in the URL.</p>
<h3 id="edit-a-message-and-compare-answers">Edit a message and compare answers
</h3>
<div class="w-full aspect-4/3">
      <video
        src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/edit-branches-edit.mp4"
        class="w-full h-full object-contain"
        title="Editing an earlier message and resending it starts a new branch, and the sibling navigation control appears on the response"
        controls></video>
    </div>
<p>Editing a message now creates a new conversation <strong>branch</strong>.
When a user edits and resends an earlier message, shinychat forks the conversation at that point: the original question and its later messages remain on one branch, while the edited question begins another. Users can move between the answers with the branch controls in the message.</p>
<div class="panel-tabset">
<ul id="tabset-5" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-5-1">Branch 1</a></li>
<li><a href="#tabset-5-2">Branch 2</a></li>
</ul>
<div id="tabset-5-1">
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/edit-branches-original.png" data-fig-alt="The original conversation with an assistant response showing a 1 / 2 sibling navigation control." />
</div>
<div id="tabset-5-2">
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/edit-branches-new.png" data-fig-alt="The same conversation after editing a message, with the new branch&#39;s response selected and the sibling navigation control showing 2 / 2." />
</div>
</div>
<p>Branches help when a prompt is almost right or when a model takes an unhelpful direction, and they make comparing answers easy without starting over.
And they are part of the saved conversation, so users return to their place in the conversation after a reload.</p>
<h2 id="add-content-and-controls">Add content and controls
</h2>
<p>When chat is part of a larger application, your users still need access to filters, settings, sources, and results.
<code>page_chat()</code> gives you a place to put those alongside the conversation: a drawer for results, toolbars for controls, and offcanvas panels for settings you would rather keep off screen.</p>
<h3 id="artifact-drawer">Artifact drawer
</h3>
<p><code>chat_drawer()</code> gives you a place to show previews, rendered reports, tables, plots, or other bits of Shiny UI next to your chat.
Your users can keep the conversation visible while they inspect a result.</p>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/complete-app-nav-drawer.png" data-fig-alt="The research assistant app with the Research assistant and Sources navigation pages in the header, the conversation in the main region, and the artifact drawer open beside the chat showing a bar chart of penguin counts." />
<p>See the <a href="https://posit-dev.github.io/shinychat/r/reference/chat_drawer.html" target="_blank" rel="noopener">drawer documentation for R</a> or <a href="https://posit-dev.github.io/shinychat/py/api/chat_drawer.html" target="_blank" rel="noopener">Python</a> for the full API.
The <a href="#complete-application">complete application example</a> combines a drawer with the rest of the application layout.</p>
<h3 id="toolbars">Toolbars
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/toolbars-home.png" data-fig-alt="A close-up of a research assistant chat. Part of the conversation is visible beside the global toolbar&#39;s Refresh, Help, and Answer settings buttons, with a response style selector below the chat input." />
<p>Your app may need a Help button that works on every page, while the chat home needs an action that&rsquo;s only relevant when you&rsquo;re looking at the conversation. <code>page_chat()</code> gives each action a home through scoped <a href="https://opensource.posit.co/blog/2026-05-26_introducing-toolbars">toolbars</a>, built on the toolbar components that bslib and Shiny shipped earlier this year.</p>
<p>If you want an action to follow users through the whole app &mdash; pass it to <code>toolbar_global</code>. Put chat-home actions in <code>toolbar</code> in <code>page_chat()</code>, and give a <a href="#complete-application">secondary page</a> its own <code>toolbar</code> through <code>chat_nav_panel()</code>. <code>toolbar_input</code> puts related actions below the message box.</p>
<p>See the <a href="https://posit-dev.github.io/shinychat/r/articles/get-started.html" target="_blank" rel="noopener">R get started guide</a> or the <a href="https://posit-dev.github.io/shinychat/py/page-chat.html" target="_blank" rel="noopener">Python Page chat guide</a> for the full toolbar API.</p>
<h3 id="offcanvas-panels">Offcanvas panels
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/toolbars-offcanvas.png" data-fig-alt="The research assistant app with the Answer settings offcanvas open along the right edge, showing a target length slider and a citations checkbox beside the conversation." />
<p><code>page_chat()</code> pairs <a href="https://opensource.posit.co/blog/2026-08-04_shiny-r-1-14-python-1-7">offcanvas panels</a> with secondary content, such as an answer-length slider or citation setting, and a toolbar button can open an <strong>Answer settings</strong> panel from any page:</p>
<h3 id="complete-application">Complete application
</h3>
<p>As your app grows, <code>page_chat()</code> can grow around the conversation. You can add secondary pages and a sidebar for filters or other app UI and the application menu keeps those options available on narrow screens.</p>
<p>The following example brings the toolbars, sidebar, navigation, and drawer together.</p>
<details class="callout callout-tip" role="note" aria-label="Tip">
<summary class="callout-header">
<span class="callout-title">A complete <code>page_chat()</code> example</span>
</summary>
<div class="callout-body">
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-6" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-6-1">R</a></li>
<li><a href="#tabset-6-2">Python</a></li>
</ul>
<div id="tabset-6-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">ui</span> <span class="o">&lt;-</span> <span class="nf">page_chat</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">  <span class="s">&#34;Research assistant&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">  <span class="n">id</span> <span class="o">=</span> <span class="s">&#34;chat&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">  <span class="n">toolbar</span> <span class="o">=</span> <span class="n">bslib</span><span class="o">::</span><span class="nf">toolbar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">bslib</span><span class="o">::</span><span class="nf">toolbar_input_button</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">      <span class="s">&#34;clear_chat&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="s">&#34;Clear conversation&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="n">icon</span> <span class="o">=</span> <span class="n">bsicons</span><span class="o">::</span><span class="nf">bs_icon</span><span class="p">(</span><span class="s">&#34;arrow-counterclockwise&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="p">),</span>
</span></span><span class="line"><span class="cl">  <span class="n">toolbar_global</span> <span class="o">=</span> <span class="n">bslib</span><span class="o">::</span><span class="nf">toolbar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">bslib</span><span class="o">::</span><span class="nf">toolbar_input_button</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">      <span class="s">&#34;help&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="s">&#34;Help&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="n">icon</span> <span class="o">=</span> <span class="n">bsicons</span><span class="o">::</span><span class="nf">bs_icon</span><span class="p">(</span><span class="s">&#34;question-circle&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="p">),</span>
</span></span><span class="line"><span class="cl">  <span class="n">sidebar</span> <span class="o">=</span> <span class="nf">chat_sidebar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">tags</span><span class="o">$</span><span class="nf">p</span><span class="p">(</span><span class="s">&#34;Use filters to focus the results.&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">history</span> <span class="o">=</span> <span class="kc">FALSE</span>
</span></span><span class="line"><span class="cl">  <span class="p">),</span>
</span></span><span class="line"><span class="cl">  <span class="n">pages_navbar</span> <span class="o">=</span> <span class="nf">list</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="nf">chat_nav_panel</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">      <span class="s">&#34;Sources&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="n">tags</span><span class="o">$</span><span class="nf">p</span><span class="p">(</span><span class="s">&#34;Sources selected during this session appear here.&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">      <span class="n">toolbar</span> <span class="o">=</span> <span class="n">bslib</span><span class="o">::</span><span class="nf">toolbar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">        <span class="n">bslib</span><span class="o">::</span><span class="nf">toolbar_input_button</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">          <span class="s">&#34;refresh_sources&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">          <span class="s">&#34;Refresh&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">          <span class="n">icon</span> <span class="o">=</span> <span class="n">bsicons</span><span class="o">::</span><span class="nf">bs_icon</span><span class="p">(</span><span class="s">&#34;arrow-repeat&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">        <span class="p">)</span>
</span></span><span class="line"><span class="cl">      <span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="p">),</span>
</span></span><span class="line"><span class="cl">  <span class="n">drawer</span> <span class="o">=</span> <span class="nf">chat_drawer</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">tags</span><span class="o">$</span><span class="nf">p</span><span class="p">(</span><span class="s">&#34;Select a result to inspect it here.&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">title</span> <span class="o">=</span> <span class="s">&#34;Latest result&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">open</span> <span class="o">=</span> <span class="kc">FALSE</span>
</span></span><span class="line"><span class="cl">  <span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span></span></span></code></pre></div></div>
</div>
<div id="tabset-6-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">faicons</span> <span class="kn">import</span> <span class="n">icon_svg</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shiny</span> <span class="kn">import</span> <span class="n">ui</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shinychat</span> <span class="kn">import</span> <span class="n">chat_drawer</span><span class="p">,</span> <span class="n">chat_nav_panel</span><span class="p">,</span> <span class="n">chat_sidebar</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">shinychat.express</span> <span class="kn">import</span> <span class="n">page_chat</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">page_chat</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;Research assistant&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nb">id</span><span class="o">=</span><span class="s2">&#34;chat&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">toolbar</span><span class="o">=</span><span class="n">ui</span><span class="o">.</span><span class="n">toolbar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">        <span class="n">ui</span><span class="o">.</span><span class="n">toolbar_input_button</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">            <span class="nb">id</span><span class="o">=</span><span class="s2">&#34;clear_chat&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="n">label</span><span class="o">=</span><span class="s2">&#34;Clear conversation&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="n">icon</span><span class="o">=</span><span class="n">icon_svg</span><span class="p">(</span><span class="s2">&#34;arrow-counterclockwise&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">        <span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">toolbar_global</span><span class="o">=</span><span class="n">ui</span><span class="o">.</span><span class="n">toolbar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">        <span class="n">ui</span><span class="o">.</span><span class="n">toolbar_input_button</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">            <span class="nb">id</span><span class="o">=</span><span class="s2">&#34;help&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="n">label</span><span class="o">=</span><span class="s2">&#34;Help&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="n">icon</span><span class="o">=</span><span class="n">icon_svg</span><span class="p">(</span><span class="s2">&#34;question-circle&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">        <span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">sidebar</span><span class="o">=</span><span class="n">chat_sidebar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">        <span class="n">ui</span><span class="o">.</span><span class="n">p</span><span class="p">(</span><span class="s2">&#34;Use filters to focus the results.&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">        <span class="n">history</span><span class="o">=</span><span class="kc">False</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">pages_navbar</span><span class="o">=</span><span class="p">[</span>
</span></span><span class="line"><span class="cl">        <span class="n">chat_nav_panel</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">            <span class="s2">&#34;Sources&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="n">ui</span><span class="o">.</span><span class="n">p</span><span class="p">(</span><span class="s2">&#34;Sources selected during this session appear here.&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">            <span class="n">toolbar</span><span class="o">=</span><span class="n">ui</span><span class="o">.</span><span class="n">toolbar</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">                <span class="n">ui</span><span class="o">.</span><span class="n">toolbar_input_button</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">                    <span class="nb">id</span><span class="o">=</span><span class="s2">&#34;refresh_sources&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">                    <span class="n">label</span><span class="o">=</span><span class="s2">&#34;Refresh&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">                    <span class="n">icon</span><span class="o">=</span><span class="n">icon_svg</span><span class="p">(</span><span class="s2">&#34;arrow-repeat&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">                <span class="p">)</span>
</span></span><span class="line"><span class="cl">            <span class="p">),</span>
</span></span><span class="line"><span class="cl">        <span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">drawer</span><span class="o">=</span><span class="n">chat_drawer</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">        <span class="n">ui</span><span class="o">.</span><span class="n">p</span><span class="p">(</span><span class="s2">&#34;Select a result to inspect it here.&#34;</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">        <span class="n">title</span><span class="o">=</span><span class="s2">&#34;Latest result&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">        <span class="nb">open</span><span class="o">=</span><span class="kc">False</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="p">),</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span></span></span></code></pre></div></div>
</div>
</div>
<p>Users see the <strong>Clear conversation</strong> button while they chat, the <strong>Help</strong> button on every page, a <strong>Sources</strong> page with its own <strong>Refresh</strong> toolbar, and a <strong>Latest result</strong> drawer beside the conversation.</p>
</div>
</details>
<h2 id="show-how-the-model-reached-an-answer">Show how the model reached an answer
</h2>
<p>Understanding how an LLM arrived at an answer is just as &mdash; if not more &mdash; important than getting the answer from the model.
A response can include ordinary text, thinking content, web activity, citations, tool calls, tool results, and custom UI.
shinychat works hard to make the model&rsquo;s work visible and presents each part in a way that helps users understand the answer and what produced it.</p>
<h3 id="keep-tool-calls-readable">Keep tool calls readable
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/tool-calls-collapsed.png" data-fig-alt="A sales assistant conversation where two SQL queries and a schema read appear as compact activity rows above the answer." />
<p>Tool calls are now shown as compact activity rows instead of letting them take over the conversation, refining the <a href="https://opensource.posit.co/blog/2025-11-20_shinychat-tool-ui">tool-call cards shinychat introduced last year</a>.
By default, related calls are grouped together into a single row, and users can still expand a group, open an individual call, and inspect the request and result when they need more detail.</p>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/tool-calls-expanded.png" data-fig-alt="The grouped tool-call row expanded to show the two SQL queries with row count and result previews." />
<p>Opening an individual call shows the request and the result in a card:</p>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/tool-calls-result.png" data-fig-alt="The SQL query expanded to a card showing the full tool call arguments and the query result as a small table." />
<p>Grouping keeps the answer readable, and the request and result stay one click away.
To customize grouping or register tools, see <a href="https://posit-dev.github.io/shinychat/r/articles/tool-ui.html" target="_blank" rel="noopener">Tool UI in shinychat for R</a>, <a href="https://shiny.posit.co/py/docs/genai-tools.html" target="_blank" rel="noopener">Tools in Shiny for Python</a>, <a href="https://ellmer.tidyverse.org/articles/tool-calling.html" target="_blank" rel="noopener">tool/function calling in ellmer</a>, or <a href="https://posit-dev.github.io/chatlas/get-started/tools.html" target="_blank" rel="noopener">tool calling in chatlas</a>.</p>
<h3 id="show-citations-for-web-search-and-fetch">Show citations for web search and fetch
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/citations-popover.png" data-fig-alt="An assistant response where each cited claim is underlined and a pill reading Internal report +1 marks the message&#39;s sources, with the citation popover open just below the pill showing the source name, a link, the supporting passage, and controls to move between the message&#39;s two citations." />
<p>Many LLM providers offer built-in web search and web fetch tools that let your agent search the web, and their APIs return citations when the model uses that content in a reply. shinychat now displays those citations automatically.</p>
<p>For example, here&rsquo;s how to register Claude&rsquo;s tools with an ellmer or chatlas client:</p>
<div class="panel-tabset" data-tabset-group="language">
<ul id="tabset-7" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-7-1">R</a></li>
<li><a href="#tabset-7-2">Python</a></li>
</ul>
<div id="tabset-7-1">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">ellmer</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">client</span> <span class="o">&lt;-</span> <span class="nf">chat_anthropic</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">client</span><span class="o">$</span><span class="nf">register_tool</span><span class="p">(</span><span class="nf">claude_tool_web_search</span><span class="p">())</span>
</span></span><span class="line"><span class="cl"><span class="n">client</span><span class="o">$</span><span class="nf">register_tool</span><span class="p">(</span><span class="nf">claude_tool_web_fetch</span><span class="p">())</span></span></span></code></pre></div></div>
</div>
<div id="tabset-7-2">
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">chatlas</span> <span class="kn">import</span> <span class="n">ChatAnthropic</span><span class="p">,</span> <span class="n">tool_web_fetch</span><span class="p">,</span> <span class="n">tool_web_search</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">client</span> <span class="o">=</span> <span class="n">ChatAnthropic</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">kwargs</span><span class="o">=</span><span class="p">{</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;default_headers&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">            <span class="s2">&#34;anthropic-beta&#34;</span><span class="p">:</span> <span class="s2">&#34;web-fetch-2025-09-10&#34;</span>
</span></span><span class="line"><span class="cl">        <span class="p">}</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">client</span><span class="o">.</span><span class="n">register_tool</span><span class="p">(</span><span class="n">tool_web_search</span><span class="p">())</span>
</span></span><span class="line"><span class="cl"><span class="n">client</span><span class="o">.</span><span class="n">register_tool</span><span class="p">(</span><span class="n">tool_web_fetch</span><span class="p">())</span></span></span></code></pre></div></div>
</div>
</div>
<p>When this client is used with <code>chat_server()</code>, citations are connected directly to the portions of the assistant&rsquo;s response that they support.</p>
<p>Custom retrieval applications, like the RAG systems you can build with <a href="https://ragnar.tidyverse.org/" target="_blank" rel="noopener">ragnar</a> or <a href="https://opensource.posit.co/blog/2026-04-14_rag-with-raghilda">raghilda</a>, can use the same citation UI by prompting the assistant to use a <code>&lt;shiny-aside&gt;</code> tag to attach a source to a claim.</p>
<h3 id="stream-responses-and-show-thinking">Stream responses and show thinking
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/thinking-collapsed.png" data-fig-alt="An assistant response with a collapsed panel reading Thought for 4s between the user&#39;s question and the answer." />
<p>With <code>chat_server()</code> in R or <code>Chat(client=...)</code> in Python, shinychat streams responses and shows supported thinking content in a collapsible panel.
Users can cancel a slow response with the stop button or the Escape key, and the partial response stays in the conversation.</p>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/streaming-stop.png" data-fig-alt="While a response streams in, the send button at the right of the chat input becomes a red stop button." />
<h2 id="add-files-and-shortcuts">Add files and shortcuts
</h2>
<h3 id="attach-files">Attach files
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/attachments-plot.png" data-fig-alt="A plot attached to the chat input as a thumbnail chip above the prompt Explain this plot, with the attach button at the left of the input." />
<p>File attachments are now supported in shinychat! Your users can send images, PDFs, and text files through a file picker, drag and drop, or paste, and shinychat sends each file to the model alongside the user&rsquo;s message. When you use <code>chat_server()</code> in R or <code>Chat(client=...)</code> in Python, your app gets that support for free.</p>
<h3 id="add-slash-commands">Add slash commands
</h3>
<img src="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/slash-commands-palette.png" data-fig-alt="The slash command palette open above the chat input, listing /help, /search, and /clear with short descriptions." />
<p>You can now register chat shortcuts, or <em>slash commands</em>, with <a href="https://posit-dev.github.io/shinychat/r/reference/chat_server.html" target="_blank" rel="noopener"><code>chat$slash_command()</code></a> in R or <a href="https://posit-dev.github.io/shinychat/py/api/Chat.html" target="_blank" rel="noopener"><code>@chat.slash_command()</code></a> in Python. The command palette appears when users type <code>/</code>, and they serve as a way to trigger server-side code, inject context or additional prompting, or even just take an action in your app, all from the chat input.</p>
<p>Check out the <a href="https://posit-dev.github.io/shinychat/r/" target="_blank" rel="noopener">shinychat for R</a> or <a href="https://posit-dev.github.io/shinychat/py/" target="_blank" rel="noopener">shinychat for Python</a> documentation for details.</p>
<h2 id="more-shinychat-powered-apps">More shinychat-powered apps
</h2>
<p>The next release of <a href="https://posit-dev.github.io/querychat/" target="_blank" rel="noopener">querychat</a> will bring these chat features to data applications, including conversation history, attachments, tool displays, and citations.
It will introduce a page-first <code>querychat_app()</code> workflow and a new <code>page()</code> API for adding querychat to an existing Shiny page.</p>
<p><a href="https://posit-dev.github.io/btw/news/index.html#btw-150" target="_blank" rel="noopener">btw 1.5.0</a> already uses shinychat 0.5.0 to give <code>btw_app()</code> a complete coding assistant for your R projects.
It adds conversation history, a <code>page_chat()</code> layout, and slash commands to an assistant that can use your R session, project files, and package documentation.</p>
<h2 id="a-few-changes-for-existing-apps">A few changes for existing apps
</h2>
<p>Existing <code>chat_ui()</code> applications remain supported when chat shares a page with other top-level content. When the conversation should fill the application instead, choose <code>page_chat()</code> and use it as the outermost page container; nesting it inside another page layout breaks the full-window layout and history experience.</p>
<p>In R, <code>chat_mod_ui()</code> and <code>chat_mod_server()</code> are soft-deprecated in favor of pairing <code>chat_ui()</code> and <code>chat_server()</code> by ID. In both languages, a startup message no longer seeds a conversation when history is enabled; use a greeting or append messages through the chat object instead.</p>
<p>The release also protects users from unsafe model-authored Markdown, shows an error when a response fails before streaming starts, and preserves tool results, citations, attachments, and other rich content when users return to a conversation.</p>
<p>With <code>page_chat()</code>, <code>chat_server()</code> or <code>Chat(client=...)</code>, and the history options, you can now give your users a complete chat application: saved conversations they can return to, messages they can edit into new branches, greetings and suggestions to start from, and responses with visible tool calls, citations, and thinking.</p>
<p>Read the <a href="https://posit-dev.github.io/shinychat/r/" target="_blank" rel="noopener">shinychat for R documentation</a> or the <a href="https://posit-dev.github.io/shinychat/py/" target="_blank" rel="noopener">shinychat for Python documentation</a> to explore the examples.
For the complete list of changes, see the <a href="https://github.com/posit-dev/shinychat/blob/main/pkg-r/NEWS.md" target="_blank" rel="noopener">R release notes</a> and the <a href="https://github.com/posit-dev/shinychat/blob/main/pkg-py/CHANGELOG.md" target="_blank" rel="noopener">Python changelog</a>.</p>
<h2 id="acknowledgements">Acknowledgements
</h2>
<p>We thank everyone who contributed to these releases, for opening issues,
submitting pull requests, and providing feedback:
<a href="https://github.com/bastianolea" target="_blank" rel="noopener">@bastianolea</a>,
<a href="https://github.com/bianchenhao" target="_blank" rel="noopener">@bianchenhao</a>,
<a href="https://github.com/christophsax" target="_blank" rel="noopener">@christophsax</a>,
<a href="https://github.com/cpsievert" target="_blank" rel="noopener">@cpsievert</a>,
<a href="https://github.com/crissthiandi" target="_blank" rel="noopener">@crissthiandi</a>,
<a href="https://github.com/elnelson575" target="_blank" rel="noopener">@elnelson575</a>,
<a href="https://github.com/gadenbuie" target="_blank" rel="noopener">@gadenbuie</a>,
<a href="https://github.com/Harshit28j" target="_blank" rel="noopener">@Harshit28j</a>,
<a href="https://github.com/JamesHWade" target="_blank" rel="noopener">@JamesHWade</a>,
<a href="https://github.com/jcheng5" target="_blank" rel="noopener">@jcheng5</a>,
<a href="https://github.com/jlxAtNovozymes" target="_blank" rel="noopener">@jlxAtNovozymes</a>,
<a href="https://github.com/jnhyeon" target="_blank" rel="noopener">@jnhyeon</a>,
<a href="https://github.com/jose-c-milliman" target="_blank" rel="noopener">@jose-c-milliman</a>,
<a href="https://github.com/kaipingyang" target="_blank" rel="noopener">@kaipingyang</a>,
<a href="https://github.com/lucasrod16" target="_blank" rel="noopener">@lucasrod16</a>,
<a href="https://github.com/markmcd" target="_blank" rel="noopener">@markmcd</a>,
<a href="https://github.com/nbenn" target="_blank" rel="noopener">@nbenn</a>,
<a href="https://github.com/parmsam" target="_blank" rel="noopener">@parmsam</a>,
<a href="https://github.com/schloerke" target="_blank" rel="noopener">@schloerke</a>,
<a href="https://github.com/shea-parkes" target="_blank" rel="noopener">@shea-parkes</a>,
<a href="https://github.com/simonpcouch" target="_blank" rel="noopener">@simonpcouch</a>,
<a href="https://github.com/slupczynskim" target="_blank" rel="noopener">@slupczynskim</a>,
<a href="https://github.com/thisisnic" target="_blank" rel="noopener">@thisisnic</a>,
<a href="https://github.com/wlandau" target="_blank" rel="noopener">@wlandau</a>, and
<a href="https://github.com/xx02al" target="_blank" rel="noopener">@xx02al</a>.</p>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>If you&rsquo;re new to LLM apps with Shiny, <a href="https://opensource.posit.co/blog/2025-09-15_shiny-side-of-llms-part-3">Build Your First LLM App with Shiny</a> walks through the process from the beginning in detail.&#160;<a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-15_shinychat-r-0.5.0-python-0.7.1/images/og-header.png" length="78200" type="image/png" />
    </item>
    <item>
      <title>ellmer 0.5.0</title>
      <link>https://opensource.posit.co/blog/2026-09-14_ellmer-0-5-0/</link>
      <pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-14_ellmer-0-5-0/</guid>
      <dc:creator>Nic Crane</dc:creator><description><![CDATA[<p>We are happy to announce that <a href="https://ellmer.tidyverse.org" target="_blank" rel="noopener">ellmer</a> 0.5.0 is now available on CRAN! ellmer is an R package that makes it easy to work with large language models directly from R. It supports a wide variety of providers (including OpenAI, Anthropic, Google, AWS Bedrock, Azure, Snowflake, Databricks, Posit, and many more), makes it easy to extract structured data, and lets the model call R functions via tool calling.</p>
<p>You can install the latest version from CRAN with</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">install.packages</span><span class="p">(</span><span class="s">&#34;ellmer&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
<p>This blog post covers the major changes in this release: a lifecycle update, updates to how you can work with files with ellmer, new ways to update model price data, returning citations when using web search tools, and new hooks for developers building on ellmer&rsquo;s tool loop.</p>
<p>The full list of changes can be found in the <a href="https://github.com/tidyverse/ellmer/releases/tag/v0.5.0" target="_blank" rel="noopener">release notes</a>.</p>
<h2 id="lifecycle">Lifecycle
</h2>
<p><code>chat_github()</code> and <code>models_github()</code> are now defunct, since GitHub Models has been retired.</p>
<p>We&rsquo;ve tightened up what a tool can return: a string, an atomic vector, a JSON string, or a <code>Content</code> object. Returning anything else, like a data frame or a list, now gives a deprecation warning. Previously ellmer converted these to JSON for you, but any problem with the conversion surfaced long after your function had finished, and it was easy to forget that the model can only read the result, not compute with it. For a data frame, convert it yourself:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">get_weather</span> <span class="o">&lt;-</span> <span class="nf">tool</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">  <span class="kr">function</span><span class="p">(</span><span class="n">cities</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="n">df</span> <span class="o">&lt;-</span> <span class="nf">weather_api</span><span class="p">(</span><span class="n">cities</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="n">jsonlite</span><span class="o">::</span><span class="nf">toJSON</span><span class="p">(</span><span class="n">df</span><span class="p">,</span> <span class="n">dataframe</span> <span class="o">=</span> <span class="s">&#34;columns&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="p">},</span>
</span></span><span class="line"><span class="cl">  <span class="kc">...</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span></span></span></code></pre></div></div>
<h2 id="new-features">New features
</h2>
<h3 id="sending-files-to-the-model">Sending files to the model
</h3>
<p>There are now two ways to give the model a file. New <code>content_document_file()</code> and <code>content_document_url()</code> send text-based documents like CSV, Markdown, and code files, just as <code>content_pdf_file()</code> and <code>content_image_file()</code> already do for PDFs and images. The contents go inline with your message, so this works with every provider:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">penguins</span> <span class="o">&lt;-</span> <span class="nf">tempfile</span><span class="p">(</span><span class="n">fileext</span> <span class="o">=</span> <span class="s">&#34;.csv&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">readr</span><span class="o">::</span><span class="nf">write_csv</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">  <span class="nf">data.frame</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">penguin</span> <span class="o">=</span> <span class="nf">c</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">      <span class="s">&#34;Waddlesworth&#34;</span><span class="p">,</span> <span class="s">&#34;Flipper McGee&#34;</span><span class="p">,</span> <span class="s">&#34;Turbo Tuxedo&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="s">&#34;Captain Blubber&#34;</span><span class="p">,</span> <span class="s">&#34;Sir Slidesalot&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">race_time_seconds</span> <span class="o">=</span> <span class="nf">c</span><span class="p">(</span><span class="m">43.2</span><span class="p">,</span> <span class="m">38.7</span><span class="p">,</span> <span class="m">31.9</span><span class="p">,</span> <span class="m">45.1</span><span class="p">,</span> <span class="m">36.4</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="p">),</span>
</span></span><span class="line"><span class="cl">  <span class="n">penguins</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">&lt;-</span> <span class="nf">chat_anthropic</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span><span class="s">&#34;Who won the race?&#34;</span><span class="p">,</span> <span class="nf">content_document_file</span><span class="p">(</span><span class="n">penguins</span><span class="p">))</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Based on the race times, **Turbo Tuxedo** won the race with the fastest time of</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; **31.9 seconds**.</span></span></span></code></pre></div></div>
<p>Inline contents are re-sent with every turn, which adds up over a long conversation, especially with a large file. For those cases, <code>chat$file_upload()</code> sends the file to the provider once and returns a reference you pass to <code>$chat()</code> instead. Because the file isn&rsquo;t re-sent with every message, this also reduces your token usage and costs:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">&lt;-</span> <span class="nf">chat_google_gemini</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">race</span> <span class="o">&lt;-</span> <span class="n">chat</span><span class="o">$</span><span class="nf">file_upload</span><span class="p">(</span><span class="n">penguins</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span><span class="s">&#34;Who won the race?&#34;</span><span class="p">,</span> <span class="n">race</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; **Turbo Tuxedo** won the race with the fastest time of **31.9 seconds**.</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span><span class="s">&#34;And who came last?&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; **Captain Blubber** came in last with the slowest time of **45.1 seconds**.</span></span></span></code></pre></div></div>
<p>You can manage your uploads with <code>chat$file_list()</code>, <code>$file_get()</code>, <code>$file_download()</code>, and <code>$file_delete()</code>. File management works with <code>chat_openai()</code>, <code>chat_anthropic()</code>, and <code>chat_google_gemini()</code>, and replaces the now-deprecated <code>claude_file_upload()</code> and <code>google_upload()</code>.</p>
<h3 id="citations">Citations
</h3>
<p>When a model answers using a built-in web search or fetch tool, the provider usually reports which sources back the answer. ellmer 0.5.0 captures citations from Claude, Google, and OpenAI.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">&lt;-</span> <span class="nf">chat_anthropic</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">register_tool</span><span class="p">(</span><span class="nf">claude_tool_web_search</span><span class="p">())</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span><span class="s">&#34;What are the current stable versions of R and Python? Look them up, one line each, no commentary.&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; R: R version 4.6.1 (Happy Hop) has been released on 2026-06-24[1]</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Python: Python 3.14.7 / 5 August 2026[2]</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Sources</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; [1] R: The R Project for Statistical Computing: https://www.r-project.org/</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; [2] Python (programming language):</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; https://en.wikipedia.org/wiki/Python_(programming_language)</span></span></span></code></pre></div></div>
<p>Citations are also kept in the chat history and included in streamed output, so apps built on ellmer can display them as well.</p>
<h3 id="counting-tokens-and-keeping-prices-current">Counting tokens and keeping prices current
</h3>
<p>Two additions make it easier to know what a conversation will cost before you commit to it. <code>Chat$token_count()</code> asks the provider how many tokens some input would use, without actually sending it to the model.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">&lt;-</span> <span class="nf">chat_anthropic</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">prompt</span> <span class="o">&lt;-</span> <span class="s">&#34;Tell me a joke about an R programmer&#34;</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">token_count</span><span class="p">(</span><span class="n">prompt</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; [1] 19</span></span></span></code></pre></div></div>
<p>Token counting is currently supported by <code>chat_anthropic()</code>, <code>chat_openai()</code>, <code>chat_google_gemini()</code>, <code>chat_google_vertex()</code>, and <code>chat_posit()</code>.</p>
<p>Providers change their prices more often than we release ellmer. You can now update in between releases with <code>models_update_prices()</code>, which downloads the latest pricing data from GitHub and caches it locally.</p>
<h2 id="other-improvements">Other improvements
</h2>
<ul>
<li>Default models have been updated across providers: <code>chat_anthropic()</code>, <code>chat_aws_bedrock()</code>, <code>chat_databricks()</code>, <code>chat_posit()</code>, and <code>chat_snowflake()</code> now use Claude Sonnet 5; <code>chat_openai()</code> and <code>chat_openrouter()</code> use GPT 5.6 Terra; and <code>chat_google_gemini()</code> and <code>chat_google_vertex()</code> use Gemini 3.7 Flash. We update the default models regularly, so if you&rsquo;d prefer to pin your code to a specific model, you should specify it using the <code>model</code> parameter.</li>
<li><code>chat_aws_bedrock()</code> now supports Bedrock Mantle, the newer endpoint that serves models like Claude Mythos and the GPT-5 family through the Anthropic Messages and OpenAI Responses APIs, rather than only the Converse API. ellmer picks the right API from the model name, so this should just work. If you&rsquo;re using a model it doesn&rsquo;t recognize, you can set the new <code>api</code> argument yourself.</li>
<li>You can now stream structured output. <code>Chat$stream()</code> and <code>$stream_async()</code> gain a <code>type</code> argument, which works the same way as in <code>$chat_structured()</code>, for providers that support it.</li>
</ul>
<h3 id="developer-updates">Developer updates
</h3>
<p>Two new features help developers building on ellmer&rsquo;s tool loop.</p>
<p>Inside a tool, <code>tool_context()</code> returns the request that triggered it and the conversation so far, so a tool can make decisions, not just the model. Here, a query tool stops after three calls:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">run_query</span> <span class="o">&lt;-</span> <span class="nf">tool</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">  <span class="kr">function</span><span class="p">(</span><span class="n">sql</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="c1"># count_tool_results() is a stand-in for your own helper</span>
</span></span><span class="line"><span class="cl">    <span class="kr">if</span> <span class="p">(</span><span class="nf">count_tool_results</span><span class="p">(</span><span class="nf">tool_context</span><span class="p">()</span><span class="o">$</span><span class="n">turns</span><span class="p">)</span> <span class="o">&gt;=</span> <span class="m">3</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">      <span class="nf">tool_reject</span><span class="p">(</span><span class="s">&#34;Query budget used up. Answer with what you have.&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl">    <span class="n">jsonlite</span><span class="o">::</span><span class="nf">toJSON</span><span class="p">(</span><span class="n">DBI</span><span class="o">::</span><span class="nf">dbGetQuery</span><span class="p">(</span><span class="n">con</span><span class="p">,</span> <span class="n">sql</span><span class="p">),</span> <span class="n">dataframe</span> <span class="o">=</span> <span class="s">&#34;columns&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">  <span class="p">},</span>
</span></span><span class="line"><span class="cl">  <span class="n">name</span> <span class="o">=</span> <span class="s">&#34;run_query&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">  <span class="n">description</span> <span class="o">=</span> <span class="s">&#34;Run a SQL query&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">  <span class="n">arguments</span> <span class="o">=</span> <span class="nf">list</span><span class="p">(</span><span class="n">sql</span> <span class="o">=</span> <span class="nf">type_string</span><span class="p">())</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span></span></span></code></pre></div></div>
<p><code>Chat</code> also gains <code>$on_request_start()</code> and <code>$on_request_end()</code>, which fire before and after every request to the model, including each round of the tool loop. For example, to time each request:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">chat</span> <span class="o">&lt;-</span> <span class="nf">chat_anthropic</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">register_tool</span><span class="p">(</span><span class="n">run_query</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">started</span> <span class="o">&lt;-</span> <span class="kc">NULL</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">on_request_start</span><span class="p">(</span><span class="nf">\</span><span class="p">(</span><span class="n">turns</span><span class="p">)</span> <span class="n">started</span> <span class="o">&lt;&lt;-</span> <span class="nf">Sys.time</span><span class="p">())</span>
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">on_request_end</span><span class="p">(</span><span class="nf">\</span><span class="p">(</span><span class="n">turn</span><span class="p">)</span> <span class="nf">message</span><span class="p">(</span><span class="s">&#34;Request took &#34;</span><span class="p">,</span> <span class="nf">round</span><span class="p">(</span><span class="nf">Sys.time</span><span class="p">()</span> <span class="o">-</span> <span class="n">started</span><span class="p">,</span> <span class="m">1</span><span class="p">),</span> <span class="s">&#34;s&#34;</span><span class="p">))</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">chat</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span><span class="s">&#34;Find the mean of every column in mtcars, one query at a time.&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Request took 3.6s</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ◯ [tool call] run_query(sql = &#34;SELECT * FROM mtcars LIMIT 5&#34;)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ● #&gt; [{&#34;mpg&#34;:21,&#34;cyl&#34;:6,&#34;disp&#34;:160,&#34;hp&#34;:110,&#34;drat&#34;:3.9,&#34;wt&#34;:2.62,&#34;qsec&#34;:16.46,…</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Now I&#39;ll compute the mean of each column one query at a time, as requested.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Request took 2.7s</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ◯ [tool call] run_query(sql = &#34;SELECT AVG(mpg) AS mean_mpg FROM mtcars&#34;)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ● #&gt; [{&#34;mean_mpg&#34;:20.0906}]</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Request took 1.9s</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ◯ [tool call] run_query(sql = &#34;SELECT AVG(cyl) AS mean_cyl FROM mtcars&#34;)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ● #&gt; [{&#34;mean_cyl&#34;:6.1875}]</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Request took 2s</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ◯ [tool call] run_query(sql = &#34;SELECT AVG(disp) AS mean_disp FROM mtcars&#34;)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; ■ #&gt; Error: Tool call rejected. Query budget used up. Answer with what you</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; have.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; It looks like the query budget has been used up, so I can only report the means</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; I was able to compute before being cut off:</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; | Column | Mean |</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; |--------|------|</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; | mpg | 20.0906 |</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; | cyl | 6.1875 |</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Request took 3.4s</span></span></span></code></pre></div></div>
<p><code>$on_request_start()</code> also receives the turns about to be sent, so an agent can compact its history with <code>chat$set_turns()</code> before the context window fills up.</p>
<h2 id="acknowledgements">Acknowledgements
</h2>
<p>A big thanks to the 73 people who helped make this release possible by filing issues, contributing code, and asking questions: <a href="https://github.com/1beb" target="_blank" rel="noopener">@1beb</a>, <a href="https://github.com/abiyug" target="_blank" rel="noopener">@abiyug</a>, <a href="https://github.com/aclink88" target="_blank" rel="noopener">@aclink88</a>, <a href="https://github.com/AdaemmerP" target="_blank" rel="noopener">@AdaemmerP</a>, <a href="https://github.com/alesanGreat" target="_blank" rel="noopener">@alesanGreat</a>, <a href="https://github.com/Apollo7777777" target="_blank" rel="noopener">@Apollo7777777</a>, <a href="https://github.com/arnavchauhan7" target="_blank" rel="noopener">@arnavchauhan7</a>, <a href="https://github.com/arunrajes" target="_blank" rel="noopener">@arunrajes</a>, <a href="https://github.com/atheriel" target="_blank" rel="noopener">@atheriel</a>, <a href="https://github.com/awunderground" target="_blank" rel="noopener">@awunderground</a>, <a href="https://github.com/bakaburg1" target="_blank" rel="noopener">@bakaburg1</a>, <a href="https://github.com/bastianolea" target="_blank" rel="noopener">@bastianolea</a>, <a href="https://github.com/bshor" target="_blank" rel="noopener">@bshor</a>, <a href="https://github.com/cerebrixos" target="_blank" rel="noopener">@cerebrixos</a>, <a href="https://github.com/CoryMcCartan" target="_blank" rel="noopener">@CoryMcCartan</a>, <a href="https://github.com/cpsievert" target="_blank" rel="noopener">@cpsievert</a>, <a href="https://github.com/D-M4rk" target="_blank" rel="noopener">@D-M4rk</a>, <a href="https://github.com/dareneiri" target="_blank" rel="noopener">@dareneiri</a>, <a href="https://github.com/debruine" target="_blank" rel="noopener">@debruine</a>, <a href="https://github.com/diegomsg" target="_blank" rel="noopener">@diegomsg</a>, <a href="https://github.com/diegoperoni" target="_blank" rel="noopener">@diegoperoni</a>, <a href="https://github.com/dipterix" target="_blank" rel="noopener">@dipterix</a>, <a href="https://github.com/earthcli" target="_blank" rel="noopener">@earthcli</a>, <a href="https://github.com/etiennebacher" target="_blank" rel="noopener">@etiennebacher</a>, <a href="https://github.com/feddelegrand7" target="_blank" rel="noopener">@feddelegrand7</a>, <a href="https://github.com/FrancescoMonti-source" target="_blank" rel="noopener">@FrancescoMonti-source</a>, <a href="https://github.com/frankiethull" target="_blank" rel="noopener">@frankiethull</a>, <a href="https://github.com/gadenbuie" target="_blank" rel="noopener">@gadenbuie</a>, <a href="https://github.com/hadley" target="_blank" rel="noopener">@hadley</a>, <a href="https://github.com/hectorgray" target="_blank" rel="noopener">@hectorgray</a>, <a href="https://github.com/hopessugar" target="_blank" rel="noopener">@hopessugar</a>, <a href="https://github.com/hswerdfe" target="_blank" rel="noopener">@hswerdfe</a>, <a href="https://github.com/JamesHWade" target="_blank" rel="noopener">@JamesHWade</a>, <a href="https://github.com/jamesinottawa" target="_blank" rel="noopener">@jamesinottawa</a>, <a href="https://github.com/jcheng5" target="_blank" rel="noopener">@jcheng5</a>, <a href="https://github.com/jcrodriguez1989" target="_blank" rel="noopener">@jcrodriguez1989</a>, <a href="https://github.com/jeroenjanssens" target="_blank" rel="noopener">@jeroenjanssens</a>, <a href="https://github.com/JosiahParry" target="_blank" rel="noopener">@JosiahParry</a>, <a href="https://github.com/jrosell" target="_blank" rel="noopener">@jrosell</a>, <a href="https://github.com/kaipingyang" target="_blank" rel="noopener">@kaipingyang</a>, <a href="https://github.com/karawoo" target="_blank" rel="noopener">@karawoo</a>, <a href="https://github.com/kbenoit" target="_blank" rel="noopener">@kbenoit</a>, <a href="https://github.com/kchou496" target="_blank" rel="noopener">@kchou496</a>, <a href="https://github.com/klin333" target="_blank" rel="noopener">@klin333</a>, <a href="https://github.com/kolabearafk" target="_blank" rel="noopener">@kolabearafk</a>, <a href="https://github.com/ksr-zguo" target="_blank" rel="noopener">@ksr-zguo</a>, <a href="https://github.com/lazasaurus-ai" target="_blank" rel="noopener">@lazasaurus-ai</a>, <a href="https://github.com/lionel-" target="_blank" rel="noopener">@lionel-</a>, <a href="https://github.com/MLiedgens" target="_blank" rel="noopener">@MLiedgens</a>, <a href="https://github.com/n8layman" target="_blank" rel="noopener">@n8layman</a>, <a href="https://github.com/nbenn" target="_blank" rel="noopener">@nbenn</a>, <a href="https://github.com/neil-bray" target="_blank" rel="noopener">@neil-bray</a>, <a href="https://github.com/nrineausanofi" target="_blank" rel="noopener">@nrineausanofi</a>, <a href="https://github.com/ntentes" target="_blank" rel="noopener">@ntentes</a>, <a href="https://github.com/omorante" target="_blank" rel="noopener">@omorante</a>, <a href="https://github.com/petzi53" target="_blank" rel="noopener">@petzi53</a>, <a href="https://github.com/rajabzadehalidip" target="_blank" rel="noopener">@rajabzadehalidip</a>, <a href="https://github.com/rempsyc" target="_blank" rel="noopener">@rempsyc</a>, <a href="https://github.com/Sade154" target="_blank" rel="noopener">@Sade154</a>, <a href="https://github.com/sarahsdao" target="_blank" rel="noopener">@sarahsdao</a>, <a href="https://github.com/scjohannes" target="_blank" rel="noopener">@scjohannes</a>, <a href="https://github.com/simonpcouch" target="_blank" rel="noopener">@simonpcouch</a>, <a href="https://github.com/Sirhubi007" target="_blank" rel="noopener">@Sirhubi007</a>, <a href="https://github.com/SokolovAnatoliy" target="_blank" rel="noopener">@SokolovAnatoliy</a>, <a href="https://github.com/sounkou-bioinfo" target="_blank" rel="noopener">@sounkou-bioinfo</a>, <a href="https://github.com/stefanlinner" target="_blank" rel="noopener">@stefanlinner</a>, <a href="https://github.com/t-kalinowski" target="_blank" rel="noopener">@t-kalinowski</a>, <a href="https://github.com/Tazinho" target="_blank" rel="noopener">@Tazinho</a>, <a href="https://github.com/thisisnic" target="_blank" rel="noopener">@thisisnic</a>, <a href="https://github.com/thoov08" target="_blank" rel="noopener">@thoov08</a>, <a href="https://github.com/trangdata" target="_blank" rel="noopener">@trangdata</a>, <a href="https://github.com/WvdH-Novus3" target="_blank" rel="noopener">@WvdH-Novus3</a>, and <a href="https://github.com/xmarquez" target="_blank" rel="noopener">@xmarquez</a>.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-14_ellmer-0-5-0/featured.jpg" length="385161" type="image/jpeg" />
    </item>
    <item>
      <title>Positron September Release Highlights</title>
      <link>https://opensource.posit.co/blog/2026-09-09_positron-2026-09-release/</link>
      <pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-09_positron-2026-09-release/</guid>
      <dc:creator>Julia Silge</dc:creator><description><![CDATA[<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p><a href="https://positron.posit.co" target="_blank" rel="noopener">Positron</a> is Posit&rsquo;s new, next-generation IDE for data science. Positron is designed to be an extensible, polyglot tool for exploring data and reproducible authoring in Python, R, and more.</p>
</div>
</div>
<p>Welcome back to another edition of our monthly Positron updates! Each month we share highlights from our <a href="https://positron.posit.co/release-notes" target="_blank" rel="noopener">latest release</a> and useful resources. <a href="https://opensource.posit.co/blog/2026-08-13_positron-2026-08-release">Last release</a> we told you about new Data Connections sources, a round of polish for inline output in Quarto documents, and help installing missing packages. This milestone brings a redesigned welcome page, a first version of Import Data, new Posit Assistant features, expanded Data Connections, package vulnerability scanning, and a more responsive Console.</p>
<h2 id="welcome-page-refresh-and-interpreter-setup">Welcome page refresh and interpreter setup
</h2>
<p>We redesigned the Positron welcome page. It now leads with an environment setup card that checks whether Python and R are ready to use, with actions to help resolve any problem it finds. The Positron badge and name now come with a <strong>Help</strong> button that opens the Help pane, and a banner links to the walkthroughs, including a new &ldquo;Get Started with Positron&rdquo; walkthrough that covers the Positron panes, keyboard shortcuts, built-in extensions, and Git.</p>
<img src="https://opensource.posit.co/blog/2026-09-09_positron-2026-09-release/welcome-page.gif" data-fig-align="center" data-fig-alt="The redesigned Positron Welcome page, showing an Environment setup card with Python checks (3 of 4 passed) including a Create Python Environment button, and R already set up successfully." />
<p>The environment setup theme continues into interpreter selection itself. When you select a Python managed by your operating system or a package manager, Positron now offers to create a virtual environment for your workspace instead of installing packages into that shared interpreter. With no folder open, the environment Positron creates now lives at <code>~/.virtualenvs/positron</code> rather than <code>~/.venv</code>, so it stays in the interpreter picker after a restart, and Positron asks before creating an environment in your home directory. We also removed the confusing startup notification that reported no interpreters were found while linking to documentation saying no setup was needed, and interpreters installed under <code>/opt/python</code> are now labeled <code>Global</code> instead of <code>Unknown</code> in the interpreter picker.</p>
<h2 id="import-data">Import Data
</h2>
<p>Before we started work this month, importing data via the Positron UI was our most upvoted feature request and this release now delivers a first version. When you view a CSV or TSV file in the Data Explorer, a new <strong>Import Data</strong> button in the action bar opens a dialog. The dialog shows the code to load that file into a data frame. You can copy the code, or click <strong>Import</strong> to run it in the console, starting a session if one is not already running.</p>
<img src="https://opensource.posit.co/blog/2026-09-09_positron-2026-09-release/import-data-excel.gif" data-fig-align="center" data-fig-alt="An Excel spreadsheet open in the Data Explorer with an Import Data button in the action bar, next to a Python console session ready to run the generated import code." />
<p>Import Data supports CSV and TSV files in Python with pandas and in R with the readr package, as well as Excel workbooks and Parquet files (with readxl and nanoparquet in R). The generated code can reproduce the filters and sorts you have applied in the Data Explorer, and it names the file by a workspace-relative path when the file is inside your workspace, so the code is easier to share and rerun. You can open the dialog from the Data Explorer, the File menu, the Variables pane, or the File Explorer context menu.</p>
<h2 id="posit-assistant">Posit Assistant
</h2>
<p>This release brings a new <strong>Agent Layout</strong> that opens <a href="https://pos.it/assistant" target="_blank" rel="noopener">Posit Assistant</a> in the editor area with a compact Session pane, giving you more visibility into the agent&rsquo;s actions as it works alongside your code. Configuring language model providers also gets a redesign. Our new Configure LLM Providers modal groups providers by connection state, so you can see at a glance which providers are ready to use. If you have trouble with the new dialog and need to switch back, set <a href="positron://settings/assistant.newProviderModal"><code>assistant.newProviderModal</code></a> to <code>false</code>.</p>
<img src="https://opensource.posit.co/blog/2026-09-09_positron-2026-09-release/provider-modal-dialog.png" data-fig-align="center" data-fig-alt="The Configure LLM Providers modal in Positron, listing connected providers (Posit AI Pass, Anthropic, GitHub Copilot) and additional model providers available to connect (Amazon Bedrock, Microsoft Foundry, OpenAI)." />
<p>You can now configure multiple custom providers, each with its own name, type, endpoint, credential, and model list; the previous single &ldquo;Custom Provider&rdquo; option is now called &ldquo;OpenAI Compatible&rdquo; to better match what it actually does. Amazon Bedrock users get a smoother experience as well. An expired AWS SSO session can be renewed right from the provider modal instead of requiring <code>aws sso login</code> in a terminal, and the AWS profile and region can now be set in the configuration dialog rather than only through environment variables or a hand-edited <code>providers.json</code>. Speaking of which, <code>providers.json</code> now accepts comments, and your comments survive edits that Positron makes to the file.</p>
<h2 id="data-connections">Data Connections
</h2>
<p>The Data Connections preview keeps growing. A new ODBC data connection driver lets you browse any database with an installed ODBC driver in the Connections pane and open it in the Data Explorer. Data sources already configured on your machine appear automatically. Databricks gains OAuth sign-in on desktop and reads <code>DATABRICKS_TOKEN</code>, <code>DATABRICKS_HOST</code>, and <code>DATABRICKS_CONFIG_FILE</code> credentials managed by Posit Workbench automatically. A <strong>Disconnect</strong> option in the context menu closes a connection and any Data Explorers opened from it, and each data connection driver now gets its own log output channel.</p>
<p>Smaller improvements round out the preview. You can now set <a href="positron://settings/dataConnections.enabled"><code>dataConnections.enabled</code></a> per workspace, so a repository can turn on the Connections pane for anyone who opens it. The tree is shallower so table and column names get more of the panel&rsquo;s width, and a new <a href="positron://settings/dataConnections.tree.indent"><code>dataConnections.tree.indent</code></a> setting controls the indentation. DuckDB connections now default to read-only, so the Connections pane and a Python or R session can have the same database open at once; when a lock conflict does happen, the error now explains that another session has locked the database.</p>
<h2 id="package-security-vulnerabilities">Package security vulnerabilities
</h2>
<p>The Packages pane now shows known security vulnerabilities (Common Vulnerabilities and Exposures scoring) for installed Python and R packages, so you can see at a glance whether something in your environment has a known CVE. The data comes from your environment&rsquo;s own Posit Package Manager repository when it has one, and from the public Posit instance otherwise, so what you see reflects the same package source your organization already governs.</p>
<img src="https://opensource.posit.co/blog/2026-09-09_positron-2026-09-release/packages-pane-cve.png" data-fig-align="center" data-fig-alt="The Packages pane showing the tornado package&#39;s Security tab with three known vulnerabilities listed by severity, each with a CVSS score, description, and the version where it was fixed." />
<h2 id="a-more-responsive-console">A more responsive Console
</h2>
<p>Console code submission is now faster, always shows visual feedback, and can be canceled while a completeness check is in flight; the Console no longer waits indefinitely with no feedback when a kernel is slow or unreachable. The Console breaks multi-statement input into complete expressions and executes them statement by statement for languages that support it. A new <a href="positron://settings/console.promptWhenIncomplete"><code>console.promptWhenIncomplete</code></a> setting runs submitted code immediately without a completeness check.</p>
<p>A few more Console fixes are worth knowing about. The Console no longer takes focus at runtime startup when you are working in another view or editor, <strong>Interrupt</strong> stays visible after switching between busy consoles, and session names now ellipsize to fit as the console tab list narrows.</p>
<h2 id="whats-coming-next">What&rsquo;s coming next
</h2>
<ul>
<li>posit::conf(2026) is next week! Our team will have several sessions on Positron, and there is still time to <a href="https://conf.posit.co/2026/" target="_blank" rel="noopener">register</a> to join us virtually from anywhere in the world.</li>
<li>Meet Posit at <a href="https://posit.co/events/cdao-government-2026" target="_blank" rel="noopener">CDAO Government</a> on September 22-23 in Washington, D.C. Stop by our booth to talk data modernization and where agentic AI fits in government.</li>
</ul>
<div class="callout callout-tip" role="note" aria-label="Tip">
<div class="callout-header">
<span class="callout-title">Tip</span>
</div>
<div class="callout-body">
<p><a href="https://positron.posit.co/download" target="_blank" rel="noopener">Download Positron</a> to try out the new features and improvements in this release!</p>
</div>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-09_positron-2026-09-release/featured.svg" length="93678" type="image/svg&#43;xml" />
    </item>
    <item>
      <title>AI Newsletter: You probably don&#39;t want to fine-tune</title>
      <link>https://opensource.posit.co/blog/2026-09-04_ai-newsletter/</link>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-04_ai-newsletter/</guid>
      <dc:creator>Sara Altman</dc:creator>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<div class="callout callout-tip" role="note" aria-label="Tip">
<div class="callout-header">
<span class="callout-title"><strong>Subscribe to the AI Newsletter!</strong></span>
</div>
<div class="callout-body">
<p>The AI newsletter is published as an RSS feed. Follow it in your favorite reader:</p>
<p><a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener noreferrer" class="btn-shortcode inline-flex mb-5 mr-5 items-center px-4 py-3 text-sm leading-5 gap-2 rounded-lg bg-blue-400 !text-white font-semibold align-middle hover:bg-blue-500 transition no-underline">Subscribe via RSS</a></p>
<p><strong>Want the newsletter as an email?</strong> Paste the feed URL, <a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener">https://opensource.posit.co/tags/ai-newsletter/index.xml</a>, into a free RSS-to-email service such as <a href="https://blogtrottr.com/" target="_blank" rel="noopener">Blogtrottr</a>, <a href="https://feedrabbit.com/" target="_blank" rel="noopener">Feedrabbit</a>, or <a href="https://follow.it/" target="_blank" rel="noopener">Follow.it</a>, and each new issue will arrive in your inbox.</p>
</div>
</div>
<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p>This week, we are busy preparing for posit::conf, so we&rsquo;re running an abbreviated version of a post Simon published on his personal blog. You can read the full version <a href="https://simonpcouch.com/blog/2026-09-03-fine-tune" target="_blank" rel="noopener">here</a>.</p>
</div>
</div>
<p>Recently, there&rsquo;s been a lot of talk about fine-tuning models. The reasoning usually goes something like this:</p>
<ul>
<li>There&rsquo;s some task that needs to be done regularly.</li>
<li>Today&rsquo;s frontier models can do it quite reliably, but it&rsquo;s expensive and the bills are starting to rack up.</li>
<li>Cheaper models can&rsquo;t quite do the task reliably, but a fine-tuned variant of a small, open-weights model might be able to just as cheaply.</li>
<li>Therefore, you should fine-tune models for tasks that your organization does.</li>
</ul>
<p>I&rsquo;ve seen <a href="https://seldo.com/posts/2026-is-the-year-of-fine-tuned-small-models/" target="_blank" rel="noopener">a</a> <a href="https://fermisense.com/when-machines-take-the-wheel/" target="_blank" rel="noopener">number</a> <a href="https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency" target="_blank" rel="noopener">of</a> <a href="https://shopify.engineering/sidekicks-continual-learning-loop" target="_blank" rel="noopener">blog</a> <a href="https://turbopuffer.com/blog/reinforcement-learning-sid-ai" target="_blank" rel="noopener">posts</a> cited in support of this idea.</p>
<p>It makes sense that fine-tuning is appealing! AI is getting expensive, many players in the space cannot make the privacy guarantees we&rsquo;d hope, and just as you start to rely on one proprietary model, it&rsquo;s phased out in favor of a newer release. That said, my reaction is that this approach seems more engineering-intensive and, likely, more expensive than alternative approaches. I&rsquo;ll try to make that case here.</p>
<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p>When I say you <em>probably</em> don&rsquo;t want to fine-tune, I mean that I totally understand there are some valid use cases here. Some of the linked blog posts are themselves valid cases. I can see fine-tuning making sense for some <em>very</em> high-volume workloads and/or for asynchronous workloads (where it doesn&rsquo;t matter if a response comes back in a second or a day).</p>
</div>
</div>
<h2 id="what-is-fine-tuning">What is fine-tuning?
</h2>
<p>Before we go further, it&rsquo;s probably worth quickly outlining what I mean by fine-tuning. In short, fine-tuning means training an existing model further to improve its performance on a particular task.</p>
<p>Large language models contain huge arrays of parameters. Input, usually in the form of text, is turned into arrays of numbers, and those arrays get multiplied a bunch of times to form the output, which is then typically converted back into text. Those multiplications are quite computationally intensive, and the array of parameters is itself quite large. For example, the smallest models that are <a href="https://simonpcouch.com/blog/2026-09-02-local-agents-3/" target="_blank" rel="noopener">beginning to be able to</a> reliably complete basic agentic work contain 8 billion parameters (in short, &ldquo;8B&rdquo;).</p>
<p>Fine-tuning changes a model&rsquo;s behavior, either by updating the existing parameters, or by training a <a href="https://huggingface.co/learn/llm-course/en/chapter11/4" target="_blank" rel="noopener">smaller set of additional parameters</a> that work alongside the originals. The goal of fine-tuning is, broadly, to increase the performance of a model on a given task while minimizing the impact on the model&rsquo;s ability to do other tasks.</p>
<h2 id="fine-tuning-is-hard">Fine-tuning is hard
</h2>
<p>It is very, very difficult to successfully fine-tune a model. I don&rsquo;t claim that it&rsquo;s impossible, but it is a substantial engineering effort, requiring the careful attention of dedicated scientists with access to today&rsquo;s frontier models across several weeks, just for the first edition of the model. To show why, let&rsquo;s revisit each step of that process.</p>
<p><strong>What model should you start with?</strong> In short, you want the smallest possible model that has the ability to learn to do the task reliably while retaining sufficient general intelligence. It is very hard to predict when or how capabilities will emerge during fine-tuning without first fine-tuning, meaning that the choice will need to be revisited several times once you&rsquo;ve made a first go at the remaining steps.</p>
<p><strong>What data gets used for training?</strong> Once you&rsquo;ve chosen a model, you need to find some training data to fine-tune it with. What should you use? Ideally, you would use the real data that represents the task and its inputs. However, in order to train on that real data, you typically need the user&rsquo;s consent, and it&rsquo;s likely that you won&rsquo;t have that consent.</p>
<p>And then, when you do have user consent and choose to train on that data, you now have a mandate not to overfit. This is because if the model you&rsquo;re fine-tuning internalizes the data you&rsquo;re training on and is able to (even hazily) recollect it, you&rsquo;ve now exposed that data to any users of the model.</p>
<p>The other approach, then, is synthetic data, created by asking a model to generate a bunch of scenarios that resemble the real inputs and then demonstrate how to carry out the task in those scenarios. But this type of synthetic data has its own problems: notably, the models creating the data have their own set of tics that make the data unrepresentative of real-world data. The <a href="https://www.404media.co/elias-thorne-chatbots-llms-chatgpt-lighthouse-keeper-story/" target="_blank" rel="noopener">Elias in the Lighthouse</a> effect is a notable example of this phenomenon. A wide variety of models will, when asked to write a story, write about a lighthouse keeper named Elias Thorne. On its own, the story of Elias Thorne might be informative training data. However, 1,000 stories about a lighthouse keeper Elias Thorne are not.</p>
<p><strong>How do you nudge the weights?</strong> Let&rsquo;s assume you do have a diverse, representative set of training data to work with. You&rsquo;ll now need to decide how, mechanistically, to change the behavior of the model&mdash;do you need a full fine-tune, or will a <a href="https://huggingface.co/learn/llm-course/en/chapter11/4" target="_blank" rel="noopener">LoRA</a> be sufficient?</p>
<p>Next, you&rsquo;ll need to tune a set of 5-10 hyperparameters. Notably, because they are hyperparameters, there is no generally good &ldquo;magic number&rdquo; for them, and instead you&rsquo;ll need to try a bunch of values and see what works. You&rsquo;d make some guesses, see what happens, then generate some hypotheses on what a given change to those parameters might do, then try out a different number and see if it has the desired effect. Every time you&rsquo;re figuring out what to try next, you&rsquo;ll be reading a <em>ton</em> of test cases and trying to observe general failure modes exhibited in this intermediate draft of your fine-tuned model. That might change the initial choice of model you&rsquo;re fine-tuning, and it might change which subsets of the training data you&rsquo;re exposing to the training process.</p>
<p><strong>How do you measure success?</strong> You need a reliable way to tell whether each version of your model is actually improving. That is especially tricky for free-text tasks, because two answers can be equally good even when they differ syntactically. Because of this, you&rsquo;ll probably want an LLM-as-a-judge system, where another model compares the responses to a target and grades according to a rubric you&rsquo;ve supplied. LLM-as-a-judge systems are themselves quite hard to design correctly, and you&rsquo;ll be reading a bunch of that judge&rsquo;s grading transcripts, too.</p>
<p>It&rsquo;s also worth mentioning that <strong>getting your fine-tuned model to score well on your own benchmark is not the hard part</strong>. Many of these fine-tuning writeups include a plot along these lines:</p>
<img src="https://opensource.posit.co/blog/2026-09-04_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/plot-finetune-benchmark-1.png" style="width:100.0%" data-fig-align="center" data-fig-alt="A scatter plot of evaluation performance against cost per million tokens (log scale). Unlabeled grey synthetic points loosely follow a rising trend. Labeled model points include Fable 5.1, GPT 5.6 Sol, Sonnet 5, Gemma 4 26B, Llama 4 8B, and Phi 5 mini. A red point labeled &#39;Our fine-tune&#39; sits in the upper left, cheap yet scoring above the frontier models." />
<p>When I see that a single-digit-billion-parameter fine-tuned model scores better than Fable 5.1 or GPT 5.6 Sol on an evaluation, I interpret that as evidence that the evaluation is not meaningful. In my own experience, it is not (comparatively) hard to get a small model to score very well on any given benchmark. What&rsquo;s much more difficult is preserving the broad intelligence of a model while driving the evaluation score up. A meaningful evaluation can do both at once, measuring task performance under a realistic, broad distribution of possible task configurations. It is very hard to author meaningful evaluations, especially for models as capable as those that exist today.</p>
<p>But let&rsquo;s say you&rsquo;ve made it this far! You found a good model to start from, a diverse set of training data that you&rsquo;ve obtained consent to train on, and a meaningful way to measure progress. Now, it&rsquo;s time to put it in production.</p>
<h2 id="fine-tuning-is-expensive">Fine-tuning is expensive
</h2>
<p>Fine-tuning can seem appealing because the actual fine-tuning part of the process can be relatively inexpensive, likely on the order of a few dollars to a few hundred dollars.<sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup></p>
<p>However, actually deploying the model to do the task you trained it to do is likely to be substantially more expensive.</p>
<h3 id="because-hosting-is-expensive">&hellip;because hosting is expensive
</h3>
<p>Hosting a model yourself is more expensive than using a similarly capable model hosted by someone else unless you have extraordinary volume. Frontier-model providers like Anthropic and OpenAI serve extraordinarily large volumes, keeping their compute almost always near max capacity. It&rsquo;s likely that the same won&rsquo;t be true for your fine-tuned, self-hosted model.</p>
<p>As an example, let&rsquo;s say I host Gemma 4 26B A4B on a single H100. I can rent a high-availability <a href="https://lambda.ai/instances" target="_blank" rel="noopener">H100</a> <a href="https://fireworks.ai/pricing" target="_blank" rel="noopener">GPU</a> <a href="https://www.baseten.co/pricing/" target="_blank" rel="noopener">instance</a> at $0.0833 a minute, or $44,000 a year. Assuming a coding-agent-like workload,<sup id="fnref:2"><a href="#fn:2" class="footnote-ref" role="doc-noteref">2</a></sup> that would buy about 96 billion <a href="https://platform.claude.com/docs/en/about-claude/pricing#:~:text=%2475%20/%20MTok-,Claude%20Sonnet%205,%2410%20/%20MTok,-Claude%20Sonnet%204.6" target="_blank" rel="noopener">Sonnet 5</a> tokens, and Sonnet 5 is a much more capable model. To break even, this fine-tuned model would need to serve about 182,000 tokens per minute, 24/7, throughout the year.</p>
<p>It might be possible to serve that many tokens. The hard part is finding enough demand for one task to keep the model busy, and that demand needs to be relatively smooth. If the model sits idle, you&rsquo;re still paying for it. If demand spikes, you either rent a second GPU or accept worse performance for users.</p>
<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p>I&rsquo;m not arguing against self-hosting in general. Self-hosting (either literally on your organization&rsquo;s own hardware or by renting hardware and serving open-weights models on it with open source software) can be quite cost-effective at sufficient scale. If you know that a <em>lot</em> of traffic will go through that endpoint, do the same napkin math shown above and see if you can save some money. What I&rsquo;m particularly arguing against is that it&rsquo;s a good idea to self-host a model <em>that can only do one thing</em>. Unless you serve an extraordinary amount of traffic that does that task specifically, the payoff is not there.</p>
</div>
</div>
<h3 id="and-you-must-host">&hellip;and you must host
</h3>
<p>The other suggestion here is that, if the model is small enough, users who need to do the task can just download the weights and run it on their laptop or a dedicated workstation.</p>
<p>The effectiveness of this argument depends on the kind of task to be done. However, it&rsquo;s worth considering the potential for this to be a very unpleasant user experience compared to just using a model that someone else hosts. For example, let&rsquo;s say my colleague does some task for an hour every week and there&rsquo;s some model small enough to run on my laptop that&rsquo;s capable of doing the task. In order to use the model for that task, my colleague would need to download the 4GB or 8GB or 100GB or whatever of weights. That user would also need to be capable of configuring the serving of the model, and they&rsquo;d need to make that happen every time they started the task. If that model was accessed through a tool that the user was using regularly anyway&mdash;Claude Code pointed at a local ollama model, for instance&mdash;they&rsquo;d need to remember to switch the model over in the application&rsquo;s settings. If that model was accessed through a different interface, the user would need to use a different interface than they normally use LLMs with for that task specifically. Neither of these are good UX.</p>
<p>Even if the software for locally serving models got <em>much</em> more pleasant than it currently is, it&rsquo;s hard to compete with the experience of pay-as-you-go for the Everything Tool that wraps the Everything Model.</p>
<h2 id="your-fine-tune-will-quickly-fall-behind">Your fine-tune will quickly fall behind
</h2>
<p>Let&rsquo;s say you successfully fine-tune a model based on some fictional model Kuen 3. Then, the following week, Kuen 3.5 is released. 3 of the 100 most capable LLM scientists on this planet worked on it. It&rsquo;s almost as good as your fine-tune on that specific task, and it&rsquo;s also broadly capable at a very broad array of tasks.</p>
<p>Fine-tuning is not a boat that is lifted by the rising tide of broader AI progress&mdash;in order to take advantage of the Kuen 3.5 release with your fine-tune, you&rsquo;d need to restart that process of choosing parameters and observing fine-tuning runs. The underlying architecture of the model may have changed, and thus the approaches that you used to fine-tune Kuen 3 might not work for Kuen 3.5 on your first try.</p>
<h2 id="what-you-should-do-instead">What you should do instead
</h2>
<p>Instead, the boat that is lifted by a rising tide in this context is plain old prompt engineering. (POPE, as <a href="https://github.com/jcheng5" target="_blank" rel="noopener">Joe Cheng</a> calls it.) Choose a model that&rsquo;s available to you, that fits your price point, and seems broadly capable across a wide variety of public benchmarks. (Extra points if the vibes on the model from people you trust are good, and extra points if someone else is serving it.)</p>
<p>Then, put together a short prompt telling the model how to do the task, perhaps inside of some coding agent harness like Claude Code (or, hey, <a href="https://assistant.posit.co/" target="_blank" rel="noopener">Posit Assistant</a>), and see what it does. Adjust the prompt to tell it how to do things correctly that it tends to trip up on, and iterate from there. If you&rsquo;ve already put together an evaluation as part of your fine-tuning process, you could even reuse that! My colleague Sara and I have written about prompt engineering in the past if you&rsquo;re interested in learning more, once about the more specific task of POPE for <a href="https://posit.co/blog/custom-chat-app" target="_blank" rel="noopener">teaching LLMs about R packages</a> but in relatively generalizable ways, and once focused on <a href="https://opensource.posit.co/blog/2026-07-03_ai-newsletter/" target="_blank" rel="noopener">deciding between the popular ways to deliver context to coding agents</a>.</p>
<h2 id="recent-past-newsletters">Recent past newsletters
</h2>
<ul>
<li><a href="https://opensource.posit.co/blog/2026-08-14_ai-newsletter">How to choose a model</a></li>
<li><a href="https://opensource.posit.co/blog/2026-07-31_ai-newsletter">Keep track of your data exploration with Posit Assistant&rsquo;s EDA log</a></li>
</ul>
<p><a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener noreferrer" class="btn-shortcode inline-flex mb-5 mr-5 items-center px-4 py-3 text-sm leading-5 gap-2 rounded-lg bg-blue-400 !text-white font-semibold align-middle hover:bg-blue-500 transition no-underline">Subscribe via RSS</a></p>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>Various technical details might mean this is an order of magnitude or two off. Regardless, the larger point stands that a single fine-tuning run is not the expensive part of deploying a fine-tuned model.&#160;<a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
<li id="fn:2">
<p>Meaning around 90% of tokens are cached input, 9% of tokens are uncached input, and 1% of tokens are output.&#160;<a href="#fnref:2" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-04_ai-newsletter/images/hero.png" length="182577" type="image/png" />
    </item>
    <item>
      <title>vitals 0.4.0</title>
      <link>https://opensource.posit.co/blog/2026-09-03_vitals-0-4-0/</link>
      <pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-09-03_vitals-0-4-0/</guid>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<p>I&rsquo;m as amped as Ella Langley&rsquo;s Gibson to share that <a href="https://vitals.tidyverse.org/" target="_blank" rel="noopener">vitals</a> 0.4.0 is now on CRAN! vitals implements a large language model evaluation toolkit for R, and this release contains several exciting features.</p>
<p>To install the newest release, run the following in R:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">install.packages</span><span class="p">(</span><span class="s">&#34;vitals&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
<p>The package includes two new helpers, <a href="https://vitals.tidyverse.org/reference/agent_solvers.html" target="_blank" rel="noopener"><code>claude_code()</code></a> and <a href="https://vitals.tidyverse.org/reference/agent_solvers.html" target="_blank" rel="noopener"><code>codex()</code></a>, which allow
you to compare your own ellmer-built agents with leading coding agents. This release also ships another new helper, <a href="https://vitals.tidyverse.org/reference/vitals_log_read.html" target="_blank" rel="noopener"><code>vitals_log_read()</code></a>, which supports reading log files back into tibbles, including columns of resumable ellmer Chats. Finally, the release includes several performance improvements; log files are much smaller, and the log viewer that reads them is now substantively faster.</p>
<p>To read the full list of changes, see the <a href="https://vitals.tidyverse.org/news/index.html#vitals-040" target="_blank" rel="noopener">changelog</a>.</p>
<h2 id="agent-solvers">Agent solvers
</h2>
<p>vitals is a port of <a href="https://inspect.aisi.org.uk/" target="_blank" rel="noopener">Inspect</a>, a well-adopted Python framework for LLM eval from Posit&rsquo;s own JJ Allaire. One of the concepts that vitals borrows from Inspect is the concept of a &ldquo;solver,&rdquo; or the LLM-powered system that sets out to solve some task. The simplest solver is just the LLM itself, with no system prompt or tools, like what you&rsquo;d get from running <code>chat_anthropic()</code> from ellmer. Solvers can gain all sorts of prompts and tools, which allows vitals users to test the effect of a change in their prompt or the addition of a new tool.</p>
<p>In the last year or so, the dominant interface to solvers in Inspect has become &ldquo;agent solvers&rdquo;: interfaces to the popular coding agents Claude Code and Codex. You call the helper <code>claude_code()</code> or <code>codex()</code>, and Inspect will proxy traffic through the real coding agent harness.</p>
<p>vitals now has first-class support for these two helpers, allowing users to compare their own agents built with ellmer to popular coding agents like Claude Code and Codex. You provide the set of tasks and grading guidance, and vitals will take care of the communication with Inspect.</p>
<h2 id="read-eval-logs-back-into-ellmer-chats">Read eval logs back into ellmer Chats
</h2>
<p>One of the big annoyances I&rsquo;ve had in my own usage of vitals is log storage. So that users can use Inspect&rsquo;s log viewer directly, we write evaluation logs to a JSON format that Inspect can read.^[1] However, I often want to write R code against the original R objects—ellmer Chats especially—that the logs were generated from. Loading in the ellmer Chats would especially be helpful for inspecting (ha!) the conversation histories in the same interface that users of the ellmer application would see.</p>
<p>Because of this, I&rsquo;ve often saved <em>both</em> the JSON logs and <code>.rda</code> logs, the latter of which contain the ellmer Chats. These files are large on their own, and it feels even more silly passing around duplicates of them.</p>
<p>The new release of vitals introduces <code>vitals_log_read()</code>, which reads an eval log file back into a tibble of samples. (It&rsquo;s almost exactly what you&rsquo;d get if you ran the <code>get_samples()</code> method on a vitals Task object.) That tibble includes reconstructed solver (and, for model-graded scorers, scorer) chats as ellmer Chat objects. For some providers, the chats will even be resumable; you can load in a solver from a JSON file into an R session and ask that solver a question yourself.</p>
<h2 id="performance-improvements">Performance improvements
</h2>
<p>The long and the short of this section is just to say that:</p>
<ol>
<li>Logs will take up less storage space than they did before. Roughly, logs written with the new vitals version will be 4x smaller than before, and the magnitude of savings increases with the complexity of the log.</li>
<li>We now display logs (with <code>vitals_view()</code>) <em>much</em> more quickly. The log viewer should feel very snappy for almost all uses of the package.</li>
</ol>
<p>I&rsquo;m really excited to have this release on CRAN! Take it for a spin and let me know if you run into issues on the <a href="https://github.com/tidyverse/vitals" target="_blank" rel="noopener">package repository</a>.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-09-03_vitals-0-4-0/featured.png" length="414699" type="image/png" />
    </item>
    <item>
      <title>Neural networks in Orbital for Python 0.6.0: PyTorch straight to your database</title>
      <link>https://opensource.posit.co/blog/2026-08-17_pyorbital-0-6-0/</link>
      <pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-08-17_pyorbital-0-6-0/</guid>
      <dc:creator>Alessandro Molina</dc:creator><description><![CDATA[<p>Over the past couple of months I&rsquo;ve been teaching Orbital to speak PyTorch.</p>
<p>Orbital&rsquo;s whole pitch is that a trained model becomes SQL, so a database can run predictions on its own, with no Python process anywhere near it. Until 0.6.0, &ldquo;trained model&rdquo; meant scikit-learn: pipelines, trees, linear models, all <code>.fit()</code> in Python and then turned into a <code>SELECT</code> statement. It never covered what a lot of teams are actually training now: PyTorch models, not scikit-learn pipelines.</p>
<p>Orbital 0.6.0 closes that gap. A <code>torch.nn.Sequential</code> network, trained exactly the way you already train it, now compiles to the same kind of SQL a linear regression would.</p>
<p>No ONNX Runtime. No model server. No separate inference service to keep alive next to the database. Just a query.</p>
<p>Why does that work at all, for a framework Orbital was never written for? Because of a decision made long before PyTorch was ever on the table.</p>
<h2 id="why-this-isnt-a-bolt-on">Why this isn&rsquo;t a bolt-on
</h2>
<p>Orbital was never really a scikit-learn tool. Underneath, it converts a scikit-learn pipeline to ONNX (Open Neural Network Exchange, a standard graph format for trained models) using the <code>skl2onnx</code> library, then walks that graph node by node to produce SQL. Scikit-learn was always one hop removed from what Orbital actually translates.</p>
<p>Adding PyTorch meant taking the same hop from a different starting point. <code>torch.onnx.export</code> turns a <code>torch.nn.Sequential</code> model into that same kind of graph. Feed it into the translator that already existed, and the translator doesn&rsquo;t know or care whether the graph came from PyTorch or scikit-learn.</p>
<p>The proof is in how little new code that took. The entire engine for running a feed-forward network in SQL is three small classes: a <code>Gemm</code> translator for <code>Linear</code> layers, <code>ReLU</code>, <code>Sigmoid</code>. Everything else (the translator base class, the optimizer, the per-dialect SQL compiler) already existed, built earlier for trees and linear models.</p>
<p>When I first thought of support for PyTorch, I put it this way: <em>&ldquo;the underlying translation works on ONNX graphs&hellip; the same value proposition orbital already provides for scikit-learn models can apply as it is to pytorch networks exported to ONNX.&rdquo;</em> That sentence turned out to be the whole implementation plan.</p>
<p>If you want the fuller picture of how a graph becomes SQL, parser, then translator, then optimizer, the <a href="https://posit-dev.github.io/orbital/learnmore/" target="_blank" rel="noopener">architecture docs</a> walk through all three stages in order.</p>
<p>That architecture is also why I keep calling this <strong>multiple frameworks</strong>, not two frameworks. <code>Relu</code>, <code>Sigmoid</code>, <code>Tanh</code>, and <code>Softmax</code> are single translators, not one per framework. Scikit-learn&rsquo;s <code>MLPClassifier</code> and <code>MLPRegressor</code> reach them through <code>MatMul</code> and <code>Add</code>, PyTorch&rsquo;s <code>nn.Sequential</code> reaches the exact same translators through <code>Gemm</code>. One implementation, two entry points, and no reason it has to stop at two.</p>
<p>Scikit-learn and PyTorch are both real and shipping today, which is enough on its own to call this &ldquo;multiple frameworks.&rdquo; But the dependency story is already moving that direction: <a href="https://github.com/posit-dev/orbital/issues/113" target="_blank" rel="noopener">issue #113</a> proposes turning scikit-learn itself into an optional dependency, the same way PyTorch already is, so the core stops assuming any particular framework at all.</p>
<h2 id="making-it-actually-usable">Making it actually usable
</h2>
<p>Neural networks are layered, and every neuron in layer two reads every output of layer one. If Orbital just inlines those outputs at each place they&rsquo;re read, instead of naming them once, that repetition compounds from one layer to the next. Two layers doubles the inlined text. Five layers is a different order of magnitude.</p>
<p>That&rsquo;s not theoretical: <a href="https://github.com/posit-dev/orbital/issues/115" target="_blank" rel="noopener">the issue that tracked the fix</a> measured scikit-learn&rsquo;s own default <code>MLPClassifier(hidden_layer_sizes=(100,))</code>, the first thing anyone reaches for, at 53MB of generated SQL and roughly 894 seconds just to generate it. A hundred neurons in one hidden layer, and the query was already unusable.</p>
<p>PyTorch&rsquo;s <code>Gemm</code> translator never had this problem. It has called <code>preserve()</code>, materializing its output as a real SQL column, since the day it was written. Scikit-learn&rsquo;s <code>MLPClassifier</code> and <code>MLPRegressor</code> compile through different ONNX ops though: <code>MatMul</code> then <code>Add</code>, because that&rsquo;s what <code>skl2onnx</code> emits, not <code>Gemm</code>. Neither of those translators called <code>preserve()</code> at all.</p>
<p>The fix, <code>Optimizer.preserve_referenced_outputs()</code>, runs after every single node in the translation loop, for every translator, not just <code>MatMul</code> and <code>Add</code>, and checks how many times that node&rsquo;s output is actually referenced downstream. Referenced more than once, it gets materialized as a named column. Referenced once or not at all, it stays inlined, no extra column, no extra noise.</p>
<p>None of this is new machinery either. Tree ensembles already lean on the same trick: <code>preserve()</code> materializes per-tree votes, or the whole ensemble&rsquo;s aggregated vote so it isn&rsquo;t re-emitted everywhere it&rsquo;s read, as real SQL columns. <code>preserve_referenced_outputs()</code> generalizes that same idea automatically, for every translator, whether it&rsquo;s part of an ordinary pipeline or a neural network.</p>
<p>Here&rsquo;s what that fix was worth, measured on three shapes while it was being built:</p>
<table>
  <thead>
      <tr>
          <th>Network</th>
          <th>Before</th>
          <th>After</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><code>20→64→64→1</code> (Orbital&rsquo;s own deep-network scaling test)</td>
          <td>12.4MB, ~185s to generate</td>
          <td>234KB</td>
      </tr>
      <tr>
          <td><code>MLPClassifier(hidden_layer_sizes=(100,))</code>, 3-class (scikit-learn&rsquo;s own default)</td>
          <td>53MB, ~894s to generate</td>
          <td>46.7KB</td>
      </tr>
      <tr>
          <td><code>MLPRegressor(hidden_layer_sizes=(32, 32))</code></td>
          <td>1.75MB, ~21s to generate</td>
          <td>57.6KB</td>
      </tr>
  </tbody>
</table>
<p>Generation time collapsed just as hard: the <code>20→64→64→1</code> network above went from about 185 seconds to about 3.6 seconds. Running the resulting SQL got faster too, if less dramatically: on 200,000 rows, an <code>MLP(32,32)</code> query dropped from 0.43s to 0.35s, and an <code>MLP(100,100)</code> from 3.58s to 3.31s.</p>
<p>Same hyperparameters. Same defaults everyone actually reaches for. The difference is entirely in how the SQL gets built, not in what the network computes.</p>
<h2 id="what-it-can-do-today">What it can do today
</h2>
<p>Neural network support in 0.6.0 covers binary classification, multiclass classification, and regression, for both scikit-learn and PyTorch. Five new or updated examples in the repo prove it out: <code>pytorch_fraud_detector.py</code>, <code>pytorch_maintenance_classifier.py</code>, <code>pytorch_demand_regressor.py</code>, <code>pipeline_mlp_classifier.py</code>, <code>pipeline_mlp_regressor.py</code>.</p>
<p>The one worth walking through is the fraud detector, since it&rsquo;s the shape most teams actually need: a handful of numeric features, a binary &ldquo;is this fraud&rdquo; output, trained the same way this kind of model always is.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="n">FEATURES</span> <span class="o">=</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;amount&#34;</span><span class="p">:</span> <span class="n">orbital</span><span class="o">.</span><span class="n">types</span><span class="o">.</span><span class="n">DoubleColumnType</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;hour&#34;</span><span class="p">:</span> <span class="n">orbital</span><span class="o">.</span><span class="n">types</span><span class="o">.</span><span class="n">DoubleColumnType</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;v1&#34;</span><span class="p">:</span> <span class="n">orbital</span><span class="o">.</span><span class="n">types</span><span class="o">.</span><span class="n">DoubleColumnType</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;v2&#34;</span><span class="p">:</span> <span class="n">orbital</span><span class="o">.</span><span class="n">types</span><span class="o">.</span><span class="n">DoubleColumnType</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">model</span> <span class="o">=</span> <span class="n">torch</span><span class="o">.</span><span class="n">nn</span><span class="o">.</span><span class="n">Sequential</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">torch</span><span class="o">.</span><span class="n">nn</span><span class="o">.</span><span class="n">Linear</span><span class="p">(</span><span class="nb">len</span><span class="p">(</span><span class="n">FEATURES</span><span class="p">),</span> <span class="mi">16</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">torch</span><span class="o">.</span><span class="n">nn</span><span class="o">.</span><span class="n">ReLU</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl">    <span class="n">torch</span><span class="o">.</span><span class="n">nn</span><span class="o">.</span><span class="n">Linear</span><span class="p">(</span><span class="mi">16</span><span class="p">,</span> <span class="mi">8</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">torch</span><span class="o">.</span><span class="n">nn</span><span class="o">.</span><span class="n">ReLU</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl">    <span class="n">torch</span><span class="o">.</span><span class="n">nn</span><span class="o">.</span><span class="n">Linear</span><span class="p">(</span><span class="mi">8</span><span class="p">,</span> <span class="mi">1</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">torch</span><span class="o">.</span><span class="n">nn</span><span class="o">.</span><span class="n">Sigmoid</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># ... train model normally: Adam, BCELoss, a plain training loop ...</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">orbital_pipeline</span> <span class="o">=</span> <span class="n">orbital</span><span class="o">.</span><span class="n">parse_pytorch_model</span><span class="p">(</span><span class="n">model</span><span class="p">,</span> <span class="n">FEATURES</span><span class="p">)</span></span></span></code></pre></div></div>
<p>From that one <code>orbital_pipeline</code>, two engines:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="n">duckdb_sql</span> <span class="o">=</span> <span class="n">orbital</span><span class="o">.</span><span class="n">export_sql</span><span class="p">(</span><span class="s2">&#34;transactions&#34;</span><span class="p">,</span> <span class="n">orbital_pipeline</span><span class="p">,</span> <span class="n">dialect</span><span class="o">=</span><span class="s2">&#34;duckdb&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">postgres_sql</span> <span class="o">=</span> <span class="n">orbital</span><span class="o">.</span><span class="n">export_sql</span><span class="p">(</span><span class="s2">&#34;transactions&#34;</span><span class="p">,</span> <span class="n">orbital_pipeline</span><span class="p">,</span> <span class="n">dialect</span><span class="o">=</span><span class="s2">&#34;postgres&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
<p>I ran both, against a real DuckDB and a real Postgres, not just read the generated text. Same four test transactions, three ways to compute a prediction (PyTorch itself, the DuckDB query, the Postgres query), and all three agree to within 1e-5. Both queries land around 10KB, not the tens of megabytes a naive translation would have produced before 0.6.0&rsquo;s optimizer fix.</p>
<p>That works because <code>Sigmoid</code> and <code>ReLU</code> both compile to plain arithmetic: <code>EXP</code>, a division, a <code>CASE WHEN</code>. Every SQL engine has those. It&rsquo;s not an accident which activation this example uses, either: <code>Tanh</code> compiles to a native <code>TANH()</code> call instead, and not every dialect implements that the same way, so it&rsquo;s the one activation in Orbital&rsquo;s NN support with an actual portability caveat attached.</p>
<p>DuckDB and Postgres are two of the three dialects <a href="https://posit-dev.github.io/orbital/learnmore/" target="_blank" rel="noopener">Orbital actively tests in CI</a>. SQLite is the third. But that&rsquo;s a testing choice, not an architecture boundary: Orbital doesn&rsquo;t write dialect-specific SQL at all. Translation ends at ibis. <code>export_sql</code> just hands the finished expression to whichever of ibis&rsquo;s own backend compilers matches the dialect you ask for, and ibis ships about twenty of those: Snowflake, BigQuery, Trino, MySQL, and so on.</p>
<h2 id="limits-honestly">Limits, honestly
</h2>
<p>What Orbital 0.6.0 handles is feed-forward, fully connected networks: stacks of <code>Linear</code> layers with <code>ReLU</code>, <code>Sigmoid</code>, <code>Tanh</code>, or <code>Softmax</code> in between. That already covers real use cases people put into production: fraud scoring, churn, demand forecasting, risk models. None of those need a CNN or a transformer.</p>
<p>What it doesn&rsquo;t do yet is exactly what that shape excludes: no convolutions, no recurrence, no attention, no embedding layers for categorical features. If your model needs any of those, Orbital isn&rsquo;t there yet.</p>
<p>SQL size still grows with the network. A few hidden layers of 64 to 128 neurons land comfortably in the KB range. Wider or deeper than that, it&rsquo;s worth checking the generated SQL against whatever statement-size limit your engine has, before you deploy it.</p>
<p>That headroom exists at all thanks to <code>Optimizer.preserve_referenced_outputs()</code>, and that mechanism helps every translator in Orbital, not just neural networks. It&rsquo;s a big enough story on its own that I&rsquo;ll get back to in a future post.</p>
<h2 id="try-it">Try it
</h2>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">pip install orbital<span class="o">[</span>pytorch<span class="o">]</span></span></span></code></pre></div></div>
<p>From there, the <a href="https://posit-dev.github.io/orbital/getstarted/" target="_blank" rel="noopener">getting-started guide</a> walks through this same fraud-detector shape end to end, and the <a href="https://github.com/posit-dev/orbital/tree/main/examples" target="_blank" rel="noopener">examples directory</a> has five more, covering classification and regression in both frameworks.</p>
<p>Orbital speaks PyTorch now. Scikit-learn too. Same query either way.</p>
<p>There are more frameworks already in the works. I won&rsquo;t name them here, only that the architecture was built for exactly this.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-08-17_pyorbital-0-6-0/featured.png" length="1070756" type="image/png" />
    </item>
    <item>
      <title>AI Newsletter: How to choose a model</title>
      <link>https://opensource.posit.co/blog/2026-08-14_ai-newsletter/</link>
      <pubDate>Fri, 14 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-08-14_ai-newsletter/</guid>
      <dc:creator>Sara Altman</dc:creator>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-body">
<p><strong>Subscribe to the AI Newsletter!</strong></p>
<p>The AI newsletter is published as an RSS feed. Follow it in your favorite reader:</p>
<p><a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener noreferrer" class="btn-shortcode inline-flex mb-5 mr-5 items-center px-4 py-3 text-sm leading-5 gap-2 rounded-lg bg-blue-400 !text-white font-semibold align-middle hover:bg-blue-500 transition no-underline">Subscribe via RSS</a></p>
<p><strong>Want the newsletter as an email?</strong> Paste the feed URL — <a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener">https://opensource.posit.co/tags/ai-newsletter/index.xml</a> — into a free RSS-to-email service such as <a href="https://blogtrottr.com/" target="_blank" rel="noopener">Blogtrottr</a>, <a href="https://feedrabbit.com/" target="_blank" rel="noopener">Feedrabbit</a>, or <a href="https://follow.it/" target="_blank" rel="noopener">Follow.it</a>, and each new issue will arrive in your inbox.</p>
</div>
</div>
<br>
<p>How do you know which model to use and when? Often, it&rsquo;s not a question of which model is best, but of which model suits your task and needs at a given time. You might switch between models for different projects (package development vs. data analysis), or even within a single project (planning vs. implementation). Different tasks require a different mix of cost, token usage, speed, intelligence, and capabilities.</p>
<p>If you want our most durable, high-level advice: <strong>start with the most expensive model you have access to from either OpenAI or Anthropic.</strong><sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup> Then, once you have a sense of the &ldquo;ceiling,&rdquo; try less expensive models and see how they compare. Developing a feel for what&rsquo;s possible with LLMs will help you make better decisions about trade-offs.</p>
<p>That said, we&rsquo;ll try to tackle this question more thoroughly in this post.</p>
<h2 id="the-model-landscape">The model landscape
</h2>
<p>Currently, Anthropic and OpenAI set the bar for AI capabilities. Google is often discussed as a third frontier lab, though its current model lineup is less competitive.</p>
<p>Anthropic and OpenAI each release a &ldquo;family&rdquo; of models: Claude and GPT, respectively. The most capable models in both families are also the slowest and most expensive. Conversely, the least capable models are the cheapest and quickest. Other labs tend to follow this same pattern, releasing a set of models with different trade-offs along the cost-performance curve.</p>
<p>Models within a given family tend to share a similar shape of intelligence, with related capabilities (relative to model size), shortcomings, and idiosyncrasies. For example, Claude Fable 5, Claude Opus 5, and Claude Sonnet 5 often use the same turns of phrase and are quite good at writing code and debugging it. Models from different families can have different shapes of intelligence even when their benchmark scores and prices are very similar. For example, GPT-5.6 Terra and Claude Sonnet 5 are priced similarly and comparably capable at agentic coding, but Terra doesn&rsquo;t &ldquo;see&rdquo; data visualizations as well as Sonnet, while Sonnet doesn&rsquo;t communicate as clearly as Terra.</p>
<p><strong>Other labs release models that score nearly as high as models from Anthropic and OpenAI on benchmarks. However, these evaluation scores can be deceiving.</strong> Labs can now train models to optimize for benchmarks (&ldquo;benchmaxxing&rdquo;). These models score well on benchmark-shaped tasks, which tend to be highly autonomous and &ldquo;tricky,&rdquo; but can fail to generalize to real-world tasks, which often involve more ambiguous requests. <a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart" target="_blank" rel="noopener">Kimi K3</a> and <a href="https://z.ai/blog/glm-5.2" target="_blank" rel="noopener">GLM 5.2</a> offer a counterexample: they score well on benchmarks, and we&rsquo;ve also found them exceptionally well-rounded and intuitive. We recently <a href="https://opensource.posit.co/blog/2026-08-10_kimi-k3-glm-5-2-posit-ai/" target="_blank" rel="noopener">introduced both models to Posit AI</a>.</p>
<p><strong>In <a href="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/" target="_blank" rel="noopener">previous</a> <a href="https://opensource.posit.co/blog/2026-06-19_ai-newsletter/" target="_blank" rel="noopener">newsletters</a> and <a href="https://posit.co/blog/introducing-bluffbench" target="_blank" rel="noopener">blog posts</a>, we&rsquo;ve shared results from targeted evaluations. Here, instead, we offer an approximate and entirely vibes-based comparison of these model families&rsquo; characteristics.</strong></p>
<p><div class="not-prose"><figure>
    <img class="h-auto max-w-full rounded-lg"
      src="https://opensource.posit.co/blog/2026-08-14_ai-newsletter/images/model-landscape.png"
      alt="A comparison of Anthropic, OpenAI, and Google Gemini across agentic coding, vision, image generation, intuitiveness, cost effectiveness, latency, and communication style. The horizontal scale runs from a lower relative level on the left to a higher relative level on the right; higher is not necessarily better."  title="Although imprecise, vibes are an important part of evaluating models." 
      loading="lazy"
    ><figcaption class="text-sm text-center text-gray-500">Although imprecise, vibes are an important part of evaluating models.</figcaption>
  </figure></div>
</p>
<p>For data science applications broadly, you can loosely think of the relevant score as the average of the agentic coding and vision scores we&rsquo;ve assigned here. Beyond writing R and Python code, models need to be able to interpret plots accurately and faithfully.</p>
<h2 id="find-the-model-that-fits-your-task">Find the model that fits your task
</h2>
<p>Much of model choice is constrained by the models you have access to. Your organization may only allow a certain provider, or you might not want to pay multiple (possibly expensive!) subscriptions just to have access to all the top models.</p>
<p>If you do have your pick, however, here are some quick guidelines, partially shaped by what models are currently available through <a href="https://docs.posit.co/posit-ai/user/" target="_blank" rel="noopener">Posit AI</a>.</p>
<p><strong>The best open-weights model, especially for data analysis:</strong> <em>Kimi K3</em>.</p>
<p>As Simon wrote in the <a href="https://opensource.posit.co/blog/2026-08-10_kimi-k3-glm-5-2-posit-ai/" target="_blank" rel="noopener">Kimi K3 and GLM 5.2 in Posit AI announcement post</a>, &ldquo;Kimi K3 is currently the most capable open weights model out there. In our internal testing, it feels somewhere between Opus 5 and Fable 5, and is notably well-rounded compared to other open weights releases.&rdquo;</p>
<p>It also ranks near the top in <a href="https://github.com/posit-dev/bluffbench2" target="_blank" rel="noopener">bluffbench2</a> (13.46%, compared with 16.35% for the tied top scorers, Gemini 3.5 Flash and Claude Fable 5), which evaluates models&rsquo; abilities to spot subtle data quality issues in visualizations.<sup id="fnref:2"><a href="#fn:2" class="footnote-ref" role="doc-noteref">2</a></sup></p>
<p><strong>A highly autonomous model for a complex coding or data task when cost and speed aren&rsquo;t a concern:</strong> <em>Claude Opus 5</em>, <em>Claude Fable 5</em>, or <em>GPT-5.6 Sol</em>.</p>
<p>These are the top-of-the-line models from Anthropic and OpenAI. They are expensive and relatively slow, but can be worth using for ambitious or highly autonomous work.</p>
<p><strong>A strong open-weights model for coding when you don&rsquo;t need vision:</strong> <em>GLM 5.2</em> (from <a href="https://z.ai/model-api" target="_blank" rel="noopener">Z.ai</a>).</p>
<p><a href="https://opensource.posit.co/blog/2026-08-10_kimi-k3-glm-5-2-posit-ai/" target="_blank" rel="noopener">&ldquo;GLM 5.2 excels at agentic coding and less so at data analysis tasks.&rdquo;</a> It is much less expensive than the proprietary models it resembles for coding tasks, but it <a href="https://docs.z.ai/guides/llm/glm-5.2" target="_blank" rel="noopener">does not natively support image inputs</a>.</p>
<p><strong>A middle-tier generalist for coding or data analysis:</strong> <em>Claude Sonnet 5</em> or <em>GPT-5.6 Terra</em>. Both are capable across coding and routine data analysis, support vision, and are less expensive than their respective labs&rsquo; higher-tier models.</p>
<p><strong>Good plot or image interpretation:</strong> One of the <em>Gemini 3.x Flash</em> models (<a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/" target="_blank" rel="noopener">3.7 was released on August 13</a>).</p>
<p>This model series is particularly strong at vision and has performed well on <a href="https://github.com/posit-dev/bluffbench" target="_blank" rel="noopener">bluffbench</a> (3.5 Flash) and <a href="https://github.com/posit-dev/bluffbench2" target="_blank" rel="noopener">bluffbench2</a> (3.5 and 3.6 Flash).</p>
<p><strong>Fast answers from an Anthropic model for a task that is not particularly complex:</strong> <em>Claude Haiku 4.5</em>.</p>
<p><strong>Fast answers from an open-weights model for a task that is not particularly complex:</strong> <a href="https://posit.co/blog/gemma-4-new-budget-focused-model-posit-ai" target="_blank" rel="noopener"><em>Gemma 4</em></a>.</p>
<h2 id="assorted-notes-from-august-2026">Assorted notes from August 2026
</h2>
<p>In late summer 2026, a few developments feel notable, but, as with much in the AI world, who knows how long these observations will hold.</p>
<ul>
<li>
<p><strong>Google Gemini currently does not have any models that perform near the frontier.</strong> With the releases of Gemini 2.5 Pro (June 2025) and the Gemini 3 series (November 2025), Google seemed positioned as a third frontier lab. However, it&rsquo;s been quite a while since they released a frontier model, and they&rsquo;re now meaningfully behind. As of the time of writing, Google says Gemini 3.5 Pro is still testing with partners. Meanwhile, <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/" target="_blank" rel="noopener">the company has begun training Gemini 4 and says it is excited by the progress</a>.</p>
</li>
<li>
<p>For a year or so, it seemed like Anthropic was solidly ahead of OpenAI in agentic coding. However, <strong>since the Claude 4.6 releases, it has become less clear that Anthropic is meaningfully ahead of OpenAI</strong>. For one, Anthropic&rsquo;s current high-end models use newer tokenizers that, <a href="https://platform.claude.com/docs/en/about-claude/models/migration-guide#what-changed-5" target="_blank" rel="noopener">according to Anthropic, produce roughly 35% more tokens for the same text than their predecessors</a>. This means the same listed price per token does not necessarily translate to the same cost for comparable text. Further, in our experience, the Claude series has become increasingly token-hungry and difficult to communicate with. At the same time, OpenAI&rsquo;s models have a notably clear, concise communication style compared with the Claude 5 series. Anthropic is still likely ahead on autonomous, long-horizon coding, but OpenAI no longer feels behind for day-to-day software engineering and data science.<sup id="fnref:3"><a href="#fn:3" class="footnote-ref" role="doc-noteref">3</a></sup></p>
</li>
<li>
<p><strong>There are now a number of balanced, well-rounded open-weights models relatively close to the closed-weights frontier.</strong> Kimi K3 and GLM 5.2, in particular, combine strong capabilities with a more pleasant, intuitive feel than their predecessors. While earlier open-weights models were just as close to the frontier in benchmark scores, some newer releases are notably more well-rounded and respond to prompts about as effectively as proprietary models. These releases are also priced at a steep discount compared with the proprietary models they most resemble.</p>
</li>
</ul>
<h2 id="recent-past-newsletters">Recent past newsletters
</h2>
<ul>
<li><a href="https://opensource.posit.co/blog/2026-07-31_ai-newsletter/">Keep track of your data exploration with Posit Assistant&rsquo;s EDA log</a></li>
<li><a href="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/">Which models are best at spotting subtle data quality problems?</a></li>
</ul>
<br>
<br>
<p><a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener noreferrer" class="btn-shortcode inline-flex mb-5 mr-5 items-center px-4 py-3 text-sm leading-5 gap-2 rounded-lg bg-blue-400 text-white font-semibold align-middle hover:bg-blue-500 transition no-underline">Subscribe via RSS</a></p>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>By &ldquo;access to,&rdquo; we mean either the most expensive model you can afford or the most expensive model that your organization allows you to use.&#160;<a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
<li id="fn:2">
<p>Kimi K3 supports <code>low</code>, <code>high</code>, and <code>max</code> reasoning levels. This run used <code>high</code>, its middle setting.&#160;<a href="#fnref:2" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
<li id="fn:3">
<p>Notably, many of our colleagues are still using Claude Opus 4.6 as their daily driver. Despite being less capable than newer high-end Claude models on especially long-horizon work, the model is capable of day-to-day software engineering and is cheaper in practice. For example, <a href="https://platform.claude.com/docs/en/about-claude/pricing" target="_blank" rel="noopener">Opus 4.6 and Opus 5 have the same listed per-token price</a>, but Opus 4.6 predates <a href="https://platform.claude.com/docs/en/about-claude/models/migration-guide#migrating-from-claude-opus-46" target="_blank" rel="noopener">the newer tokenizer that can produce up to roughly 35% more tokens for the same text</a>. Many of our colleagues also find Opus 4.6 easier to communicate with.&#160;<a href="#fnref:3" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-08-14_ai-newsletter/images/hero.png" length="468856" type="image/png" />
    </item>
    <item>
      <title>Positron August Release Highlights</title>
      <link>https://opensource.posit.co/blog/2026-08-13_positron-2026-08-release/</link>
      <pubDate>Thu, 13 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-08-13_positron-2026-08-release/</guid>
      <dc:creator>Julia Silge</dc:creator><description><![CDATA[<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p><a href="https://positron.posit.co" target="_blank" rel="noopener">Positron</a> is Posit&rsquo;s new, next-generation IDE for data science. Positron is designed to be an extensible, polyglot tool for exploring data and reproducible authoring in Python, R, and more.</p>
</div>
</div>
<p>Welcome back to another edition of our monthly Positron updates! Each month we share highlights from our <a href="https://positron.posit.co/release-notes" target="_blank" rel="noopener">latest release</a> and useful resources. <a href="https://opensource.posit.co/blog/2026-07-13_positron-2026-07-release">Last release</a> we told you about several major features that came out of preview to general availability, including the <a href="https://opensource.posit.co/blog/2026-07-29_positron-jupyter-notebook-editor-ga">new notebook editor</a>. This milestone we are excited to share new functionality for SQL support, reproducible authoring, helping you know when packages are missing, and more.</p>
<h2 id="data-connections">Data Connections
</h2>
<p>Data Connections is our new way to work with SQL and database-like resources in Positron, from local files and database servers to cloud data warehouses. It is currently available as a preview feature, and you can enable it with the <a href="positron://settings/dataConnections.enabled"><code>dataConnections.enabled</code></a> setting. This release more than doubles the number of data sources you can reach. Amazon Redshift, Snowflake, Databricks, and Posit Connect pins join the existing DuckDB, PostgreSQL, and SQLite support.</p>
<img src="https://opensource.posit.co/blog/2026-08-13_positron-2026-08-release/data-connections.gif" data-fig-align="center" data-fig-alt="Browsing the schemas and tables of a DuckDB connection in the Data Connections panel, then opening a table in the Data Explorer to see its column profiles and data." />
<p>The panel itself is more capable as well. <strong>Refresh</strong> and <strong>Refresh All</strong> reload the tree while preserving what you have expanded, briefly highlighting the rows that were reloaded. Open connections now show an indicator. Collapsing a connection in the UI keeps you connected to your data source, while anything you&rsquo;ve previewed with the Data Explorer stays still open. When you remove a connection, Positron asks for confirmation and reports how many Data Explorers will close with it.</p>
<p>Data Connections is still an experimental preview, and your feedback continues to shape it. Tell us which databases and warehouses you need, and anything confusing, missing, or broken, in the <a href="https://github.com/posit-dev/positron/discussions/14571" target="_blank" rel="noopener">Data Connections discussion</a>.</p>
<h2 id="inline-output-for-quarto">Inline output for Quarto
</h2>
<p><a href="https://positron.posit.co/quarto-inline-output" target="_blank" rel="noopener">Inline output for <code>.qmd</code> documents</a> came out of preview last release, and this release brings you a substantial round of polish for this way of working. Be aware that the Quarto settings have moved into a dedicated <code>quarto.*</code> namespace with its own group in the Settings editor. The previous <code>positron.quarto.*</code> keys still work but are deprecated, and Positron will prompt you to update your settings.</p>
<p>Before a kernel starts, the kernel status names the interpreter it will start and offers an explicit <strong>Start Kernel</strong> action. When a cell fails, <strong>Fix</strong> and <strong>Explain</strong> buttons send the error to <a href="https://assistant.posit.co/" target="_blank" rel="noopener">Posit Assistant</a>, matching the Positron notebook experience. The editor also scrolls to reveal output as it is produced, which you can turn off with the new <a href="positron://settings/quarto.inlineOutput.autoScroll"><code>quarto.inlineOutput.autoScroll</code></a> setting.</p>
<p>Output renders more faithfully as well. The editor gutter now shows which statement is currently executing and per-statement progress. Positron draws images at your display&rsquo;s pixel ratio, so plots are sharp on retina screens, and Python figures now respect the <code>fig-width</code> and <code>fig-height</code> cell options.</p>
<img src="https://opensource.posit.co/blog/2026-08-13_positron-2026-08-release/inline-output-plot-metadata.png" data-fig-align="center" data-fig-alt="A Quarto document open in Positron with a Python cell that sets the fig-width and fig-height options, and the resulting matplotlib scatter plot rendered inline below the cell at that size." />
<p>HTML widgets no longer stick in the editor corner when you scroll past them, or trap scrolling instead of letting the document scroll. HTML widgets no longer render as raw HTML after a reload, and collapsed output no longer springs back open when its cell re-runs. Running code in a Quarto document also pins the editor tab now, so Positron does not silently close the document and its session when you open another file.</p>
<h2 id="ai-model-providers">AI model providers
</h2>
<p>Positron now reads AI model provider configuration from a single <code>providers.json</code> file rather than a scattered set of settings. The new release will migrate your existing configuration automatically when you start it, and deprecates the <code>authentication.*</code> and <code>positron.assistant.provider.*.enable</code> settings in favor of it. Two new commands give you direct access: <em>Open AI Provider Settings (JSON)</em> opens <code>providers.json</code> from the Command Palette, and <em>Migrate Provider Settings to providers.json</em> runs the migration on demand.</p>
<h2 id="install-missing-packages">Install missing packages
</h2>
<p>Positron now notices when your code refers to a package you do not have installed and offers to install it for you. The prompt appears for packages referenced by your scripts and notebooks in both R and Python.</p>
<img src="https://opensource.posit.co/blog/2026-08-13_positron-2026-08-release/missing-package.gif" data-fig-align="center" data-fig-alt="A Shiny app in Positron showing a missing package button in the editor toolbar. Clicking it installs bslib in the console, and the app then runs with its bubble chart in the Viewer pane." />
<p>A <code>library()</code> or <code>import</code> call for something missing becomes a single click instead of an error you have to go resolve by hand.</p>
<h2 id="performance-and-memory">Performance and memory
</h2>
<p>We continue to invest in the memory footprint, performance, and reliability of Positron. Several components now load only when they are actually needed, and turning off <a href="positron://settings/ai.enabled"><code>ai.enabled</code></a> now means Positron never loads some heavy AI-related components at all. We fixed a cluster of long-standing reliability problems around session restarts and lifecycles.</p>
<p>Startup and editing are faster as well. R and Python kernels start much faster on Windows systems with aggressive antivirus software. The interpreter session picker appears immediately instead of waiting for interpreter discovery to finish. We also fixed slow typing, formatting, and saving in long R and Python files.</p>
<h2 id="whats-coming-next">What&rsquo;s coming next
</h2>
<ul>
<li>Join <a href="https://posit.co/workflow-demo/ai-governance-workbench" target="_blank" rel="noopener">our upcoming webinar</a> on August 26 to learn about AI governance in Posit Workbench.</li>
<li>We are looking forward to posit::conf(2026) next month, where our team will have several sessions on Positron. <a href="https://conf.posit.co/2026/" target="_blank" rel="noopener">Register now</a> to join us in person in Houston or virtually from anywhere in the world.</li>
</ul>
<div class="callout callout-tip" role="note" aria-label="Tip">
<div class="callout-header">
<span class="callout-title">Tip</span>
</div>
<div class="callout-body">
<p><a href="https://positron.posit.co/download" target="_blank" rel="noopener">Download Positron</a> to try out the new features and improvements in this release!</p>
</div>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-08-13_positron-2026-08-release/featured.svg" length="91423" type="image/svg&#43;xml" />
    </item>
    <item>
      <title>Kimi K3 and GLM 5.2 are now in Posit AI</title>
      <link>https://opensource.posit.co/blog/2026-08-10_kimi-k3-glm-5-2-posit-ai/</link>
      <pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-08-10_kimi-k3-glm-5-2-posit-ai/</guid>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<p>We&rsquo;re super stoked to share that two more open weights models—Kimi K3 and GLM 5.2—are now available as part of <a href="https://posit.ai/" target="_blank" rel="noopener">Posit AI</a>. Both of these models are substantially cheaper than the proprietary models they feel most similar to, albeit a bit rougher on the edges. Notably, as well, these models are served much more quickly than Anthropic models; we&rsquo;ve been seeing these models stream almost twice as many tokens per second in our internal testing, and working with them is a qualitatively different feel. <a href="https://docs.posit.co/posit-ai/user/faq/#privacy-data-storage" target="_blank" rel="noopener">As with the other models</a> made available in Posit AI, <strong>your conversation histories will not be stored unless you choose to opt-in to data retention at sign-up.</strong></p>
<p>Kimi K3 is currently the most capable open weights model out there. In our internal testing, it feels somewhere between Opus 5 and Fable 5, and is notably well-rounded compared to other open weights releases. It&rsquo;s priced at the same price-per-token as Claude Sonnet 5.<sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup></p>
<p>GLM 5.2 excels at agentic coding and less so at data analysis tasks. Because the model is not vision-capable, it cannot &lsquo;see&rsquo; plots and thus feels less capable for data science work. Its per-token pricing is also similar to Haiku 4.5, but feels something like Opus 4.6 or Sonnet 5 for agentic coding tasks.</p>
<h2 id="pricing">Pricing
</h2>
<p>The models in Posit AI today, at a glance:</p>
<table>
  <thead>
      <tr>
          <th style="text-align: left">Model</th>
          <th style="text-align: right">Cached input</th>
          <th style="text-align: right">Input</th>
          <th style="text-align: right">Cache write</th>
          <th style="text-align: right">Output</th>
          <th style="text-align: right">Context length</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td style="text-align: left">Claude Opus 5</td>
          <td style="text-align: right">$0.55</td>
          <td style="text-align: right">$5.50</td>
          <td style="text-align: right">$6.875</td>
          <td style="text-align: right">$27.50</td>
          <td style="text-align: right">1M</td>
      </tr>
      <tr>
          <td style="text-align: left">Claude Sonnet 5</td>
          <td style="text-align: right">$0.33</td>
          <td style="text-align: right">$3.30</td>
          <td style="text-align: right">$4.125</td>
          <td style="text-align: right">$16.50</td>
          <td style="text-align: right">1M</td>
      </tr>
      <tr>
          <td style="text-align: left"><strong>Kimi K3</strong></td>
          <td style="text-align: right"><strong>$0.33</strong></td>
          <td style="text-align: right"><strong>$3.30</strong></td>
          <td style="text-align: right"><strong>$0.00</strong></td>
          <td style="text-align: right"><strong>$16.50</strong></td>
          <td style="text-align: right"><strong>250K</strong><sup id="fnref:2"><a href="#fn:2" class="footnote-ref" role="doc-noteref">2</a></sup></td>
      </tr>
      <tr>
          <td style="text-align: left"><strong>GLM 5.2</strong></td>
          <td style="text-align: right"><strong>$0.154</strong></td>
          <td style="text-align: right"><strong>$1.54</strong></td>
          <td style="text-align: right"><strong>$0.00</strong></td>
          <td style="text-align: right"><strong>$4.84</strong></td>
          <td style="text-align: right"><strong>256K</strong></td>
      </tr>
      <tr>
          <td style="text-align: left">Claude Haiku 4.5</td>
          <td style="text-align: right">$0.11</td>
          <td style="text-align: right">$1.10</td>
          <td style="text-align: right">$1.375</td>
          <td style="text-align: right">$5.50</td>
          <td style="text-align: right">200K</td>
      </tr>
      <tr>
          <td style="text-align: left">Gemma 4 26B A4B</td>
          <td style="text-align: right">$0.033</td>
          <td style="text-align: right">$0.33</td>
          <td style="text-align: right">$0.00</td>
          <td style="text-align: right">$1.65</td>
          <td style="text-align: right">100K</td>
      </tr>
  </tbody>
</table>
<p><strong>The cost savings are greater than the per-token pricing differences alone might suggest.</strong> For one, the Claude 5 series models use a tokenizer that results in substantially more tokens (~35%) per word than Kimi K3&rsquo;s or GLM 5.2&rsquo;s tokenizer. Also, users are not billed at a higher rate for Cache writes than &rsquo;normal&rsquo; input tokens; all input tokens are written to the cache by default, but we can&rsquo;t make a guarantee that you&rsquo;ll hit the cache after any specific delay between requests. In practice, we&rsquo;ve seen that the cache efficiency of conversations with these deployments is slightly lower than with Claude models.</p>
<h2 id="get-started">Get started
</h2>
<p>To get started, open up <a href="https://assistant.posit.co/" target="_blank" rel="noopener">Posit Assistant</a> and update when prompted! Then, select your model of choice under the Posit AI model provider. If you&rsquo;re not already a Posit AI subscriber, you can learn more <a href="https://posit.ai/" target="_blank" rel="noopener">here</a>.</p>
<p>It&rsquo;s worth saying that we suspect these models will rotate somewhat regularly in the service; as the months go by, we plan to introduce support for new models and deprecate others.</p>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>Claude Sonnet 5 is currently priced at a promotional $2/$10, but will be back to its usual $3/$15 in a few weeks. The pricing is the same <em>after</em> Sonnet 5&rsquo;s promotional pricing ends.&#160;<a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
<li id="fn:2">
<p>While Kimi K3 technically supports a 1M-token context window, we&rsquo;ve limited it to 250K in Posit AI to ensure we have enough capacity.&#160;<a href="#fnref:2" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-08-10_kimi-k3-glm-5-2-posit-ai/featured.png" length="200054" type="image/png" />
    </item>
    <item>
      <title>AI Newsletter: EDA log in Posit Assistant</title>
      <link>https://opensource.posit.co/blog/2026-07-31_ai-newsletter/</link>
      <pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-07-31_ai-newsletter/</guid>
      <dc:creator>Sara Altman</dc:creator>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-body">
<p><strong>Subscribe to the AI Newsletter!</strong></p>
<p>The AI newsletter is now published as an RSS feed. Follow it in your favorite reader:</p>
<p><a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener noreferrer" class="btn-shortcode inline-flex mb-5 mr-5 items-center px-4 py-3 text-sm leading-5 gap-2 rounded-lg bg-blue-400 !text-white font-semibold align-middle hover:bg-blue-500 transition no-underline">Subscribe via RSS</a></p>
<p><strong>Want the newsletter as an email?</strong> Paste the feed URL — <a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener">https://opensource.posit.co/tags/ai-newsletter/index.xml</a> — into a free RSS-to-email service such as <a href="https://blogtrottr.com/" target="_blank" rel="noopener">Blogtrottr</a>, <a href="https://feedrabbit.com/" target="_blank" rel="noopener">Feedrabbit</a>, or <a href="https://follow.it/" target="_blank" rel="noopener">Follow.it</a>, and each new issue will arrive in your inbox.</p>
</div>
</div>
<br>
[Posit Assistant](https://assistant.posit.co/) in Positron now includes an EDA log feature to help you keep track of exploratory analysis done with the agent.
<p><div class="not-prose"><figure>
    <img class="h-auto max-w-full rounded-lg"
      src="https://opensource.posit.co/blog/2026-07-31_ai-newsletter/images/eda-log-zoom.png"
      alt="Screenshot of Positron. On the left, Posit Assistant has analyzed a dataset of U.S. language speakers, showing a bar chart and written findings. On the right, the EDA Log opens in an editor tab titled &ldquo;ACS language speakers&rdquo;: a table with Area, Status, and Notes columns lists three areas—&ldquo;Dataset structure &amp; quality&rdquo; and &ldquo;Top languages nationwide&rdquo; marked Explored, and &ldquo;Language coverage across states&rdquo; marked Partial—each with bullet-point findings and an arrow that links back to the conversation, followed by a &ldquo;Next steps&rdquo; section of suggested directions." 
      loading="lazy"
    >
  </figure></div>
</p>
<p>The log summarizes findings for different areas of exploration and keeps track of next steps. To use the log, run the <code>/eda-log</code> slash command after you&rsquo;ve started the EDA process.</p>
<h3 id="why-we-made-this">Why we made this
</h3>
<p>Exploratory data analysis, the open-ended orientation to your data that often comes before anything else, can be a branching, nonlinear process. There are many questions you can ask of your data, and new areas of inquiry can open with each question you ask. Because of this, it is often hard to keep track of what you&rsquo;ve looked into, where that code lives, and what you want to explore next.</p>
<p>Historically, the EDA process was limited by how quickly you could write code and interpret the output. Coding agents like Posit Assistant lift the first of those constraints. They can carry out EDA far faster than you can on your own, which can exacerbate the issue of keeping track of what you&rsquo;ve explored.</p>
<p>This speed also introduces a new problem: the point of EDA is typically for you, the human, to understand your data, but coding agents can produce output faster than you can absorb it. If the agent completes an analysis but you haven&rsquo;t understood the insights in the data, the exploration process hasn&rsquo;t really happened.</p>
<p>Posit Assistant already has various features that tackle this problem, including an exploratory mode of interaction where it runs shorter turns and stops more frequently to involve the user.</p>
<p>The EDA log is another lightweight tool for the same goal. It keeps a running summary of what you and Posit Assistant have explored, helping your understanding keep pace with the agent&rsquo;s and giving you a clearer picture of what&rsquo;s already been done.</p>
<h3 id="details">Details
</h3>
<p>Here&rsquo;s what the EDA log looks like in action:</p>
<script src="https://fast.wistia.com/player.js" async></script>
<script src="https://fast.wistia.com/embed/bu9ch5gqvx.js" async type="module"></script>
<style>wistia-player[media-id='bu9ch5gqvx']:not(:defined) { background: center / contain no-repeat url('https://fast.wistia.com/embed/medias/bu9ch5gqvx/swatch'); display: block; filter: blur(5px); padding-top:60.42%; }</style>
<p><wistia-player media-id="bu9ch5gqvx" aspect="1.6551724137931034"></wistia-player></p>
<p>At a high level:</p>
<ul>
<li>When you run <code>/eda-log</code>, Posit Assistant will create a log for the exploration done in the conversation so far. The log then opens in the editor area in Positron.</li>
<li>The underlying logs are stored as YAML files in <code>.posit/assistant/eda-logs/</code>, next to where plans are stored.</li>
<li>Posit Assistant is instructed to loosely keep the log up to date as the conversation progresses, but you can also manually trigger an update at any time with the &ldquo;Refresh&rdquo; button.</li>
<li>Clicking the arrow next to an area scrolls you back to the spot in the conversation where that insight originated, so you can revisit the code and context that produced it.</li>
<li>Suggested next steps appear as clickable text. Clicking one sends it to Posit Assistant as your next message.</li>
<li>The creation of an EDA log is always user-triggered. Posit Assistant will never create one on its own.</li>
<li>The feature is currently only in Positron, but will come to RStudio soon.</li>
</ul>
<h2 id="recent-past-newsletters">Recent past newsletters
</h2>
<ul>
<li><a href="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/">Which models are best at spotting data quality problems?</a></li>
<li><a href="https://opensource.posit.co/blog/2026-07-03_ai-newsletter/">How to choose between AGENTS.md, skills, and MCP servers</a></li>
</ul>
<br>
<br>
<p><a href="https://opensource.posit.co/tags/ai-newsletter/index.xml" target="_blank" rel="noopener noreferrer" class="btn-shortcode inline-flex mb-5 mr-5 items-center px-4 py-3 text-sm leading-5 gap-2 rounded-lg bg-blue-400 text-white font-semibold align-middle hover:bg-blue-500 transition no-underline">Subscribe via RSS</a></p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-07-31_ai-newsletter/images/hero.png" length="954735" type="image/png" />
    </item>
    <item>
      <title>AI Newsletter: LLMs often miss subtle visual artifacts in data visualizations</title>
      <link>https://opensource.posit.co/blog/2026-07-17_ai-newsletter/</link>
      <pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-07-17_ai-newsletter/</guid>
      <dc:creator>Sara Altman</dc:creator>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<p>Imagine that you receive some patient data and load it into R or Python for the first time. You make a couple plots to get a sense of the data before coming across this one:</p>
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/artifact-plot-1.png" data-fig-align="center" width="768" />
<p>Huh. It mostly looks normal, except there&rsquo;s a few points perfectly aligned with what looks to be a &ldquo;fitted&rdquo; line. You dig into it a bit more, and realize that the rows from one study site have their cholesterol values imputed. You set them to <code>NA</code> and go along your way.</p>
<p>Would today&rsquo;s frontier LLMs catch such an oddity? We designed <a href="https://github.com/posit-dev/bluffbench2" target="_blank" rel="noopener">an LLM evaluation</a> to help us answer this question. As it turns out, LLMs mostly miss these sorts of artifacts:</p>
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/results-plot-1.png" data-fig-align="center" data-fig-alt="A bar plot showing scores for several frontier models. The two leaders, Gemini 3.5 Flash and Claude Fable 5, score in the mid teens. Models from OpenAI cluster at the bottom, never eclipsing 10%." width="768" />
<p>During exploratory or open-ended data analysis, Posit assistant <a href="https://opensource.posit.co/blog/2026-06-08_comparing-posit-assistant-and-claude-code/#specialized-data-analysis-capabilities" target="_blank" rel="noopener">&ldquo;only runs a few bits of code at a time, then summarizes what it found and suggests next steps&rdquo;</a>. This is motivated by our stance that, for now, a data scientist should mostly keep pace with and understand what the agent is doing when analyzing data. This stance was initially informed by our observation that last year&rsquo;s frontier models <a href="https://posit.co/blog/introducing-bluffbench" target="_blank" rel="noopener">tended to see what they expected to see</a> when visualizing data. While <a href="https://opensource.posit.co/blog/2026-06-19_ai-newsletter/" target="_blank" rel="noopener">LLMs have since become much better at interpreting counterintuitive plots</a>, bluffbench2 shows they still lag behind human data scientists in interpreting data visualizations. As such, we are still cautious on the prospect of highly autonomous data agents.</p>
<h2 id="how-the-eval-works">How the eval works
</h2>
<p>The eval harness is a relatively generic coding agent harness, similar to that of Claude Code or Posit Assistant. The agent has a tool to run R code in a persistent REPL and some vague prompting about data analysis:</p>
<blockquote>
<p>You are an AI assistant embedded in the user&rsquo;s data science IDE. You can read and modify files in the user&rsquo;s workspace and execute R code in their active session, including rendering plots. Prioritize correctness and clear communication&hellip;</p>
</blockquote>
<p>In each sample, the agent first carries out a few &ldquo;lull&rdquo; turns, making a couple plots and tables unrelated to the eval. Short user messages like &ldquo;load in the csv in this folder&rdquo; are decorated with &ldquo;System Reminders&rdquo; and other noise like that injected by popular agent harnesses.</p>
<div style="display: flex; flex-direction: column; gap: 8px; padding: 20px; max-width: 100%; margin: 20px auto;">
<div style="align-self: flex-end; background-color: #e8f3fc; padding: 12px 18px; border-radius: 18px 18px 4px 18px; max-width: 70%;">
take a look at <code>dat</code> in my env
</div>
<div style="align-self: flex-start; background-color: white; padding: 12px 18px; border-radius: 18px 18px 18px 4px; max-width: 70%; border: 1px solid #e0e0e0;">
<em>Tool: Run R code</em>
</div>
<div style="align-self: flex-end; background-color: #e8f3fc; padding: 12px 18px; border-radius: 18px 18px 4px 18px; max-width: 70%;">
<em>Tool result</em>
</div>
<div style="align-self: flex-start; background-color: white; padding: 12px 18px; border-radius: 18px 18px 18px 4px; max-width: 70%; border: 1px solid #e0e0e0;">
Looks like <code>dat</code> is a data frame of...
</div>
<div style="align-self: flex-end; background-color: #e8f3fc; padding: 12px 18px; border-radius: 18px 18px 4px 18px; max-width: 70%;">
<span style="display: block; margin-bottom: 8px; font-family: monospace; font-size: 0.8em; opacity: 0.55;">&lt;system-reminder&gt;<br>Your to-do list is currently empty. If you are working on tasks that would benefit from tracking progress, consider creating to-dos. This is just a gentle reminder - ignore if not applicable.<br>&lt;/system-reminder&gt;</span>
summarize <code>$cholesterol</code>
</div>
<div style="align-self: flex-start; background-color: white; padding: 12px 18px; border-radius: 18px 18px 18px 4px; max-width: 70%; border: 1px solid #e0e0e0;">
<em>Tool: Run R code</em>
</div>
</div>
<p>After a few turns, the agent is asked to produce a data visualization that includes a subtle visual artifact that could feasibly result from a real data-generating process. The artifacts span a range of realistic data quality issues: stuck sensors, bad joins, points imputed onto a line, swapped columns, pseudoreplication, differing units, etc.</p>
<div style="display: flex; flex-direction: column; gap: 8px; padding: 20px; max-width: 100%; margin: 20px auto;">
<div style="align-self: flex-end; background-color: #e8f3fc; padding: 12px 18px; border-radius: 18px 18px 4px 18px; max-width: 70%;">
plot bmi vs cholesterol
</div>
<div style="align-self: flex-start; background-color: white; padding: 12px 18px; border-radius: 18px 18px 18px 4px; max-width: 70%; border: 1px solid #e0e0e0;">
<em>Tool: Run R code</em>
</div>
<div style="align-self: flex-end; background-color: #e8f3fc; padding: 12px 18px; border-radius: 18px 18px 4px 18px; max-width: 70%;">
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/images/labs-thumb.png" width="220" style="border-radius: 12px; display: block;">
</div>
</div>
<p>If the agent mentions the artifact in its follow-up response, it receives a full point. If the agent does not mention the artifact, it can also receive a half point by mentioning it in response to a follow-up user message along the lines of &ldquo;what do you see in the plot?&rdquo; If the agent never mentions the artifact, it is graded as incorrect.</p>
<h2 id="designing-the-eval">Designing the eval
</h2>
<p>Once we understood the mechanism behind bluffbench, implementing the eval was relatively straightforward. bluffbench demonstrates the degree to which an LLM will ignore evidence shown in a plot in favor of its expectations. So, to implement a given sample, we&rsquo;d just think of some situation that would elicit a strong prior and then subvert it. For example, a dataset called <code>doug_firs</code> with variables <code>height</code> and <code>circumference</code>; one might expect that, as height increases, so does circumference. So, instead, we did a transformation under the hood that made the relationship parabolic.</p>
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/trees-plot-1.png" data-fig-align="center" data-fig-alt="Two scatterplots side by side, both with circumference on the x axis and height on the y axis. The left, labeled &#39;Original Plot&#39;, shows height rising with circumference, a positive trend. The right, labeled &#39;Tampered Plot&#39;, shows height rising then falling as circumference increases, an inverted-U shape." width="768" />
<p>A year ago, triggering this prior was enough to frequently &rsquo;trick&rsquo; the current frontier LLMs.</p>
<p>Slipping a plotted artifact past today&rsquo;s LLMs is much harder. Any human could ace bluffbench, but only an attentive data analyst would excel at bluffbench2.</p>
<p>In our early work on a successor to bluffbench, we started off with trying to elicit priors in the same way as bluffbench did, but in more realistic, longer-context scenarios. We were surprised to find that the same mechanism broadly doesn&rsquo;t seem to trick today&rsquo;s models even in these more realistic settings.<sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup> We then tried a &lsquo;reverse bluffbench&rsquo;, where we let the model being evaluated in on the trick, asking it to carry out the transformation itself and then look at the plotted result which was tampered with to show the original relationship. We anticipated that this stronger prior (&ldquo;I did a thing with an obvious effect&rdquo;) might cause the models to miss the (re)manipulation, but models reliably noted that the plot looked as if it hadn&rsquo;t been manipulated.</p>
<p>As such, there isn&rsquo;t a similar &rsquo;trick&rsquo; in bluffbench2 per se. The transcripts read like relatively normal data analysis sessions and the plotted artifacts are designed to plausibly result from real data-generating processes. Instead, the eval elicits 1) the &lsquo;shape&rsquo; of LLMs&rsquo; vision being different than humans&rsquo; and 2) the model&rsquo;s tendencies to perform progress, simulating a data analysis moving along smoothly.</p>
<p>Today&rsquo;s frontier models are in the mid-teens at best; the top scores belong to Claude Fable 5 and Gemini 3.5 Flash at 16%. That said, we&rsquo;d caution folks from interpreting the current scores on this eval as &lsquo;LLMs don&rsquo;t see plots well.&rsquo; The plotted artifacts are actually quite subtle, and when they&rsquo;re made even a bit more marked, models tend to call them out consistently.</p>
<p>For example, the previous version of the scatterplot had a slightly more dense cluster of points:</p>
<div class="panel-tabset">
<ul id="tabset-1" class="panel-tabset-tabby">
<li><a data-tabby-default href="#tabset-1-1">Previous</a></li>
<li><a href="#tabset-1-2">Current</a></li>
</ul>
<div id="tabset-1-1">
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/previous-labs-plot-1.png" data-fig-align="center" data-fig-alt="A scatterplot of BMI versus cholesterol with a dense run of roughly fifty points falling exactly on a straight line through the noisy cloud." width="768" />
</div>
<div id="tabset-1-2">
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/current-labs-plot-1.png" data-fig-align="center" data-fig-alt="A scatterplot of BMI versus cholesterol with a sparse run of about thirty points falling exactly on a straight line through the noisy cloud, subtler than the previous version." width="768" />
</div>
</div>
<p>Opus 4.8 (medium) consistently got this sample right in the previous iteration.</p>
<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p>The fact that this was the case&mdash;that models would call out more marked artifacts reliably&mdash;gave us confidence that our grading setup was reasonable. In other words, it does indeed seem like models are struggling with these tasks because their vision is not capable enough to &lsquo;see&rsquo; the plotted artifact rather than a behavioral tendency to not mention those artifacts when they do see them.</p>
</div>
</div>
<h2 id="exploring-the-evals-results">Exploring the eval&rsquo;s results
</h2>
<p>At least for now, there&rsquo;s a loosely linear relationship between the cost to run the eval and the resulting score:</p>
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/cost-plot-1.png" data-fig-align="center" data-fig-alt="A scatterplot of score against total cost for each frontier model, colored by lab. The two leaders, Gemini 3.5 Flash and Claude Fable 5, sit highest at around the mid teens, while the OpenAI models sit low regardless of cost. Higher spend does not buy a higher score." width="768" />
<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p>Given that Gemini 3.5 Flash is so much cheaper than Claude Fable 5 per-token ($1.50/$9 per mTok I/O vs. $10/$50), it&rsquo;s surprising that the eval was so expensive to run for Gemini 3.5 Flash. This is primarily driven by cache (in)efficiency; the harness is implemented against Gemini&rsquo;s <code>generateContent</code> API, which makes it difficult to make use of discounted cached input pricing compared to Anthropic and OpenAI&rsquo;s APIs. Implementing and switching to Gemini&rsquo;s newer Interactions API would push the Flash 3.5 point to the left.</p>
</div>
</div>
<p>One of the most interesting learnings from examining the logs is a behavioral one. Even though we never request that LLMs introduce modeled results to plots, like fitted lines and confidence intervals with <code>geom_smooth(method = &quot;lm&quot;, se = TRUE)</code>, they sometimes do so anyway. For example:</p>
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/smooth-example-1.png" data-fig-align="center" data-fig-alt="The BMI versus cholesterol scatterplot with a straight fitted line and a shaded confidence-interval ribbon laid over the points, an overlay the model added on its own." width="768" />
<p>In general, adding modeled results to data visualizations without first looking at data is bad practice; it makes it hard to see the data itself. In the eval, adding a modeled result like this seems to substantially lower the chances that the model will notice the plotted artifact:</p>
<img src="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/index.markdown_strict_files/figure-markdown_strict/smooth-plot-1.png" data-fig-align="center" data-fig-alt="A dumbbell plot, one row per model, comparing accuracy on artifact plots the model drew with a geom_smooth() overlay versus without. For nearly every model the &#39;with overlay&#39; point sits well to the left of the &#39;without&#39; point; Claude Fable 5 falls from about a quarter correct to zero, and Gemini 3.5 Flash from about a quarter to under a tenth." width="768" />
<h2 id="more-bluffbench">More bluffbench
</h2>
<p>If you&rsquo;d like to learn more about the bluffbench set of evals, take a look at these past posts:</p>
<ul>
<li><a href="https://posit.co/blog/introducing-bluffbench" target="_blank" rel="noopener"><strong>Introducing bluffbench</strong></a>: Writeup of the the original eval.</li>
<li><a href="https://posit.co/blog/llm-plot-interpretation" target="_blank" rel="noopener"><strong>LLMs interpret plots well, until expectations interfere</strong></a>: In-depth post on why models at the time didn&rsquo;t perform well on bluffbench, as well as various interventions we tried to improve performance.</li>
<li><a href="https://opensource.posit.co/blog/2026-06-19_ai-newsletter/" target="_blank" rel="noopener"><strong>LLMs are getting much better at interpreting counterintuitive plots</strong></a>: In spring 2026, bluffbench scores suddenly jumped as models improved.</li>
<li><a href="https://skaltman.github.io/scipy-2026/" target="_blank" rel="noopener"><em><strong>It&rsquo;s (still) very bad to be wrong</strong></em></a>: Slides from our recent SciPy 2026 talk on building agents for correct, transparent, and reproducible data analysis in light of the bluffbench results.</li>
</ul>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>This somewhat alleviated the fear that models had just memorized the bluffbench eval setup, mentioned in <a href="https://opensource.posit.co/blog/2026-06-19_ai-newsletter/" target="_blank" rel="noopener">our previous post</a>.&#160;<a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-07-17_ai-newsletter/images/hero.png" length="114420" type="image/png" />
    </item>
    <item>
      <title>Tips for managing your Python &amp; R environments in Positron</title>
      <link>https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/</link>
      <pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/</guid>
      <dc:creator>Cindy Tong</dc:creator>
      <dc:creator>Brice Stacey</dc:creator><description><![CDATA[<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p><a href="https://positron.posit.co" target="_blank" rel="noopener">Positron</a> is the Posit next-generation IDE for data science. Positron is an extensible, polyglot tool for exploring data and reproducible authoring in Python, R, and more.</p>
</div>
</div>
<p>If you work across Python and R or juggle multiple projects with different dependency needs, environment management is probably one of the most challenging parts of your workflow. <a href="https://positron.posit.co/" target="_blank" rel="noopener">Positron</a> comes with several out-of-the-box ways to help simplify environment management. Here are a few tips to streamline your workflow.</p>
<h2 id="1-understand-how-positron-discovers-your-environment">1: Understand how Positron discovers your environment
</h2>
<p>Positron does not just look at your system PATH. For Python, it actively discovers venv, uv, pyenv, and conda environments. If Positron is not seeing your project&rsquo;s virtual environment, you can add custom search locations via the <a href="positron://settings/python.interpreters.include"><code>python.interpreters.include</code></a> setting, or trigger a manual rescan with <em>Interpreter: Discover All Interpreters</em>. Learn more about <a href="https://positron.posit.co/python-installations.html#python-installation-discovery" target="_blank" rel="noopener">Python discovery in the Positron documentation</a>.</p>
<p>For R, discovery works differently. Positron consults various sources to build the list of R interpreters. These include your PATH, R root folders based on specific operating systems, well-known executable locations, and on Windows the registry. You can customize your R discovery through a few settings including <a href="positron://settings/positron.r.customRootFolders"><code>positron.r.customRootFolders</code></a> and <a href="positron://settings/positron.r.customBinaries"><code>positron.r.customBinaries</code></a>. Learn more about <a href="https://positron.posit.co/r-installations.html#customizing-r-discovery" target="_blank" rel="noopener">R discovery in the Positron documentation</a>.</p>
<h2 id="2-use-the-interpreter-selector">2. Use the Interpreter Selector
</h2>
<p>To begin your first session, click &ldquo;Start Session&rdquo; in the top right corner and select your preferred R or Python interpreter. Positron can run multiple R and Python interpreter sessions at once, but only one is ever the active session at a given moment. The Interpreter Selector always shows you the active session and its status (idle, busy, or shut down) and you can use it to switch between or start additional sessions.</p>
<img src="https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/active-interpreter-session.png" width="100%" data-fig-align="center" data-fig-alt="Use the Interpreter Selector to change your active session" />
<h2 id="3-install-python-with-uv">3. Install Python with uv
</h2>
<p>If Positron does not find a usable Python on your machine, it will offer to install Python via uv to help streamline your setup. If you prefer to manage the installation yourself, you can disable uv with the <a href="positron://settings/python.allowUvPythonInstall"><code>python.allowUvPythonInstall</code></a> setting. Check out our <a href="https://opensource.posit.co/blog/2026-07-08_positron-uv/" target="_blank" rel="noopener">blog post exploring on-demand Python installation in Positron</a> or explore configurations in our <a href="https://positron.posit.co/python-installations.html#troubleshooting" target="_blank" rel="noopener">Python installation documentation</a>.</p>
<img src="https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/install-python-with-uv.png" width="100%" data-fig-align="center" data-fig-alt="Install Python via uv" />
<h2 id="4-manage-packages-with-the-packages-pane">4. Manage packages with the Packages Pane
</h2>
<p>The Packages Pane in Positron lets you manage the packages installed in your active session. Whether you use pip, uv, conda, pak, base R, or renv, you can browse installed packages and search package repositories. You can also track outdated packages and install, update, or uninstall packages without leaving Positron or writing any code.</p>
<img src="https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/packages-pane.gif" width="100%"  data-fig-align="center" data-fig-alt="Manage Python and R packages in the Packages Pane" />
<h2 id="5-view-package-documentation">5. View package documentation
</h2>
<p>If you need to learn more about a specific package, the Packages Pane has buttons to navigate to the source documentation or link to the package&rsquo;s website. You can also pull up the <a href="https://positron.posit.co/help-pane.html" target="_blank" rel="noopener">Help Pane</a> for any reference in the Console using the <code>?</code> operator, including both packages and functions.</p>
<img src="https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/packages-pane-help.png" width="100%" data-fig-align="center" data-fig-alt="View source documentation for specific packages in the Help Pane" />
<h2 id="6-start-projects-with-new-folder-from-template">6. Start projects with New Folder from Template
</h2>
<p>The New Folder from Template flow helps you start new projects faster. Instead of running multiple setup commands you can make a few selections and Positron helps you set up an environment directory, version control, directory structure, and an interpreter instance. Learn more about the <a href="https://positron.posit.co/folder-templates.html" target="_blank" rel="noopener">Python and R templates</a> available in the documentation.</p>
<img src="https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/new-folder-from-template.png" width="100%" data-fig-align="center" data-fig-alt="New Folder from Template helps you start new projects faster with ready to use environments" />
<h2 id="7-make-use-of-the-command-palette">7. Make use of the Command Palette
</h2>
<p>We recommend taking advantage of the following commands when managing your environments:</p>
<ul>
<li><em>Interpreter: Discover All Interpreters</em>: Find new interpreters</li>
<li><em>Interpreter: Select Session</em>: Select a running interpreter session</li>
<li><em>Python: Create Environment</em>: Create a new virtual environment</li>
</ul>
<h2 id="8-let-posit-assistant-help-manage-your-environment">8. Let Posit Assistant help manage your environment
</h2>
<p><a href="https://assistant.posit.co/docs/features/context-management/" target="_blank" rel="noopener">Posit Assistant</a> has knowledge of your R and Python session including the language, version, names, and types of variables in your environment. You can prompt Posit Assistant to set up environments and troubleshoot issues you run into.</p>
<img src="https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/assistant-installed-packages.png" width="100%" data-fig-align="center" data-fig-alt="Posit Assistant in Positron showing a summary of packages installed in the user's Python environment, organized by category." />
<p>Have an idea for how we can improve environment management in Positron? We would love to hear from you in a <a href="https://github.com/posit-dev/positron/discussions" target="_blank" rel="noopener">discussion on GitHub</a>.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-07-15_positron-environment-management-tips/featured.svg" length="75593" type="image/svg&#43;xml" />
    </item>
    <item>
      <title>Positron July Release Highlights</title>
      <link>https://opensource.posit.co/blog/2026-07-13_positron-2026-07-release/</link>
      <pubDate>Mon, 13 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-07-13_positron-2026-07-release/</guid>
      <dc:creator>Julia Silge</dc:creator><description><![CDATA[<div class="callout callout-note" role="note" aria-label="Note">
<div class="callout-header">
<span class="callout-title">Note</span>
</div>
<div class="callout-body">
<p><a href="https://positron.posit.co" target="_blank" rel="noopener">Positron</a> is Posit&rsquo;s new, next-generation IDE for data science. Positron is designed to be an extensible, polyglot tool for exploring data and reproducible authoring in Python, R, and more.</p>
</div>
</div>
<p>Welcome back to another edition of our monthly Positron updates! Each month we share highlights from our <a href="https://positron.posit.co/release-notes" target="_blank" rel="noopener">latest release</a> and useful resources. <a href="https://opensource.posit.co/blog/2026-06-08_positron-2026-06-release/" target="_blank" rel="noopener">Last release</a> we told you that several major features were on track to leave preview in July, and the time is now here! The new Notebook editor, the Packages pane, and Posit Assistant are all now generally available.</p>
<h2 id="positron-notebook-editor">Positron Notebook Editor
</h2>
<p>Positron&rsquo;s <a href="https://positron.posit.co/positron-notebook-editor" target="_blank" rel="noopener">new Notebook editor</a> is now the default experience for Jupyter (<code>.ipynb</code>) files. This release brings a long list of additions, including split-pane editing, cell tag management, and executing a line or selection within a cell with <kbd>Cmd/Ctrl+Shift+Enter</kbd>. You can export notebooks to Quarto, Python, or R, and inline PDF rendering makes it easier to work with generated output. If you are coming from JupyterLab, you will find familiar keyboard shortcuts, and we have improved output fidelity for Mermaid diagrams, htmlwidgets, and ipywidgets.</p>
<img src="https://opensource.posit.co/blog/2026-07-13_positron-2026-07-release/positron-notebook-editor.png" data-fig-align="center" data-fig-alt="The Notebook editor in Positron, showing a Jupyter notebook with a Mermaid diagram." />
<p>Posit has built on and for the Jupyter ecosystem for years now, and we are excited to have recently <a href="https://opensource.posit.co/blog/2026-06-25_posit-joins-jupyter-foundation/" target="_blank" rel="noopener">joined the Jupyter Foundation</a>.</p>
<h2 id="packages-pane">Packages pane
</h2>
<p>The <a href="https://positron.posit.co/packages-pane" target="_blank" rel="noopener">Packages pane</a> has also come out of preview. It gives you a live view of the R and Python packages installed in your active session, so you can search, install, update, and remove packages, and jump to their documentation, without leaving Positron.</p>
<img src="https://opensource.posit.co/blog/2026-07-13_positron-2026-07-release/packages-pane.gif" data-fig-align="center" data-fig-alt="The Packages pane in Positron, showing installed Python packages with a detail editor." />
<p>This release also makes the pane more capable. Clicking a package opens a detail editor with its overview, metadata, and actions, similar to the Extensions pane. Installing a package searches results as you type, and each row has a button that opens the package&rsquo;s website. Update indicators appear immediately on a new session from a cached snapshot, and <strong>Update All Packages</strong> reports exactly what changed. Package operations keep your environment consistent too, resolving against a workspace <code>requirements.txt</code> for Python and updating the <code>renv.lock</code> snapshot for R.</p>
<h2 id="posit-assistant">Posit Assistant
</h2>
<p><a href="https://pos.it/assistant" target="_blank" rel="noopener">Posit Assistant</a>, our unified, data-science-focused approach to AI assistance, is now out of preview and generally available; Posit Assistant fully replaces the legacy Positron Assistant. We know the names are similar, and Joe Cheng <a href="https://opensource.posit.co/blog/2026-06-11_history-of-posit-data-science-agents/" target="_blank" rel="noopener">recently walked through a bit of the story behind how our AI agent tools have evolved</a>.</p>
<p>This release also gives you finer control over AI in Positron. The new <a href="positron://settings/ai.enabled"><code>ai.enabled</code></a> setting turns off every Positron AI feature at once, and administrators can enforce it; <a href="positron://settings/notebook.ai.enabled"><code>notebook.ai.enabled</code></a> does the same for notebooks specifically. The set of language model providers keeps growing, with DeepSeek joining as an experimental provider and Microsoft Foundry reaching general availability. The configuration modal now shows every provider by default with preview and experimental badges.</p>
<h2 id="data-explorer">Data Explorer
</h2>
<p>The <a href="https://positron.posit.co/data-explorer" target="_blank" rel="noopener">Data Explorer</a> can now open Excel workbooks directly, with no code required, without first needing to load it via Python or R. Sort, filter, and profile columns, switch between worksheets, and toggle whether the first row holds column names. An <strong>Open in Excel</strong> button opens the workbook in your native spreadsheet application.</p>
<p>This release broadens what you can open in the Data Explorer overall. Backed by a native DuckDB engine, it now also previews compressed CSV, TSV, and Parquet files. Also, a new <strong>Open in Data Explorer</strong> code action, from the editor lightbulb or <kbd>Cmd+.</kbd>, opens the data frame under your cursor in R, Python, and Quarto files, so you can jump straight from your code to exploring your data.</p>
<img src="https://opensource.posit.co/blog/2026-07-13_positron-2026-07-release/open-in-data-explorer-code-action.gif" data-fig-align="center" data-fig-alt="The Open in Data Explorer code action in Positron, showing a lightbulb menu in a Python file with the option to open a data frame in the Data Explorer." />
<h2 id="r-language-intelligence">R language intelligence
</h2>
<p>Last release we introduced within-file symbol resolution for R; this release expands it across files. Go to Definition, Find References, and Rename Symbol now work across the packages and scripts in your workspace, and diagnostics and workspace symbols react to external file changes so your language intelligence stays in sync as your project evolves.</p>
<h2 id="whats-coming-next">What&rsquo;s coming next
</h2>
<ul>
<li>New in preview this release, Data Connections lets you browse the schemas, tables, views, and indexes of a database, open tables in the Data Explorer, and generate connection code. It currently supports DuckDB, PostgreSQL, and SQLite. Learn how to try it out and tell us what you think in the <a href="https://github.com/posit-dev/positron/discussions/14571" target="_blank" rel="noopener">Data Connections discussion post</a>: which databases and warehouses you need, whether the connection setup is clear, and anything confusing, missing, or broken.</li>
<li>We are looking forward to posit::conf(2026) in September, where our team will have several sessions on Positron. <a href="https://conf.posit.co/2026/" target="_blank" rel="noopener">Register now</a> to join us in person in Houston or virtually from anywhere in the world.</li>
</ul>
<div class="callout callout-tip" role="note" aria-label="Tip">
<div class="callout-header">
<span class="callout-title">Tip</span>
</div>
<div class="callout-body">
<p><a href="https://positron.posit.co/download" target="_blank" rel="noopener">Download Positron</a> to try out the new features and improvements in this release!</p>
</div>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-07-13_positron-2026-07-release/featured.svg" length="88493" type="image/svg&#43;xml" />
    </item>
    <item>
      <title>mcptools 1.0.0</title>
      <link>https://opensource.posit.co/blog/2026-07-06_mcptools-1-0-0/</link>
      <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-07-06_mcptools-1-0-0/</guid>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<p>The first major release of <a href="https://posit-dev.github.io/mcptools/" target="_blank" rel="noopener">mcptools</a>, an R SDK for the Model Context Protocol, is now on CRAN! This release includes several notable features: fetching tools from remote authenticated servers, deploying MCP servers on Posit Connect, and support for images and other rich content types.</p>
<p>To install the package, run:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">install.packages</span><span class="p">(</span><span class="s">&#34;mcptools&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
<p>To demo these new features, I&rsquo;ll deploy an R function that returns a picture to an MCP server on Posit Connect. Then, in another R session, I&rsquo;ll connect to that server and ask an ellmer chat to take a look at the picture and tell me what it sees.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">mcptools</span><span class="p">)</span></span></span></code></pre></div></div>
<h2 id="images-in-tool-results">Images in tool results
</h2>
<p>mcptools supports &ldquo;both directions&rdquo; of MCP. In one direction, users can deploy R functions as MCP servers. In the other direction, users can fetch tools from third-party MCP servers as R functions. <strong>mcptools now supports both serving and fetching tools that return images.</strong></p>
<p>As an example of a function that returns an image, let&rsquo;s consider a tool <code>fetch_reference_image()</code>:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="nf">library</span><span class="p">(</span><span class="n">ellmer</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">tools</span> <span class="o">&lt;-</span> <span class="nf">list</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">  <span class="n">fetch_reference_image</span> <span class="o">=</span> <span class="nf">tool</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="kr">function</span><span class="p">()</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">      <span class="nf">content_image_url</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">        <span class="s">&#34;https://simonpcouch.com/blog/2026-04-16-local-agents-2/featured.png&#34;</span>
</span></span><span class="line"><span class="cl">      <span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="n">name</span> <span class="o">=</span> <span class="s">&#34;fetch_reference_image&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">description</span> <span class="o">=</span> <span class="s">&#34;Fetch the reference image.&#34;</span>
</span></span><span class="line"><span class="cl">  <span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span></span></span></code></pre></div></div>
<p>Running that function:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">tools</span><span class="o">$</span><span class="nf">fetch_reference_image</span><span class="p">()</span></span></span></code></pre></div></div>
<figure>
<img src="https://simonpcouch.com/blog/2026-04-16-local-agents-2/featured.png" alt="A brown and white Border Collie on a deck, looking attentively at the camera, with wire railings and blurred greenery in the background." />
<figcaption aria-hidden="true">A brown and white Border Collie on a deck, looking attentively at the camera, with wire railings and blurred greenery in the background.</figcaption>
</figure>
<p>Without MCP, I can ask a model to look at the image and tell me what it sees:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">ch</span> <span class="o">&lt;-</span> <span class="nf">chat_claude</span><span class="p">(</span><span class="s">&#34;Be brief.&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">ch</span><span class="o">$</span><span class="nf">register_tool</span><span class="p">(</span><span class="n">tools[[1]]</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">ch</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span><span class="s">&#34;What do you see in the reference image?&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; I can see a beautiful Border Collie dog in the reference</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; image. The dog has:</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - A chocolate brown and white coat with distinctive</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   coloring</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - Amber/brown eyes</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - Alert, perked ears</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - A white blaze down the center of its face</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - A pink/brown nose</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - Long, fluffy fur typical of the breed</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; The dog appears to be outdoors, positioned near what</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; looks like a wooden post or railing with wire fencing</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; visible in the background. There&#39;s greenery and trees</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; visible in the blurred background. The dog has an</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; attentive, intelligent expression that&#39;s characteristic</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; of Border Collies.</span></span></span></code></pre></div></div>
<p>In the next two sections, I&rsquo;ll deploy this function on Posit Connect and then fetch it from a fresh R session so that future ellmer chats can fetch this same image.</p>
<h2 id="mcp-servers-on-posit-connect">MCP Servers on Posit Connect
</h2>
<p>mcptools has now adopted plumber2&rsquo;s <code>_server.yml</code> open standard. This means that, <strong>to deploy an MCP server on Posit Connect, you just need to add a <code>_server.yml</code> file in your project root</strong> that looks like this:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yml" data-lang="yml"><span class="line"><span class="cl"><span class="nt">engine</span><span class="p">:</span><span class="w"> </span><span class="l">mcptools</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">tools</span><span class="p">:</span><span class="w"> </span><span class="l">tools.R</span></span></span></code></pre></div></div>
<p><code>tools.R</code> (or whatever you choose to name the file) is a file that defines ellmer tools and returns them in a list. In my case, <code>tools.R</code> looks exactly like the code chunk defining <code>tools</code> above. I can then run <code>rsconnect::deployAPI(&quot;.&quot;, contentCategory = &quot;mcp&quot;)</code> to deploy the tools from <code>tools.R</code> as an authenticated MCP server on Posit Connect.</p>
<h2 id="fetch-tools-as-r-functions-from-authenticated-mcp-servers">Fetch tools as R functions from authenticated MCP servers
</h2>
<p>Perhaps the biggest gap in mcptools before this release was that <code>mcp_tools()</code> did not support remote, authenticated MCP servers. Up to this point, I had written in documentation that folks ought to use <code>npx mcp-remote</code>, which converts remote MCP servers (like Slack, Confluence, or really any of the most well-adopted third-party MCP servers) into local ones. That meant that, even though mcptools only implemented the local half of the protocol, mcptools users could connect to remotely hosted MCP servers.</p>
<p>This is undesirable for a few reasons. For one, mcptools should &ldquo;just work&rdquo; without users having to install software from sources other than CRAN. Further, installing code via <code>npx</code> is particularly problematic; the node package registry has been the source of <a href="https://www.axios.com/2026/03/31/north-korean-hackers-implicated-in-major-supply-chain-attack" target="_blank" rel="noopener">a</a> <a href="https://arstechnica.com/security/2026/06/dozens-of-red-hat-packages-backdoored-through-its-offical-npm-channel/" target="_blank" rel="noopener">number</a> <a href="https://www.theregister.com/cyber-crime/2026/05/18/shai-hulud-copycat-hits-another-npm-package/5242180" target="_blank" rel="noopener">of</a> <a href="https://www.stepsecurity.io/blog/mini-shai-hulud-is-back-a-self-spreading-supply-chain-attack-hits-the-npm-ecosystem" target="_blank" rel="noopener">particularly</a> <a href="https://unit42.paloaltonetworks.com/monitoring-npm-supply-chain-attacks/" target="_blank" rel="noopener">concerning</a> supply chain attacks recently.</p>
<p><strong><code>mcp_tools()</code> now natively supports fetching tools from authenticated, third-party MCP servers.</strong> For example, a configuration that used to look like:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="p">{</span>
</span></span><span class="line"><span class="cl">  <span class="nt">&#34;mcpServers&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;connect&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">      <span class="nt">&#34;command&#34;</span><span class="p">:</span> <span class="s2">&#34;npx&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="nt">&#34;args&#34;</span><span class="p">:</span> <span class="p">[</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;mcp-remote&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;&lt;my_deployed_connect_listing&gt;&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;--header&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;Authorization: Key ${CONNECT_API_KEY}&#34;</span>
</span></span><span class="line"><span class="cl">      <span class="p">]</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl">  <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span></span></span></code></pre></div></div>
<p>Now looks like:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="p">{</span>
</span></span><span class="line"><span class="cl">  <span class="nt">&#34;mcpServers&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;connect&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">      <span class="nt">&#34;url&#34;</span><span class="p">:</span> <span class="s2">&#34;&lt;my_deployed_connect_listing&gt;&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">      <span class="nt">&#34;headers&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">        <span class="nt">&#34;Authorization&#34;</span><span class="p">:</span> <span class="s2">&#34;Key ${CONNECT_API_KEY}&#34;</span>
</span></span><span class="line"><span class="cl">      <span class="p">}</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl">  <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span></span></span></code></pre></div></div>
<p>I&rsquo;ll save that latter configuration as <code>config.json</code>. Then, I&rsquo;ll provide it to <code>mcp_tools()</code>:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">tools_fetched</span> <span class="o">&lt;-</span> <span class="nf">mcp_tools</span><span class="p">(</span><span class="n">config</span> <span class="o">=</span> <span class="s">&#34;config.json&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nf">class</span><span class="p">(</span><span class="n">tools_fetched[[1]]</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; [1] &#34;ellmer::ToolDef&#34; &#34;function&#34;        &#34;S7_object&#34;</span></span></span></code></pre></div></div>
<p>Now I can register the fetched tools with a chat in a new R session, and it has access to the same image:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-r" data-lang="r"><span class="line"><span class="cl"><span class="n">ch_new</span> <span class="o">&lt;-</span> <span class="nf">chat_claude</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="n">ch_new</span><span class="o">$</span><span class="nf">register_tools</span><span class="p">(</span><span class="n">tools_fetched</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">ch_new</span><span class="o">$</span><span class="nf">chat</span><span class="p">(</span><span class="s">&#34;What&#39;s in the reference image?&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; The reference image features a **Border Collie** dog.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Here are some details:</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - **Coloring**: The dog has a striking **brown</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   (chocolate) and white** coat, with a distinctive white</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   stripe running down the center of its face.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - **Expression**: It has an alert and attentive look,</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   with beautiful **amber/brown eyes**.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - **Tongue**: Its tongue is slightly visible, giving it a</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   cute appearance.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - **Setting**: The dog appears to be on a **deck or</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   porch**, with cable/wire railings visible, and a lush</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   green, leafy background suggesting an outdoor, wooded</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   area.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; - **Accessories**: It appears to be wearing a small</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;   **collar tag**.</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt;</span>
</span></span><span class="line"><span class="cl"><span class="c1">#&gt; Overall, it&#39;s a beautiful and expressive dog photo! 🐕</span></span></span></code></pre></div></div>
<p>Taken together, the changes in this release should allow R users to do much more with MCP! For a more complete list of changes in this release, see the package <a href="https://posit-dev.github.io/mcptools/news/index.html" target="_blank" rel="noopener">changelog</a>.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-07-06_mcptools-1-0-0/featured.png" length="874368" type="image/png" />
    </item>
    <item>
      <title>AI Newsletter: AGENTS.md vs Skills vs MCP servers</title>
      <link>https://opensource.posit.co/blog/2026-07-03_ai-newsletter/</link>
      <pubDate>Fri, 03 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-07-03_ai-newsletter/</guid>
      <dc:creator>Sara Altman</dc:creator>
      <dc:creator>Simon Couch</dc:creator><description><![CDATA[<h2 id="prompt-agentsmd-skill-or-mcp-server">Prompt, AGENTS.md, skill, or MCP server?
</h2>
<p>There are a variety of ways to give a coding agent new information or abilities, but it can be confusing when to use each one. Generally, with coding agents like Claude Code or Posit Assistant, you can:</p>
<ul>
<li><strong>Write prompts</strong> in the chat (i.e., normal usage).</li>
<li>Add information to either a project-level or user-level <strong><code>CLAUDE.md</code> or <code>AGENTS.md</code>.</strong> The contents of a directory&rsquo;s <code>CLAUDE.md</code> or <code>AGENTS.md</code> are included in the agent&rsquo;s system prompt in every session in that directory. You can also create user-level versions that apply to every session.</li>
<li>Write a <strong>skill</strong> or use an existing one. Skills are packaged instructions that can include both text and code. The agent loads a skill only when it&rsquo;s relevant.</li>
<li>Add an <strong>MCP server.</strong> An <a href="https://modelcontextprotocol.io/docs/getting-started/intro" target="_blank" rel="noopener">MCP (Model Context Protocol)</a> server provides an agent with access to otherwise hard-to-find context, mostly through <a href="https://modelcontextprotocol.io/docs/learn/server-concepts#tools" target="_blank" rel="noopener">tools</a>, using a standardized interface.</li>
</ul>
<p>This list is roughly ordered from most straightforward to most complicated.<sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup></p>
<p>So when do you use one over the other? There are two axes that might matter for your decision.</p>
<p>The first axis is reusability: do you want the agent to perform the task or access the information just once, or many times? The more often you or others will reuse something, the more it&rsquo;s worth encoding somewhere more permanent.</p>
<p>The second axis is reach. Prompting, <code>CLAUDE.md</code>/<code>AGENTS.md</code> files, and skills all provide guidance on how to best make use of the existing context and tools. MCP servers can provide the agent with entirely new tools (in the <a href="https://ellmer.tidyverse.org/articles/tool-calling.html" target="_blank" rel="noopener">agent tools</a> sense), granting the agent access to hard-to-reach information.</p>
<p><div class="not-prose"><figure>
    <img class="h-auto max-w-full rounded-lg"
      src="https://opensource.posit.co/blog/2026-07-03_ai-newsletter/images/diagram.excalidraw.svg"
      alt="Decision tree for choosing how to customize a coding agent. The first question asks whether this is a one-time thing or something you&rsquo;ll need the agent to do similarly in the future. A one-time thing leads to normal prompting. If you&rsquo;ll need it again, the next question is when you want the agent to do the task: every session in the project leads to a CLAUDE.md or AGENTS.md file, while &ldquo;as needed&rdquo; leads to a further question. That question asks whether the agent has the capability to do the task with the tools it already has: if yes, write a skill; if no, use an MCP server."  title="A decision tree for choosing between prompting, AGENTS.md, skills, and MCP servers." 
      loading="lazy"
    ><figcaption class="text-sm text-center text-gray-500">A decision tree for choosing between prompting, AGENTS.md, skills, and MCP servers.</figcaption>
  </figure></div>
</p>
<p><strong>MCP servers often seem like the solution when you need to grant an agent access to an outside system, but they aren&rsquo;t always necessary.</strong> In many cases, what seems like a task for an MCP server can actually be solved by a command-line interface (CLI) tool, or a CLI plus a skill that tells the agent how to use it. For example, GitHub has an <a href="https://github.com/github/github-mcp-server" target="_blank" rel="noopener">MCP server</a>, but the <a href="https://cli.github.com/" target="_blank" rel="noopener"><code>gh</code></a> CLI does roughly the same thing.</p>
<p>The GitHub MCP server works by providing the agent with new tools, whereas the <code>gh</code> CLI takes advantage of the agent&rsquo;s existing bash tool that lets it run arbitrary shell commands. The skill plus CLI option is therefore generally preferable from a simplicity standpoint, but also from a token standpoint: adding the GitHub MCP server would inject tens of thousands of tokens of tool definitions into every request, whereas a skill that tells the agent to use <code>gh</code> costs almost nothing and loads only when it&rsquo;s relevant.</p>
<p>However, some information sources are hard to reach via the command line, because no CLI exists, the one that does isn&rsquo;t fully featured, or it&rsquo;s difficult for the agent to use. In some of these cases, the same sources can be accessed more effectively via MCP servers, such as design tools like Figma, knowledge repositories like Notion or Confluence, or issue trackers like Linear or Jira.</p>
<h2 id="posit-news">Posit news
</h2>
<h3 id="posit-assistant-in-positron">Posit Assistant in Positron
</h3>
<p>As of the <a href="https://positron.posit.co/download.html#release-notes" target="_blank" rel="noopener">June release</a> of Positron, Posit Assistant is now the default experience in Positron. Positron Assistant will be deprecated starting in the 2026.07 release (release date July 6).</p>
<p>We understand that the names are confusing! Our hope is that the transition state will be over soon and the confusion will lessen. If you want to understand why we gave different assistants very similar names, read this blog post from Posit CTO Joe Cheng: <a href="https://opensource.posit.co/blog/2026-06-11_history-of-posit-data-science-agents/" target="_blank" rel="noopener">A brief and biased history of Posit data science agents</a>.</p>
<p>Posit Assistant works with the same providers as Positron Assistant. If you have a working provider setup with Positron Assistant, you&rsquo;ll be able to use that same setup with Posit Assistant. You can read more about available providers <a href="https://assistant.posit.co/docs/downloads/positron/" target="_blank" rel="noopener">here</a>.</p>
<h3 id="package-updates">Package updates
</h3>
<ul>
<li><a href="https://opensource.posit.co/blog/2026-07-01_raghilda-0-2-0/" target="_blank" rel="noopener">raghilda v0.2</a>, a Python package for Retrieval Augmented Generation, is now on PyPI. This release, among other things, broadens support for crawling sites.</li>
<li><a href="https://opensource.posit.co/blog/2026-06-22_debrief-0-1-0/" target="_blank" rel="noopener">debrief</a>, an R package for LLM-friendly profiling, is now on CRAN. debrief turns profvis profiling output into text-based summaries, allowing AI agents to optimize R code more effectively.</li>
</ul>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>One interface we didn&rsquo;t mention here is &lsquo;custom agents,&rsquo; a concept popularized by GitHub Copilot and now appearing under various names. These bundle some combination of prompts and tools (sometimes gathered via MCP). We&rsquo;d reach for the options mentioned above first, which are open standards that are broadly supported across most agent platforms and thus can be shared and migrated more easily.&#160;<a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</div>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-07-03_ai-newsletter/images/hero.png" length="83964" type="image/png" />
    </item>
    <item>
      <title>raghilda `v0.2`: crawl APIs, PostgreSQL, and more</title>
      <link>https://opensource.posit.co/blog/2026-07-01_raghilda-0-2-0/</link>
      <pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://opensource.posit.co/blog/2026-07-01_raghilda-0-2-0/</guid>
      <dc:creator>Rich Iannone</dc:creator>
      <dc:creator>Tomasz Kalinowski</dc:creator><description><![CDATA[<p>When we <a href="https://opensource.posit.co/blog/2026-04-14_rag-with-raghilda/" target="_blank" rel="noopener">introduced
raghilda</a>
in April, the package handled the core RAG workflow: read documents,
chunk them, embed them into a store, and retrieve relevant chunks at
query time. raghilda <code>v0.2</code> broadens that scope considerably. The
release adds a structured crawl and ingest API that replaces the manual
read-and-upsert loop with a pipeline built around caching, concurrency,
and composable crawlers (including a <code>CloudflareCrawler</code> that can index
JavaScript-rendered sites without running a local headless browser). It
also ships a PostgreSQL store backend and NVIDIA NIM embedding support.</p>
<p>This post walks through the major additions. The package&rsquo;s fundamentals
have not changed (stores, chunkers, retrievers, and the pattern for
connecting to chatlas all work as before), but the surface area for
building and maintaining stores in production has grown substantially.</p>
<h2 id="the-crawl-and-ingest-api">The crawl and ingest API
</h2>
<p>The largest change in <code>v0.2</code> is a new API for crawling sources and
ingesting them into a store. In <code>v0.1</code>, building a store meant calling
<code>read_as_markdown()</code> on individual URLs, chunking each result, and
upserting them one by one. That works for a handful of pages, but it
becomes unwieldy for larger collections, where you also want caching (to
avoid re-fetching unchanged content) and concurrency (to finish in
minutes rather than hours).</p>
<p>The new API introduces a clean separation between crawling and storage.
On the crawl side, a crawler object produces <code>MarkdownDocument</code> objects
from a defined scope. On the store side, <code>store.ingest()</code> consumes those
documents lazily, applies an optional preparation step (typically
chunking), and writes them to the store with configurable parallelism.
The pipeline looks like this:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.chunker</span> <span class="kn">import</span> <span class="n">MarkdownChunker</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.crawl</span> <span class="kn">import</span> <span class="n">CrawlScope</span><span class="p">,</span> <span class="n">WebCrawler</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.embedding</span> <span class="kn">import</span> <span class="n">EmbeddingOpenAI</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.store</span> <span class="kn">import</span> <span class="n">DuckDBStore</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">crawler</span> <span class="o">=</span> <span class="n">WebCrawler</span><span class="p">(</span><span class="n">cache_dir</span><span class="o">=</span><span class="s2">&#34;.cache/crawl&#34;</span><span class="p">,</span> <span class="n">max_workers</span><span class="o">=</span><span class="mi">4</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">scope</span> <span class="o">=</span> <span class="n">CrawlScope</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">roots</span><span class="o">=</span><span class="p">[</span><span class="s2">&#34;https://example.com/docs&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">include_patterns</span><span class="o">=</span><span class="p">[</span><span class="s2">&#34;https://example.com/docs/**&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">depth</span><span class="o">=</span><span class="mi">2</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">documents</span> <span class="o">=</span> <span class="n">crawler</span><span class="o">.</span><span class="n">markdown_documents</span><span class="p">(</span><span class="n">scope</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">store</span> <span class="o">=</span> <span class="n">DuckDBStore</span><span class="o">.</span><span class="n">create</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">location</span><span class="o">=</span><span class="s2">&#34;raghilda.duckdb&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">embed</span><span class="o">=</span><span class="n">EmbeddingOpenAI</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl">    <span class="n">overwrite</span><span class="o">=</span><span class="kc">True</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">summary</span> <span class="o">=</span> <span class="n">store</span><span class="o">.</span><span class="n">ingest</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">documents</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">prepare</span><span class="o">=</span><span class="n">MarkdownChunker</span><span class="p">(</span><span class="n">chunk_size</span><span class="o">=</span><span class="mi">1000</span><span class="p">)</span><span class="o">.</span><span class="n">chunk</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">max_workers</span><span class="o">=</span><span class="mi">4</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nb">print</span><span class="p">(</span><span class="n">summary</span><span class="p">)</span></span></span></code></pre></div></div>
<pre><code>IngestSummary(inserted=142, replaced=0, skipped=0)
</code></pre>
<p>The <code>CrawlScope</code> dataclass defines the traversal policy: root URLs,
include/exclude patterns, depth limits, and page count caps. The crawler
handles the mechanics of fetching and converting pages to Markdown,
while the scope tells it where to go. This separation means you can
change the backend (swap <code>WebCrawler</code> for <code>CloudflareCrawler</code>, for
instance) without redefining the scope, and vice versa.</p>
<p>Three concrete crawlers ship with <code>v0.2</code>. <code>DirectoryCrawler</code> walks local
file trees and converts supported formats to Markdown. <code>WebCrawler</code>
fetches pages over HTTP using <code>requests</code> and converts them locally.
<code>CloudflareCrawler</code> delegates both fetching and rendering to
Cloudflare&rsquo;s Browser Rendering API (the right tool for sites that load
content through JavaScript). All three implement the same interface:
<code>origins()</code> to discover pages, <code>fetch_raw()</code> and <code>fetch_markdown()</code> for
single-page access, and <code>markdown_documents()</code> for the full pipeline.</p>
<p>Caching is built into the crawler layer. When you pass <code>cache_dir=True</code>
(or an explicit path), each crawler stores fetched content and converted
Markdown in a flat directory of files with metadata sidecars. On
subsequent runs, cached entries are reused if they are still fresh. For
<code>WebCrawler</code> and <code>CloudflareCrawler</code>, freshness is controlled by
<code>cache_stale_after=</code>, a <code>timedelta</code> that defines how long a cached entry
remains valid. This makes interrupted workflows resumable without any
explicit checkpoint logic: rerun the script and the cache supplies
everything that was already fetched, while only new or stale pages
trigger network requests.</p>
<p>Concurrency operates on both sides of this boundary independently. The
crawler can fetch and convert pages in parallel (controlled by
<code>max_workers=</code> on the crawler constructor), and <code>store.ingest()</code> can
write to the store concurrently (controlled by its own <code>max_workers=</code>
argument). For <code>WebCrawler</code>, the breadth-first frontier is explored
concurrently while preserving stable output order, so results come back
in a consistent sequence regardless of which pages respond first.</p>
<h2 id="cloudflarecrawler"><code>CloudflareCrawler</code>
</h2>
<p>The <code>CloudflareCrawler</code> moves the work of a crawl off the local machine.
Instead of fetching pages and converting them to Markdown with local
processes, it hands both jobs to Cloudflare&rsquo;s Browser Rendering API, so
a long crawl over a large site runs on Cloudflare&rsquo;s distributed
infrastructure rather than competing for local CPU and bandwidth. For
collections large enough that concurrent local requests become the
bottleneck, this is the primary reason to reach for it: the slow,
sustained part of building a store happens remotely, and what returns is
ready-to-chunk Markdown. The same arrangement resolves a problem that
defeats a plain HTTP fetch, because the API renders each page in a real
browser, executing JavaScript and waiting for the DOM to settle before
extracting content. Sites built with React, Vue, or Angular, which an
ordinary request reduces to an empty shell, are therefore handled
without extra configuration or a locally installed headless browser.</p>
<p>The usage looks almost identical to <code>WebCrawler</code>, because both share the
same crawl interface. The key difference is that the constructor takes
Cloudflare credentials instead of an HTTP session, and Cloudflare&rsquo;s
infrastructure handles the rendering remotely (so there is no need to
install Playwright, Selenium, or any other local headless browser):</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">import</span> <span class="nn">os</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.crawl</span> <span class="kn">import</span> <span class="n">CloudflareCrawler</span><span class="p">,</span> <span class="n">CrawlScope</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">crawler</span> <span class="o">=</span> <span class="n">CloudflareCrawler</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">account_id</span><span class="o">=</span><span class="n">os</span><span class="o">.</span><span class="n">environ</span><span class="p">[</span><span class="s2">&#34;CLOUDFLARE_ACCOUNT_ID&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">api_token</span><span class="o">=</span><span class="n">os</span><span class="o">.</span><span class="n">environ</span><span class="p">[</span><span class="s2">&#34;CLOUDFLARE_API_TOKEN&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">cache_dir</span><span class="o">=</span><span class="kc">True</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">render</span><span class="o">=</span><span class="kc">True</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">max_workers</span><span class="o">=</span><span class="mi">4</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">scope</span> <span class="o">=</span> <span class="n">CrawlScope</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">roots</span><span class="o">=</span><span class="p">[</span><span class="s2">&#34;https://my-spa-docs.example.com/&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">depth</span><span class="o">=</span><span class="mi">2</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">include_patterns</span><span class="o">=</span><span class="p">[</span><span class="s2">&#34;https://my-spa-docs.example.com/**&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">limit</span><span class="o">=</span><span class="mi">500</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">documents</span> <span class="o">=</span> <span class="n">crawler</span><span class="o">.</span><span class="n">markdown_documents</span><span class="p">(</span><span class="n">scope</span><span class="p">)</span></span></span></code></pre></div></div>
<p>Iterating the result performs the crawl lazily, yielding one
<code>MarkdownDocument</code> per page. Each document exposes the page&rsquo;s <code>origin</code>
alongside its rendered Markdown <code>content</code>, so a quick pass confirms that
the JavaScript-rendered pages came back with real text rather than the
empty shells a plain HTTP fetch would have produced:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">for</span> <span class="n">doc</span> <span class="ow">in</span> <span class="n">documents</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">    <span class="nb">print</span><span class="p">(</span><span class="n">doc</span><span class="o">.</span><span class="n">origin</span><span class="p">,</span> <span class="sa">f</span><span class="s2">&#34;(</span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">doc</span><span class="o">.</span><span class="n">content</span><span class="p">)</span><span class="si">:</span><span class="s2">,</span><span class="si">}</span><span class="s2"> chars)&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
<pre><code>https://my-spa-docs.example.com/ (3,214 chars)
https://my-spa-docs.example.com/guide/install (5,902 chars)
https://my-spa-docs.example.com/guide/config (8,477 chars)
https://my-spa-docs.example.com/api/reference (12,043 chars)
</code></pre>
<p>The <code>render=True</code> default tells Cloudflare to execute JavaScript before
extracting content. For server-rendered sites where JavaScript execution
is unnecessary, setting <code>render=False</code> reduces crawl time and API usage.
The <code>source=</code> parameter controls how pages are discovered: <code>&quot;all&quot;</code> (the
default) combines multiple discovery methods, <code>&quot;sitemap&quot;</code> reads from the
site&rsquo;s <code>sitemap.xml</code>, <code>&quot;crawl&quot;</code> follows links from the rendered DOM, and
<code>&quot;urls&quot;</code> processes only the explicitly provided roots.</p>
<p>For stores that need regular updates, the <code>modified_since=</code> parameter
restricts the crawl to pages modified after a given Unix timestamp,
keeping refresh jobs lightweight. Combined with the crawl cache and the
store&rsquo;s own deduplication (identical documents are not re-embedded), an
incremental update script can run daily without redundant work:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">import</span> <span class="nn">time</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">datetime</span> <span class="kn">import</span> <span class="n">timedelta</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">one_week_ago</span> <span class="o">=</span> <span class="nb">int</span><span class="p">(</span><span class="n">time</span><span class="o">.</span><span class="n">time</span><span class="p">())</span> <span class="o">-</span> <span class="p">(</span><span class="mi">7</span> <span class="o">*</span> <span class="mi">24</span> <span class="o">*</span> <span class="mi">60</span> <span class="o">*</span> <span class="mi">60</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">crawler</span> <span class="o">=</span> <span class="n">CloudflareCrawler</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">account_id</span><span class="o">=</span><span class="n">os</span><span class="o">.</span><span class="n">environ</span><span class="p">[</span><span class="s2">&#34;CLOUDFLARE_ACCOUNT_ID&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">api_token</span><span class="o">=</span><span class="n">os</span><span class="o">.</span><span class="n">environ</span><span class="p">[</span><span class="s2">&#34;CLOUDFLARE_API_TOKEN&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">cache_dir</span><span class="o">=</span><span class="kc">True</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">cache_stale_after</span><span class="o">=</span><span class="n">timedelta</span><span class="p">(</span><span class="n">days</span><span class="o">=</span><span class="mi">1</span><span class="p">),</span>
</span></span><span class="line"><span class="cl">    <span class="n">modified_since</span><span class="o">=</span><span class="n">one_week_ago</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span></span></span></code></pre></div></div>
<p>The tradeoff is cost: <code>CloudflareCrawler</code> requires a Cloudflare account
with Browser Rendering access. For static HTML sites where a plain HTTP
fetch returns the full content, <code>WebCrawler</code> remains the simpler and
free option. Both crawlers share the same interface, so switching
between them requires only a constructor change.</p>
<h2 id="postgresql-store">PostgreSQL store
</h2>
<p>raghilda <code>v0.1</code> shipped with three store backends: DuckDB (local,
zero-config), ChromaDB, and OpenAI Vector Stores. <code>v0.2</code> adds
<code>PostgreSQLStore</code>, backed by <code>psycopg2</code> and <code>pgvector</code>. This is the
natural choice for production deployments where the store needs to be
shared across services, or where you already have PostgreSQL
infrastructure.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.store</span> <span class="kn">import</span> <span class="n">PostgreSQLStore</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.embedding</span> <span class="kn">import</span> <span class="n">EmbeddingOpenAI</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">store</span> <span class="o">=</span> <span class="n">PostgreSQLStore</span><span class="o">.</span><span class="n">create</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">connection</span><span class="o">=</span><span class="s2">&#34;postgresql://user:pass@localhost:5432/mydb&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">embed</span><span class="o">=</span><span class="n">EmbeddingOpenAI</span><span class="p">(),</span>
</span></span><span class="line"><span class="cl">    <span class="n">name</span><span class="o">=</span><span class="s2">&#34;docs_store&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">overwrite</span><span class="o">=</span><span class="kc">True</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span></span></span></code></pre></div></div>
<p>The store supports full-text search via PostgreSQL&rsquo;s native
<code>tsvector</code>/<code>tsquery</code> with a pre-computed column and GIN index, vector
similarity search via pgvector with HNSW indexes (supporting cosine, L2,
and inner product distance metrics), and combined retrieval that merges
both result sets with deoverlap support.</p>
<p>Retrieval uses the same interface as every other backend: a single
<code>retrieve()</code> call returns a ranked list of chunks, each carrying its
similarity score under <code>metrics</code> and its heading-hierarchy <code>context</code>.
Running a query against a store populated with the raghilda
documentation returns the most relevant chunks first:</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="n">results</span> <span class="o">=</span> <span class="n">store</span><span class="o">.</span><span class="n">retrieve</span><span class="p">(</span><span class="s2">&#34;How are vector indexes configured?&#34;</span><span class="p">,</span> <span class="n">top_k</span><span class="o">=</span><span class="mi">2</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">for</span> <span class="n">r</span> <span class="ow">in</span> <span class="n">results</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">    <span class="nb">print</span><span class="p">(</span><span class="sa">f</span><span class="s2">&#34;Score: </span><span class="si">{</span><span class="n">r</span><span class="o">.</span><span class="n">metrics</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span><span class="o">.</span><span class="n">value</span><span class="si">:</span><span class="s2">.4f</span><span class="si">}</span><span class="s2">&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="nb">print</span><span class="p">(</span><span class="n">r</span><span class="o">.</span><span class="n">context</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="nb">print</span><span class="p">(</span><span class="n">r</span><span class="o">.</span><span class="n">text</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="nb">print</span><span class="p">(</span><span class="s2">&#34;---&#34;</span><span class="p">)</span></span></span></code></pre></div></div>
<pre><code>Score: 0.5142
# PostgreSQL store
Vector similarity search runs through pgvector with HNSW indexes,
supporting cosine, L2, and inner product distance metrics. Call
build_index() after ingestion to create them.
---
Score: 0.4417
# PostgreSQL store &gt; Combined retrieval
Combined retrieval merges the vector and full-text result sets and
applies deoverlap to drop redundant overlapping chunks, returning a
single ranked list from one retrieve() call.
---
</code></pre>
<p>Attributes work as expected: scalar types map to columns, struct types
map to JSONB, and attribute filters can query into JSONB fields using
the <code>-&gt;&gt;</code> operator. The <code>build_index()</code> method creates HNSW indexes
after ingestion, and the <code>vss_index=</code> parameter on <code>create()</code> controls
the default index type.</p>
<p>Connection strings are accepted directly in <code>create()</code> and <code>connect()</code>,
so you can point the store at any PostgreSQL instance with pgvector
installed. If the pgvector extension is missing, the store raises an
informative error rather than failing cryptically on the first vector
operation.</p>
<h2 id="nvidia-nim-embeddings">NVIDIA NIM embeddings
</h2>
<p>The embedding layer gains a new provider: <code>EmbeddingNVIDIA</code>, which
connects to NVIDIA&rsquo;s OpenAI-compatible embedding API. The default model
is <code>nvidia/llama-nemotron-embed-1b-v2</code>, a compact embedding model
suitable for retrieval workloads. The provider reads its API key from
the <code>NVIDIA_API_KEY</code> environment variable.</p>
<div class="code-block"><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">raghilda.embedding</span> <span class="kn">import</span> <span class="n">EmbeddingNVIDIA</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">embedding</span> <span class="o">=</span> <span class="n">EmbeddingNVIDIA</span><span class="p">()</span></span></span></code></pre></div></div>
<p>One notable feature of NVIDIA&rsquo;s embedding API is differentiated input
types: queries and documents are embedded with different prefixes
(<code>&quot;query&quot;</code> and <code>&quot;passage&quot;</code>), which can improve retrieval quality for
asymmetric search where the query is short and the documents are long.
raghilda handles this distinction automatically when the store calls the
embedding provider during ingestion and retrieval.</p>
<p>The provider includes built-in rate limit handling with exponential
backoff. NVIDIA&rsquo;s 429 responses carry no <code>Retry-After</code> header or rate
limit metadata, so backoff is the only viable strategy. For users who
want to keep embedding computation entirely local, NVIDIA NIM can also
be self-hosted, in which case you point the provider at your local
endpoint with a <code>base_url=</code> override.</p>
<h2 id="improved-duckdb-error-messages">Improved DuckDB error messages
</h2>
<p>A smaller but practical improvement: <code>DuckDBStore</code> now raises clear,
actionable errors when BM25 retrieval is attempted before the index has
been built, or after writes have made the index stale. Previously, this
produced a cryptic <code>CatalogException</code> from DuckDB about a missing
<code>match_bm25</code> function. The new message tells you exactly what to do:</p>
<pre><code>RuntimeError: DuckDBStore retrieval requires a current BM25 index.
Call `store.build_index(&quot;bm25&quot;)` after inserting or updating documents
and before calling `retrieve_bm25()` or `retrieve()`.
</code></pre>
<p>The store now tracks BM25 freshness internally: the index is marked
stale after any <code>upsert()</code> call and marked current after
<code>build_index(&quot;bm25&quot;)</code>. This tracking happens off the retrieval hot path,
so there is no per-query overhead. HNSW indexes are unaffected because
DuckDB maintains them across writes automatically.</p>
<h2 id="why-even-use-raghilda">Why even use raghilda?
</h2>
<p>raghilda is a retrieval library, not an orchestration framework. Larger
projects like LangChain and LlamaIndex offer composable retrieval
components too, but they also ship agent runtimes, chain abstractions,
prompt management, and memory systems. If all you need is the retrieval
pipeline (crawl, chunk, embed, store, retrieve), raghilda gives you that
without the surrounding framework. The API surface is small: plain
dataclasses, iterators, and direct function calls. There are fewer
layers of indirection between your code and the underlying operations,
which makes the pipeline easier to debug and reason about.</p>
<p>raghilda <code>v0.2</code> makes that focused scope practical at scale. The crawl
API adds caching and concurrency while keeping each step a separate,
inspectable call. The storage layer lets you start with a local DuckDB
file and move to PostgreSQL or OpenAI Vector Stores later without
changing retrieval code. And every backend provides hybrid retrieval
(semantic search, BM25, and attribute filtering combined in a single
<code>retrieve()</code> call) out of the box, without assembling separate retriever
classes or configuring a pipeline graph.</p>
<h2 id="getting-started">Getting started
</h2>
<p>raghilda <code>v0.2</code> is available now on PyPI (<code>pip install raghilda</code>). The
<a href="https://posit-dev.github.io/raghilda/" target="_blank" rel="noopener">raghilda documentation site</a>
covers all of the features described here in more detail. The <a href="https://posit-dev.github.io/raghilda/user-guide/getting-started.html" target="_blank" rel="noopener">Getting
Started</a>
guide walks through building a store from scratch, and the <a href="https://posit-dev.github.io/raghilda/user-guide/crawling-and-ingestion.html" target="_blank" rel="noopener">Crawling and
Ingestion</a>
guide covers the new crawl API in depth. A dedicated
<a href="https://posit-dev.github.io/raghilda/user-guide/cloudflare-crawler.html" target="_blank" rel="noopener">CloudflareCrawler</a>
guide explains browser rendering, page discovery, caching, and
incremental updates. The <a href="https://github.com/posit-dev/raghilda" target="_blank" rel="noopener">GitHub
repository</a> has the source, issue
tracker, and full changelog. If you run into problems or have feature
requests, open an issue there.</p>
]]></description>
      <enclosure url="https://opensource.posit.co/blog/2026-07-01_raghilda-0-2-0/assets/raghilda-updated.png" length="1954675" type="image/png" />
    </item>
  </channel>
</rss>
