<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Zakir Hossen — Writing</title>
    <link>https://zakirhossen.com/writing/</link>
    <atom:link href="https://zakirhossen.com/writing/rss.xml" rel="self" type="application/rss+xml" />
    <description>Notes from building software solo: MCP servers, AI coding agents and the APIs behind them.</description>
    <language>en</language>
    <lastBuildDate>Wed, 07 Oct 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>Claude Code Review: The Four Built-In Options and How I Review AI-Written Code</title>
      <link>https://zakirhossen.com/writing/claude-code-review/</link>
      <guid isPermaLink="true">https://zakirhossen.com/writing/claude-code-review/</guid>
      <description>Claude Code reviews code four ways: /code-review, ultrareview, managed Code Review and GitHub Actions. What each one does, and the review rules I use.</description>
      <pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Zakir Hossen</dc:creator>
      <category>claude code</category>
      <category>code review</category>
      <category>ai coding agents</category>
      <content:encoded><![CDATA[<p>I build my products with Claude Code every day, and I ship alone. Nobody
else reads the diffs on my own products. So review is the step I cannot skip, and
it is also the step where an agent is most tempted to tell me what I want to
hear.</p>
<p>“Claude Code review” can mean two things: the review features built into
Claude Code, or how you review the code Claude Code writes. This post covers
both. First the four built-in options and what each one is for, checked
against Anthropic’s documentation in October 2026. Then the rules I follow,
which came from things that went wrong.</p>
<h2 id="the-four-built-in-options">The four built-in options</h2>
<table>
<thead>
<tr>
<th>Option</th>
<th>Where it runs</th>
<th>What it reviews</th>
<th>Who can use it</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>/code-review</code></td>
<td>Your machine, as a background subagent</td>
<td>Your current diff, a branch, a path or a PR</td>
<td>Any Claude Code user</td>
</tr>
<tr>
<td><code>/code-review ultra</code> (ultrareview)</td>
<td>Anthropic’s cloud, many agents</td>
<td>Your branch or a GitHub PR</td>
<td>claude.ai sign-in; research preview</td>
</tr>
<tr>
<td>Code Review (managed)</td>
<td>Anthropic’s cloud, via a GitHub App</td>
<td>PRs in the repos you pick, automatically or on request</td>
<td>Team and Enterprise; research preview</td>
</tr>
<tr>
<td>Claude Code GitHub Actions</td>
<td>Your own CI runners</td>
<td>Whatever your workflow tells it to</td>
<td>Anyone with an API key or subscription token</td>
</tr>
</tbody>
</table>
<p>Two related commands are worth knowing. <code>/security-review</code> checks the diff
between your branch and the default branch on <code>origin</code> for security problems
such as injection, auth issues and data exposure. <code>/simplify</code> looks for
cleanup only: reuse, simplification, efficiency, and whether the change sits
at the right level of abstraction. The docs are explicit that
<code>/simplify</code> does not look for correctness bugs. If you want bugs found, use
<code>/code-review</code>.</p>
<h3 id="1-code-review-in-your-terminal">1. /code-review, in your terminal</h3>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="text"><code><span class="line"><span>/code-review</span></span></code></pre>
<p>With no target, it reviews your branch’s commits ahead of its upstream plus
any uncommitted changes. You can also pass a PR number, a branch, a path or a
range such as <code>main...my-feature</code>. <code>/review</code> is an alias.</p>
<p>The parts that matter in practice:</p>
<ul>
<li><strong>Effort levels.</strong> <code>/code-review low</code> through <code>/code-review max</code>. At <code>low</code>
and <code>medium</code> it reports only the findings it is most sure of, so you see
fewer false alarms. <code>high</code> to <code>max</code> cover more ground and may include
findings it is less sure of. If you do not type a level, it reuses the last
one you typed, even from an earlier session.</li>
<li><strong>It runs in the background</strong> as a subagent with its own context window, so
it does not fill up the conversation you are working in.</li>
<li><strong><code>--fix</code> applies the findings</strong> to your working tree. When the review
runs in the background (the default), those edits happen outside your
session’s checkpoints, so <code>/rewind</code> will not undo them. Commit before you
run it, and use git to back out.</li>
<li><strong><code>--comment</code> posts the findings</strong> on a GitHub pull request as inline
comments.</li>
<li><strong>It follows your <code>CLAUDE.md</code>.</strong> It does not read <code>REVIEW.md</code> (more on that
below). If you want a local review to enforce a rule, the rule has to be in
<code>CLAUDE.md</code>.</li>
</ul>
<p>This is the one to run on every change before you commit. It needs no setup
and no GitHub App.</p>
<h3 id="2-ultrareview-in-the-cloud">2. Ultrareview, in the cloud</h3>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="text"><code><span class="line"><span>/code-review ultra</span></span>
<span class="line"><span>/code-review ultra 1234</span></span></code></pre>
<p>Ultrareview sends the review to a cloud sandbox where a larger group of
reviewer agents works on it in parallel. Anthropic’s docs say every finding
it reports is independently reproduced and verified, which is the main
difference from a local review. It reviews your branch against the default
branch, or a GitHub PR if you pass its number. Your terminal stays free while
it runs.</p>
<p>Limits worth knowing before you plan around it:</p>
<ul>
<li>It needs a claude.ai sign-in. It does not run on Amazon Bedrock, Google
Cloud’s Agent Platform or Microsoft Foundry, or for organisations with Zero
Data Retention.</li>
<li>It is a research preview. Pro and Max accounts get three free runs, once.
After that it uses paid usage credits.</li>
<li>Claude never starts it on its own. You have to type it.</li>
<li>For CI, there is a <code>claude ultrareview</code> subcommand that waits for the
findings and prints them.</li>
</ul>
<p>Use it for the changes where a missed bug is expensive: billing, auth,
migrations, anything that deletes data.</p>
<h3 id="3-code-review-the-managed-github-service">3. Code Review, the managed GitHub service</h3>
<p>This one is for teams. An organisation Owner installs the Claude GitHub App
and picks which repositories to review. From then on, reviews start when a PR
opens, on every push, or only when someone comments <code>@claude review</code>,
depending on the setting for each repository.</p>
<p>Several agents look at the diff and the code around it, a verification step
filters out false positives, and the results arrive as inline comments
tagged by severity:</p>
<ul>
<li><strong>Important</strong>: a bug that should be fixed before merging.</li>
<li><strong>Nit</strong>: worth fixing, not blocking.</li>
<li><strong>Pre-existing</strong>: a real bug that this PR did not introduce.</li>
</ul>
<p>It never approves or blocks a PR. The check run always finishes as neutral,
so if you want findings to block a merge, you read them in your own CI.
Anthropic’s docs put the average review at about 20 minutes and $15 to $25
in usage, billed separately from plan usage (October 2026). Reviewing on every
push multiplies that by the number of pushes, so manual mode with
<code>@claude review</code> on the PRs that matter is the cheaper setting.</p>
<h3 id="4-claude-code-github-actions">4. Claude Code GitHub Actions</h3>
<p>If you want Claude in your own CI instead of Anthropic’s service, run
<code>/install-github-app</code> inside Claude Code. It installs the GitHub App, stores
your key as a repository secret, and pushes a branch with the workflow file.
GitHub then opens with a pull request ready for you to create. Merge it, and
you can mention <code>@claude</code> in a PR or issue to get a review or a change.
This gives you the most control and the most setup.</p>
<h2 id="claudemd-or-reviewmd">CLAUDE.md or REVIEW.md</h2>
<p>The managed Code Review reads two files from your repository:</p>
<ul>
<li><strong><code>CLAUDE.md</code></strong> is general project guidance. Code Review treats a new
violation of it as a Nit, and it also flags a PR that makes a statement in
<code>CLAUDE.md</code> out of date.</li>
<li><strong><code>REVIEW.md</code></strong> is only for review. It goes straight to the agents that find
and verify issues, so rules there land more reliably than the same rules in
a long <code>CLAUDE.md</code>.</li>
</ul>
<p>Remember the catch from earlier: the local <code>/code-review</code> reads <code>CLAUDE.md</code>
and <strong>not</strong> <code>REVIEW.md</code>. If you use both, rules that must apply everywhere go
in <code>CLAUDE.md</code>.</p>
<p>Here is the kind of <code>REVIEW.md</code> I would write for a Laravel and Inertia app.
Most of the rules come from the instructions I keep for my own Laravel
projects:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="markdown"><code><span class="line"><span style="color:#79B8FF;font-weight:bold"># Review instructions</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## What Important means here</span></span>
<span class="line"><span style="color:#E1E4E8">Important = breaks production behaviour, leaks data, or cannot be rolled</span></span>
<span class="line"><span style="color:#E1E4E8">back. Style and naming are Nit at most. Report at most five Nits.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Always check</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Index, unique and foreign key names stay under 64 characters. SQLite</span></span>
<span class="line"><span style="color:#E1E4E8">  ignores the limit locally; MySQL rejects the migration in production.</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Controllers called by Inertia forms return a redirect, never</span></span>
<span class="line"><span style="color:#E1E4E8">  response()-&gt;json().</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Every query that reads customer data is scoped to the current team.</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> A test's assertions match its name. Read the assertion, not the title.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Do not report</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Anything CI already enforces: formatting, lint, type errors.</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Generated files and lockfiles.</span></span></code></pre>
<p>Keep it short. The docs warn that a long <code>REVIEW.md</code> dilutes the rules that
matter most.</p>
<h2 id="how-i-review-code-an-agent-wrote">How I review code an agent wrote</h2>
<p>The tools above find bugs in code. They do not decide what ships. These are
the rules I follow for that part.</p>
<h3 id="run-the-gate-before-any-review">Run the gate before any review</h3>
<p>Build, type check, tests. Do this before asking any reviewer, human or
agent, to read the diff. A review of code that does not compile wastes
everyone’s time, and a model will happily review it anyway.</p>
<h3 id="do-not-let-the-author-approve-its-own-work">Do not let the author approve its own work</h3>
<p>The session that wrote the code already believes it is right. In my
experience, it is more likely to say its own diff looks fine. So the review
should come from somewhere else. When I run an implementation plan, a fresh
reviewer agent, with its own context, checks each task before the next one
starts. A different model, such as Codex, goes one step further: different
training, different blind spots. I compared the two in
<a href="https://zakirhossen.com/writing/claude-code-vs-codex/">Claude Code vs Codex</a>.</p>
<h3 id="reproduce-every-finding-before-you-act-on-it">Reproduce every finding before you act on it</h3>
<p>A finding is a claim, not a fact. Before I change code because a reviewer
said so, I run the case that is supposed to fail. Some findings are real.
Some describe code paths that cannot happen.</p>
<p>The same rule applies to a reviewer saying a failure is “pre-existing”.
Prove it: run the same test on the commit before the change. If you cannot
show it failing there, it belongs to this change.</p>
<h3 id="read-what-a-test-asserts-not-what-it-is-called">Read what a test asserts, not what it is called</h3>
<p>This one hurt. On Schedule &amp; Chill, a test named “hands the real user model
to the tools” passed for the whole time the feature it described was broken.
It checked a different way of getting the user from the one the code
actually used, so it could not fail. The name said one thing and the
assertion tested another. The details are in
<a href="https://zakirhossen.com/writing/laravel-mcp/">my Laravel MCP post</a>.</p>
<p>When an agent writes both the code and the test, the test was written by
something that already believed the code was right. Read the assertion line.</p>
<h3 id="separate-can-hurt-money-or-customers-from-everything-else">Separate “can hurt money or customers” from everything else</h3>
<p>I sort each finding into two groups before fixing anything:</p>
<ol>
<li><strong>Blocks the merge.</strong> Wrong data, lost data, a broken payment, a security
hole, a user who cannot finish what they came to do.</li>
<li><strong>Follow-up.</strong> Naming, structure, a cleaner way to write it.</li>
</ol>
<p>Group 1 gets fixed now. Group 2 gets fixed if it is cheap, or written down if
it is not. Without this split, a review turns into twenty equal-looking
comments and the one that matters gets lost.</p>
<h3 id="check-facts-by-hand">Check facts by hand</h3>
<p>Code reviewers check code. They do not check whether a sentence on your
site is still true. A free tool page on JuggleHire told users to “export the
scorecard as a PDF”. The tool only offered copy and a <code>.txt</code> download. On a
pricing comparison page, five of eleven competitor prices had changed since
the page was written. Neither is a code problem, so a code review would not
catch either.</p>
<p>Prices, limits, feature claims and tutorial steps get checked against the
real source before they ship: the live pricing page, the real tool, the
current docs.</p>
<h3 id="let-a-human-own-the-merge">Let a human own the merge</h3>
<p>On my own products, I am that human. On a shared codebase where another
engineer reviews my PRs, the agent does the first pass and the human decides.
An agent never approves and never merges there. Having access to the repo is
not the same as having approval.</p>
<h2 id="which-one-to-use">Which one to use</h2>
<ul>
<li><strong>Every change, before you commit:</strong> <code>/code-review</code> at <code>medium</code>, after your
build and tests pass.</li>
<li><strong>A change that touches money, auth or data:</strong> add <code>/security-review</code> and,
if you have access, <code>/code-review ultra</code>.</li>
<li><strong>A team that already reviews on GitHub:</strong> the managed Code Review in manual
mode, with a short <code>REVIEW.md</code>.</li>
<li><strong>You want it in your own CI:</strong> GitHub Actions through <code>/install-github-app</code>.</li>
</ul>
<p>And whichever you pick, the last reviewer is you. The built-in reviewers are
good at “does this code do what it says”. Checking that what it says is true
is still your job.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Laravel MCP: What Three Production Servers Taught Me</title>
      <link>https://zakirhossen.com/writing/laravel-mcp/</link>
      <guid isPermaLink="true">https://zakirhossen.com/writing/laravel-mcp/</guid>
      <description>How to build an MCP server with the official laravel/mcp package, and six problems I had to fix on my three production Laravel MCP servers.</description>
      <pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate>
      <dc:creator>Zakir Hossen</dc:creator>
      <category>laravel</category>
      <category>mcp</category>
      <category>ai agents</category>
      <content:encoded><![CDATA[<p>I run three MCP servers built with <code>laravel/mcp</code>, the official Laravel
package. Each one sits inside a product I already had:</p>
<ul>
<li><strong>Schedule &amp; Chill</strong>, a social media scheduler. Its MCP server is public:
22 tools, reachable with an API token or over OAuth 2.1.</li>
<li><strong>ShipTell</strong>, a support and changelog tool. Its server covers changelogs,
the inbox, the knowledge base and the roadmap.</li>
<li><strong>JuggleHire</strong>, a hiring platform. A small internal server, read-only,
that my own agents use to look up customers.</li>
</ul>
<p>The package makes the first version easy. This post covers that part
quickly, then spends most of its time on six problems I had to fix after the
first version worked. They are not in the docs, and most of them are not
bugs in the package. They come from running a Laravel app for a caller that
is a model, not a person.</p>
<p>One warning first: <code>laravel/mcp</code> changes fast. It reached 1.0 in September
2026. My servers still run 0.6 and 0.9. The API changed between those two,
and it changed again in 1.0. The code below follows the
<a href="https://laravel.com/framework/docs/mcp">current Laravel docs</a>. Check the docs for
your version before copying anything.</p>
<h2 id="the-setup-in-four-commands">The setup in four commands</h2>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="bash"><code><span class="line"><span style="color:#B392F0">composer</span><span style="color:#9ECBFF"> require</span><span style="color:#9ECBFF"> laravel/mcp</span></span>
<span class="line"><span style="color:#B392F0">php</span><span style="color:#9ECBFF"> artisan</span><span style="color:#9ECBFF"> vendor:publish</span><span style="color:#79B8FF"> --tag=ai-routes</span></span>
<span class="line"><span style="color:#B392F0">php</span><span style="color:#9ECBFF"> artisan</span><span style="color:#9ECBFF"> make:mcp-server</span><span style="color:#9ECBFF"> SocialServer</span></span>
<span class="line"><span style="color:#B392F0">php</span><span style="color:#9ECBFF"> artisan</span><span style="color:#9ECBFF"> make:mcp-tool</span><span style="color:#9ECBFF"> GetAccounts</span></span></code></pre>
<p>The second command creates <code>routes/ai.php</code>. That is where servers are
registered, and there are two ways to do it:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="php"><code><span class="line"><span style="color:#6A737D">// routes/ai.php</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> App\Mcp\Servers\SocialServer</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Facades\Mcp</span><span style="color:#E1E4E8">;</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D">// Over HTTP, for remote clients. Protected like any other route.</span></span>
<span class="line"><span style="color:#79B8FF">Mcp</span><span style="color:#F97583">::</span><span style="color:#B392F0">web</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">'/mcp'</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">SocialServer</span><span style="color:#F97583">::class</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    -&gt;</span><span style="color:#B392F0">middleware</span><span style="color:#E1E4E8">([</span><span style="color:#9ECBFF">'auth:sanctum'</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">'throttle:mcp'</span><span style="color:#E1E4E8">]);</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D">// As an Artisan command (stdio), for a client on your own machine.</span></span>
<span class="line"><span style="color:#79B8FF">Mcp</span><span style="color:#F97583">::</span><span style="color:#B392F0">local</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">'social'</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">SocialServer</span><span style="color:#F97583">::class</span><span style="color:#E1E4E8">);</span></span></code></pre>
<p><code>throttle:mcp</code> needs a rate limiter called <code>mcp</code>. Define it in a service
provider with <code>RateLimiter::for('mcp', ...)</code>, like any other named limiter.</p>
<h2 id="a-tool-written-for-a-model">A tool, written for a model</h2>
<p>Here is a tool in the shape I use most. It lists the accounts a user can post
to, so the model fetches real IDs instead of making them up:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="php"><code><span class="line"><span style="color:#F97583">&lt;?</span><span style="color:#79B8FF">php</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">namespace</span><span style="color:#B392F0"> App\Mcp\Tools</span><span style="color:#E1E4E8">;</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Illuminate\Contracts\JsonSchema\JsonSchema</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Request</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Response</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Server\Attributes\Description</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Server\Attributes\Name</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Server\Tool</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Server\Tools\Annotations\IsReadOnly</span><span style="color:#E1E4E8">;</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">#[</span><span style="color:#79B8FF">Name</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">'get_accounts'</span><span style="color:#E1E4E8">)]</span></span>
<span class="line"><span style="color:#E1E4E8">#[</span><span style="color:#79B8FF">Description</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">'List the social accounts this user can post to. Call this before scheduling anything, and use only the ids it returns.'</span><span style="color:#E1E4E8">)]</span></span>
<span class="line"><span style="color:#E1E4E8">#[</span><span style="color:#79B8FF">IsReadOnly</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#F97583">class</span><span style="color:#B392F0"> GetAccounts</span><span style="color:#F97583"> extends</span><span style="color:#B392F0"> Tool</span></span>
<span class="line"><span style="color:#E1E4E8">{</span></span>
<span class="line"><span style="color:#F97583">    public</span><span style="color:#F97583"> function</span><span style="color:#B392F0"> schema</span><span style="color:#E1E4E8">(</span><span style="color:#79B8FF">JsonSchema</span><span style="color:#E1E4E8"> $schema)</span><span style="color:#F97583">:</span><span style="color:#F97583"> array</span></span>
<span class="line"><span style="color:#E1E4E8">    {</span></span>
<span class="line"><span style="color:#F97583">        return</span><span style="color:#E1E4E8"> [];</span></span>
<span class="line"><span style="color:#E1E4E8">    }</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">    public</span><span style="color:#F97583"> function</span><span style="color:#B392F0"> handle</span><span style="color:#E1E4E8">(</span><span style="color:#79B8FF">Request</span><span style="color:#E1E4E8"> $request)</span><span style="color:#F97583">:</span><span style="color:#79B8FF"> Response</span></span>
<span class="line"><span style="color:#E1E4E8">    {</span></span>
<span class="line"><span style="color:#E1E4E8">        $accounts </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> $request</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">user</span><span style="color:#E1E4E8">()</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">socialAccounts</span><span style="color:#E1E4E8">()</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">get</span><span style="color:#E1E4E8">();</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">        return</span><span style="color:#79B8FF"> Response</span><span style="color:#F97583">::</span><span style="color:#B392F0">structured</span><span style="color:#E1E4E8">([</span></span>
<span class="line"><span style="color:#9ECBFF">            'count'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $accounts</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">count</span><span style="color:#E1E4E8">(),</span></span>
<span class="line"><span style="color:#9ECBFF">            'accounts'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $accounts</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">map</span><span style="color:#E1E4E8">(</span><span style="color:#F97583">fn</span><span style="color:#E1E4E8"> ($account) =&gt; [</span></span>
<span class="line"><span style="color:#9ECBFF">                'id'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $account</span><span style="color:#F97583">-&gt;</span><span style="color:#E1E4E8">id,</span></span>
<span class="line"><span style="color:#9ECBFF">                'platform'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $account</span><span style="color:#F97583">-&gt;</span><span style="color:#E1E4E8">platform,</span></span>
<span class="line"><span style="color:#9ECBFF">                'name'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $account</span><span style="color:#F97583">-&gt;</span><span style="color:#E1E4E8">account_name,</span></span>
<span class="line"><span style="color:#9ECBFF">                'health'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $account</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">needsReconnect</span><span style="color:#E1E4E8">() </span><span style="color:#F97583">?</span><span style="color:#9ECBFF"> 'needs_reconnect'</span><span style="color:#F97583"> :</span><span style="color:#9ECBFF"> 'ok'</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">            ])</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">all</span><span style="color:#E1E4E8">(),</span></span>
<span class="line"><span style="color:#E1E4E8">        ]);</span></span>
<span class="line"><span style="color:#E1E4E8">    }</span></span>
<span class="line"><span style="color:#E1E4E8">}</span></span></code></pre>
<p>Three details in there matter more than they look:</p>
<ol>
<li><strong>The description tells the model what to do next.</strong> “Call this before
scheduling anything” is an instruction. The description is the main prompt
text you control in a tool, so use it.</li>
<li><strong><code>health</code>, not <code>is_active</code>.</strong> On Schedule &amp; Chill, <code>is_active</code> stayed
true for an account whose refresh token had died. Anything that trusted
it, a model included, would schedule posts into a channel that could not
publish. Give the model the field you would check yourself.</li>
<li><strong><code>#[IsReadOnly]</code></strong> tells the client the tool changes nothing. It is a
hint, not a guarantee, but it costs one line.</li>
</ol>
<p>Register the tool in the server’s <code>$tools</code> array. The server class also takes
an <code>#[Instructions('...')]</code> attribute. Mine says which tool to call first and
which actions cannot be undone. The client receives it when it connects.</p>
<h2 id="what-i-had-to-fix">What I had to fix</h2>
<h3 id="1-it-worked-locally-and-returned-404-in-production">1. It worked locally and returned 404 in production</h3>
<p>This one cost me several hours on JuggleHire.</p>
<p>The MCP server worked on my machine. In production, every request returned
404. The route did not exist.</p>
<p>The cause was <code>composer.json</code>. I had never required <code>laravel/mcp</code> myself. It
was installed only because Laravel Boost, a dev dependency, pulled it in.
Production installs with <code>composer install --no-dev</code>, so the package was not
there, its service provider never booted, and <code>routes/ai.php</code> was never
loaded.</p>
<p>The fix was one line, <code>composer require laravel/mcp</code>. The lesson is wider:</p>
<ul>
<li>Run <code>composer why laravel/mcp</code>. If the only answer is a <code>require-dev</code>
package, production will not have it.</li>
<li><strong>A 404 means the route is not registered. A 401 means it is.</strong> Hit the
endpoint without a token. If you get 404, stop debugging auth.</li>
</ul>
<h3 id="2-local-and-web-are-two-different-servers-in-practice">2. Local and web are two different servers in practice</h3>
<p><code>Mcp::local</code> and <code>Mcp::web</code> can point at the same class, but they do not run
in the same place. The local one is an Artisan command on your machine, so it
reads your local <code>.env</code> and your local database. The web one runs on your
server against production.</p>
<p>On JuggleHire that meant the local server answered with seed data while the
production endpoint was broken, and “it works for me” was true and
meaningless at the same time. Test the transport your users will use. If
you keep both, name them so nobody confuses them.</p>
<h3 id="3-mcp-routes-in-routeswebphp-need-a-csrf-exception">3. MCP routes in routes/web.php need a CSRF exception</h3>
<p>The package loads <code>routes/ai.php</code> without the <code>web</code> middleware group, so
there is no CSRF check on those routes. On Schedule &amp; Chill the MCP routes
live in <code>routes/web.php</code> instead. That puts them inside the <code>web</code> group, and
a POST from an AI client carries no CSRF token, so without an exception
Laravel rejects it with a 419.</p>
<p>If your MCP routes live in <code>routes/web.php</code>, exclude them:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="php"><code><span class="line"><span style="color:#6A737D">// bootstrap/app.php</span></span>
<span class="line"><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">withMiddleware</span><span style="color:#E1E4E8">(</span><span style="color:#F97583">function</span><span style="color:#E1E4E8"> (</span><span style="color:#79B8FF">Middleware</span><span style="color:#E1E4E8"> $middleware)</span><span style="color:#F97583">:</span><span style="color:#F97583"> void</span><span style="color:#E1E4E8"> {</span></span>
<span class="line"><span style="color:#E1E4E8">    $middleware</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">validateCsrfTokens</span><span style="color:#E1E4E8">(</span><span style="color:#B392F0">except</span><span style="color:#E1E4E8">: [</span></span>
<span class="line"><span style="color:#9ECBFF">        'mcp'</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#9ECBFF">        'mcp-oauth'</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#9ECBFF">        'oauth/register'</span><span style="color:#E1E4E8">, </span><span style="color:#6A737D">// dynamic client registration, from Mcp::oauthRoutes()</span></span>
<span class="line"><span style="color:#E1E4E8">    ]);</span></span>
<span class="line"><span style="color:#E1E4E8">})</span></span></code></pre>
<p>Or keep them in <code>routes/ai.php</code>, which is what it is for.</p>
<h3 id="4-with-oauth-the-tools-saw-a-different-user-than-the-middleware">4. With OAuth, the tools saw a different user than the middleware</h3>
<p>Schedule &amp; Chill serves the same server twice. <code>/mcp</code> takes a Sanctum token
that people paste into a config file. <code>/mcp-oauth</code> uses Passport, so that
claude.ai, Claude Desktop and other clients can offer a “Connect” button.
The Laravel docs recommend Passport for exactly this reason: OAuth 2.1 is
what the MCP spec documents. <code>Mcp::oauthRoutes()</code> registers the discovery
documents and client registration routes the clients probe for.</p>
<p>On that OAuth path, my setup needed a middleware that swapped the identity
model Passport resolved for the real <code>User</code>. It called
<code>$request-&gt;setUserResolver()</code>, and the test for it passed.</p>
<p>Every tool call still failed. In <code>laravel/mcp</code>, <code>Request::user()</code> resolves the
user through the <strong>auth manager</strong>, not through the HTTP request. The <code>api</code>
guard still held the old identity model, so the first relation call died. The
fix was to also call <code>Auth::guard('api')-&gt;setUser($user)</code>.</p>
<p>The test passed because it asserted <code>request()-&gt;user()</code>, which the tools
never call. Two rules came out of it:</p>
<ul>
<li>When a framework has two ways to get the current user, test the one the
caller actually uses.</li>
<li><strong>A successful handshake proves nothing.</strong> <code>initialize</code> returned 200, the
consent screen rendered, tokens were issued. Only a real tool call that
touches a relation shows whether anything behind the door works.</li>
</ul>
<h3 id="5-the-401-had-no-www-authenticate-header">5. The 401 had no WWW-Authenticate header</h3>
<p>The MCP authorization spec gives a server two ways to point a client to its
OAuth metadata: a <code>WWW-Authenticate</code> header on the 401, or a well-known URL
the client has to try on its own. Clients read the header first. Without
it, they have to guess, and some just report that the server does not
support OAuth.</p>
<p><code>Mcp::web()</code> adds a middleware for that header, <code>AddWwwAuthenticateHeader</code>.
But Laravel’s middleware priority list runs authentication earlier, so the
401 was rendered before that middleware could add the header. Clients got a
bare 401.</p>
<p><code>laravel/mcp</code> 1.0 fixes this inside the package. On 0.9 and earlier, move
the middleware ahead of authentication yourself:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="php"><code><span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Illuminate\Contracts\Auth\Middleware\AuthenticatesRequests</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">use</span><span style="color:#79B8FF"> Laravel\Mcp\Server\Middleware\AddWwwAuthenticateHeader</span><span style="color:#E1E4E8">;</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">$middleware</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">prependToPriorityList</span><span style="color:#E1E4E8">(</span></span>
<span class="line"><span style="color:#B392F0">    before</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">AuthenticatesRequests</span><span style="color:#F97583">::class</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#B392F0">    prepend</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">AddWwwAuthenticateHeader</span><span style="color:#F97583">::class</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">);</span></span></code></pre>
<p>Use the <code>AuthenticatesRequests</code> <strong>interface</strong> as the anchor, not the
<code>Authenticate</code> class. The priority list holds the interface. Naming the class
matches nothing, and the middleware is quietly added to the end of the list,
which looks like the fix and changes nothing. Check with <code>curl -i</code> that the
header is really there.</p>
<h3 id="6-a-capped-list-looked-like-the-whole-list">6. A capped list looked like the whole list</h3>
<p>JuggleHire’s internal server has a tool that lists customers. It returned at
most 50 rows and did not say so, and production has far more customers than
that. An agent that counted the rows got 50, and nothing in the response told
it the real number was much higher. I ended up writing a warning into my own
notes: never quote that number as a total. The fix belongs in the tool, not
in my notes.</p>
<p>A model cannot know a list was cut unless you tell it. Any tool that limits
results should return the total and say the list is partial:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="php"><code><span class="line"><span style="color:#F97583">return</span><span style="color:#79B8FF"> Response</span><span style="color:#F97583">::</span><span style="color:#B392F0">structured</span><span style="color:#E1E4E8">([</span></span>
<span class="line"><span style="color:#9ECBFF">    'total'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $total,</span></span>
<span class="line"><span style="color:#9ECBFF">    'returned'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $customers</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">count</span><span style="color:#E1E4E8">(),</span></span>
<span class="line"><span style="color:#9ECBFF">    'truncated'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $total </span><span style="color:#F97583">&gt;</span><span style="color:#E1E4E8"> $customers</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">count</span><span style="color:#E1E4E8">(),</span></span>
<span class="line"><span style="color:#9ECBFF">    'customers'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> $customers</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">all</span><span style="color:#E1E4E8">(),</span></span>
<span class="line"><span style="color:#E1E4E8">]);</span></span></code></pre>
<p>The same rule applies to anything the server decides for the model. On
Schedule &amp; Chill, scheduling tools return the time we understood in UTC, in
local time and with the timezone name, so an agent can check our reading of
its input. It cannot look at a calendar to confirm.</p>
<h2 id="errors-are-instructions">Errors are instructions</h2>
<p>The Laravel docs say it directly: on validation failure, AI clients act on
the error messages you give them. So write messages a model can act on:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="php"><code><span class="line"><span style="color:#E1E4E8">$request</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">validate</span><span style="color:#E1E4E8">([</span></span>
<span class="line"><span style="color:#9ECBFF">    'scheduled_at'</span><span style="color:#F97583"> =&gt;</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">'required'</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">'date'</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">'after:now'</span><span style="color:#E1E4E8">],</span></span>
<span class="line"><span style="color:#E1E4E8">], [</span></span>
<span class="line"><span style="color:#9ECBFF">    'scheduled_at.after'</span><span style="color:#F97583"> =&gt;</span><span style="color:#9ECBFF"> 'scheduled_at must be in the future. Send an ISO-8601 time after the current time, with a timezone offset.'</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">]);</span></span></code></pre>
<p>On Schedule &amp; Chill, the tools that act on posts and uploads return
errors as a JSON object with <code>code</code>, <code>problem</code>, <code>cause</code> and <code>fix</code>. The <code>fix</code>
field is the one that matters. A model that reads it can correct itself
without asking the human.</p>
<p>One more thing about errors: the error path must never throw. Error messages
often include outside text, such as an account name or a platform’s error
body. On Schedule &amp; Chill one invalid UTF-8 byte in that text made
<code>json_encode</code> throw, and the whole tool call failed instead of returning the
error. Clean outside strings before you encode them.</p>
<h2 id="testing-it">Testing it</h2>
<p>The package has two built-in ways to test, and both are worth using.</p>
<p><strong>Unit tests</strong> call a tool directly on the server class:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="php"><code><span class="line"><span style="color:#B392F0">it</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">'lists the user</span><span style="color:#79B8FF">\'</span><span style="color:#9ECBFF">s accounts'</span><span style="color:#E1E4E8">, </span><span style="color:#F97583">function</span><span style="color:#E1E4E8"> () {</span></span>
<span class="line"><span style="color:#E1E4E8">    $user </span><span style="color:#F97583">=</span><span style="color:#79B8FF"> User</span><span style="color:#F97583">::</span><span style="color:#B392F0">factory</span><span style="color:#E1E4E8">()</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">create</span><span style="color:#E1E4E8">();</span></span>
<span class="line"><span style="color:#79B8FF">    SocialAccount</span><span style="color:#F97583">::</span><span style="color:#B392F0">factory</span><span style="color:#E1E4E8">()</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">for</span><span style="color:#E1E4E8">($user)</span><span style="color:#F97583">-&gt;</span><span style="color:#B392F0">create</span><span style="color:#E1E4E8">([</span><span style="color:#9ECBFF">'account_name'</span><span style="color:#F97583"> =&gt;</span><span style="color:#9ECBFF"> 'Acme on LinkedIn'</span><span style="color:#E1E4E8">]);</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF">    SocialServer</span><span style="color:#F97583">::</span><span style="color:#B392F0">actingAs</span><span style="color:#E1E4E8">($user)</span></span>
<span class="line"><span style="color:#F97583">        -&gt;</span><span style="color:#B392F0">tool</span><span style="color:#E1E4E8">(</span><span style="color:#79B8FF">GetAccounts</span><span style="color:#F97583">::class</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">        -&gt;</span><span style="color:#B392F0">assertOk</span><span style="color:#E1E4E8">()</span></span>
<span class="line"><span style="color:#F97583">        -&gt;</span><span style="color:#B392F0">assertSee</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">'Acme on LinkedIn'</span><span style="color:#E1E4E8">);</span></span>
<span class="line"><span style="color:#E1E4E8">});</span></span></code></pre>
<p><strong>The MCP Inspector</strong> connects to a running server so you can call tools by
hand: <code>php artisan mcp:inspector mcp</code> for the web server at <code>/mcp</code>, or
<code>php artisan mcp:inspector social</code> for the local one.</p>
<p>Then do the test that finds the real problems. Connect an actual client and
give it a plain-language task that needs three or four tool calls, without
naming the tools. In Claude Code:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="bash"><code><span class="line"><span style="color:#B392F0">claude</span><span style="color:#9ECBFF"> mcp</span><span style="color:#9ECBFF"> add</span><span style="color:#79B8FF"> --transport</span><span style="color:#9ECBFF"> http</span><span style="color:#9ECBFF"> social</span><span style="color:#9ECBFF"> https://your-app.test/mcp</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#79B8FF">  --header</span><span style="color:#9ECBFF"> "Authorization: Bearer YOUR_TOKEN"</span></span></code></pre>
<p>Watch which tool it calls first, which argument it guesses, and which
description sends it the wrong way. That found more problems in my servers
than all the schema work did.</p>
<h2 id="who-this-is-not-for">Who this is not for</h2>
<p>If your code decides which endpoint to call and the model only writes text,
you do not need an MCP server. A normal API call is simpler. I wrote about
where that line sits in <a href="https://zakirhossen.com/writing/mcp-vs-api/">MCP vs API</a>. For the
framework-neutral version of this post, with the TypeScript SDK and the
mistakes every MCP server makes, see
<a href="https://zakirhossen.com/writing/how-to-build-an-mcp-server/">how to build an MCP server</a>.</p>
<p>If you are already on Laravel and your users work in Claude, Cursor or
another MCP client, the package is the shortest path I know. You are not
building a second app. You are exposing the one you have, with its policies,
validation and services intact. The work is in what this post covers: making
it behave when the caller is a model.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Claude Code vs Codex: What Actually Changed After Six Months of Both</title>
      <link>https://zakirhossen.com/writing/claude-code-vs-codex/</link>
      <guid isPermaLink="true">https://zakirhossen.com/writing/claude-code-vs-codex/</guid>
      <description>I use Claude Code and Codex daily on production Laravel and TypeScript work. Where each one wins, where benchmarks stop helping, and my workflow.</description>
      <pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Zakir Hossen</dc:creator>
      <category>claude code</category>
      <category>codex</category>
      <category>ai coding agents</category>
      <content:encoded><![CDATA[<p>I ship software alone. There is no team to catch my mistakes, so the tools that
write code with me get judged on one thing: how much of what they produce
survives to production without me rewriting it.</p>
<p>I have run both Claude Code and OpenAI’s Codex CLI daily for about six months
across Laravel, Inertia/React and TypeScript codebases. This is what I actually
found, including the parts where the tool I use more lost.</p>
<h2 id="the-short-version">The short version</h2>
<table>
<thead>
<tr>
<th></th>
<th>Claude Code</th>
<th>Codex CLI</th>
</tr>
</thead>
<tbody>
<tr>
<td>Where it fits</td>
<td>Long, multi-file work you want to steer</td>
<td>Bounded tasks you want done and returned</td>
</tr>
<tr>
<td>Interaction</td>
<td>Interactive, shows reasoning, stops to ask</td>
<td>More autonomous, sandboxed, review at the end</td>
</tr>
<tr>
<td>Token appetite</td>
<td>Higher</td>
<td>Noticeably lower for the same task</td>
</tr>
<tr>
<td>Best single use</td>
<td>Building a feature across many files</td>
<td>Reviewing a diff someone else wrote</td>
</tr>
<tr>
<td>Where it frustrates me</td>
<td>Cost, and it will over-build if you let it</td>
<td>Less steerable once it has started</td>
</tr>
</tbody>
</table>
<p>If you only take one thing: <strong>they are not really competing for the same job.</strong>
I stopped choosing between them around month three and started using both,
which I explain at the end.</p>
<h2 id="be-careful-with-the-benchmark-numbers">Be careful with the benchmark numbers</h2>
<p>Search for this comparison and you will find SWE-bench Verified scores quoted
to one decimal place, usually both above 96%. I am not going to repeat those as
if I verified them, because I did not, and neither did most of the posts
quoting them.</p>
<p>Two things are worth knowing about that:</p>
<ol>
<li><strong>The gap between the top agents on those benchmarks is now smaller than
the gap between a good prompt and a lazy one.</strong> When two tools are within a
point of each other, the benchmark has stopped being the deciding factor for
your work.</li>
<li><strong>Benchmarks measure “did the test pass”, not “would you merge this”.</strong>
Those are different questions, and the second one is the one that costs me
time.</li>
</ol>
<p>So the rest of this is what I observed, framed as observation rather than
measurement. Where I say “faster” or “cheaper” I mean in my work, on my
codebases, which is a sample of one.</p>
<h2 id="where-claude-code-wins-for-me">Where Claude Code wins for me</h2>
<p><strong>Multi-file work where the plan matters more than the code.</strong> The thing I do
most is change a feature that touches a controller, a React page, a migration,
a config file and three tests. Claude Code holds that shape in its head. It
reads the surrounding code and matches it — my naming, my comment density, the
way I structure a Laravel domain folder — rather than writing generically
correct code that reads like it came from somewhere else.</p>
<p><strong>It tells you when it disagrees.</strong> More than once it has stopped and said the
approach I asked for would break something else, and it was right. That is
worth real money when nobody else is reviewing your work.</p>
<p><strong>Skills, subagents and hooks.</strong> This is the part that changed how I work, not
just how fast I type. I keep a set of skill files that encode things I have
learned the hard way — how my deploy pipeline works, which database gotchas
bite me when SQLite passes locally and MySQL fails in production, what my
commit messages should look like. Claude Code loads the relevant one on its
own. Codex has no real equivalent that I have found.</p>
<p><strong>Worktrees.</strong> I often run several sessions on the same repository. Claude Code
handles git worktrees well enough that I can genuinely parallelise, which was
not true six months ago.</p>
<h2 id="where-codex-wins-for-me">Where Codex wins for me</h2>
<p><strong>It uses far fewer tokens for the same result.</strong> This is the difference I feel
most in the bill. On a bounded task — “add validation to this endpoint and
update the test” — Codex gets there with a fraction of the back-and-forth.
Claude Code will read more files than it strictly needed to.</p>
<p><strong>Review is its best mode.</strong> This surprised me. Handing Codex a diff and asking
what is wrong with it produces tighter, less agreeable feedback than asking
Claude Code the same thing. Claude Code is more likely to tell me the diff
looks good. Codex is more likely to find the thing I missed.</p>
<p><strong>Sandboxed autonomy.</strong> When I genuinely do not want to watch — a mechanical
refactor across forty files, a dependency bump — letting it run and reviewing
at the end is the right shape, and Codex is built for that shape.</p>
<h2 id="where-both-of-them-still-fail">Where both of them still fail</h2>
<p>Neither tool checks whether what it wrote is <em>true</em>.</p>
<p>This is the failure that has cost me the most. A tool page on one of my
products had a step in its instructions saying “export the result as a PDF”.
There was no PDF export. The code offered clipboard copy and a <code>.txt</code>
download, and the prose next to it had been confidently wrong for months. Both
agents had read that page. Neither flagged it, because prose that contradicts
the component beside it is not a syntax error.</p>
<p>The same class of thing bit me on competitor pricing: five of eleven prices on
a comparison page had drifted since they were written, and no amount of code
review catches a number that is simply out of date.</p>
<p><strong>So the rule I ended up with: agents are good at “does this run” and bad at
“is this still true”.</strong> Anything factual — a price, a limit, a capability
claim, a step in a tutorial — gets verified by me against the actual thing,
every time. I now write tests that fail when a claim goes stale, which is the
only version of this that survives me forgetting.</p>
<p>The second shared failure is over-building. Ask for a fix, get a fix plus a
refactor plus three new tests plus a README section nobody asked for. Claude
Code does this more than Codex does. The prompt that works is boring and
specific: say what to change, say what not to touch.</p>
<h2 id="what-i-actually-do-now">What I actually do now</h2>
<p>I stopped picking one:</p>
<ol>
<li><strong>Claude Code writes the feature.</strong> It gets the multi-file work, the
planning, and anything where matching the existing codebase matters.</li>
<li><strong>Codex reviews the diff before it merges.</strong> Different model, different
training, genuinely different opinions — it catches things the author
missed, in a way that asking the author to re-read its own work does not.
The full routine, including Claude Code’s own review commands, is in
<a href="https://zakirhossen.com/writing/claude-code-review/">Claude Code review</a>.</li>
<li><strong>I verify every factual claim myself.</strong> Prices, limits, capabilities,
tutorial steps. Neither agent does this and both will state a wrong thing
confidently.</li>
</ol>
<p>That third step is not optional and it is the one people skip.</p>
<h2 id="which-should-you-pick-if-you-can-only-pick-one">Which should you pick, if you can only pick one</h2>
<ul>
<li><strong>Building something new, or working alone with nobody reviewing you</strong> —
Claude Code. The steerability and the skills system compound over months in a
way that raw speed does not.</li>
<li><strong>Working in a team with review already in place, or cost-sensitive</strong> —
Codex. You already have humans catching mistakes; you want throughput, and it
is cheaper per task.</li>
<li><strong>Mostly reviewing rather than writing</strong> — Codex, without much hesitation.</li>
</ul>
<p>And if the honest answer is that you have not tried either seriously for more
than a week: pick either one and use it on real work for a month. The
difference between the two is much smaller than the difference between using
one properly and using it as autocomplete.</p>
<hr>
<p><em>I build <a href="https://zakirhossen.com/projects/">software products</a> solo under Lomeyo, LLC — a hiring
platform, a social scheduler, and a few smaller things. Everything above comes
out of that work. Three of them have MCP servers built in Laravel, and
<a href="https://zakirhossen.com/writing/laravel-mcp/">here is what building those taught me</a>.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>How to Get a Grok API Key and Build a Bot With It</title>
      <link>https://zakirhossen.com/writing/grok-api-key/</link>
      <guid isPermaLink="true">https://zakirhossen.com/writing/grok-api-key/</guid>
      <description>The console path to an xAI API key, why the OpenAI SDK works unchanged, and a working Telegram bot — plus the billing detail that catches people out.</description>
      <pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Zakir Hossen</dc:creator>
      <category>grok</category>
      <category>xai</category>
      <category>api</category>
      <category>bots</category>
      <content:encoded><![CDATA[<p>Getting a Grok API key takes about three minutes. Understanding what you are
being billed for takes slightly longer, and that is the part people get wrong.</p>
<p>This covers both, then builds a small bot that actually runs.</p>
<h2 id="getting-the-key">Getting the key</h2>
<ol>
<li>Go to <strong>console.x.ai</strong>. This is the developer console, and it is separate
from the Grok consumer app — you do <strong>not</strong> need an X Premium subscription
to use the API.</li>
<li>Sign up or sign in. There is a short onboarding asking what you are building
and asking you to accept the terms.</li>
<li>Open <strong>API Keys</strong> in the sidebar, then <strong>Create API Key</strong>.</li>
<li>Copy it immediately. Like most providers, the full key is shown once.</li>
</ol>
<p>That is the whole flow. If you are being asked to pay for X Premium, you are in
the consumer product, not the developer console.</p>
<h3 id="before-you-write-any-code-check-two-things-in-the-console">Before you write any code, check two things in the console</h3>
<ul>
<li><strong>Your credit balance and what it is.</strong> xAI runs promotional credits for new
accounts, and has at times offered additional monthly credits in exchange for
opting into a data-sharing programme. <strong>Both the amounts and the terms have
changed more than once</strong> — read what your own console says rather than
trusting any blog post, including this one. There is no permanent free tier;
after credits, it is pay per token.</li>
<li><strong>Whether the data-sharing option is on.</strong> If extra credits are attached to
sharing your API traffic for training, that is a real trade. For a hobby bot
it is probably fine. For anything touching customer data it is a decision you
should make deliberately, not one you should discover later.</li>
</ul>
<h2 id="store-it-properly">Store it properly</h2>
<p>The single most common mistake with a fresh API key is committing it.</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="bash"><code><span class="line"><span style="color:#6A737D"># .env  — and make sure .env is in .gitignore</span></span>
<span class="line"><span style="color:#E1E4E8">XAI_API_KEY</span><span style="color:#F97583">=</span><span style="color:#9ECBFF">xai-...</span></span></code></pre>
<p>If you have already pushed one to a public repository, rotate it in the console
rather than deleting the commit. Scrapers find keys in public git history
within minutes, and the commit is not the only copy.</p>
<h2 id="the-useful-part-it-speaks-openai">The useful part: it speaks OpenAI</h2>
<p>The thing that makes the Grok API quick to adopt is that it is
OpenAI-compatible. You do not need a new SDK. You point the OpenAI client at a
different base URL:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#F97583">from</span><span style="color:#E1E4E8"> openai </span><span style="color:#F97583">import</span><span style="color:#E1E4E8"> OpenAI</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">client </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> OpenAI(</span></span>
<span class="line"><span style="color:#FFAB70">    api_key</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">os.environ[</span><span style="color:#9ECBFF">"XAI_API_KEY"</span><span style="color:#E1E4E8">],</span></span>
<span class="line"><span style="color:#FFAB70">    base_url</span><span style="color:#F97583">=</span><span style="color:#9ECBFF">"https://api.x.ai/v1"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">response </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> client.chat.completions.create(</span></span>
<span class="line"><span style="color:#FFAB70">    model</span><span style="color:#F97583">=</span><span style="color:#9ECBFF">"grok-4"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#FFAB70">    messages</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">[</span></span>
<span class="line"><span style="color:#E1E4E8">        {</span><span style="color:#9ECBFF">"role"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"system"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"content"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"You are terse. Answer in one sentence."</span><span style="color:#E1E4E8">},</span></span>
<span class="line"><span style="color:#E1E4E8">        {</span><span style="color:#9ECBFF">"role"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"user"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"content"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"Why is the sky blue?"</span><span style="color:#E1E4E8">},</span></span>
<span class="line"><span style="color:#E1E4E8">    ],</span></span>
<span class="line"><span style="color:#E1E4E8">)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF">print</span><span style="color:#E1E4E8">(response.choices[</span><span style="color:#79B8FF">0</span><span style="color:#E1E4E8">].message.content)</span></span></code></pre>
<p>Node is the same idea:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="js"><code><span class="line"><span style="color:#F97583">import</span><span style="color:#E1E4E8"> OpenAI </span><span style="color:#F97583">from</span><span style="color:#9ECBFF"> "openai"</span><span style="color:#E1E4E8">;</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">const</span><span style="color:#79B8FF"> client</span><span style="color:#F97583"> =</span><span style="color:#F97583"> new</span><span style="color:#B392F0"> OpenAI</span><span style="color:#E1E4E8">({</span></span>
<span class="line"><span style="color:#E1E4E8">  apiKey: process.env.</span><span style="color:#79B8FF">XAI_API_KEY</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">  baseURL: </span><span style="color:#9ECBFF">"https://api.x.ai/v1"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">});</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">const</span><span style="color:#79B8FF"> res</span><span style="color:#F97583"> =</span><span style="color:#F97583"> await</span><span style="color:#E1E4E8"> client.chat.completions.</span><span style="color:#B392F0">create</span><span style="color:#E1E4E8">({</span></span>
<span class="line"><span style="color:#E1E4E8">  model: </span><span style="color:#9ECBFF">"grok-4"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">  messages: [{ role: </span><span style="color:#9ECBFF">"user"</span><span style="color:#E1E4E8">, content: </span><span style="color:#9ECBFF">"Give me one fact about Bangladesh."</span><span style="color:#E1E4E8"> }],</span></span>
<span class="line"><span style="color:#E1E4E8">});</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">console.</span><span style="color:#B392F0">log</span><span style="color:#E1E4E8">(res.choices[</span><span style="color:#79B8FF">0</span><span style="color:#E1E4E8">].message.content);</span></span></code></pre>
<p>Two caveats on that compatibility:</p>
<ul>
<li><strong>Model names are xAI’s, not OpenAI’s.</strong> Check the console or the models
endpoint for what is currently available rather than guessing a version
string. Model IDs change.</li>
<li><strong>Compatible does not mean identical.</strong> The common path — chat completions,
streaming, tool calling — works. Provider-specific extras on either side do
not always map. If something behaves oddly, that seam is the first place to
look.</li>
</ul>
<h2 id="building-a-bot">Building a bot</h2>
<p>The most common thing people want a Grok key for is a bot. Here is a Telegram
one that works, in about forty lines.</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#F97583">import</span><span style="color:#E1E4E8"> os</span></span>
<span class="line"><span style="color:#F97583">import</span><span style="color:#E1E4E8"> httpx</span></span>
<span class="line"><span style="color:#F97583">from</span><span style="color:#E1E4E8"> openai </span><span style="color:#F97583">import</span><span style="color:#E1E4E8"> OpenAI</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF">XAI</span><span style="color:#F97583"> =</span><span style="color:#E1E4E8"> OpenAI(</span><span style="color:#FFAB70">api_key</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">os.environ[</span><span style="color:#9ECBFF">"XAI_API_KEY"</span><span style="color:#E1E4E8">], </span><span style="color:#FFAB70">base_url</span><span style="color:#F97583">=</span><span style="color:#9ECBFF">"https://api.x.ai/v1"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#79B8FF">TG</span><span style="color:#F97583"> =</span><span style="color:#F97583"> f</span><span style="color:#9ECBFF">"https://api.telegram.org/bot</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">os.environ[</span><span style="color:#9ECBFF">'TELEGRAM_TOKEN'</span><span style="color:#E1E4E8">]</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF">SYSTEM</span><span style="color:#F97583"> =</span><span style="color:#9ECBFF"> "You are a helpful assistant in a group chat. Keep replies under 60 words."</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> reply</span><span style="color:#E1E4E8">(chat_id: </span><span style="color:#79B8FF">int</span><span style="color:#E1E4E8">, text: </span><span style="color:#79B8FF">str</span><span style="color:#E1E4E8">) -&gt; </span><span style="color:#79B8FF">None</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">    httpx.post(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{TG}</span><span style="color:#9ECBFF">/sendMessage"</span><span style="color:#E1E4E8">, </span><span style="color:#FFAB70">json</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">{</span><span style="color:#9ECBFF">"chat_id"</span><span style="color:#E1E4E8">: chat_id, </span><span style="color:#9ECBFF">"text"</span><span style="color:#E1E4E8">: text})</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> answer</span><span style="color:#E1E4E8">(prompt: </span><span style="color:#79B8FF">str</span><span style="color:#E1E4E8">) -&gt; </span><span style="color:#79B8FF">str</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">    res </span><span style="color:#F97583">=</span><span style="color:#79B8FF"> XAI</span><span style="color:#E1E4E8">.chat.completions.create(</span></span>
<span class="line"><span style="color:#FFAB70">        model</span><span style="color:#F97583">=</span><span style="color:#9ECBFF">"grok-4"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#FFAB70">        messages</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">[</span></span>
<span class="line"><span style="color:#E1E4E8">            {</span><span style="color:#9ECBFF">"role"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"system"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"content"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">SYSTEM</span><span style="color:#E1E4E8">},</span></span>
<span class="line"><span style="color:#E1E4E8">            {</span><span style="color:#9ECBFF">"role"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"user"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"content"</span><span style="color:#E1E4E8">: prompt},</span></span>
<span class="line"><span style="color:#E1E4E8">        ],</span></span>
<span class="line"><span style="color:#E1E4E8">    )</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#E1E4E8"> res.choices[</span><span style="color:#79B8FF">0</span><span style="color:#E1E4E8">].message.content</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> main</span><span style="color:#E1E4E8">() -&gt; </span><span style="color:#79B8FF">None</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">    offset </span><span style="color:#F97583">=</span><span style="color:#79B8FF"> 0</span></span>
<span class="line"><span style="color:#F97583">    while</span><span style="color:#79B8FF"> True</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#6A737D">        # long-poll Telegram; 30s timeout keeps this cheap</span></span>
<span class="line"><span style="color:#E1E4E8">        r </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> httpx.get(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{TG}</span><span style="color:#9ECBFF">/getUpdates"</span><span style="color:#E1E4E8">, </span><span style="color:#FFAB70">params</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">{</span><span style="color:#9ECBFF">"offset"</span><span style="color:#E1E4E8">: offset, </span><span style="color:#9ECBFF">"timeout"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">30</span><span style="color:#E1E4E8">}, </span><span style="color:#FFAB70">timeout</span><span style="color:#F97583">=</span><span style="color:#79B8FF">35</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> update </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> r.json().get(</span><span style="color:#9ECBFF">"result"</span><span style="color:#E1E4E8">, []):</span></span>
<span class="line"><span style="color:#E1E4E8">            offset </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> update[</span><span style="color:#9ECBFF">"update_id"</span><span style="color:#E1E4E8">] </span><span style="color:#F97583">+</span><span style="color:#79B8FF"> 1</span></span>
<span class="line"><span style="color:#E1E4E8">            message </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> update.get(</span><span style="color:#9ECBFF">"message"</span><span style="color:#E1E4E8">, {})</span></span>
<span class="line"><span style="color:#E1E4E8">            text </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> message.get(</span><span style="color:#9ECBFF">"text"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> text:</span></span>
<span class="line"><span style="color:#F97583">                continue</span></span>
<span class="line"><span style="color:#E1E4E8">            reply(message[</span><span style="color:#9ECBFF">"chat"</span><span style="color:#E1E4E8">][</span><span style="color:#9ECBFF">"id"</span><span style="color:#E1E4E8">], answer(text))</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">if</span><span style="color:#79B8FF"> __name__</span><span style="color:#F97583"> ==</span><span style="color:#9ECBFF"> "__main__"</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">    main()</span></span></code></pre>
<p>That is deliberately the simplest thing that works. Three changes turn it into
something you would leave running:</p>
<ol>
<li><strong>Keep conversation history per chat</strong>, capped. Without it every message is
context-free. With it uncapped, your token bill grows with the conversation
forever — keep the last N turns, not all of them.</li>
<li><strong>Rate limit per user.</strong> A bot in an open group is a bot anyone can spend
your credits on. This is the failure mode that turns a hobby project into a
bill.</li>
<li><strong>Handle errors.</strong> Wrap the API call. A 429 or a 500 should mean “try again
shortly”, not a crashed process and a silent bot.</li>
</ol>
<h2 id="the-billing-detail-that-catches-people">The billing detail that catches people</h2>
<p>You are billed per token, input and output, and <strong>input includes everything you
send</strong> — the system prompt, the conversation history, and any documents you
paste in. A long system prompt is not free; you pay for it on every single
call.</p>
<p>For a bot, the practical consequences are:</p>
<ul>
<li>A 500-word system prompt costs you on every message, forever. Keep it short.</li>
<li>Unbounded chat history is the most common source of a surprising bill.</li>
<li>Set a spend alert in the console on day one, not after the first bad week.</li>
</ul>
<h2 id="when-grok-is-the-right-pick">When Grok is the right pick</h2>
<p>Honestly: the API is quick to adopt because it is OpenAI-compatible, and its
distinctive asset is real-time access to what is happening on X. If your
product needs current, public conversation — a monitoring tool, a trend bot,
anything where “what are people saying right now” is the question — that is a
genuine differentiator.</p>
<p>If you need general reasoning or coding, the OpenAI-compatible base URL means
you can benchmark it against alternatives by changing one line. Do that with
your own prompts on your own task rather than trusting anyone’s benchmark
table, including mine.</p>
<p>If your bot needs to take actions and not only answer, that is where tool
calling and MCP come in. I wrote about
<a href="https://zakirhossen.com/writing/mcp-vs-api/">when an MCP server helps and when a plain API call is better</a>.</p>
<hr>
<p><em>I write about the tools I actually ship with — see
<a href="https://zakirhossen.com/writing/">more writing</a>, or <a href="https://zakirhossen.com/projects/">what I am building</a>.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>How to Build an MCP Server (And the Five Things That Broke Mine)</title>
      <link>https://zakirhossen.com/writing/how-to-build-an-mcp-server/</link>
      <guid isPermaLink="true">https://zakirhossen.com/writing/how-to-build-an-mcp-server/</guid>
      <description>Build a production MCP server: tool design, transports, OAuth, and the five failures that only show up once a real model starts calling it.</description>
      <pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Zakir Hossen</dc:creator>
      <category>mcp</category>
      <category>tutorial</category>
      <category>ai agents</category>
      <category>laravel</category>
      <content:encoded><![CDATA[<p>Building an MCP server that connects is easy. There is a quickstart in every
SDK and it works in about ten minutes.</p>
<p>Building one a model uses <em>correctly</em> took me considerably longer, and the gap
between those two things is what this post is about. I run MCP servers on four
products now. Everything below is a mistake I made and had to fix.</p>
<h2 id="the-ten-minute-version">The ten-minute version</h2>
<p>Every MCP server is the same three things:</p>
<ol>
<li><strong>A list of tools</strong>, each with a name, a description and a JSON schema for
its arguments.</li>
<li><strong>A handler per tool</strong> that does the work and returns a result.</li>
<li><strong>A transport</strong> — stdio for something running on the user’s machine, HTTP
for something running on yours.</li>
</ol>
<p>Here is the smallest useful shape, in TypeScript, using v2 of the official
SDK (<code>@modelcontextprotocol/server</code>), the stable line since July 2026:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="ts"><code><span class="line"><span style="color:#F97583">import</span><span style="color:#E1E4E8"> { McpServer } </span><span style="color:#F97583">from</span><span style="color:#9ECBFF"> "@modelcontextprotocol/server"</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">import</span><span style="color:#E1E4E8"> { StdioServerTransport } </span><span style="color:#F97583">from</span><span style="color:#9ECBFF"> "@modelcontextprotocol/server/stdio"</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#F97583">import</span><span style="color:#79B8FF"> *</span><span style="color:#F97583"> as</span><span style="color:#E1E4E8"> z </span><span style="color:#F97583">from</span><span style="color:#9ECBFF"> "zod/v4"</span><span style="color:#E1E4E8">;</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">const</span><span style="color:#79B8FF"> server</span><span style="color:#F97583"> =</span><span style="color:#F97583"> new</span><span style="color:#B392F0"> McpServer</span><span style="color:#E1E4E8">({ name: </span><span style="color:#9ECBFF">"my-server"</span><span style="color:#E1E4E8">, version: </span><span style="color:#9ECBFF">"1.0.0"</span><span style="color:#E1E4E8"> });</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">server.</span><span style="color:#B392F0">registerTool</span><span style="color:#E1E4E8">(</span></span>
<span class="line"><span style="color:#9ECBFF">  "list_accounts"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">  {</span></span>
<span class="line"><span style="color:#E1E4E8">    description:</span></span>
<span class="line"><span style="color:#9ECBFF">      "List the social accounts this user can post to. Call this before scheduling anything so you use real account ids."</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">    inputSchema: z.</span><span style="color:#B392F0">object</span><span style="color:#E1E4E8">({</span></span>
<span class="line"><span style="color:#E1E4E8">      platform: z.</span><span style="color:#B392F0">string</span><span style="color:#E1E4E8">().</span><span style="color:#B392F0">optional</span><span style="color:#E1E4E8">().</span><span style="color:#B392F0">describe</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">"Only accounts on this platform, e.g. linkedin"</span><span style="color:#E1E4E8">),</span></span>
<span class="line"><span style="color:#E1E4E8">    }),</span></span>
<span class="line"><span style="color:#E1E4E8">  },</span></span>
<span class="line"><span style="color:#F97583">  async</span><span style="color:#E1E4E8"> ({ </span><span style="color:#FFAB70">platform</span><span style="color:#E1E4E8"> }) </span><span style="color:#F97583">=&gt;</span><span style="color:#E1E4E8"> {</span></span>
<span class="line"><span style="color:#F97583">    const</span><span style="color:#79B8FF"> accounts</span><span style="color:#F97583"> =</span><span style="color:#F97583"> await</span><span style="color:#E1E4E8"> db.accounts.</span><span style="color:#B392F0">findMany</span><span style="color:#E1E4E8">({ where: platform </span><span style="color:#F97583">?</span><span style="color:#E1E4E8"> { platform } </span><span style="color:#F97583">:</span><span style="color:#E1E4E8"> {} });</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#E1E4E8"> {</span></span>
<span class="line"><span style="color:#E1E4E8">      content: [{ type: </span><span style="color:#9ECBFF">"text"</span><span style="color:#E1E4E8">, text: </span><span style="color:#79B8FF">JSON</span><span style="color:#E1E4E8">.</span><span style="color:#B392F0">stringify</span><span style="color:#E1E4E8">(accounts) }],</span></span>
<span class="line"><span style="color:#E1E4E8">    };</span></span>
<span class="line"><span style="color:#E1E4E8">  },</span></span>
<span class="line"><span style="color:#E1E4E8">);</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">await</span><span style="color:#E1E4E8"> server.</span><span style="color:#B392F0">connect</span><span style="color:#E1E4E8">(</span><span style="color:#F97583">new</span><span style="color:#B392F0"> StdioServerTransport</span><span style="color:#E1E4E8">());</span></span></code></pre>
<p>The SDK turns the Zod schema into the JSON Schema the model sees, and it
rejects arguments that do not match before your handler runs. Two notes from
the SDK docs: over stdio, stdout is the protocol channel, so log with
<code>console.error</code>, never <code>console.log</code>. And if you are still on v1
(<code>@modelcontextprotocol/sdk</code>), the imports and some method names differ, so
follow the v1 docs rather than this snippet.</p>
<p>That runs. Now here is what goes wrong.</p>
<h2 id="1-the-model-invents-ids">1. The model invents IDs</h2>
<p>This was the first thing to break and it is the one I would warn everyone
about.</p>
<p>My scheduling tool took <code>social_account_ids</code>. A developer using the REST
version fetches the accounts once and hardcodes them. A model does not. Given a
tool that wants account IDs, it will confidently supply account IDs that do not
exist. The schema validates. The call either fails oddly or hits the wrong
account.</p>
<p><strong>The fix is a tool, not documentation.</strong> Add <code>list_accounts</code> — or whatever the
equivalent is in your domain — and say in the <em>description</em> of every tool that
consumes an ID that the model should call it first. Descriptions are prompt
text. They are the only place you get to instruct the model.</p>
<h2 id="2-error-messages-written-for-humans">2. Error messages written for humans</h2>
<p><code>422 Unprocessable Entity</code> is fine for a developer with docs open. To a model
it is a dead end, and a model at a dead end guesses.</p>
<p>Compare:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="plaintext"><code><span class="line"><span>// Useless to a model</span></span>
<span class="line"><span>{ "error": "Validation failed", "code": 422 }</span></span>
<span class="line"><span></span></span>
<span class="line"><span>// Useful</span></span>
<span class="line"><span>{ "error": "scheduled_at must be in the future. You sent 2026-01-01T10:00:00Z,</span></span>
<span class="line"><span>   which is in the past. Current time is 2026-09-09T14:22:00Z. Send a timestamp</span></span>
<span class="line"><span>   after that." }</span></span></code></pre>
<p>The second one tells the model what to do next. Every error your MCP server
returns should read like an instruction, and should include the value that was
wrong and what a correct one looks like.</p>
<h2 id="3-silently-dropped-arguments">3. Silently dropped arguments</h2>
<p>This one cost me a real production mistake, so it gets its own section.</p>
<p>Most MCP servers — mine included, at first — accept an options object and
ignore keys they do not recognise. The call returns success. The feature you
thought you enabled was never enabled, because you guessed the key name and the
server threw it away without saying anything.</p>
<p>If a caller passes <code>platform_options: { threadReply: true }</code> and your server
expects <code>thread_reply</code>, and you silently drop unknown keys, the post publishes
with the default and reports success. Nothing in the logs says otherwise.</p>
<p><strong>Fix it in two places.</strong> Reject unknown keys loudly, and ship a
<code>get_capabilities</code> tool that returns the exact option keys each surface
accepts, so the model can ask instead of guessing. Guessing is the default
behaviour of every model and you cannot prompt it away.</p>
<h2 id="4-publishing-is-irreversible-and-models-are-not-careful">4. Publishing is irreversible and models are not careful</h2>
<p>Anything your server can do that cannot be undone needs a safe default.</p>
<p>For a scheduler that means: the natural tool is <code>schedule_post</code> with a future
timestamp, not <code>publish_now</code>. The review window costs nothing and cancel undoes
anything wrong. <code>publish_now</code> exists but it is the explicit choice, not the
path of least resistance. The whole flow from the user’s side, including who
holds the OAuth tokens and how to stop a bad post, is in
<a href="https://schedulenchill.com/blog/can-ai-agents-post-on-social-media">this walkthrough of AI agents posting to social media</a>
on the Schedule &amp; Chill blog.</p>
<p>Apply this generally. Ask which of your tools are irreversible, then make the
reversible version the obvious one — in the tool name, in the description, and
in what happens when a required field is missing.</p>
<h2 id="5-remote-mcp-means-oauth-and-oauth-is-most-of-the-work">5. Remote MCP means OAuth, and OAuth is most of the work</h2>
<p>stdio is easy: the server runs on the user’s machine as their user, so
authentication is “it already is them”.</p>
<p>Remote is a different product. Once your MCP server runs on your infrastructure
and serves many users over HTTP, you need:</p>
<ul>
<li>A real OAuth flow, including the consent screen and the callback.</li>
<li>Token storage and refresh, per user, that survives your deploys.</li>
<li>Scoping, so a token cannot reach another tenant’s data.</li>
<li>A way for the user to revoke it.</li>
</ul>
<p>None of that is MCP-specific and all of it is required. Budget for it. If you
already have OAuth for your product, you are most of the way there; if you do
not, the MCP server is not the small project it looked like.</p>
<h2 id="what-laravel-developers-should-know">What Laravel developers should know</h2>
<p>If you are on Laravel — I mostly am — the ecosystem got much better. The
official Laravel MCP package lets you register a server as a route and define
tools as classes, which means your existing authorisation, validation and
service layer apply unchanged. That is the main advantage: you are not building
a parallel application, you are exposing the one you have.</p>
<p>The trap is the same as everywhere else. Your form requests were written to
return validation errors to a human looking at a form. Read them again as if a
model is the reader.</p>
<p>I wrote the full Laravel version separately:
<a href="https://zakirhossen.com/writing/laravel-mcp/">Laravel MCP: what three production servers taught me</a>
covers the package setup and six failures that only showed up in production.</p>
<h2 id="the-test-that-actually-tells-you-it-works">The test that actually tells you it works</h2>
<p>Not a unit test. Connect a real client — Claude Desktop or Claude Code — and
give it a task in plain language that requires three or four tool calls in
sequence, without telling it which tools to use.</p>
<p>Watch what it calls. You will find:</p>
<ul>
<li>The tool it reached for first, which is rarely the one you expected.</li>
<li>The argument it guessed instead of fetching.</li>
<li>The description that was ambiguous enough to send it the wrong way.</li>
</ul>
<p>That fifteen-minute exercise found more real problems in my servers than any
amount of schema work. Do it before you ship, then again after every tool you
add.</p>
<h2 id="the-short-checklist">The short checklist</h2>
<ul>
<li>One tool that lists the real IDs, and descriptions telling the model to call
it first.</li>
<li>Errors that say what to do next, with the bad value and a correct example.</li>
<li>Unknown arguments rejected loudly, plus a <code>get_capabilities</code> tool.</li>
<li>The reversible action is the obvious one; irreversible actions are explicit.</li>
<li>Remote means OAuth, scoping and revocation — plan for it as its own project.</li>
<li>Test by watching a real model use it cold.</li>
</ul>
<hr>
<p><em>More on what this protocol is and is not:
<a href="https://zakirhossen.com/writing/mcp-vs-api/">MCP vs API</a> — why an MCP server sits on top of your API
rather than replacing it, and when building one is a waste of time.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>MCP vs API: What the Difference Actually Is, From Someone Who Shipped Both</title>
      <link>https://zakirhossen.com/writing/mcp-vs-api/</link>
      <guid isPermaLink="true">https://zakirhossen.com/writing/mcp-vs-api/</guid>
      <description>MCP does not replace your API. It sits on top of one. What changes when you add an MCP server, what does not, and when building one is a waste of time.</description>
      <pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Zakir Hossen</dc:creator>
      <category>mcp</category>
      <category>api</category>
      <category>ai agents</category>
      <content:encoded><![CDATA[<p>I have shipped MCP servers on four of my own products. Every one of them sits
on top of a REST API I had already built. Not one of them replaced it.</p>
<p>That is the whole answer, and most explanations of this bury it. But it leaves
the useful question open: if MCP does not replace your API, what does it
actually change, and is it worth building?</p>
<h2 id="the-one-sentence-version">The one-sentence version</h2>
<p><strong>An API is an interface for a developer who reads documentation. MCP is an
interface for a model that cannot.</strong></p>
<p>Same functions underneath. Different consumer, and the consumer is what changes
the design.</p>
<h2 id="what-is-actually-different">What is actually different</h2>
<p>Put the two side by side and the differences are narrow but they matter.</p>
<table>
<thead>
<tr>
<th></th>
<th>REST API</th>
<th>MCP server</th>
</tr>
</thead>
<tbody>
<tr>
<td>Who reads the interface</td>
<td>A developer, once, at build time</td>
<td>A model, every session, at runtime</td>
</tr>
<tr>
<td>How capabilities are found</td>
<td>You read docs and write code</td>
<td>The client asks the server what it can do</td>
</tr>
<tr>
<td>When the tool list is fixed</td>
<td>Design time — you ship code</td>
<td>Runtime — the server answers on connect</td>
</tr>
<tr>
<td>What a bad call costs</td>
<td>A 400 you see in your logs</td>
<td>A model that quietly does the wrong thing</td>
</tr>
<tr>
<td>Error messages written for</td>
<td>A developer debugging</td>
<td>A model deciding what to do next</td>
</tr>
<tr>
<td>Transport</td>
<td>HTTP, your choice of shape</td>
<td>JSON-RPC over stdio or HTTP</td>
</tr>
</tbody>
</table>
<p>The row that matters most is the third one. On the HN thread about this,
someone reduced the whole protocol to: <em>tools can be added at runtime instead
of at design time</em>. That is closer to the truth than most long-form
explanations, mine included.</p>
<h2 id="the-part-that-surprised-me-building-one">The part that surprised me building one</h2>
<p>I assumed an MCP server was a thin wrapper over my existing controllers. Take
the endpoint, describe it in a schema, done.</p>
<p>That is technically true and produces a bad server.</p>
<p>Here is why. My scheduler’s REST API has an endpoint that takes
<code>social_account_ids</code>. A developer integrating it reads the docs, calls the
accounts endpoint once, hardcodes the IDs, and never thinks about it again.</p>
<p>A model does not do that. Given a tool that takes account IDs, a model will
<strong>invent</strong> account IDs. They look plausible. The call succeeds against the
schema and posts to the wrong channel, or fails with an error the model then
tries to “fix” by inventing different IDs.</p>
<p>So the MCP server needed things the REST API never did:</p>
<ul>
<li><strong>A tool whose only job is “tell me what accounts exist”</strong>, so the model
fetches real IDs instead of guessing. On REST this is a documentation
problem. On MCP it is a tool.</li>
<li><strong>Errors written as instructions, not as status.</strong> <code>422 Unprocessable Entity</code> tells a developer to go read the docs. It tells a model nothing. The
message has to say what to do instead.</li>
<li><strong>A read-before-write tool.</strong> The model needs to see the existing queue
before adding to it, or it duplicates work a human already did.</li>
<li><strong>A safe default that is not “publish”.</strong> Publishing is irreversible.
Scheduling is not. Every agent-created post gets a future timestamp so a
human has a window to cancel.</li>
</ul>
<p>None of that is protocol. It is interface design for a consumer that reasons
instead of reading. That is the actual work, and it is why “just wrap your API”
produces something that technically connects and practically misbehaves.</p>
<p>If you build on Laravel, I wrote up
<a href="https://zakirhossen.com/writing/laravel-mcp/">what broke on my Laravel MCP servers</a>, including the
failures that only showed up in production.</p>
<h2 id="what-mcp-does-not-give-you">What MCP does not give you</h2>
<p>Worth saying plainly, because the marketing around this is loose:</p>
<ul>
<li><strong>It is not faster.</strong> It is another layer over the same HTTP call.</li>
<li><strong>It does not remove the need for an API.</strong> You still built one. MCP calls it.</li>
<li><strong>It does not make a model reliable.</strong> It gives the model a well-described
door. It will still walk through the wrong one sometimes.</li>
<li><strong>It is not required for an AI feature.</strong> If your code decides what to call
and the model only writes text, you want a normal API call, not MCP.</li>
</ul>
<h2 id="when-to-build-one-and-when-not-to">When to build one, and when not to</h2>
<p>Build an MCP server when <strong>the model decides which action to take</strong>. That is
the whole test.</p>
<p>Concretely, it is worth it when:</p>
<ul>
<li>Your users already work inside an MCP-capable client — Claude Desktop, Claude
Code, Cursor and similar — and want your product reachable from there without
switching windows.</li>
<li>The set of useful actions is larger than one, and which one to take depends
on context the model has and your code does not.</li>
<li>Your API is stable. An MCP server over an API you are still redesigning means
maintaining two moving interfaces.</li>
</ul>
<p>Skip it when:</p>
<ul>
<li>Your code picks the endpoint and the model only generates content. Use the
API. Adding MCP here is architecture for its own sake.</li>
<li>You have one action. A single-tool MCP server is a lot of protocol for one
HTTP call.</li>
<li>You do not have an API yet. Build that first. The MCP server is a projection
of it, and projecting something that does not exist yet means designing both
at once, badly.</li>
</ul>
<h2 id="the-honest-summary">The honest summary</h2>
<p>MCP is not an API replacement and treating it as one is how people end up
disappointed. It is a second interface, aimed at a different consumer, over
work you already did.</p>
<p>The value is not the protocol — it is that the runtime tool discovery lets a
model use your product without anyone writing an integration first. Whether
that is worth building depends entirely on whether models are actually the ones
deciding what to call in your product.</p>
<p>For me, on a scheduler where the whole point is that an AI tool writes and
queues the post, it clearly was. On a hiring platform where a human clicks the
buttons, it clearly was not. JuggleHire has no MCP server for its customers.
The only one there is a small, read-only server my own agents use to look up
customer records, and it passes the same test: a model decides what to call.</p>
<p>If posting is your use case too, the Schedule &amp; Chill blog walks through
<a href="https://schedulenchill.com/blog/can-ai-agents-post-on-social-media">how an AI agent gets a post onto a social account</a>,
over MCP and over a REST API.</p>
<hr>
<p><em>If you want the practical side, I wrote up
<a href="https://zakirhossen.com/writing/how-to-build-an-mcp-server/">how to build an MCP server</a> — the
structure, the mistakes, and the parts the tutorials skip.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
