Case Study
Evaluating and Deploying a Third-Party Claude Plugin Suite
The problem
A developer at an outside SEO agency built four Claude Code plugins (SEO, blog, ads, and an anti-slop editing pass) and shared them with us as a favor. Ownership described them to our marketing lead as a full "marketing suite."
Three things were wrong with that picture, and nobody had checked:
- The plugins had never been reviewed for security or licensing before being handed to a business user.
- They were built for Claude Code, Anthropic's agent environment with sub-agents, hooks and network access. When run from the plain chat interface, roughly 60 to 70 percent of that functionality is inert. Getting real value out of them meant putting the marketing lead on Claude Code, not the chat box.
- Two weeks earlier, the marketing lead had asked in writing for an application: a shared dashboard with budget, analytics, a calendar and post scheduling. The plugins are skills. Against her eight stated needs, one was covered well, one partially, six not at all.
Left alone, this would have been an underwhelming handoff that made a genuinely good tool look like a broken promise.
What I built
Evaluation (one day)
- Read the plugin source. Pinned dependencies with vulnerability notes, security test fixtures, hash-verified installs, no phone-home behavior. Verdict: well built, and built by someone who takes supply chain seriously.
- Caught the license: internal use by the agency only. Cleared org distribution directly with the developer before anything shipped.
- Mapped the platform gap and the expectation gap: what the plugins do against what the marketing lead asked for.
- Wrote ownership a plain-language memo before the handoff: give marketing the SEO scan and an ads audit now, treat the dashboard as a separate build I'd own, and learn from the developer which skills carry the value.
Deployment (about ten days)
- Fixed package upload failures on all four plugins.
- Installed the suite on our office Mac and ran live audits against the company website with the blog and SEO plugins to confirm real output before anyone else touched them.
- Set up Google Analytics and Search Console credentials, so the SEO plugin pulls live search data instead of guessing.
- Installed the three working plugins on my own Windows machine first, then on the marketing lead's laptop. That meant getting each executable through the company's application allowlisting tool, coordinating with IT, and staging fixed packages on a USB drive.
- Found and fixed four platform-specific bugs (below), reported all of them to the developer, and kept the local patches tracked so they drop out when upstream fixes land.
- Set the marketing lead up in the Claude Code desktop app, so she runs the plugins in the environment they were built for without ever opening a terminal.
- Wrote a setup guide and a user manual for her.
- Designed the workflow with a human gate: analytics find underperforming pages, the plugins draft the fix, the marketing lead reviews and publishes. Nothing auto-publishes. The ads plugin was set to draft-only.
- Locked down the organization's plugin catalog so staff only see what's been reviewed.
What was mine
The evaluation, the memo to ownership, every install and workaround, the documentation, and the workflow design were mine. Specifically:
- The decision to stop the handoff and re-set expectations first. The easy move was to install the plugins and let the marketing lead find out. I chose to tell ownership what the tools were and weren't before she touched them.
- Not shipping a disabled safeguard. The ads installer refused to run because a security check failed (details below). I could have switched the check off and moved on. I patched it locally to keep working, kept the original, refused to put that patch on anyone else's machine, and got the real fix from the developer.
- The review gate. The developer's loop ran analytics to fix to publish. I made it end at a draft for human review. Marketing content goes out under the company's name; an agent doesn't get to publish it.
- Draft-only on ads. Ad spend is money. The plugin proposes; a person commits.
- Tracking local patches against upstream instead of forking, so we stay on the developer's releases.
What came from others: the plugins and the install path were the developer's work. Ownership set a $2,000 test ad budget. IT handled the allowlisting approvals.
Problems worth knowing about
- A failed integrity check on the ads installer. The installer verifies that the dependency list it's about to install from is the exact file the developer security-reviewed, by comparing a fingerprint (a SHA-256 hash). The pinned fingerprint didn't match the file in the package. Two explanations: someone tampered with the list, or the developer regenerated it and forgot to update the constant. I checked everything I could from inside the package: the list correctly described every file it covered, and all 106 package hashes matched. That pointed to a stale number, not tampering, but it can't be proven from the outside, because a tampered list would look equally consistent. So I patched the constant to keep working on my machine, backed up the original, sent the developer both fingerprints, and held the Windows install until he answered. He confirmed it was stale, shipped fixed packages, and added a test so it can't recur.
- An installer looking in the wrong folder. The SEO installer expected its launcher in
bin/and the package shipped it inscripts/. A one-line copy, but the Windows installer did the same thing, so it went in the report to the developer. - Python version pinning. The office Mac had Python 3.14 as default; the plugins needed 3.12. Fixed by putting 3.12 first on the path for that one install command, without changing anything else on the machine.
- Three Windows-specific bugs nobody had hit, because nobody had run the Windows path end to end:
- The SEO plugin had no Windows launcher at all. Its launcher was a bash script copied as-is. I wrote a
.cmdequivalent. - The blog installer picked up Windows' Microsoft Store Python stub instead of real Python, so every call failed silently and the dependency step skipped. I installed the dependencies into their own environment under the plugin folder, which also kept them inside the paths our security software allows.
- The ads lock file was one character short:
fonttoolswhere a downstream package requiredfonttools[woff]. In hash-verified install mode that's a hard failure. Fixed the pin, no hash or version changes, and flagged it as a likely typo.
- The SEO plugin had no Windows launcher at all. Its launcher was a bash script copied as-is. I wrote a
- Application allowlisting. Company laptops block any executable that hasn't been individually approved. A plugin suite with several binaries meant a per-file approval loop with IT before anything ran. Slow, but the right control to have in place.
- A coverage gap the vendor didn't know about. The ads plugin supports eleven ad platforms. It doesn't cover Google Local Services Ads, which is where our phone leads actually come from. Flagged to the developer.
Results
Not measured: content output or ad performance. The ads plugin's developer token was still pending approval at the time of writing. The anti-slop plugin has no Windows installer and was deferred.
What I'd do differently
- Ask the requester what they need before evaluating what was offered. I found the marketing lead's written request after I'd started reviewing the plugins. Reading it first would have reframed the whole evaluation on day one.
- Separate credentials for the marketing lead. She runs on a Google service account shared with me. It was faster, and it's wrong: access should trace to one person, and I'd give her her own from the start.