“Does this code follow good patterns?” is a question every tech lead asks and almost nobody answers well. Manual architecture reviews are slow, and asking an AI assistant to “review the architecture” usually gives you a generic list of GoF pattern names sprinkled over your code.
So I wrote an AI skill that does it properly:
dev-pattern-compliance-audit. It audits a codebase (or one module) against the Gang of Four design patterns and
Martin Fowler’s Patterns of Enterprise Application Architecture (PEAA), and produces a report you can actually act on.
What is an agent skill?
A skill is a folder with a SKILL.md file: frontmatter plus instructions that an AI coding agent loads when the task
matches its description. It can also bundle reference documents and helper scripts. The skill works with agents
such as OpenCode and Claude Code, and the repository is a small catalog you can install skills from.
Instead of re-explaining “how I want an architecture review done” in every prompt, you encode the method once.
The core idea: compliance is not “more patterns”
The most important part of the skill is not a checklist, it is a stance. A pattern is an answer to a specific force. Applied without that force, it is over-engineering. Fowler himself says a Domain Model is not always the right choice, and a Service Layer is unnecessary for a single-interface app.
So every candidate goes through a fit test:
- Force: is the problem the pattern solves actually present (a type code branching that keeps growing, the same query duplicated, two clients needing the same operation)?
- Evidence: can you point at code (
path:line) that shows it? - Cost/benefit: is the fix proportionate to the size of the code and how often it changes?
Each finding gets a verdict: well applied, misapplied, missing, over-engineered, or anti-pattern. A report that says “mostly fine, here are 4 things worth fixing and 6 patterns that are well applied” is more useful, and more trusted, than a 60-item wish list.
How the audit works
The skill follows a ten-step workflow:
- Scope: detect the stack, read
AGENTS.md,CLAUDE.md, ADRs. A documented deliberate deviation is not a defect. - Scan for leads: a Python script (
scan_candidates.py) does a fast regex-based inventory: Singleton shapes,instanceof/switchchains, anemic entities,newof collaborators inside services, layering leaks, god classes, single-implementation interfaces, float money, and more. Its output is leads, not findings. - Map the architecture top-down (PEAA first): domain logic style (Transaction Script, Table Module, Domain Model), data source style (Active Record, Data Mapper, Repository), distribution, concurrency, and base patterns. This is decided per entity, not once for the whole codebase.
- API contract check: more on this below.
- Credit what the framework gives you: Spring Data is a Repository, JPA is Data Mapper plus Unit of Work,
@Transactionalis a proxy, and so on. Reporting “missing Repository” in a Spring Data project is a false positive. - Evaluate at class level (GoF): classify each candidate, matching the language’s idiom. In modern Java, sealed interfaces and pattern-matching
switchoften replace Visitor, and lambdas replace many Strategy classes. - Cross-checks: catch defects where a pattern is applied here but skipped there, or where code disagrees with the schema, tests, or its own comments.
- Verify, rate, deduplicate: severity (High/Medium/Low) plus confidence (Confirmed vs Likely). Ten
instanceofchains are one missing-polymorphism finding with ten locations. - Write the report: saved as
docs/pattern-compliance-<date>.md. - Offer hand-off: only for findings you accept, as an OpenSpec change or plain ticket list. The audit itself is read-only.
No hallucinated line numbers
The most common way an AI-generated review goes wrong is citing UserService.java:142 when the file has 90 lines.
One bad citation makes the reader doubt everything else.
The skill ships check_citations.py, which fails on any path:line reference past the end of the file. It cannot tell
whether the line says what the report claims, so the skill still requires re-reading the code, but it catches the
cheapest class of errors mechanically.
OpenAPI as the API contract
Version 1.2 added first-class OpenAPI/Swagger support. When a spec exists, it is treated as the process-boundary contract: PEAA’s Data Transfer Object plus Remote Facade. The skill then checks:
- Source of truth: contract-first (committed spec drives codegen) or code-first (Springdoc, Nest Swagger, FastAPI export)? Two sources with no single published artifact is a finding.
- Drift: paths in code missing from the spec, diverging fields, documented status codes that are never returned, loose
additionalPropertieshiding a real schema. - Generated clients: if a TypeScript frontend calls an API that has a spec, the default correct shape is a generated client (orval, openapi-typescript, kubb, and similar), wired into
package.jsonscripts and CI. A hand-maintainedapi.tsnext to a spec is a Medium finding. - Leaking entities: persistence entities used directly as API schemas.
There are exceptions too. It will not demand OpenAPI plus codegen for an internal module with a single client.
What the report looks like
Every finding includes the catalog and pattern name, verdict, severity, confidence, path:line evidence, why it matters
for this code, and a proportionate recommendation with the smallest first step. Two sections keep it honest:
- Well applied: so the team does not “fix” working design.
- Deliberately not recommended: patterns considered and rejected because the force is absent, which pre-empts “why didn’t you suggest X?”.
For big repositories the skill has an execution budget: read all entry points and 3-5 vertical slices, sample the rest, cap the main list at roughly 15 findings, and say clearly what was read and what was sampled.
Try it
The skill lives in the catalog at github.com/inver/ai-skills. Add the catalog with the agent-sync CLI:
npx skillfish add https://github.com/inver/ai-skills
Or copy skills/dev-pattern-compliance-audit/ into your agent’s skills directory, then ask something like:
- “Audit this module for GoF and PEAA pattern compliance.”
- “Is our OpenAPI spec really the source of truth for the frontend client?”
- “Which patterns are we using wrong?”
The scanner is strongest on Java/Kotlin/Spring, good on TypeScript/Nest/Prisma/React and Python/FastAPI/Django, and best-effort elsewhere. The report says so in its method section.
The skill is Apache 2.0 licensed. Issues and pull requests are welcome, especially new framework mappings and scanner heuristics for stacks I do not use daily.