AI & Engineering Judgment / A developer’s perspective
Your AI Agent Sounds Right. Do You Know Why?
The explanation is clear. The diff is small. The tests are green. You click Accept. If someone asks which requirement those tests proved, can you answer?
The patch comes with a story you want to believe
Imagine reviewing a report-download feature in a multi-tenant application. An agent spots a tenant comparison inside an access policy and proposes removing it. Its explanation is tidy: the request middleware already selects the tenant, the permission check remains, and two existing tests pass. One condition disappears. The code looks simpler.
“The tenant check is redundant because middleware already establishes the tenant. I kept the download permission check, and all existing tests pass.”
This is an illustrative proposal, not a recorded agent response or a discovered Semitexa vulnerability. It is useful because every part sounds plausible. Middleware can establish context. Duplicate checks can create maintenance work. Passing tests are valuable. The missing step is showing that those facts justify this particular deletion.
A tenant context identifies the execution. A report has its own tenant identity. To remove the comparison, you need evidence that every caller supplying a report has already enforced the relationship between the two. The agent’s explanation has turned that obligation into an assumption.
When an answer replaces your judgment
Steven D. Shaw and Gideon Nave use cognitive surrender to describe adopting AI output with minimal scrutiny in their research paper Thinking—Fast, Slow, and Artificial. Their Wharton interview discusses experiments involving logic and reasoning tasks. Those results offer a reason to examine our habits; they do not measure software review or the effectiveness of Semitexa.
For a developer, the practical warning sign is accepting a conclusion because the agent supplied a convincing explanation, while losing track of the evidence needed to support it. A reassuring test summary can become part of that explanation. Even an ultimately correct patch can be accepted through a process that leaves you unable to judge the next one.
Delegating work is useful. You can let an agent find callers, propose code, and run checks while retaining responsibility for the requirement and the acceptance decision. The review needs an inspectable path from the requested behavior to the result. Semitexa helps make that path available through explicit boundaries, structural queries, runtime observations, and recorded verification.
Give the requirement a sentence and a place in code
For this example, the requirement is deliberately narrow: a person may download a report only when the person and report belong to the same tenant and the person has download permission. We model it in an ordinary PHP policy:
final class ReportAccessPolicy
{
public function canDownload(
string $actorTenant,
string $reportTenant,
bool $hasDownloadPermission,
): bool {
return $actorTenant === $reportTenant
&& $hasDownloadPermission;
}
}
The proposed replacement is:
return $hasDownloadPermission;
You can now compare two concrete statements. The requirement asks for tenant agreement and permission. The replacement checks permission. A claim that tenant agreement is established elsewhere needs a source location, a description of the callers it covers, and a test at that boundary.
This small policy is illustrative application code, not a built-in Semitexa authorization API. A real feature may need subject-specific grants, privileged cross-tenant operations, revoked access, and other rules. Giving each decision an explicit owner makes those requirements reviewable. The DDD policy walkthrough explores that responsibility, while the tenant-context guide explains why carrying context and enforcing data access remain separate concerns.
Two passing tests can describe the wrong confidence
Suppose the existing suite covers a permitted download and a denied download within one tenant. Both implementations satisfy those two examples. Expand the cases before accepting the deletion:
| Case | Required result | Original policy | Proposed patch |
|---|---|---|---|
| Same tenant, permission granted | Allow | Allow | Allow |
| Same tenant, permission absent | Deny | Deny | Deny |
| Different tenant, permission granted | Deny | Deny | Allow: regression |
The limited suite reports 2/2 passes for the original and proposed policies. With the cross-tenant case included, the original passes 3/3 and the proposal passes 2/3. Restoring the comparison passes all three. The failure was absent from the limited suite’s questions.
What would make removing the comparison defensible?
You would need a boundary that actually establishes tenant agreement for every supported caller, plus verification that callers cannot supply a report outside it. A test showing that one normal request has a tenant context would leave the report-access question unanswered. If the guarantee belongs upstream, make that contract explicit and exercise its rejection case.
For this article, a retained PHP harness executed nine checks across the original policy, the limited-suite result, the detected regression, and the restored policy. It performs no HTTP request, database query, middleware execution, or real report download. It verifies this predicate and the coverage example; integration checks would still be required for a deployed feature.
Use the graph to question “already handled elsewhere”
A claim about another layer is a claim about structure. Start with the actual policy class and ask who uses it. Semitexa’s ai:review-graph:query --usages follows indexed relationships; ai:review-graph:impact helps inspect the reach of a change. Follow the callers back toward their entry points, then read the code that is supposed to establish the guarantee.
The visual Project Graph in the Observatory gives a developer a view of dependencies, dependents, source links, and recorded traces in a development environment. The CLI impact guide shows the command workflow. These views let you examine the same structural evidence an agent uses.
A graph edge explains a relationship extracted from code. It cannot establish that a permission predicate is correct or that all runtime paths were indexed. Check freshness and coverage, investigate dynamic references, and treat an empty result according to what the map actually includes. The illustrative policy above lives in an isolated evidence script; it is not a discovered application route shown in the graph.
The useful review question becomes specific: which caller establishes tenant agreement for the report passed here, and what makes that promise hold for the other callers? You can ask the agent to retrieve that evidence, then assess it yourself.
Read what ran, what passed, and what remains unknown
Structural inspection and runtime observation answer different questions. In a development environment, the Observatory and recorded request traces can show an executed path and help connect it to source. A trace of an allowed request is evidence about that execution. To evaluate the removed tenant condition, reproduce the denied cross-tenant case through the relevant application boundary as well.
Semitexa’s ai:verify selects syntax, lint, structure, static-analysis, and test checks from the supplied file scope and available tooling. Read the selected targets and their results. A passing lint check establishes something different from a permission test. A skipped test or incomplete analyzer contributes no passing result to the access requirement.
The counterexample in this article was executed separately with php var/docs/cognitive-surrender-example.php. We do not infer its execution from an ai:verify summary. In your own project, run the relevant behavioral test explicitly when the selected checks do not cover it. The AI-native investigation demonstrates how focused checks fit into a broader examination of behavior.
Keep the decision and supporting evidence in ai:work and ai:trace so another session can inspect them. These task artifacts preserve what was recorded; their contents still need evidence. Task traces are also distinct from runtime request traces: one records the investigation, the other records an execution. An unsupported conclusion remains unsupported when saved.
Make acceptance an outcome you can explain
An inspectable review
Keep a reason for accepting the patch.
- 01 / REQUIREName the invariant
Write the behavior that must remain true.
Same tenant + permission - 02 / INSPECTFollow the claim
Find the rule’s owner and relevant callers.
Source + graph - 03 / CHALLENGETry the boundary
Run a case that would expose the deletion.
Other tenant + permission - 04 / DECIDERecord the reason
Connect the acceptance to observed evidence.
Result + remaining limits
A useful request to an agent is: “Show which requirement this change preserves, where the existing guarantee is enforced, and a case that would fail if your assumption were wrong.” That produces a reviewable proposal. You can delegate discovery and execution while assessing the reasoning against source and results.
Scale the review to the consequence of the change. A spelling correction and a removed access condition need different evidence. Focus attention on behavior, contracts, and affected callers so the review remains practical. If the agent cannot substantiate a key assumption, keep that uncertainty visible and continue the investigation.
Keep exercising the part you want to retain
Repeated delegation becomes more useful when you can still formulate requirements, recognize a counterexample, and explain an acceptance decision. Reading every generated line is rarely enough on its own. Understanding where the decision lives and checking its boundaries gives the reading a purpose.
Semitexa makes those activities easier to carry out: explicit request and domain boundaries, a shared structural map, observable executions, and durable verification records put evidence within reach. This article demonstrates an engineering workflow; it does not claim that Semitexa has been experimentally shown to prevent cognitive surrender. The developer still has to use the evidence and judge what it establishes.
For our proposed patch, the decision is clear: retain the tenant comparison because the stated requirement includes it, the cross-tenant case fails without it, and the proposed upstream guarantee has not been established. That is a reason another developer can challenge, reproduce, or extend.
Try the same exercise with your next AI-generated change. Before accepting it, finish this sentence in your own words: “I accept this change because this evidence shows that the requirement still holds.”