Security Depth Is A Huge Advantage for Software Engineer

Anthropic's 2026 Agentic Coding Trends Report says engineers can now work across frontend, backend, databases, and infrastructure, areas where they previously lacked expertise, because "AI fills in knowledge gaps while humans provide oversight and direction." Anthropic sells coding agents, so read that as a vendor's view. The surveys below point the same way.
When agents fill knowledge gaps across the stack, narrow specializations lose their edge, and the scarce skill is the judgment to tell whether what an agent produced is safe. A developer with strong security skills carries that judgment inside the work they already do.
This post goes through the evidence that stack specializations are blurring, the ways generated code fails that security depth catches, where that depth multiplies a developer's existing stack, and how to build it without changing jobs.
Agents make breadth inexpensive
Three data points on where the boundaries are going:
In the 2026 AI Engineering Survey, 81% of respondents say roles in the product organization are blurring (44% significantly, 37% somewhat).
In the Stack Overflow 2026 survey, 56% of respondents use AI or agents to generate code in an area unfamiliar to them.
Addy Osmani expects AI tools to "augment generalists more," and describes the profile that holds up as a T shape: "Deep expertise in one or two areas (the vertical stroke), broad familiarity with many others (the horizontal stroke)."
The surveys show roles blurring. They do not show specialists disappearing, and Osmani's T shape is the version of specialization that survives: breadth from the agent, depth from you. The useful question is which depth keeps paying when an agent writes the first draft. Security qualifies because generated code fails in security-shaped ways.
Generated code often lacks security controls unless instructed
The OWASP Top 10:2025 calls itself "a standard awareness document for developers," and broken access control sits at the top. GitHub's Octoverse 2025 shows the same shift in code scanning: broken access control overtook injection as the most common CodeQL alert, flagged in more than 151,000 repositories (up 172% year over year). GitHub attributes much of it to "misconfigured permissions in CI/CD pipelines and AI-generated scaffolds that skip critical auth checks."
Two studies show how that happens:
Perry et al. (Stanford, ACM CCS 2023, 47 participants). Developers with an AI assistant "wrote significantly less secure code" and "were also more likely to believe they wrote secure code." On the SQL task, 36% of the AI group wrote injectable code, against 7% of the control group. The sample is small, so read it as a study and keep the mechanism: confidence rose while quality fell.
Veracode's 2026 GenAI Code Security Report (a vendor study, run on raw models with no security prompting or review). The average security pass rate was 56%, essentially unchanged from 55% the year before, so about 44% of tasks still introduced a known vulnerability while syntax correctness is near 100%. Model size made no difference, and coding-specialized models (51%) scored no higher than general-purpose ones (52%).
Developers already work with this gap in mind. In the Stack Overflow 2026 survey, 48% say they trust AI output only when they can easily verify it. Verification is the oversight half of Anthropic's sentence, and for security it means reading generated code and finding the missing authorization check or the string-concatenated query. That takes knowing both the framework and the attack.
Security depth multiplies the stack you already have
Security depth adds to what a developer already does in each stack:
| Your stack | What security depth adds |
|---|---|
| Backend, web, API | Reviewing authorization layers, catching missing checks in generated scaffolds, threat modeling a feature before an agent builds it |
| Platform, infra, SRE | Reviewing the IAM and pipeline permissions that agents and builds inherit |
| CI/CD, build, tooling | Dependency pinning, artifact signing, and limiting which secrets a build or an agent can reach |
| Data or backend with logging experience | Writing detections as code, with tests and pipelines |
| ML or LLM application developer | Setting tool boundaries and agent permissions, with prompt injection and excessive agency from the OWASP Top 10 for LLM Applications as the threat list |
| Systems, C or C++ | Memory-safety migration and fuzzing |
The last row has the clearest measured result. Android's share of memory-safety vulnerabilities fell from 76% in 2019 to 24% in 2024, and below 20% in 2025. Google's explanation is one sentence: "The problem is overwhelmingly with new code." The drop followed a change in how developers write new components, with memory-safe languages as the default. Memory-safety bugs account for about 70% of Microsoft CVEs (2006 to 2018) and about 70% of Chromium's high-severity bugs, according to CISA and Chromium's own analysis, and that figure describes those large C and C++ codebases. In those codebases, the decision with the largest security effect was an engineering decision about new code.
Agents find bugs faster than people can verify them
AI is taking over specific tasks: alert triage, summaries, and finding and patching known bug classes. The evidence shows a new bottleneck forming behind the discovery step.
HackerOne 2025: valid AI-related vulnerability reports rose 210% and prompt injection reports rose 540%. AI adds new attack surface to secure.
Triage fatigue: discovery got cheap in 2026 and validation did not. Two examples show where the queue forms.
Project Glasswing. Anthropic's May 2026 update says maintainers face "a deluge of low-quality, AI-generated bug reports," that several are "severely capacity constrained," and that some asked it to slow its disclosures "because they need more time to design patches." Its disclosure ledger names independent human triage and review as "the rate limiting step." VulnCheck's September read of that ledger counted 26,153 findings, of which 2,736 (10.5%) had reached the ledger and 202 (0.8%) were marked fixed. Where both ratings exist, Claude scored 91.5% of findings high or critical and maintainers scored 51.3% that way. VulnCheck disputes parts of Anthropic's true-positive claim and reports that Anthropic's dashboard counts 421 fixes where the ledger lists 202, so treat the exact rates as contested. The queue is the consistent part: Anthropic's own summary is that progress "is limited by how quickly we can verify, disclose, and patch" what AI finds.
Google's open source reward program. On October 1, 2026, Google stopped accepting product vulnerability submissions to its Open Source Software Vulnerability Reward Program, citing "a significant rise in automated submissions, the vast majority of which are not valid." Supply-chain reports and reports already filed stay open, and Google committed to an update in Q1 2027. In early September the Go project had added a section on LLM-generated reports to its security policy: reporters should review and filter the output first, and those who forward large amounts of unfiltered output won't be credited.
Reproducing a finding, judging reachability and severity, and writing and testing the patch are the steps maintainers are short of. Each one needs someone who can read the code and think like the attacker. That is a developer with security depth, in any stack.
Build the depth inside your current role
The route stays inside the job you have, and each step applies directly to code you own.
Learn the vocabulary attackers use. Read the OWASP Top 10:2025 and the OWASP Top 10 for LLM Applications 2025, and map each item to a bug you have seen or shipped. Read ten entries in the CISA Known Exploited Vulnerabilities catalog (1,734 entries as of October 4, 2026) for technology you use, and ask how each bug got into the code. Follow the Google Project Zero blog for root-cause analyses.
Practice on labs. PortSwigger Web Security Academy is free and has interactive labs (30 XSS, 16 SQL injection, 7 on web LLM attacks, 5 on API testing). OWASP Juice Shop is an intentionally vulnerable app you can run locally.
Run a standard against your own service. OWASP ASVS 5.0 is a checklist for what "secure" means for an application. Write down every control you can't point to in code.
Bring it into your team. Join or start a security champions program: OWASP's guidance is that "security engineers do not scale well across teams of developers," and that champions should be nominated, with management buy-in for protected time. Own threat modeling for your team's next design doc. Burn down the SAST, dependency, and secret-scanning backlog (CodeQL, Dependabot, secret scanning), since those findings tend to sit unfixed. Pair with the AppSec team when a pentest or bug bounty report hits your service.
Putting it together
Agents are making stack breadth cheap, and the work that remains for the human is oversight: deciding whether generated code, a generated patch, or a generated finding is correct and safe. Security depth is the vertical stroke of the T that applies in every stack, from access control in a web app to agent permissions in an LLM system to memory safety in C.
This week, run ASVS against one service you own and write down the three controls you can't point to in code. That list is your first threat model.
References
Developer roles and AI
Code and vulnerabilities
AI-generated code
AI and defenders
VulnCheck: The Anthropic Glasswing Receipts Are Starting to Trickle In
Google OSS VRP rules, with the product vulnerability pause notice
The Hacker News: Google pauses OSS product bug bounty rewards
Build the depth
Research for this post was assisted by AI agents.





