← All research

EXP.002 / AI-BUILT SOFTWARE SECURITY

Security Drift in Iterative AI Coding

How does repeated AI-assisted modification affect software security?

PLANNED

Abstract

A planned study of how software security changes when an AI coding agent repeatedly modifies the same program. It will track vulnerabilities that appear, disappear, or change in severity, alongside functional correctness. It does not assume that security necessarily gets worse with more iterations.

Background

AI-assisted development often involves a sequence of feature requests, bug fixes, and refactoring tasks. A change that meets a functional request may also affect security. This study asks how those effects vary across repeated modifications and between function-focused and security-focused requests.

Proposed methodology

  1. Select test applications and record their initial functionality and security state.
  2. Ask an AI coding agent to add features, make changes, and refactor the same program across multiple iterations, preserving each code snapshot and request.
  3. Compare function-focused requests with security-focused requests under documented conditions.
  4. Review security findings and run functional checks at selected checkpoints, comparing each snapshot with the baseline and earlier versions.
  5. Report improvements, regressions, unchanged findings, and limitations without assuming a direction of change.

Iteration strategy

  • Possible checkpoints include iterations 1, 3, 5, and 10. The final schedule will be defined before running the experiment.
  • Record the model or agent, requests, code changes, and evaluation conditions so each iteration can be compared.

Planned security evaluation

  • Vulnerability count and severity at each checkpoint
  • CWE categories associated with reviewed findings
  • Newly introduced and fixed vulnerabilities
  • Security regressions, including issues that return after a fix
  • Functional correctness after each modification

RESULTS

Study planned

No experimental results or benchmark values have been published.

Limitations

The study is planned. Models, test applications, and evaluation tools have not been specified, and no experimental results are available. Finding counts will require review: tool output alone does not establish whether a vulnerability is real. Any later conclusions will be limited to the tested applications and conditions.

Possible future application

Findings may inform future VibeGuard security checks for AI-built software, including checks for changes or regressions after repeated modifications. This is a possible research direction, not an implemented VibeGuard feature.