# CodeLoom AI — Complete Technical & Company Knowledge Base URL: https://codeloom-ai.com/ Status: Pre-launch Founder: Vasil Vasilev (Founder & Lead Systems Architect) Contact: founder@codeloom-ai.com ## 1. Executive Summary Modernize legacy codebases without breaking existing behavior. CodeLoom AI maps repository dependencies, generates characterization tests to lock in current behavior, and migrates legacy code incrementally through small, test-verified pull requests. CodeLoom AI is an automated codebase modernization system that migrates legacy software—outdated frameworks, deprecated runtimes, and tangled architectures—to modern architectures while preserving existing behavior. Founded by Vasil Vasilev and currently in pre-launch development, CodeLoom AI is built around a simple engineering rule: behavior preservation comes first. Every code change is verified against characterization tests generated and run against the original codebase before any source file is modified. ## 2. The Problem with Manual Refactoring & Conventional AI Coding Agents Enterprises spend heavily maintaining brittle, poorly documented legacy software. Manual refactoring takes months and diverts senior engineers from product work. - **Hidden dependency coupling**: In large or multi-repo codebases, modules depend on implicit state and undocumented side effects. Changing a low-level utility can trigger cascading failures across unrelated services. - **Undocumented edge-case behavior**: Legacy systems encode years of bug fixes and business rules that exist only in production code. Rewriting code without baseline characterization tests silently drops critical behavior. - **Unreviewable migration branches**: Long-running migration branches accumulate thousands of changed lines and drift from main. Reviewers cannot meaningfully audit massive diffs, delaying merges or letting regressions slip in. ## 3. How CodeLoom AI Works (Three-Stage Pipeline) ### Stage 01: Map the repository CodeLoom AI ingests whole codebases—including multi-repo setups—and builds a structural dependency graph from imports, call sites, and shared types. It identifies leaf modules with zero downstream dependents first, ordering the migration plan so the impact of every change is known upfront. - Parses cross-file and cross-repository dependency edges - Ranks leaf nodes first for safe, bottom-up migration ordering - Calculates blast radius before scheduling any transformation ### Stage 02: Lock in current behavior Before touching a single line of legacy code, CodeLoom AI generates characterization tests (unit and integration) that capture what the code actually does today—including edge cases and error handling—and runs them against the original implementation. - Generates unit and integration tests from existing execution paths - Runs test suite against unmodified legacy source to establish baseline - Freezes behavioral contract prior to any code modification ### Stage 03: Modernize incrementally & repair AI agents rewrite modules toward the target stack and run the locked characterization tests against every change. If a test fails, the agent inspects the failure, repairs the code, and re-verifies before opening a small, audit-ready pull request that explains what changed, why, and what verified it. - Transforms legacy syntax and patterns to the target architecture - Executes automated verify-and-repair loop when a test flags a mismatch - Produces small, human-reviewable pull requests with verification logs ## 4. Why CodeLoom AI vs. Conventional AI Coding Agents IDE copilots and prompt-driven coding agents are built to generate new features from scratch. When pointed at a tangled legacy repository, they edit files in arbitrary order, hallucinate missing context, and rewrite tests to match their own broken output. - **Execution & Migration Order** - Conventional AI Coding Agents: Edits files ad-hoc based on user prompts or limited context windows, breaking downstream callers across the repository. - CodeLoom AI: Builds a static AST dependency graph first and migrates strictly bottom-up from zero-dependency leaf modules. - **Behavioral Baseline** - Conventional AI Coding Agents: Rewrites code immediately—or writes new unit tests after editing the code, baking regressions directly into the test suite. - CodeLoom AI: Synthesizes and runs characterization tests against the untouched legacy code first, locking in real runtime behavior before any edit. - **Undocumented Quirks & Edge Cases** - Conventional AI Coding Agents: Cleans up 'messy' conditionals that actually encode years of production bug fixes, causing silent business-logic regressions. - CodeLoom AI: Treats legacy execution paths—including null quirks, rounding, and error envelopes—as an immutable contract that must pass. - **Verification & Repair** - Conventional AI Coding Agents: Relies on the developer to manually spot bugs in the IDE, run tests locally, and paste stack traces back into a chat window. - CodeLoom AI: Runs every candidate diff inside an automated sandbox gate; if a locked test fails, the repair loop patches the diff before you ever see it. - **Code Review & Auditability** - Conventional AI Coding Agents: Dumps sprawling multi-file diffs across unrelated modules that senior engineers cannot safely review or approve. - CodeLoom AI: Emits small, single-module pull requests documenting exactly what changed, why it was scheduled, and which baseline tests verified it. ## 5. System Architecture (4 Layers) ### LAYER 01: Static AST & Cross-Repo Dependency Graph Engine Parses source trees into an explicit directed graph of modules, symbols, and call sites before any transformation is planned. - Builds symbol-level import/export edges across packages and repositories - Detects circular dependency clusters and isolates strongly connected components - Produces a bottom-up topological queue starting at zero-dependency leaf modules ### LAYER 02: Characterization Test Synthesizer & Baseline Runner Captures observable behavior of the untouched legacy module by generating and executing characterization suites against the original runtime. - Synthesizes unit and integration tests covering standard paths, boundary values, and error branches - Executes the generated test suite against the unmodified legacy code to establish a green baseline - Locks the test assertions as an immutable verification gate for that migration unit ### LAYER 03: Constrained Transformation & Automated Repair Loop Rewrites a single scoped module toward the target stack and iteratively repairs diffs against the locked test suite. - Applies target architecture rules (module system, type annotations, framework lifecycle) - Runs the locked characterization suite in an isolated sandbox after each edit - Feeds assertion diffs and stack traces back into the repair agent if any test flags a mismatch ### LAYER 04: Audit-Ready Pull Request Packaging Emits human-reviewable pull requests only after all baseline characterization checks pass. - Scopes diffs to small, single-module units that engineers can review in minutes - Attaches structured documentation: What Changed, Why It Changed, and What Verified It - Integrates with standard Git code review workflows so human maintainers approve every merge ## 6. Security, Isolation & Data Governance Enterprise codebases are core intellectual property. CodeLoom AI is designed from day one so customer source code is never used to train foundation models and every test execution runs inside an isolated, ephemeral sandbox. - **Zero training on customer repositories (Data Policy)**: Customer source code, AST dependency graphs, and characterization test outputs are never used to train, fine-tune, or improve shared AI models. All model inference calls are stateless. - **Ephemeral sandboxed test runners (Execution Boundary)**: Baseline characterization suites and candidate migration diffs execute inside network-restricted, ephemeral containers that are destroyed immediately after the verification run completes. - **Self-hosted VPC & on-premises runner architecture (Deployment Control)**: For regulated organizations, CodeLoom AI's graph analysis and test execution engine is designed to run entirely inside your own cloud VPC or on-premises CI infrastructure. - **No autonomous merges to protected branches (Human Governance)**: CodeLoom AI never pushes directly to main or production branches. Its sole output is a scoped, test-verified pull request subject to your existing branch protection rules and human code review. ## 7. Planned Licensing & Deployment Tiers Pre-launch notice: CodeLoom AI is not yet commercially available. Pricing figures are to be announced. ### Pilot / Module Scope (Planned Tier • Evaluation) — Pricing: To be announced For engineering teams validating behavior-preserving migration on a single service or bounded package. - Single-repository dependency graph mapping & leaf-node ordering - Automated characterization test generation for selected modules - Verify-and-repair transformation loop with full test logs - Audit-ready pull request generation for human review ### Engineering Organization (Planned Tier • Multi-Repo) — Pricing: To be announced For platform and core engineering teams migrating interconnected services, shared libraries, and frameworks. - Multi-repository dependency mapping & cross-package blast-radius analysis - Custom target architecture rulesets and internal framework conventions - Parallelized characterization test runner integration with existing CI pipelines - Batch PR scheduling ordered by topological dependency rank ### Enterprise VPC / Self-Hosted (Planned Tier • Dedicated Control) — Pricing: To be announced For regulated enterprises requiring strict source-code isolation and dedicated infrastructure deployment. - Planned deployment inside customer-controlled VPC or on-premises runners - Zero external source-code retention outside your security boundary - Custom compliance audit trails attached to every generated pull request - Dedicated architectural onboarding for custom legacy runtimes ## 8. Engineering Roadmap & Company Mission Most teams attempting to use general-purpose AI coding assistants for large-scale migrations hit the same wall: rewriting code is easy, but proving that the rewritten code preserves years of subtle production behavior is hard. CodeLoom AI flips the workflow—investing in static dependency graph analysis and automated characterization test synthesis before a single line of legacy code is modified. - **Milestone 01 • Active Development — Core AST Dependency Mapper & Characterization Harness (In Progress)**: Static cross-file dependency graph construction, leaf-first topological sorting, and automated characterization test synthesis for JavaScript/TypeScript and Python modules. - **Milestone 02 • Pre-Launch Validation — Closed-Loop Verify-and-Repair Agent & Git PR Packager (In Progress)**: Automated sandboxed test execution that feeds assertion failures and stack traces back into the transformation loop until the baseline contract passes. - **Milestone 03 • Upcoming — Design Partner Pilots & Multi-Repo VPC Runners (Planned)**: Controlled pilot deployments on bounded production repositories with engineering teams, followed by self-hosted VPC runner packaging. ## 9. Frequently Asked Questions ### Q: Is CodeLoom AI live? A: No. CodeLoom AI is currently in pre-launch development. ### Q: When will CodeLoom AI launch? A: To be announced. We are focused on building and validating our core mapping, characterization testing, and incremental migration pipeline. ### Q: What languages and frameworks will it support? A: We are validating our initial static analysis and characterization runner on JavaScript/TypeScript (Node.js, CommonJS-to-ESM, legacy React) and Python codebases, with additional enterprise runtimes to be announced. ### Q: How are changes verified? A: Before modifying any legacy code, CodeLoom AI generates unit and integration characterization tests and runs them against the original codebase to capture its current behavior. Every subsequent migration change is run against those tests. If a test fails, the system repairs the change and re-runs the suite until the tests pass before generating a pull request. ### Q: Who reviews the output? A: Your engineering team reviews every change. CodeLoom AI produces small, self-contained pull requests that explain what changed, why it changed, and which characterization tests verified the behavior, so human engineers can audit and approve them before merge. ### Q: What does "behavior preservation" mean? A: Behavior preservation means the modernized code produces the same observable outputs, side effects, and error handling as the original legacy code for the inputs covered by its characterization test suite. Instead of rewriting features from scratch, CodeLoom AI locks in what the software does today and verifies refactored code against that baseline. ### Q: Where is the product hosted and how is codebase data handled? A: CodeLoom AI is architected for ephemeral, isolated sandbox execution with zero customer code retention for model training, plus planned self-hosted VPC runner options for enterprise deployments. Full compliance details are available on our Security page.