# diniscruz.ai — every page, one file Site version: v0.1.1. Generated by admin/build/build.py. Canonical host: https://diniscruz.ai/ — every section below is one page of the site. All writing CC BY 4.0 unless noted. ============================================================================== PAGE: /index.html (markdown twin: /index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/index.html* > Dinis Cruz: founder of sgit.ai, MyFeeds.ai and The Cyber Boardroom, former OWASP Board member. Essays and research on GenAI, AppSec, threat modeling and knowledge graphs. --- AI · cyber security · semantic knowledge graphs · open source # Dinis Cruz I build security and AI products in the open, and I write down what I learn while building them. Right now that means [**sgit.ai**](https://sgit.ai), git for encrypted vaults, built for humans and AI agents. Before that came The Cyber Boardroom, MyFeeds.ai and a long run of open-source security tooling going back to the **OWASP O2 Platform**. Founder of [sgit.ai](https://sgit.ai), [sgraph.ai](https://sgraph.ai), [MyFeeds.ai](https://investor.myfeeds.ai/), [The Cyber Boardroom](https://thecyberboardroom.com), [RiskMandate.ai](https://riskmandate.ai) and [VoiceDebrief.ai](https://voicedebrief.ai). Former OWASP Board member. Based in the UK. [Read the writing →](writing/index.md)[What I am building →](building/index.md)[About me →](about/index.md) ## Latest writing Essays, research briefs and project proposals. Many were written with an LLM as co-author, and the byline says so. 103 pieces so far, all under CC BY 4.0. [All writing →](writing/index.md) ### [GenLegalAdvise Project Plan](2025/10/03/genlegaladvise-project-plan.md) *2 Oct 2025* Small businesses, freelancers, and independent consultants regularly encounter legal documents -- consulting agreements, NDAs, service terms, EULAs, data sharing… ### [Time as a Calibrator of Credibility and Trust in Information Systems](2025/10/02/time-as-a-calibrator-of-credibility-and-trust-in-information-systems.md) *2 Oct 2025* In an era of information overload and rampant misinformation, time emerges as a critical factor in determining what information we trust. Traditional approaches to… ### [Project Voice2SIEM: Turning Customer Support Audio Into Real-Time Security Events](2025/10/02/project-voice2siem-turning-customer-support-audio-into-real-time-security-events.md) *2 Oct 2025* Phone-based social engineering and vishing (voice phishing) attacks are on the rise, targeting customer support and help desk agents. Attackers impersonate customers or… ### [Next-Generation API Security Platform: Semantic Graphs, GenAI Testing & Ephemeral Environments for 2025](2025/09/01/next-generation-api-security-platform-semantic-graphs-genai-testing-ephemeral-environments-2025.md) *1 Sep 2025* Modern API security requires going beyond traditional scanning -- it must blend into the API testing lifecycle and leverage cutting-edge AI to map and probe complex… ### [Dinis Cruz's Research on API Security (2009-2025)](2025/09/01/dinis-cruz-research-on-api-security-2009-2025.md) *1 Sep 2025* API security has been a persistent theme in Dinis Cruz's work, spanning early insights in 2009--2010 through to innovative ideas in 2025. His contributions center on how… ### [LLM Workflows/Stateflow Service - Technical Brief](2025/08/23/llm-workflows-stateflow-service-technical-brief.md) *23 Aug 2025* The LLM Workflows/Stateflow Service is a proposed stateless web service for executing AI-driven workflows with well-defined, deterministic steps. It acts as a state… ## What I am building Six companies, one strategy. Everything they ship is open source, and so are their investor materials. What they sell is the running, maintained, trusted service, not the code. ### [sgit.ai](https://sgit.ai) *Now · Apache-2.0* Git for encrypted vaults. Clone, commit, branch and merge files that are encrypted before they leave your machine. The server stores ciphertext it cannot read. ### [sgraph.ai](https://sgraph.ai) *Commercial home* Where the sgit layer turns into revenue: SG/Send, the secure file-sharing service, and hosted SG/Vaults. ### [MyFeeds.ai](https://investor.myfeeds.ai/) *Semantic graphs* Role-aware cybersecurity briefings built on semantic knowledge graphs, with CISO, engineer and board views of the same news and the source attribution kept. ### [The Cyber Boardroom](https://thecyberboardroom.com) *Security & the board* An AI-powered platform for the conversation between technical security teams and the board. Also the UK company behind sgit.ai and RiskMandate.ai. ### [RiskMandate.ai](https://riskmandate.ai) *Autonomous systems* The business risk layer for autonomous systems. A named human underwrites the exposure, and the interval is the decision. ### [VoiceDebrief.ai](https://voicedebrief.ai) *In the browser* Voice recordings into transcripts and debriefs, entirely in the browser. No account, and nothing uploaded to a server. ## Research, by topic The writing grouped into the areas I keep coming back to. Each hub is a curated reading list, not a tag cloud. ### [Cyber security & threat modeling](research/cyber-security.md) *AppSec* Semantic threat models, supply-chain security, MCP and OAuth risks, API security from 2009 to 2025, and security as a board conversation. ### [Semantic knowledge graphs](research/graphs.md) *G³* Graphs of graphs of graphs, LLMs as ephemeral graph databases, evolving ontologies, and why meaning lives in the edges. ### [AI & development](research/development-and-genai.md) *Engineering* Deterministic GenAI pipelines, Iterative Flow Development, surrogate dependencies, and the joy of programming with AI. ### [The future of news](research/the-future-of-news.md) *Trust* Fact provenance, identity graphs for authors and sources, micro-payments, and fair compensation for AI crawling. ### [Europe & learning](research/europe-and-learning.md) *Sovereignty* An open-source sovereign cloud for Europe, Europe's GenAI opportunity, and generative AI in education. ### [Projects & innovation](research/projects.md) *The lab* Project briefs and MVPs: VulnAI, InsightFlow, Voice2SIEM, JSync, SupplyShield and the rest of the innovation lab. ## The record The parts of my track record that the writing here builds on. | What | Why it matters here | |---|---| | **Former OWASP Board member** | And organiser of the OWASP Summits, Lisbon 2011 and Woburn 2017: working sessions with no spectators, only participants. The [Open Security Summit](https://open-security-summit.org/) series went on to build on that format. | | **Creator of the O2 Platform** | The OWASP static-analysis engine of 2010 to 2012, and the first of a line of open-source tooling that continues in OSBot, MGraph-DB, memory_fs, Issues-FS and sgit-ai. | | **CISO and security practitioner** | Security leadership inside UK companies, and one UK company taken through to an exit. This is where the threat-modeling and "security for the board" work comes from. | | **Founder, six times over** | The companies above, all run on the same idea: [open source is a strategy, not a charity](https://open-source.sgit.ai/). Trust is what gets sold. | The long version, with the interests I should declare, is on the [about page](about/index.md). ## Elsewhere Most of what I write is posted first on LinkedIn. The code is on GitHub. The sgit.ai network is a set of focused sites that each take one argument further than a blog post can. ### [LinkedIn](https://www.linkedin.com/in/diniscruz) *Fastest route* Where the essays are first posted and discussed, and the best way to reach me. ### [GitHub](https://github.com/DinisCruz) *Code* The open-source work, including the sgit CLI, OSBot, MGraph-DB and the source of this site. ### [open-source.sgit.ai](https://open-source.sgit.ai/) *Position* My position on open source as a strategy: the licences, survivability, the history checked against its sources. ### [The sgit.ai network](https://sgit.ai/network/index.html) *The network* Sites on graphs, threat modeling, standards, risk, agent identity and more, each publishing its argument before its implementation. --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /about/index.html (markdown twin: /about/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/about/index.html* > Dinis Cruz: founder of sgit.ai, sgraph.ai, MyFeeds.ai, The Cyber Boardroom, RiskMandate.ai and VoiceDebrief.ai; former OWASP Board member and organiser of the OWASP Summits; creator of the O2 Platform; CISO and security practitioner in the UK for thirty years. The record, where I write, the interests I declare, and how to reach me. --- # Dinis Cruz I have spent my career at the point where security, software development and, more recently, generative AI meet. I have been a security practitioner, a CISO for UK companies, an OWASP leader and a founder. What I build now, I build in the open, and this site is where the writing behind it lives. ## The record | Role | What it involves | |---|---| | **Founder, [sgit.ai](https://sgit.ai)** | Encrypted vaults with git semantics: clone, commit, branch and merge files that are encrypted before they leave your machine, under Apache-2.0. It is also the hub of [a network of focused sites](https://sgit.ai/network/index.html), each of which publishes its argument before its implementation, so the commitments can be checked. | | **Founder, [sgraph.ai](https://sgraph.ai)** | Where the strategy turns into revenue: the commercial home of SG/Send, the secure file-sharing service built on the open-source sgit layer, and of hosted SG/Vaults. The code stays Apache-2.0. What is sold is the running, maintained, certified service. | | **Founder, [MyFeeds.ai](https://investor.myfeeds.ai/)** | Role-aware cybersecurity briefings built on semantic knowledge graphs, with source attribution: CISO, engineer and board views of the same news. Open source and serverless, with the seed pitch and unit economics published in the open. | | **Founder, [The Cyber Boardroom](https://thecyberboardroom.com)** | An AI-powered platform for the conversation between technical security teams and the board, bridging the two with knowledge-graph technology. Apache-2.0, with the community edition and the [investment repository](https://github.com/the-cyber-boardroom/cbr-investment) public. The Cyber Boardroom Limited, a UK-registered company, is the commercial vehicle behind sgit.ai and RiskMandate.ai. | | **Founder, [RiskMandate.ai](https://riskmandate.ai)** | The business risk layer for autonomous systems. | | **Founder, [VoiceDebrief.ai](https://voicedebrief.ai)** | Voice recordings into transcripts and debriefs, entirely in the browser, with nothing uploaded to a server. | | **Former OWASP Board member** | And organiser of the OWASP Summits, Lisbon 2011 and Woburn 2017. That working-session format, with no spectators and only participants, is the one the [Open Security Summit](https://open-security-summit.org/) series went on to build on. My current open-source work still ships under the owasp-sbot organisation. | | **Creator, the O2 Platform** | The OWASP static-analysis engine of 2010 to 2012, and the first of a line of open-source tooling that continues in the osbot and mgraph families, memory_fs, Issues-FS and sgit-ai. All of it is Apache-2.0 and on PyPI. | | **Thirty years in the UK** | A security practitioner, a CISO for UK companies and a founder, with one UK company taken through to an exit. My stated intent is to grow more UK-based companies on this technology, in the open. | ## What I write about The [writing on this site](../writing/index.md) runs from February 2024 onwards. Most of it is research briefs and project proposals, written to be used rather than just read. The themes keep recurring: - **Application security that finally works.** Threat modeling as [semantic knowledge graphs](../research/cyber-security.md), threat models as mandatory disclosures, and security as a conversation the board can take part in. - **Semantic knowledge graphs.** G³ (graphs of graphs of graphs), LLMs as ephemeral graph databases, and ontologies that evolve rather than being designed top-down. [The graphs hub →](../research/graphs.md) - **Deterministic, provable GenAI.** Outputs with provenance, small models plus code in place of large models, and data pipelines you can debug. [AI & development →](../research/development-and-genai.md) - **Trust in news.** Fact provenance, identity graphs for authors and sources, and new ways to fund journalism. [The future of news →](../research/the-future-of-news.md) - **Europe and sovereignty.** An open-source sovereign cloud, and Europe's opportunity in GenAI. [Europe & learning →](../research/europe-and-learning.md) ## Writing elsewhere | | | |---|---| | **[LinkedIn](https://www.linkedin.com/in/diniscruz)** | Where most of the essays are first posted and discussed, and the fastest route to reach me. | | **[sgit.ai articles](https://sgit.ai/articles/index.html)** | The first-person articles on the sgit.ai site: the future of news as a story vault, the SaaS apocalypse, and Fractal Semantic Graphs. The byline there links to [a sibling of this page](https://sgit.ai/about/index.html). | | **[open-source.sgit.ai](https://open-source.sgit.ai/)** | My position on open source as a strategy rather than a charity, with the licences, the investor materials published in the open, and [the interests declared there](https://open-source.sgit.ai/about/index.html). | | **[graphs.sgit.ai](https://graphs.sgit.ai/)** | The Fractal Semantic Graphs argument in full: the grammar, the worked examples and the evidence. | | **[GitHub](https://github.com/DinisCruz)** | The code, including [the sgit CLI](https://github.com/SGit-AI/SGit-AI__CLI) and [this site's source](https://github.com/DinisCruz/DinisCruz-AI__Website). | ## Interests declared I run companies whose strategy the writing here describes, and they sell the running, maintained service rather than the code. The markets many of these essays describe are markets I intend to be in. Read the arguments knowing that. The project proposals on this site were written to start conversations with specific organisations, and they say so in their titles. Many of the documents here were written with LLMs (ChatGPT Deep Research, Claude and others) as co-authors, and the byline names them. The ideas, the direction and the editing are mine. The prose is often shared work. ## Reach me, or correct me [LinkedIn](https://www.linkedin.com/in/diniscruz) is the fastest route. [The repository for this site](https://github.com/DinisCruz/DinisCruz-AI__Website) takes issues and pull requests. If you find something wrong here, I would rather know. [← Home](../index.md)[What I am building →](../building/index.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /building/index.html (markdown twin: /building/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/building/index.html* > The companies Dinis Cruz has founded (sgit.ai, sgraph.ai, MyFeeds.ai, The Cyber Boardroom, RiskMandate.ai, VoiceDebrief.ai), the open-source projects under them (OSBot, MGraph-DB, memory_fs, Issues-FS, sgit-ai), and the sgit.ai network of focused sites. --- # What I am building Six companies, one strategy: **everything they ship is open source, and so are their investor materials.** Technology is not the moat. What gets sold is trust, meaning the running, maintained and certified service and the people accountable for it. The reasoning is set out at [open-source.sgit.ai](https://open-source.sgit.ai/). ## Companies ### [sgit.ai](https://sgit.ai) *Now · Apache-2.0* Git for encrypted vaults. A vault is a unit of work (data, app, history and sources) versioned like git and handed over with a single read key. No account, no hosting, nothing for the reader to install. The server stores ciphertext it cannot read. `pip install sgit-ai`. ### [sgraph.ai](https://sgraph.ai) *Commercial home* SG/Send, the secure file-sharing service built on the sgit layer, and hosted SG/Vaults. The code stays Apache-2.0. What is sold is the running service. ### [MyFeeds.ai](https://investor.myfeeds.ai/) *Semantic graphs* Role-aware cybersecurity briefings on semantic knowledge graphs. It is 100% open source and serverless, with no vendor lock-in, and the seed pitch is published in the open. ### [The Cyber Boardroom](https://thecyberboardroom.com) *Security & the board* An AI-powered platform that bridges security teams and the board with knowledge-graph technology. The UK company behind sgit.ai and RiskMandate.ai. ### [RiskMandate.ai](https://riskmandate.ai) *Autonomous systems* The business risk layer for autonomous systems. There is no "deny" button for a risk, only a decision on how long to accept it and who underwrites it. ### [VoiceDebrief.ai](https://voicedebrief.ai) *In the browser* Voice recordings into transcripts and debriefs, entirely client-side. It asks for keys at run time and stores nothing. ## Open source The code under those companies is Apache-2.0 and published on PyPI. It descends from the OWASP O2 Platform and has been built up one reusable layer at a time. | Project | What it is | |---|---| | **[sgit-ai](https://github.com/SGit-AI/SGit-AI__CLI)** | The sgit CLI: git semantics over client-side AES-256-GCM encrypted vaults. [PyPI](https://pypi.org/project/sgit-ai/). | | **[OSBot family](https://github.com/owasp-sbot)** | The `osbot-*` libraries: a serverless Python framework for building AI applications, and the building blocks behind the startups. They are published under the owasp-sbot organisation. | | **MGraph-DB** | A memory-first graph database for GenAI and serverless workloads, publicly credited to the OWASP community. It is the engine behind much of the [graphs research](../research/graphs.md). | | **memory_fs** | A file-system abstraction used to build file-based representations of documents and standards. For example, [the GDPR as files](../2025/07/02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md). | | **Issues-FS** | Part of the same Apache-2.0 family on PyPI, alongside memory_fs. | | **OWASP O2 Platform** | Where it started: the open-source static-analysis engine of 2010 to 2012. [O2 Platform's MethodStreams (2010) →](../2025/02/11/o2-platforms-methodstreams-2010-open-source-sast-engine.md) | ## The sgit.ai network The [sgit.ai network](https://sgit.ai/network/index.html) is a set of focused sites on `*.sgit.ai`, each taking one question further than an essay can. They share a design language (this site uses it too), and each publishes its argument **before** the thing it describes exists, so the commitments can be checked afterwards. A few that pick up threads from the writing here: ### [graphs.sgit.ai](https://graphs.sgit.ai) *Graphs* A node is just a node, and meaning lives in the edges. A grammar for semantic graphs, from five rules to a full position. ### [threat-modeling.sgit.ai](https://threat-modeling.sgit.ai) *AppSec* A threat model is a claim you can check. White papers, a threat model validated against the code, and the method behind them. ### [standards.sgit.ai](https://standards.sgit.ai) *Compliance* Laws and standards as addressable graphs: the EU AI Act, GDPR, ISO/IEC 27001 and ISO 31000, with a subset method for agents. ### [risks.sgit.ai](https://risks.sgit.ai) *Risk* You cannot deny a risk. You can only say how long you accept it, and a named human underwrites it. ### [nhi.sgit.ai](https://nhi.sgit.ai) *Agents* Identity for AI agents, and why the industry only answers half the question. ### [open-source.sgit.ai](https://open-source.sgit.ai) *Strategy* Open source is a strategy, not a charity. Survivability, the licences, and a history checked against its sources. [← About](../about/index.md)[The writing →](../writing/index.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /writing/index.html (markdown twin: /writing/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/writing/index.html* # Writing > Every essay, research brief and project proposal by Dinis Cruz, newest first — 103 pieces. --- ## October 2025 - 2025-10-02 — [GenLegalAdvise Project Plan](../2025/10/03/genlegaladvise-project-plan.md): Small businesses, freelancers, and independent consultants regularly encounter legal documents -- consulting agreements, NDAs, service terms, EULAs, data sharing policies -- that are lengthy, dense… - 2025-10-02 — [Time as a Calibrator of Credibility and Trust in Information Systems](../2025/10/02/time-as-a-calibrator-of-credibility-and-trust-in-information-systems.md): In an era of information overload and rampant misinformation, time emerges as a critical factor in determining what information we trust. Traditional approaches to credibility tend to evaluate a… - 2025-10-02 — [Project Voice2SIEM: Turning Customer Support Audio Into Real-Time Security Events](../2025/10/02/project-voice2siem-turning-customer-support-audio-into-real-time-security-events.md): Phone-based social engineering and vishing (voice phishing) attacks are on the rise, targeting customer support and help desk agents. Attackers impersonate customers or executives to manipulate… ## September 2025 - 2025-09-01 — [Next-Generation API Security Platform: Semantic Graphs, GenAI Testing & Ephemeral Environments for 2025](../2025/09/01/next-generation-api-security-platform-semantic-graphs-genai-testing-ephemeral-environments-2025.md): Modern API security requires going beyond traditional scanning -- it must blend into the API testing lifecycle and leverage cutting-edge AI to map and probe complex behaviors. This white paper… - 2025-09-01 — [Dinis Cruz's Research on API Security (2009-2025)](../2025/09/01/dinis-cruz-research-on-api-security-2009-2025.md): API security has been a persistent theme in Dinis Cruz's work, spanning early insights in 2009--2010 through to innovative ideas in 2025. His contributions center on how to test and analyze web… ## August 2025 - 2025-08-23 — [LLM Workflows/Stateflow Service - Technical Brief](../2025/08/23/llm-workflows-stateflow-service-technical-brief.md): The LLM Workflows/Stateflow Service is a proposed stateless web service for executing AI-driven workflows with well-defined, deterministic steps. It acts as a state machine for orchestrating Large… - 2025-08-22 — [Personas Service - Technical LLM Brief](../2025/08/23/personas-service-technical-llm-brief.md): The Personas Service (to be deployed at personas.prod.mgraph.ai) is a stateless microservice that leverages Large Language Models (LLMs) to translate and tailor content for specific target personas.… - 2025-08-22 — [Surrogate Dependencies: Simulating Backends for Offline-First Development](../2025/08/22/surrogate-dependencies-simulating-backends-for-offline-first-development.md): Modern software development often relies on remote APIs and services, yet developers frequently need to work in environments where those backends are unavailable or unstable. Surrogate dependencies… - 2025-08-22 — [Iterative Flow Development (IFD) Methodology: JavaScript Web Application Implementation](../2025/08/22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.md): Generative AI (GenAI) is rapidly transforming software development, enabling new coding paradigms where human developers collaborate with Large Language Models (LLMs) as intelligent assistants. ## July 2025 - 2025-07-27 — [Project VulnAI: AI-Powered Vulnerability Risk Management Platform](../2025/07/27/project-vulnai-ai-powered-vulnerability-risk-management-platform.md): Organizations continue to struggle with a deluge of security vulnerabilities and alerts, but traditional vulnerability management approaches have failed to translate these findings into effective… - 2025-07-06 — [Semantic Knowledge Graphs, G³, and Sustainable AI: Aligning Innovations with ESG Objectives](../2025/07/06/semantic-knowledge-graphs-g3-and-sustainable-ai-aligning-innovations-with-esg-objectives.md): In an era when organizations are increasingly accountable for the environmental, social, and governance (ESG) impacts of their technology, innovative approaches in data and AI engineering are… - 2025-07-06 — [Graph‑Based Cloud IAM in the GenAI Agentic World](../2025/07/06/graph-based-cloud-iam-in-the-genai-agentic-world.md): The rapid rise of generative AI agents – autonomous software powered by large language models (LLMs) – is redefining how applications interact with cloud services. These AI agents can dynamically… - 2025-07-06 — [From Large Language Models to Small Models and Code: The Evolution of AI Solutions](../2025/07/06/from-large-language-models-to-small-models-and-code-the-evolution-of-ai-solutions.md): Large Language Models (LLMs) like GPT-3 and GPT-4 have revolutionized what machines can do with human language. These massive models (with hundreds of billions of parameters) stunned the world by… - 2025-07-06 — [Finding the “Good Enough” Threshold: Optimizing Risk, Creativity, and Product Decisions](../2025/07/06/finding-the-good-enough-threshold-optimizing-risk-creativity-and-product-decisions.md): Every decision in business, creativity, or engineering comes with a trade-off between striving for perfection and delivering “good enough” at the right time. There is an optimal point at which… - 2025-07-04 — [Usage-Based Billable Entities: Aligning SaaS Pricing with Customer Usage](../2025/07/04/usage-based-billable-entities-aligning-saas-pricing-with-customer-usage.md): Software-as-a-Service (SaaS) companies have traditionally relied on fixed subscription plans – charging customers a monthly or annual fee for access to a product. However, this model often misaligns… - 2025-07-04 — [The Joy of Programming in the Age of AI-Assisted Development](../2025/07/04/the-joy-of-programming-in-the-age-of-ai-assisted-development.md): In the software engineering community, the term “flow state” or being “in the zone” describes a peak experience of focus and creativity that makes programming deeply rewarding. A coder in this state… - 2025-07-04 — [From Free Scraping to Fair Compensation: Cloudflare’s GenAI Crawler Charges and the Future of News Monetization](../2025/07/04/from-free-scraping-to-fair-compensation-cloudflares-genai-crawler-charges-and-the-future-of-news-monetization.md): The rise of generative AI has upended the traditional relationship between content creators and aggregators. AI web crawlers now scrape vast amounts of text, news, and images to train large language… - 2025-07-04 — [FAQ - Evolving Semantic Graphs and Ontologies with LLMs and MGraph-DB](../2025/07/04/faq-evolving-semantic-graphs-and-ontologies-with-llms-and-mgraph-db.md): This FAQ-style white paper addresses key technical questions about Dinis Cruz’s approach to building and evolving semantic knowledge graphs using Large Language Models (LLMs) and the MGraph-DB… - 2025-07-03 — [LLM-Driven GDPR Compliance Q&A Graph – Technical Brief](../2025/07/03/llm-driven-gdpr-compliance-q-and-a-graph-technical-brief.md): This project aims to build an interactive chatbot UI that guides a user through a series of questions about their organization’s GDPR practices, dynamically builds a knowledge graph from the answers… - 2025-07-02 — [Using Memory_FS to Build a File-Based Representation of the GDPR Standard](../2025/07/02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md): This white paper presents a method for converting the General Data Protection Regulation (GDPR) text into a structured, file-based representation using the MemoryFS framework. Targeted at technical… - 2025-07-02 — [Ephemeral GenAI SIEM: A Serverless, Graph-Driven Approach to Security Event Management](../2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md): Security Information and Event Management (SIEM) tools have long been central to enterprise defense, but traditional SIEMs struggle with scale, context, and cost. They often ingest massive volumes of… ## June 2025 - 2025-06-25 — [Using Ephemeral Neo4j Instances for a Cybersecurity Risk Graph Scenario](../2025/06/25/using-ephemeral-neo4j-instances-for-a-cybersecurity-risk-graph-scenario.md): Graph thinkers – those who naturally conceptualize information as networks of nodes and relationships – can greatly benefit from on-demand graph analytics. Recently, we demonstrated how a Large… - 2025-06-25 — [Ephemeral Neo4j Instances for On-Demand Graph Analytics](../2025/06/25/ephemeral-neo4j-instances-for-on-demand-graph-analytics.md): Organizations are increasingly seeking ways to perform complex graph data transformations and analytics without the overhead of maintaining long-running database servers. Neo4j – a leading graph… - 2025-06-25 — [Data Tests for Neo4j: Bringing Automated Testing to Graph Databases](../2025/06/25/data-tests-for-neo4j-bringing-automated-testing-to-graph-databases.md): Neo4j’s flexible, schema-optional nature is a double-edged sword: it empowers rapid graph modeling but can lead to hidden data inconsistencies as the graph evolves. Data tests for Neo4j apply the… - 2025-06-22 — [Project Plan: High Street GenAI Learning Hub](../2025/06/22/project-plan-high-street-genai-learning-hub.md): In many communities, there is a need for a "third space" – a place outside of home and work where people can gather, learn, and collaborate in a non-commercial, inclusive environment. As author Zadie… - 2025-06-22 — [FIST Meets the Semantic Knowledge Graph: Aligning Fast, Inexpensive, Simple, Tiny with Dinis Cruz's G³ Approach](../2025/06/22/fist-meets-the-semantic-knowledge-graph-aligning-fast-inexpensive-simple-tiny-with-dinis-cruzs-g3-approach.md): Modern engineering projects thrive on agility, thrift, and simplicity. These values are encapsulated in the FIST framework – Fast, Inexpensive, Simple, Tiny – originally coined in the defense… - 2025-06-19 — [Using LLMs as Ephemeral Graph Databases: Empowering the Graph Thinkers in the Age of Generative AI](../2025/06/19/using-llms-as-ephemeral-graph-databases--empowering-the-graph-thinkers-in-the-age-of-generative-ai.md): Graph thinkers – those who naturally conceptualize information as networks of nodes and relationships – are poised to benefit immensely from advances in generative AI. Traditional graph databases… - 2025-06-19 — [Empowering Workshops with Custom GPTs for GenAI Training](../2025/06/19/empowering-workshops-with-custom-gpts-for-genai-training.md): GenAI (Generative AI) has rapidly emerged as a transformative technology, yet many executives and professionals struggle to find practical entry points for its adoption. Often, the barrier is not… - 2025-06-18 — [No Code Development (NCD): A Paradigm Shift Beyond 'Vibe Coding'](../2025/06/18/no-code-development--ndc--a-paradigm-shift-beyond-vibe-coding.md): In recent years, advances in generative AI – especially large language models (LLMs) like GPT-4, Codex, and others – have ushered in a new style of software creation. Instead of writing syntax… - 2025-06-18 — [Empowering the Graph Thinkers in the Age of Generative AI](../2025/06/18/empowering-the-graph-thinkers-in-the-age-of-generative-ai.md): In a world increasingly defined by complex interconnections, a unique set of individuals thrives on seeing relationships and patterns everywhere – the "graph thinkers." These are people who… - 2025-06-18 — [Comparing the EU FED Cloud vs. an Open-Source Federated Cloud Proposal](../2025/06/18/comparing-the-eu-fed-cloud-vs-an-open-source-federated-cloud-proposal.md): The EU FED Cloud is a European initiative to build a federated, sovereign cloud ecosystem with built-in compliance and security. It's not a single public cloud provider, but a framework of multiple… - 2025-06-18 — [Bridging Niklas Luhmann's Ideas with Semantic Knowledge Graphs and G³](../2025/06/18/bridging-niklas-luhmanns-ideas-with-semantic-knowledge-graphs-and-g3.md): Niklas Luhmann (1927–1998) was a renowned German sociologist and systems theorist, famous not only for his influential theory of social systems but also for his extraordinarily productive writing… - 2025-06-16 — [Workshop Plan: User-Driven Semantic Persona Graphs Powered by GenAI](../2025/06/16/workshop-plan-user-driven-semantic-persona-graphs-powered-by-genai.md): This workshop will demonstrate how to use Generative AI (GenAI) to build a personalized semantic knowledge graph for a user and then generate tailored outputs for different stakeholder personas. We… - 2025-06-15 — [The Hidden Cost of Ephemeral Testing and the Case for Automation](../2025/06/15/the-hidden-cost-of-ephemeral-testing-and-the-case-for-automation.md): The key insight is that lack of visible tests doesn't mean testing isn't happening – it means testing is happening inefficiently. Developers practicing Ephemeral Test-Driven Development (ETDD)… - 2025-06-15 — [Personal Content Rights: Protecting Individuals in the Age of Deepfakes and AI Cloning](../2025/06/15/personal-content-rights-protecting-individuals-in-the-age-of-deepfakes-and-ai-cloning.md): The rise of generative AI has enabled anyone to clone voices, faces, and writing styles with startling realism. From fabricated videos of leaders to AI-generated voice scams, this technology is being… - 2025-06-14 — [User-Driven Semantic Persona Graphs Powered by GenAI](../2025/06/14/user-driven-semantic-persona-graphs-powered-by-genai.md): Modern organizations increasingly rely on knowledge graphs to model complex relationships between business entities, risks, and controls. However, building these graphs traditionally requires… - 2025-06-13 — [Technical Briefing: Web Content Filtering Project](../2025/06/13/technical-briefing-web-content-filtering-project.md): This technical briefing outlines the architecture and design of the Web Content Filtering Project, which aims to give users fine-grained control over the content they see as they browse the web. The… - 2025-06-13 — [Follow-Up Technical Vision: Optimizations, Deployment, and Security](../2025/06/13/follow-up-technical-vision-optimizations-deployment-and-security.md): In our initial briefing, we outlined a web content filtering platform that acts as a proxy to intercept web pages and filter out disallowed content before rendering to users. The goal of such a… - 2025-06-10 — [Security Implications of the Model Context Protocol (MCP) and the Need for Robust Infrastructure](../2025/06/10/security-implications-of-the-model-context-protocol-mcp-and-the-need-for-robust-infrastructure.md): The Model Context Protocol (MCP) has emerged as a powerful open standard that enables Large Language Models (LLMs) to interface with external tools and data sources in a uniform way. By acting as a… - 2025-06-10 — [Explorers, Villagers, and Town Planners: Understanding the Generative AI Divide](../2025/06/10/explorers-villagers-and-town-planners-understanding-the-generative-ai-divide.md): In the generative AI (GenAI) community of mid-2025, a puzzling divide has emerged. On one side stand the enthusiastic explorers – pioneers who see GenAI as a technology that can be used everywhere… - 2025-06-09 — [Supercharging AppSec Threat Modeling Services with GenAI and Semantic Graphs](../2025/06/09/supercharging-appsec-threat-modeling-services-with-genai-and-semantic-graphs.md): Application security (AppSec) consulting is on the cusp of a transformation driven by Generative AI (GenAI) and semantic knowledge graphs. By integrating these technologies, AppSec services companies… - 2025-06-08 — [Using Presentations Instead of CVs in Hiring](../2025/06/08/using-presentations-instead-of-cvs-in-hiring.md): Conventional resumes or CVs often provide a flat, one-dimensional snapshot of a candidate. They are typically static lists of past job titles, dates, and bullet-point skills – which can be misleading… - 2025-06-08 — [Personalized Briefing: Semantic Knowledge Graphs – Intersection of Dinis Cruz & Kerstin Clessienne's Work](../2025/06/08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md): Kerstin Clessienne is an internationally experienced advisor in applied marketing technology and organizational strategy, who explicitly brands herself as a “Marketing Business Graph and Business AI… - 2025-06-07 — [History and Analysis of OWASP In-Person Summits](../2025/06/07/history-and-analysis-of-owasp-in-person-summits.md): The OWASP in-person summits are intensive, collaborative gatherings of the Open Web Application Security Project’s community. These summits (distinct from regular OWASP conferences) bring together… - 2025-06-07 — [GenAI Legacy Code Refactoring – Business Plan](../2025/06/07/genai-legacy-code-refactoring-business-plan.md): GenAI Legacy Code Refactoring is a SaaS company dedicated to modernizing and refactoring legacy software codebases using generative AI (GenAI), human expertise, and automated workflows. Our platform… - 2025-06-06 — [Personalised Briefing for Dan Raywood on the Future of News](../2025/06/06/personalised-briefing-for-dan-raywood-on-the-future-of-news.md): This executive briefing is a synthesis of hundreds of pages of research by Dinis Cruz (primarily from The Future of News and the Monetisation of Trust report and related documents on… - 2025-06-03 — [Proposal: Strategic AWS Partnership with Dinis Cruz’s GenAI and Graph Innovations](../2025/06/03/proposal-strategic-aws-partnership-with-dinis-cruz-genai-and-graph-innovations.md): Dinis Cruz is a seasoned cybersecurity technologist and researcher who has spent over six years building advanced solutions on Amazon Web Services (AWS). He has a proven track record of leveraging… - 2025-06-03 — [Proposal for Neo4j Collaboration with Dinis Cruz](../2025/06/03/proposal-for-neo4j-collaboration-with-dinis-cruz.md): Dinis Cruz is a seasoned cybersecurity expert and innovator with a long-standing relationship with Neo4j, marked by years of advocacy for graph technologies. He brings over two decades of experience… - 2025-06-03 — [Jira as a Graph Database – Proposal for Atlassian Executives](../2025/06/03/jira-as-a-graph-database–proposal-for-atlassian-executives.md): Dinis Cruz is a seasoned technology leader and Atlassian expert with over a decade of experience pushing Jira beyond traditional use cases. He has pioneered research and practical implementations of… - 2025-06-02 — [Linking Threat Models with Semantic Business Graphs](../2025/06/02/linking-threat-models-with-semantic-business-graphs.md): Threat modeling is traditionally a technical exercise: security architects enumerate assets, diagram data flows, identify threats/vulnerabilities, and recommend mitigations. Often, the output is a… ## May 2025 - 2025-05-30 — [Using Threat Modeling and Semantic Graphs to Secure the Digital Supply Chain](../2025/05/30/using-threat-modeling-and-semantic-graphs-to-secure-the-digital-supply-chain.md): Modern organizations rely on complex digital supply chains comprising not only their own systems, but countless third-party components and service providers. Each link in this chain – from… - 2025-05-30 — [Scaling Supply Chain Security using Threat Modeling Semantic Knowledge Graphs and Maps](../2025/05/30/scaling-supply-chain-security-using-threat-modeling-semantic-knowledge-graphs-and-maps.md): In today’s interconnected software ecosystem, digital supply chain security has become a critical challenge. Modern applications rely on countless third-party components and services, yet buyers and… - 2025-05-30 — [Graphs of Graphs of Graphs (G3) in Threat Modeling](../2025/05/30/graphs-of-graphs-of-graphs-g3-in-threat-modeling.md): Modern cybersecurity architectures are highly complex, with myriad components, third-party services, and evolving threat landscapes. Traditional threat modeling approaches – often static diagrams or… - 2025-05-29 — [Threat Models as Mandatory Disclosures: A Vision for Security Transparency](../2025/05/29/threat-models-as-mandatory-disclosures__a-vision-for-security-transparency.md): Argues that publishing threat models improves security transparency. - 2025-05-29 — [Semantic Knowledge Graphs for LLM-Driven Source Code Analysis](../2025/05/29/semantic-knowledge-graphs-for-llm-driven-source-code-analysis.md): Using knowledge graphs and large language models for analyzing source code. - 2025-05-29 — [Advancing Threat Modeling with Semantic Knowledge Graphs](../2025/05/29/advancing-threat-modeling-with-semantic-knowledge-graphs.md): Shows how semantic knowledge graphs enhance threat modeling practices. - 2025-05-27 — [LETS (Load, Extract, Transform, Save): A Deterministic and Debuggable Data Pipeline Architecture](../2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md): Outline of a deterministic, debuggable data pipeline architecture. - 2025-05-22 — [Technical Debrief: Evolution from Electron‑Based to Python‑Native Web Content Capture App](../2025/05/22/technical-debrief__evolution-from-electron-based-to-python-native-web-content-capture-app.md): Lessons learned migrating a web capture tool from Electron to Python. - 2025-05-21 — [Project: Electron-Based Web Content Capture App (with Playwright & Python)](../2025/05/21/project__electron-based-web-content-capture-app-with-playwright-and-python.md): Following the failed attempt to implement the Project: Web Content Capture Extension with Pyodide and Serverless Backend this version creates a desktop app - 2025-05-18 — [Security Debrief: OpenAI’s ChatGPT Connector GitHub App](../2025/05/18/security-debrief__openai_chatgpt_connector_gitHub_app.md): Analyzes security implications of the ChatGPT Connector GitHub app. - 2025-05-18 — [Project: Web Content Capture Extension with Pyodide and Serverless Backend](../2025/05/18/project__web-content-capture-extension-with-pyodide-and-serverless-backend.md): Prototype browser extension leveraging Pyodide and a serverless backend to capture webpage content, including a debrief on why the initial approach failed. - 2025-05-18 — [OAuth Security Concerns and Implications for the Model Context Protocol (MCP)](../2025/05/18/oauth-security-concerns-and-implications-for-the-model-context-protocol.md): Discusses OAuth risks in relation to the Model Context Protocol. - 2025-05-18 — [Briefing on Canada’s New Minister of Artificial Intelligence and Digital Innovation (vs UK and PT)](../2025/05/18/briefing-on-canada-new-minister-of-artificial-intelligence-and-digital-innovation-vs-uk-and-pt.md): Canada has recently established a new cabinet position: Minister of Artificial Intelligence and Digital Innovation. This unprecedented role signals a heightened focus on AI governance and digital… - 2025-05-07 — [Using GenAI to graph and map your company’s data](../2025/05/07/using-genai-to-graph-and-map-your-companys-data__the-grafter__may_2025.md): Presentation delivered at The Grafter meetup Chapter on the 7th Apr 2025 - 2025-05-04 — [Research - AI-Powered Customer Service Solutions for Multi-Property Airbnb Hosts](../2025/05/04/research__ai-powered-customer-service-solutions-for-multi-property-airbnb-hosts.md): Managing multiple Airbnb properties means juggling constant guest inquiries across different channels. The goal is to respond quickly and accurately 24/7 – which is where AI-powered assistants can… - 2025-05-04 — [Project GenBnB: Enhancing Airbnb Host Workflows with GenAI](../2025/05/04/project-genbnb__enhancing-airbnb-host-workflows-with-gen-ai.md): Explores using generative AI to automate property listings and customer interactions so Airbnb hosts can focus on high-value tasks. - 2025-05-03 — [Project InsightFlow: GenAI-Powered Transformation of Regulatory and News Feeds](../2025/05/03/project-insightflow__genai-powered-transformation-of-regulatory-and-news-feeds.md): Describes a semantic knowledge graph pipeline that ingests regulations and news, using generative AI to deliver personalized client newsletters. ## April 2025 - 2025-04-29 — [Fail Safe, Not Fail Big: Cyber-Security-Inspired Strategies to Prevent the Next Iberian Grid Crisis](../2025/04/29/fail-safe-not-fail-big__cyber-security-inspired-strategies-to-prevent-the-next-iberian-grid-crisis.md): This white paper was authored on April 29, 2025, one day after the Iberian power blackout incident. At the time of writing, no official root cause analysis or detailed technical explanation of the… - 2025-04-23 — [Semantic OWASP: Leveraging GenAI and Graphs to Customise and Scale Security Knowledge](../2025/04/23/semantic_owasp__leveraging_genai_and_garphs_to_customise_and_scale_security_knowledge.md): Presentation delivered at the OWASP London Chapter meeting on the 23rd Apr 2025. - 2025-04-22 — [Navigating the AI Revolution: A University Student’s Guide to Generative AI in Education](../2025/04/22/navigating-the-ai-revolution__a_university_students_guide_to_generative-ai-in-education.md): Generative Artificial Intelligence (AI) – especially large language models (LLMs) like GPT-based tools – is transforming how we learn, research, and communicate. As a university student, you are at… - 2025-04-22 — [Graph-Powered Legal Knowledge: An Open, Distributed, and GenAI-Assisted Roadmap](../2025/04/22/graph-powered-legal-knowledge__an-open-distributed-and-ai-assisted-roadmap.md): Legal systems around the world are growing in complexity and volume, yet the data of law is often trapped in static documents and siloed systems. This paper presents a roadmap for a graph-powered… - 2025-04-21 — [Strengthening Trust in News: Implementing Identity Graphs for Authors and Sources](../2025/04/21/strengthening-trust-in-news__implementing-identity-graphs-for-authors-and-sources.md): We propose a newsroom‑wide adoption of Identity Graphs: dynamic, data‑rich profiles that verify and continuously update the credentials, affiliations, and past media appearances of both authors and… - 2025-04-10 — [Project Cybersage: AI-Powered Risk Contextualization & Security Reporting](../2025/04/10/project-cybersage__ai-powered-risk-contextualization_security-reporting.md): Proposal for an AI platform that contextualizes vulnerability data and generates clear, risk-based security reports for technical and executive audiences. - 2025-04-10 — [Project Agenda: GenAI-Powered Transformation of Meetings and Documentation](../2025/04/10/project-agenda__gen-ai-powered-transformation-of-meetings-and-documentation.md): Introduces a GenAI-driven workflow that prepares meetings, generates personalized briefs and tracks action items to reduce wasted time. - 2025-04-07 — [Scaling a Solo Cybersecurity Consulting Practice - Business Plan Research](../2025/04/07/scaling-a-solo-cybersecurity-consulting-practice__business-plan-research.md): In summary, current market conditions strongly favor cybersecurity consulting and training services. A confluence of factors – relentless cyberattacks, a shortage of experienced talent, new tech… - 2025-04-06 — [Think Different, Again: Reimagining Apple’s Role in the AI Era](../2025/04/06/think-different-again__reimagining-apple-role-in-the-ai-era.md): Apple Inc. stands at a crossroads in the spring of 2025. The company that revolutionized personal computing, smartphones, and wearable technology now faces perhaps its greatest challenge: defining… - 2025-04-06 — [Enhancing Cybersecurity Event Networking with Semantic Knowledge Graphs](../2025/04/06/enhancing-cybersecurity-event-networking-with-semantic-knowledge-graphs.md): Cybersecurity conferences and trade events thrive on the connections made between attendees, vendors, and presenters. Yet too often, who meets whom is left to chance. Even as attendee networking… - 2025-04-02 — [The Future of News Monetization: Embracing Micro and Nano Payments](../2025/04/02/the-future-of-news-monetization__embracing-micro-and-nano-payments.md): News organizations worldwide face a dual challenge: sustaining revenue in the digital age while maintaining user engagement and trust. Traditional subscription and advertising models are straining… - 2025-04-02 — [Semantic OWASP: Leveraging GenAI and Graphs to Customise and Scale Security Knowledge](../2025/04/02/semantic-owasp__leveraging-genai-and-graphs-to-customise-and-scale-security-knowledge.md): See also this slides and video presented at OWASP's London Chapter Meeting - 2025-04-02 — [Maturity Models vs. Traditional Standards in Application Security](../2025/04/02/maturity-modes-vs-traditional-standards-in-application-security.md): Every organization’s security capabilities are different. Traditional standards often assume a uniform set of controls for all teams or applications, which may not reflect reality. Maturity models… - 2025-04-01 — [Intent-Based Feedback Loops in Cloud Environments](../2025/04/01/intent-based-feedback-loops-in-cloud-environments.md): Today, a successful deployment typically yields only a confirmation that resources were created or updated – it does not automatically verify that the user’s higher-level goal was achieved, nor does… - 2025-04-01 — [Europe’s Strategic Opportunity in GenAI: A Deep Dive into Six Defining Trends](../2025/04/01/europe-strategic-opportunity-in-gen-ai__a-deep-dive-into-six-defining-trends.md): A convergence of trends is reshaping the GenAI landscape – models are becoming commoditised and open, systems are shifting from centralized GPU clusters to local devices, and knowledge is moving from… ## March 2025 - 2025-03-31 — [Scaling Europe’s Regulatory Superpower: From Static Cybersecurity Standards to Semantic Graphs](../2025/03/31/scaling-europe-regulatory-superpower.md): White paper on using semantic graphs to manage EU cybersecurity standards. - 2025-03-29 — [From Top-Down to Organic Evolving Graphs, Ontologies, and Taxonomies](../2025/03/29/from-top-down-to-organic-evolving-graphs-ontologies-and-taxonomies.md): Discusses dynamic, community-driven approaches to building knowledge graphs. - 2025-03-24 — [Journalists' Challenges with Digital Content Provenance and Trust](../2025/03/24/journalists-challenges-with-digital-content-provenance-and-trust.md): Explores how semantic graphs build trust in AI-generated news. - 2025-03-03 — [LinkedIn Vault: Professional Data Preservation Service](../2025/03/03/linkedin-vault__professional-data-preservation-servic.md): Automates regular backups of LinkedIn career data to user storage. - 2025-03-01 — [Project StartLLM: UC-01: Pentesting Insights Acceleration (PIA)](../2025/03/01/project-startllm__uc-01__pentesting-insight-acceleration.md): Explores how generative AI can streamline pentesting workflows by summarizing findings and highlighting priorities for faster remediation. - 2025-03-01 — [Project StartLLM: Technical Proposal for 5x GenAI Projects](../2025/03/01/project-startllm__technical-proposal-for-5x-genai-projects.md): Strategic proposal outlining five focused generative AI projects designed to deliver quick wins and accelerate organizational adoption. ## February 2025 - 2025-02-26 — [My Journey Building a GenAI Startup: The Power of MVPs and CI Pipelines - PART 2](../2025/02/26/my-journey-building_a_genai_startup__the-power-of-mvps-and-ci-pipelines__part-2.md): Presentation delivered at the AI Security Collective meetup in London on the 26th Feb 2025 (Part 2) - 2025-02-24 — [An Open-Source Sovereign Cloud for an Open Europe: The Case for a Federated, AI-Enabled, and Multilingual Digital Infrastructure](../2025/02/24/an-open-source-sovereign-cloud-for-an-open-europe.md): The European Union (EU) stands at a critical juncture in defining its digital sovereignty. As reliance on U.S.-based cloud providers like Amazon Web Services (AWS), Microsoft Azure, and Google Cloud… - 2025-02-15 — [Project SupplyShield: GenAI-Driven Supply Chain Risk Management and Compliance](../2025/02/15/project-supplyshield__genai-driven-supply-chain-risk-management-and-compliance.md): Introduces an AI-powered third-party risk platform that uses generative models and knowledge graphs to deliver continuous supply chain compliance. - 2025-02-13 — [Scaling Kubernetes with One-Node Clusters: A Paradigm Shift in Cloud-Native Orchestration](../2025/02/13/scaling-kubernetes-with-one-node-clusters-a-paradigm-shift-in-cloud-native-orchestration.md): Kubernetes has revolutionized container orchestration, providing a powerful abstraction for managing workloads across distributed infrastructure. - 2025-02-13 — [Project Lumos: Serverless JIRA-to-GraphDB XYZ Connector](../2025/02/13/project-lumos__serverless-jira-to-graphdb-xyz-connector.md): Working Backwards plan for an open-source connector that streams JIRA data into GraphDB XYZ using a fully serverless architecture. - 2025-02-12 — [Generative AI and the Future of Learning](../2025/02/12/generative-ai-and-the-future-of-learning.md): Generative AI (GenAI) – think of tools like ChatGPT – is rapidly changing how people learn. Unlike traditional one-size-fits-all teaching, GenAI can tailor lessons to each learner in real time. - 2025-02-11 — [Portuguese as a Programming Language in the AI Era](../2025/02/11/portuguese-as-a-programming-language-in-the-AI-Era.md): As the official language of nine countries and a working language in multiple international organizations, Portuguese wields significant cultural and geopolitical influence. In economic terms, the… - 2025-02-11 — [O2 Platform's MethodStreams (2010 Open Source SAST engine)](../2025/02/11/o2-platforms-methodstreams-2010-open-source-sast-engine.md): The OWASP O2 Platform is an open-source toolkit created by Dinis Cruz as a “new paradigm” for performing and documenting web application security reviews (Dinis Cruz Blog: OWASP O2 Platform). - 2025-02-10 — [Second Stories: From Three Mile Island to Cybersecurity](../2025/02/10/second-stories__from-three-mile-island-to-cybersecurity.md): The concept of "second stories" originates from safety science and the analysis of accidents such as the Three Mile Island (TMI) nuclear incident. In the aftermath of TMI (1979), investigators… - 2025-02-09 — [Project JSync: JIRA Exporter and Synchronization System](../2025/02/09/project-jsync__jira-exporter-and-synchronization-system.md): Serverless pipeline that captures JIRA issue changes in real time and stores them in S3 and GitHub, exposing a FastAPI interface for querying the data. - 2025-02-05 — [The Future of News: Building Trust Through Fact Provenance](../2025/02/05/the-future-of-news-building-trust-through-fact-provenance.md): These principles directly align with The Cyber Boardroom: Personalized News Feed Architecture, where fact provenance underpins the platform’s ability to deliver role-based cybersecurity insights.… - 2025-02-02 — [Monetising Trust and Knowledge: How News Providers can leverage Personalised Semantic Graphs](../2025/02/02/monetising-trust-and-knowledge-for-news-providers.md): The rise of GenAI-driven content consumption is reshaping how news organisations deliver and monetise information. As platforms, search engines, and GenAI-powered assistants increasingly rely on… ## January 2025 - 2025-01-29 — [My Journey Building a GenAI Startup: The Power of MVPs and CI Pipelines - PART 1](../2025/01/29/my-journey-building_a_genai_startup__the-power-of-mvps-and-ci-pipelines__part-1.md): Presentation covering early lessons from building a GenAI startup. ## June 2024 - 2024-06-28 — [Deterministic GenAI Outputs with Provenance (OWASP EU AppSec Lisbon)](../2024/06/28/deterministic-genai-outputs-with-provenance__owasp-appsec-lisbon__gslides.md): Presentation delivered at the Global OWASP AppSec in Lisbon in 28th June 2024 ## February 2024 - 2024-02-22 — [It’s 2024 and, with GenAI, we can finally make AppSec work](../2024/02/22/its-2024-and-with-genai-we-can-finally-make-appsec-work.md): Presentation delivered at the OWASP London Chapter meeting on the 22nd Feb 2024 --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/index.html (markdown twin: /research/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/index.html* # Research Hub > Portal to articles on AI, cybersecurity, knowledge graphs, and innovation. --- Welcome to the Research Hub, where I explore the intersection of AI, cybersecurity, and knowledge management through practical projects and innovative approaches. ## Research Areas ### 🔒 [Cyber-Security & Threat Modeling](cyber-security.md) Next-generation approaches to application security, threat modeling, and security governance in the age of AI. Featuring 25+ articles on semantic threat modeling, supply chain security, AppSec innovation, and ephemeral security architectures. ### 🤖 [AI & Development](development-and-genai.md) Exploring the practical applications of generative AI in software development, from LLM security patterns to deterministic data pipelines. Includes methodologies like IFD (Iterative Flow Development) and No Code Development. ### 🕸️ [Knowledge Graphs](graphs.md) Semantic graphs, ontologies, and their applications in security, AI, and knowledge management. Learn how graphs can transform the way we organize and connect information, including ephemeral graph databases and LLMs as graph systems. ### 📰 [The Future of News](the-future-of-news.md) Building trust in journalism through provenance, identity graphs, and new monetization models. Examining how technology can restore faith in news and create sustainable business models. ### 🇪🇺 [Europe & Learning](europe-and-learning.md) Europe's AI governance, sovereign cloud opportunities, and the transformation of education through generative AI. Includes perspectives on European tech sovereignty and learning methodologies. ### 🚀 [Projects & Innovation Lab](projects.md) GenAI-driven tools, MVPs, and business ventures at the cutting edge. Over 40 active projects showcasing practical applications of emerging technologies, from VulnAI to strategic partnerships. ### 🏢 [Organizational Transformation](organizational-transformation.md) How organizations can leverage AI, semantic graphs, and modern methodologies to transform their operations, from hiring practices to meeting efficiency. ## Latest Research (Summer 2025) Check out the most recent additions to the research hub: ### August 2025 - [Iterative Flow Development (IFD) Methodology](../2025/08/22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.md) - Revolutionary approach to AI-assisted development achieving 10-20x productivity gains - [Surrogate Dependencies: Simulating Backends for Offline-First Development](../2025/08/22/surrogate-dependencies-simulating-backends-for-offline-first-development.md) - Enabling fully offline development with prerecorded API data ### July 2025 - [Ephemeral GenAI SIEM: A Serverless, Graph-Driven Approach](../2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md) - Revolutionizing security event management - [Project VulnAI: AI-Powered Vulnerability Risk Management Platform](../2025/07/27/project-vulnai-ai-powered-vulnerability-risk-management-platform.md) - Next-generation vulnerability management ### June 2025 - [User-Driven Semantic Persona Graphs Powered by GenAI](../2025/06/14/user-driven-semantic-persona-graphs-powered-by-genai.md) - Interactive knowledge graph creation - [Strategic Partnership Proposals](projects.md#partnerships) - AWS, Neo4j, and Atlassian collaboration opportunities ## Research Philosophy My research focuses on practical, implementable solutions that bridge the gap between cutting-edge technology and real-world business needs. Every project and article aims to: - **Demonstrate feasibility** through working prototypes and MVPs - **Share knowledge openly** via open-source code and detailed documentation - **Connect domains** by finding synergies between security, AI, and business - **Enable others** by providing frameworks and tools others can build upon ## Monthly Archives - [August 2025](../2025/08/index.md) - IFD methodology and offline development - [July 2025](../2025/07/index.md) - Ephemeral architectures and sustainable AI - [June 2025](../2025/06/index.md) - Semantic graphs and strategic partnerships ## Get Involved - 💬 Connect on [LinkedIn](https://linkedin.com/in/diniscruz) to discuss research topics - 🐙 Explore code on [GitHub](https://github.com/diniscruz) - 📧 Reach out for collaboration opportunities - 📚 Use and extend the open-source tools featured in the research --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/06/index.html (markdown twin: /2025/06/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/06/index.html* # June 2025 Published Materials > During June 2025, I published 29 comprehensive research documents spanning artificial intelligence, cybersecurity, organizational transformation, and strategic innovation. The month's publications… --- During June 2025, I published 29 comprehensive research documents spanning artificial intelligence, cybersecurity, organizational transformation, and strategic innovation. The month's publications show a concentrated focus on semantic knowledge graphs and their applications (10 documents including user-driven personas, LLMs as graph databases, and graph thinking), ephemeral computing and testing methodologies (3 documents on ephemeral testing and Neo4j instances), and strategic partnership proposals with major technology companies (3 documents on AWS, Neo4j, and Atlassian collaboration proposals). Additional research covered GenAI applications in business transformation (5 documents on legacy code refactoring, workshops, and no-code development), cybersecurity frameworks (4 documents on threat modeling and security implications), and various topics including OWASP history, hiring practices, news media transformation, and personal content rights, demonstrating the breadth of my research interests and expertise. My research this month particularly emphasizes the intersection of AI technologies with practical business applications, from threat modeling services enhanced by GenAI to legacy code refactoring solutions. A recurring theme is the evolution from static systems to dynamic, graph-based architectures that can adapt and scale with organizational needs. The strategic partnership proposals with AWS, Neo4j, and Atlassian highlight my vision for collaborative innovation in serverless computing, graph databases, and knowledge management systems. The later part of the month saw significant work on ephemeral Neo4j instances and data testing methodologies, showcasing practical implementations of theoretical concepts in graph database management. Throughout June, I consistently explored how semantic knowledge graphs can bridge technical complexity with business value, whether through personalized news feeds, cybersecurity risk modeling, or user-driven persona development. ## Overview Table | Date | Title | Focus Area | Key Concepts | |------|-------|------------|--------------| | 06/02 | [Linking Threat Models with Semantic Business Graphs](02/linking-threat-models-with-semantic-business-graphs.md) | Cyber Security | Semantic Business Graphs, Threat Modeling, Business Context, Knowledge Graphs, Two-Way Influence | | 06/03 | [Jira as a Graph Database – Proposal for Atlassian Executives](03/jira-as-a-graph-database_proposal-for-atlassian-executives.md) | Projects | Jira Graph Database, Issue Links, Semantic Relationships, JSync, Project Lumos | | 06/03 | [Proposal for Neo4j Collaboration with Dinis Cruz](03/proposal-for-neo4j-collaboration-with-dinis-cruz.md) | Projects | Neo4j Partnership, Serverless Graph DB, MGraph-AI, Knowledge Graphs, GraphRAG | | 06/03 | [Proposal: Strategic AWS Partnership with Dinis Cruz's GenAI and Graph Innovations](03/proposal-strategic-aws-partnership-with-dinis-cruz-genai-and-graph-innovations.md) | Projects | AWS Serverless, GenAI Integration, OSBot-AWS, MGraph-AI, Amazon Bedrock | | 06/06 | [Personalised Briefing for Dan Raywood on the Future of News](06/personalised-briefing-for-dan-raywood-on-the-future-of-news.md) | The Future of News | Micropayments, Trust-as-a-Service, Semantic Graphs, Personalized News Feeds, Content Monetization | | 06/07 | [GenAI Legacy Code Refactoring – Business Plan](07/genai-legacy-code-refactoring-business-plan.md) | Projects | Legacy Code Modernization, GenAI Automation, Bug-First Testing, Serverless Architecture, SaaS Model | | 06/07 | [History and Analysis of OWASP In-Person Summits](07/history-and-analysis-of-owasp-in-person-summits.md) | Cyber Security | OWASP Summits, Collaborative Security, Working Sessions, Community Building, Knowledge Sharing | | 06/08 | [Personalized Briefing: Semantic Knowledge Graphs – Intersection of Dinis Cruz & Kerstin Clessienne's Work](08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md) | Graphs | Semantic Knowledge Graphs, Marketing Technology, Self-Evolving Graphs, Wardley Maps, Trust Networks | | 06/08 | [Using Presentations Instead of CVs in Hiring](08/using-presentations-instead-of-cvs-in-hiring.md) | Europe and Learning | Alternative Hiring Methods, Personal Presentations, Diversity in Recruitment, Creative Assessment, Talent Discovery | | 06/09 | [Supercharging AppSec Threat Modeling Services with GenAI and Semantic Graphs](09/supercharging-appsec-threat-modeling-services-with-genai-and-semantic-graphs.md) | Cyber Security | AppSec Services, GenAI Threat Modeling, Semantic Graphs, Multi-Stakeholder Reporting, AI-Assisted Analysis | | 06/10 | [Explorers, Villagers, and Town Planners: Understanding the Generative AI Divide](10/explorers-villagers-and-town-planners-understanding-the-generative-ai-divide.md) | Development and GenAI | EVTP Framework, Wardley Maps, Innovation Adoption, GenAI Perspectives, Technology Evolution | | 06/10 | [Security Implications of the Model Context Protocol (MCP) and the Need for Robust Infrastructure](10/security-implications-of-the-model-context-protocol-mcp-and-the-need-for-robust-infrastructure.md) | Cyber Security | MCP Security, LLM Tool Integration, Prompt Injection, Identity Management, Zero Trust Architecture | | 06/13 | [Follow-Up Technical Vision: Optimizations, Deployment, and Security](13/follow-up-technical-vision-optimizations-deployment-and-security.md) | Projects | Web Content Filtering, Performance Optimization, Multi-Tenant Architecture, Serverless Deployment, Cache Strategies | | 06/13 | [Technical Briefing: Web Content Filtering Project](13/technical-briefing-web-content-filtering-project.md) | Projects | Web Content Filtering, Real-Time Modification, Proxy Architecture, LLM Processing, Personalization | | 06/14 | [User-Driven Semantic Persona Graphs Powered by GenAI](14/user-driven-semantic-persona-graphs-powered-by-genai.md) | Graphs | Persona Graphs, GenAI-Driven Q&A, Interactive Workflows, Knowledge Graph Creation, Human-in-the-Loop | | 06/15 | [Personal Content Rights: Protecting Individuals in the Age of Deepfakes and AI Cloning](15/personal-content-rights-protecting-individuals-in-the-age-of-deepfakes-and-ai-cloning.md) | Cyber Security | Personal Content Rights, Deepfake Protection, AI Ethics, Digital Identity, Watermarking | | 06/15 | [The Hidden Cost of Ephemeral Testing and the Case for Automation](15/the-hidden-cost-of-ephemeral-testing-and-the-case-for-automation.md) | Development and GenAI | Ephemeral Testing, Test Automation, ETDD, Wardley Maps, Development Velocity | | 06/16 | [Workshop Plan: User-Driven Semantic Persona Graphs Powered by GenAI](16/workshop-plan-user-driven-semantic-persona-graphs-powered-by-genai.md) | Development and GenAI | Workshop Design, GPT Pipeline, Multi-Phase Architecture, Vibe Coding, Interactive Learning | | 06/18 | [Bridging Niklas Luhmann's Ideas with Semantic Knowledge Graphs and G³](18/bridging-niklas-luhmanns-ideas-with-semantic-knowledge-graphs-and-g3.md) | Graphs | Zettelkasten Method, Knowledge Management, G³ Framework, Niklas Luhmann, Personal Knowledge Systems | | 06/18 | [Comparing the EU FED Cloud vs. an Open-Source Federated Cloud Proposal](18/comparing-the-eu-fed-cloud-vs-an-open-source-federated-cloud-proposal.md) | Development and GenAI | EU FED Cloud, Federated Architecture, Open-Source Cloud, ArQiver Platform, Data Sovereignty | | 06/18 | [Empowering the Graph Thinkers in the Age of Generative AI](18/empowering-the-graph-thinkers-in-the-age-of-generative-ai.md) | Graphs | Graph Thinking, No-Code Development, LLMs as Graph Databases, MGraph-DB, Knowledge Democratization | | 06/18 | [No Code Development (NCD): A Paradigm Shift Beyond 'Vibe Coding'](18/no-code-development--ndc--a-paradigm-shift-beyond-vibe-coding.md) | Development and GenAI | No Code Development, Vibe Coding, AI-Assisted Development, Natural Language Programming, ThreatModCon 2025 | | 06/19 | [Empowering Workshops with Custom GPTs for GenAI Training](19/empowering-workshops-with-custom-gpts-for-genai-training.md) | Development and GenAI | Custom GPTs, Workshop Training, GPT Builder, Knowledge Base Integration, Interactive Learning | | 06/19 | [Using LLMs as Ephemeral Graph Databases](19/using-llms-as-ephemeral-graph-databases--empowering-the-graph-thinkers-in-the-age-of-generative-ai.md) | Graphs | LLMs as Databases, Ephemeral Graphs, Natural Language Queries, Cybersecurity Risk Modeling, Graph CRUD Operations | | 06/22 | [FIST Meets the Semantic Knowledge Graph](22/fist-meets-the-semantic-knowledge-graph-aligning-fast-inexpensive-simple-tiny-with-dinis-cruzs-g3-approach.md) | Cyber Security | FIST Framework, Fast/Inexpensive/Simple/Tiny, Defense Acquisition, Agile Development, G³ Alignment | | 06/22 | [Project Plan: High Street GenAI Learning Hub](22/project-plan-high-street-genai-learning-hub.md) | Cyber Security | Community Learning, GenAI Education, Third Space Concept, Digital Inclusion, Social Innovation | | 06/25 | [Data Tests for Neo4j: Bringing Automated Testing to Graph Databases](25/data-tests-for-neo4j-bringing-automated-testing-to-graph-databases.md) | Graphs | Data Testing, Graph Database QA, PyTest Integration, CI/CD Pipelines, Schema Validation | | 06/25 | [Ephemeral Neo4j Instances for On-Demand Graph Analytics](25/ephemeral-neo4j-instances-for-on-demand-graph-analytics.md) | Graphs | Ephemeral Databases, Serverless Neo4j, AWS Architecture, On-Demand Analytics, Cost Optimization | | 06/25 | [Using Ephemeral Neo4j Instances for a Cybersecurity Risk Graph Scenario](25/using-ephemeral-neo4j-instances-for-a-cybersecurity-risk-graph-scenario.md) | Graphs | Risk Graph Implementation, Cypher Queries, Ephemeral Setup, Threat Modeling, Incident Response | ## Detailed Summaries ### Linking Threat Models with Semantic Business Graphs *June 2, 2025* This document proposes [Semantic Business Graphs](02/linking-threat-models-with-semantic-business-graphs.md) as a revolutionary approach to bridging security analysis with business context. The research demonstrates how organizations can map their entire business landscape – including goals, processes, organizational structures, finances, and compliance obligations – into an interconnected knowledge graph that directly links with threat modeling artifacts. This creates a living model where security analyses are always performed in light of real business impact, ensuring that cyber threats are evaluated not in isolation but in terms of their effect on customer trust, revenue, and strategic goals. The implementation involves creating "graphs of graphs" that connect multiple ontologies and taxonomies, enabling complex queries like "Which critical business services would be impacted by a Log4J vulnerability?" to be answered in seconds. The paper details how this approach creates a two-way influence between business decisions and threat models, where business changes automatically flag relevant threat models for updating, and security findings visibly link to business objectives. The research includes practical examples from a SaaS provider with GenAI capabilities, showing how the semantic graph serves different stakeholders from executives to security teams to external auditors. ### Jira as a Graph Database – Proposal for Atlassian Executives *June 3, 2025* This comprehensive proposal presents [Jira as a Graph Database](03/jira-as-a-graph-database_proposal-for-atlassian-executives.md), showcasing over a decade of Dinis Cruz's experience in pushing Jira beyond traditional use cases. The document demonstrates how Jira's built-in issue linking and customization capabilities can model complex networks of information, effectively transforming it into a powerful knowledge graph where each issue serves as a node and each issue link represents a relationship edge. The proposal includes concrete examples from implementations at Photobox Group and Holland & Barrett, where Jira was successfully used to manage risk data and workflows as graph-backed systems. The technical approach involves structuring Jira projects so that each represents a distinct node type in the graph, with custom link types encoding semantic relationships between entities. The proposal introduces several innovative tools including JSync (a Jira synchronization system that exports data for offline graph queries), Project Lumos (a serverless Jira-to-GraphDB connector), and MGraph-AI (a memory-first graph database for AI workloads). These solutions address Jira's current limitations while demonstrating how Atlassian could unlock new capabilities in advanced analytics, knowledge graphs, and AI integrations through strategic collaboration. ### Proposal for Neo4j Collaboration with Dinis Cruz *June 3, 2025* This strategic partnership proposal outlines opportunities for [Neo4j collaboration](03/proposal-for-neo4j-collaboration-with-dinis-cruz.md) based on Cruz's extensive history with graph technologies and Neo4j specifically. The document traces a timeline from 2017 to 2025, detailing Cruz's pioneering work including the use of Neo4j for GDPR data flow mapping at Photobox and his leadership in sessions on "Ideas for Graph DBs like Neo4j" at the Open Security Summit. The proposal highlights current innovations including MGraph-AI, a serverless graph database designed for AI and serverless environments, and MyFeeds.ai, which demonstrates LLM-powered semantic news feeds using knowledge graphs. The collaboration opportunities span multiple areas including serverless Neo4j implementations, knowledge graphs for GenAI applications, Jira-Neo4j integration for DevSecOps, and cybersecurity knowledge graph solutions. The proposal emphasizes The Cyber Boardroom pilot as a flagship opportunity, where Neo4j could serve as the core knowledge graph repository for a GenAI-powered platform that bridges cybersecurity expertise with corporate board decision-making. Each proposed initiative aligns with Neo4j's strategic direction in AI, cloud deployment models, and enterprise knowledge graphs. ### Proposal: Strategic AWS Partnership with Dinis Cruz's GenAI and Graph Innovations *June 3, 2025* This proposal presents a [strategic AWS partnership](03/proposal-strategic-aws-partnership-with-dinis-cruz-genai-and-graph-innovations.md) opportunity based on Cruz's six years of building advanced solutions on Amazon Web Services. The document showcases proven implementations of serverless GenAI and graph-based applications, including the open-source MGraph-AI graph database and OSBot-AWS automation toolkit, which demonstrate capabilities that AWS's current offerings don't natively provide. The research highlights how Cruz identified and filled the gap for a truly serverless graph database optimized for Lambda-based use with zero cost when idle. The collaboration proposals include co-developing a serverless graph database offering, featuring Cruz's work as AWS case studies for GenAI and serverless best practices, sponsoring open-source integration development, and joint solution offerings in the AWS Marketplace. The document details flagship projects like The Cyber Boardroom (a serverless GenAI platform for cybersecurity decision-making) and MyFeeds.ai (personalized semantic news feeds), both built entirely on AWS's serverless stack. These innovations align directly with AWS's priorities in serverless computing and GenAI, including planned integration with Amazon Bedrock. ### Personalised Briefing for Dan Raywood on the Future of News *June 6, 2025* This executive briefing synthesizes extensive research on [the future of news and monetization of trust](06/personalised-briefing-for-dan-raywood-on-the-future-of-news.md) specifically tailored for Dan Raywood, Senior Editor at SC Media UK. The document presents innovative proposals for diversifying revenue through micro and nano payments, where readers pay tiny amounts for individual pieces of content rather than committing to full subscriptions. The briefing emphasizes how modern technology has removed barriers to micropayments, enabling seamless one-click payments that align monetary incentives with truth, transparency, and trust in journalism. The research introduces the concept of "Trust-as-a-Service," where news organizations can monetize their credibility and verification capabilities through real-time Verification APIs, credibility scoring services, and expert networks. The document details how personalized news feeds powered by semantic graphs can deliver tailored content while maintaining full provenance and transparency. For SC Media UK specifically, the briefing outlines opportunities to pilot custom cybersecurity news feeds, leverage their reputation through trust metrics and verification badges, and integrate content into cross-publisher micropayment systems. ### GenAI Legacy Code Refactoring – Business Plan *June 7, 2025* This comprehensive business plan presents a [GenAI Legacy Code Refactoring](07/genai-legacy-code-refactoring-business-plan.md) SaaS company dedicated to modernizing legacy software codebases using generative AI, human expertise, and automated workflows. The document outlines a three-phase process: adding comprehensive test coverage and documentation, ensuring continuous integration pipelines run smoothly, and performing AI-assisted code refactoring under the safety net of rigorous tests. The approach addresses the universal pain point that organizations allocate 60-80% of IT budgets to maintaining legacy systems, with technical debt averaging $361k per 100,000 lines of code. The business model operates on pure SaaS with token-based pricing adding 20% markup to all model costs, ensuring profitability on every project. The solution leverages bug-first testing approaches, where known bugs have passing tests that capture their current behavior, providing meaningful safety nets for refactoring. The plan includes pragmatic success metrics measured through OKRs rather than absolute claims, human-in-the-loop services for expertise coordination, and a serverless architecture that keeps costs minimal while enabling global scalability. ### History and Analysis of OWASP In-Person Summits *June 7, 2025* This comprehensive historical analysis documents all [OWASP in-person summits](07/history-and-analysis-of-owasp-in-person-summits.md) from 2008 to 2017, providing detailed insights into how these intensive collaborative gatherings shaped the application security community. The research covers four major summits: the 2008 European Summit in Algarve (80 participants), the 2009 Washington D.C. leadership summit, the 2011 Global Summit in Lisbon (175-180 participants from 20+ countries), and the 2017 Summit at Woburn Forest, UK. Each summit is analyzed for its planning, format, key topics, outcomes, and lasting impact on OWASP's evolution. The document reveals how summit formats evolved from conference-style presentations to pure working sessions, emphasizing collaboration over passive learning. Key outcomes included the establishment of OWASP's core principles and code of ethics (2008), the creation of six global committees, governance reforms introducing member-inclusive board elections (2011), and the acceleration of projects like OWASP SAMM and the Mobile Security Testing Guide (2017). The analysis demonstrates how these summits served as catalysts for organizational change, community building, and the advancement of application security practices globally. ### Personalized Briefing: Semantic Knowledge Graphs – Intersection of Dinis Cruz & Kerstin Clessienne's Work *June 8, 2025* This personalized briefing explores the [intersection of research interests](08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md) between Dinis Cruz and Kerstin Clessienne, both thought leaders in semantic knowledge graphs. The document highlights how Clessienne, a marketing technology expert, advocates for knowledge graph-driven approaches to improve customer insight and personalization, while Cruz has developed self-evolving semantic knowledge graph systems for capturing complex knowledge and generating tailored content. Both recognize that integrating knowledge graphs with AI is a game-changer for creating truly personalized and context-aware experiences. The briefing reveals key convergences in their work: both see semantic enrichment as critical for AI to move beyond pattern-matching, both address organizational data silos through unified knowledge graphs, and both use graphs to drive personalized content and insights. Cruz's unique contribution of evolving graphs into Wardley Maps for strategic insight complements Clessienne's focus on marketing strategy and intelligent transformation. The document provides resources for deeper collaboration, including links to Cruz's research repository, MyFeeds.ai architecture details, and key articles on self-improving knowledge graphs and the progression from metadata to stories. ### Using Presentations Instead of CVs in Hiring *June 8, 2025* This document presents an innovative approach to recruitment where [candidates create presentations instead of traditional CVs](08/using-presentations-instead-of-cvs-in-hiring.md), based on Dinis Cruz's successful implementations at Glasswall and Holland & Barrett. The research demonstrates how traditional resumes provide only flat, one-dimensional snapshots that can be misleading, invite unconscious bias, and fail to capture true skills or potential. In contrast, personal slide decks allow candidates to tell the story of their professional journey, showcase actual work with visuals, and demonstrate creativity and communication skills. The case studies show remarkable success: at Glasswall, Petra Vukmirović, an emergency medicine doctor transitioning to cybersecurity, created a 10-slide presentation outlining her self-learning roadmap that led to her hiring and subsequent success in the field. At Holland & Barrett, the security team formally integrated candidate presentations into recruitment, with HR buy-in after seeing improved hiring outcomes. The approach promotes diversity by allowing candidates with non-traditional backgrounds to explain their unique paths and transferable skills, while making interviews more engaging through structured discussions around prepared content. ### Supercharging AppSec Threat Modeling Services with GenAI and Semantic Graphs *June 9, 2025* This white paper outlines how [AppSec consulting services can be transformed](09/supercharging-appsec-threat-modeling-services-with-genai-and-semantic-graphs.md) through the integration of Generative AI and semantic knowledge graphs. The document demonstrates how LLMs can rapidly produce, customize, and maintain threat models at scale, generating hundreds or thousands of security documents in hours rather than weeks. By representing threats, assets, and mitigations as nodes in a semantic graph linked to business context, static threat models become living knowledge bases that are queryable and continuously updatable. The proposed service offerings include AI-Augmented Threat Modeling Services that deliver 5x more threat scenarios than manual approaches, Knowledge Graph Integration with custom security dashboards, Multi-Stakeholder Reporting Bundles that generate tailored outputs for different audiences, and AI-Driven Secure Code Review combining LLM analysis with human expertise. The implementation leverages techniques like mass-generation of threat models (demonstrated with Google's Gemini 2.0 generating 1,000 models), customization for specific tech stacks, and open schema outputs for seamless integration. The paper emphasizes that this approach enables consultants to scale their expertise, deliver more value with less effort, and differentiate themselves in a competitive market. ### Explorers, Villagers, and Town Planners: Understanding the Generative AI Divide *June 10, 2025* This opinion piece applies Simon Wardley's [EVTP framework to explain the divide](10/explorers-villagers-and-town-planners-understanding-the-generative-ai-divide.md) in the GenAI community between enthusiastic explorers who see unlimited potential and cautious town planners who emphasize current limitations. The document explains how Explorers thrive in Genesis phases with high uncertainty, celebrating what works while treating failures as learning opportunities, while Town Planners excel in Commodity phases requiring reliability, consistency, and safety at scale. Both perspectives are valid and necessary within their appropriate contexts. The analysis reveals that Explorers focus on GenAI's magical possibilities, embracing "vibe coding" and predicting imminent disruption of traditional development, while Town Planners highlight serious concerns including hallucinations, lack of determinism, scalability issues, and security risks. The paper argues that both camps are brilliant and correct from their vantage points – Explorers measure by potential and breakthroughs, Town Planners by worst failures and limitations. The solution lies in recognizing when each mindset should dominate, fostering communication between camps, and investing in bridging solutions that gradually transform exploratory prototypes into production-ready systems. ### Security Implications of the Model Context Protocol (MCP) and the Need for Robust Infrastructure *June 10, 2025* This paper examines the [security implications of MCP](10/security-implications-of-the-model-context-protocol-mcp-and-the-need-for-robust-infrastructure.md), an open standard enabling LLMs to interface with external tools and data sources uniformly. While MCP provides a solid foundation for LLM tool integration (acting as a "USB-C for AI"), its success highlights that security is only as strong as the surrounding infrastructure. The document analyzes how MCP's strength in seamlessly bridging LLMs with real-world actions can become a massive vulnerability if traditional security controls are inadequate. The research details new attack surfaces including prompt injection exploits (like the GitHub MCP exploit discovered in May 2025), malicious tool poisoning, and rug pull attacks. Current infrastructure falls short due to coarse authorization and over-privilege, lack of agent-specific identities, authentication gaps in MCP, uniform trust for all tools, and limited observability. The paper proposes solutions including fine-grained permissions and sandboxing, strong authentication with cryptographic tool verification, runtime monitoring for "toxic flow" detection, human-in-the-loop confirmation gates, and graph-based threat modeling for continuous analysis. The conclusion emphasizes that robust security must be engineered at the system level, not relying on AI model alignment alone. ### Follow-Up Technical Vision: Optimizations, Deployment, and Security *June 13, 2025* This technical document provides detailed [optimization strategies, deployment options, and security considerations](13/follow-up-technical-vision-optimizations-deployment-and-security.md) for a web content filtering platform that uses AI to intercept and filter disallowed content. The optimization approach leverages minimal overhead in monitoring mode, sophisticated caching of content segments using hash-based deduplication, DOM structure analysis for intelligent segmentation, and parallel LLM processing with configurable cost/performance trade-offs. The system breaks pages into segments, computes unique hashes, and reuses classifications for previously seen content, dramatically reducing processing time for repeat visits. The deployment strategy offers flexibility from multi-tenant SaaS using serverless functions to dedicated cloud deployments with tenant-specific encryption, and even on-premises installations for organizations requiring complete data control. Security measures include strong tenant isolation with encrypted data segregation, safe handling of sensitive content with minimal retention policies, integrity controls to prevent filter bypass, and configurable fail-open versus fail-closed behaviors. The architecture supports customer-provided API keys for LLM services, enabling organizations to maintain control over their data while benefiting from the filtering capabilities. ### Technical Briefing: Web Content Filtering Project *June 13, 2025* This technical briefing outlines the architecture and design of the [Web Content Filtering Project](13/technical-briefing-web-content-filtering-project.md), which aims to give users fine-grained control over the content they see as they browse the web. The project's core idea is to intercept and dynamically modify web pages in real-time, allowing unwanted content to be filtered out and relevant content to be highlighted. By leveraging Large Language Models (LLMs) and knowledge graphs, the system can filter or transform content on the fly without requiring changes from the websites themselves. The implementation showcases several key design principles: personalization via semantic graphs that represent both content and user preferences as structured knowledge, minimal in-line LLM usage to ensure fast browsing after initial processing, deterministic and reproducible results through structured outputs and rigorous data schemas, and comprehensive provenance and explainability where every filtering decision can be audited. The system implements a multi-stage pipeline including page capture via proxy, HTML parsing to typed structures, DOM to graph conversion, text extraction, LLM semantic classification, persona/preference graph matching, and finally page reconstruction with filtered content. The architecture leverages open-source tools like OSBot TypeSafe and MGraph-DB, demonstrating how modern AI can enable personalized web experiences while maintaining transparency and user control. ### User-Driven Semantic Persona Graphs Powered by GenAI *June 14, 2025* This white paper presents a GenAI-powered, [user-driven workflow for creating persona-based knowledge graphs](14/user-driven-semantic-persona-graphs-powered-by-genai.md) on demand, transforming the traditionally static process of gathering information into an engaging dialogue. The system uses an interactive Q&A workflow orchestrated by an LLM that dynamically generates questions, adapting to user responses and building a semantic graph of the user's domain. Each answer the user provides is parsed into structured data that expands the graph with new nodes and edges, while the next questions are informed by the growing graph, allowing a personalized path of inquiry. The technical implementation leverages a serverless architecture using OSBot-FastAPI for API endpoints, OSBot-Utils for flow orchestration, and MGraph-DB as a memory-first graph database optimized for AI workloads. The system incorporates human-in-the-loop feedback where users validate and refine the graph's accuracy, pre-population capabilities using external data sources, and the ability to generate multiple persona-specific outputs from the same base knowledge. Applications span cybersecurity risk profiling, personalized news feeds, corporate compliance auditing, customer onboarding, and education, demonstrating how combining human input with AI assistance yields self-improving knowledge structures that are far more aligned to user reality than generic templates. ### Personal Content Rights: Protecting Individuals in the Age of Deepfakes and AI Cloning *June 15, 2025* This comprehensive white paper proposes establishing [Personal Content Rights](15/personal-content-rights-protecting-individuals-in-the-age-of-deepfakes-and-ai-cloning.md) as a new legal framework to give individuals firm control over their own digital persona in the age of AI-generated deepfakes. The document outlines how generative AI has enabled anyone to clone voices, faces, and writing styles with startling realism, leading to fraud, defamation, and non-consensual pornography on a massive scale. The proposed framework would make it illegal to create or distribute AI-generated content that imitates a real person without explicit authorization, treating digital cloning without consent as a serious rights violation akin to identity theft. Key proposals include outlawing unauthorized deepfakes with stiff penalties, establishing inalienability of persona rights (preventing wholesale selling of one's digital identity), developing robust identity verification frameworks through persona and identity graphs, mandating watermarking and provenance metadata for all AI-generated media, and preserving legitimate fair uses like parody and satire. The paper details technological solutions including watermarking standards like C2PA, deepfake detection services, and identity graph infrastructure to manage permissions at scale. Implementation strategies cover enforcement mechanisms, global adoption patterns following the "Brussels effect" model, and the need for independent oversight bodies to prevent abuse while fostering innovation in ethical AI applications. ### The Hidden Cost of Ephemeral Testing and the Case for Automation *June 15, 2025* This analysis reveals how developers practicing [Ephemeral Test-Driven Development (ETDD)](15/the-hidden-cost-of-ephemeral-testing-and-the-case-for-automation.md) "write" tests in their mind and execute them with their hands, like scribbling on a whiteboard and erasing it after each use. While this ad-hoc approach feels fast in the moment, it is deceptively expensive over time – each manual test is wasted effort that must be repeated, whereas an automated test can run endlessly at virtually no extra cost. The document demonstrates how teams that embrace capturing these checks as code see compound benefits where every new test makes the next code change safer and faster. The paper addresses the "metric trap" where teams write minimal or meaningless tests just to satisfy coverage requirements, arguing that the focus must shift from "writing tests for metrics" to "automating tests for insight." It emphasizes investing in test infrastructure and developer experience, highlighting tools like Wallaby.js and NCrunch that provide real-time continuous testing inside the IDE. The analysis incorporates Wardley Map perspectives to balance speed and quality based on project maturity, showing that while Genesis-phase experiments might justify minimal testing, Product and Commodity phases demand comprehensive test suites. Ultimately, the research proves that automated testing is not a tax on development speed but a powerful enabler of sustainable velocity. ### Workshop Plan: User-Driven Semantic Persona Graphs Powered by GenAI *June 16, 2025* This detailed [workshop plan](16/workshop-plan-user-driven-semantic-persona-graphs-powered-by-genai.md) demonstrates how to use Generative AI to build personalized semantic knowledge graphs and generate tailored outputs for different stakeholder personas. The workshop guides technical participants through a multi-phase pipeline of custom GPT-powered assistants that collaboratively transform user input into meaningful insights, showcasing how complex, context-aware solutions can be built by chaining together specialized GPT agents. The implementation involves six dedicated GPT agents: GPT 1 ingests compliance criteria and designs questionnaires, GPT 2 conducts interactive Q&A interviews, GPT 3 converts raw answers into semantic knowledge graphs, GPT 4 produces comprehensive technical reports, GPT 5 generates persona-specific briefings for different stakeholders, and GPT 6 creates UI prototypes using "vibe coding" techniques. The workshop structure includes a live demonstration of the end-to-end pipeline, hands-on creation of GPTs using provided prompt templates, and discussion of potential improvements like automating data flow between GPTs and serverless deployment. This comprehensive plan with detailed GPT configuration artifacts ensures participants gain concrete understanding of how GenAI can be orchestrated to implement complex workflows, from raw text to knowledge graphs to stakeholder-specific insights and even UI prototypes. ### Bridging Niklas Luhmann's Ideas with Semantic Knowledge Graphs and G³ *June 18, 2025* This briefing introduces [Niklas Luhmann's Zettelkasten system](18/bridging-niklas-luhmanns-ideas-with-semantic-knowledge-graphs-and-g3.md) and maps its principles to modern Semantic Knowledge Graphs and the G³ (Graphs of Graphs of Graphs) approach. Luhmann, a prolific German sociologist, achieved remarkable productivity through his analog knowledge management method – a slip-box containing 90,000 handwritten notes that served as a "second brain" and thinking partner. The document demonstrates how Luhmann's Zettelkasten was essentially a proto-knowledge graph on paper: atomic notes with unique IDs, densely linked via cross-references, emergent in structure rather than hierarchically organized, and scalable over a lifetime of use. The research draws technical parallels between Luhmann's half-century-old system and contemporary knowledge graphs, showing how his approach prefigured modern Personal Knowledge Management systems. Cruz's G³ methodology particularly resonates with Luhmann's philosophy of avoiding a single "master ontology" in favor of multiple interconnected perspectives. The paper explores how Luhmann's practices of building knowledge through incremental connections, allowing structure to emerge organically, and maintaining multiple parallel note collections align with modern approaches to semantic graphs, modularity, and federated knowledge systems. ### Comparing the EU FED Cloud vs. an Open-Source Federated Cloud Proposal *June 18, 2025* This white paper provides a comparative analysis of the [EU FED Cloud initiative versus an open-source federated cloud proposal](18/comparing-the-eu-fed-cloud-vs-an-open-source-federated-cloud-proposal.md). The EU FED Cloud is a European initiative building a federated, sovereign cloud ecosystem with built-in compliance and security, featuring the ArQiver platform that provides compliance, identity, and workflow backbone. It adopts a three-layer architecture designed to enforce zero-trust security and EU regulatory compliance by default, emphasizing data sovereignty, interoperability, and elimination of vendor lock-in. In contrast, the open-source sovereign cloud proposal calls for a European cloud infrastructure that is fully open-source and API-compatible with major clouds like AWS and Azure. This approach prioritizes maximum compatibility, allowing developers to migrate applications with minimal code changes while maintaining sovereignty and transparency. The document analyzes key differences: the EU FED Cloud introduces its own framework and APIs focused on compliance-first design, while the open-source proposal emphasizes compatibility-first with existing cloud services. Both approaches aim to empower Europe with sovereign cloud infrastructure, but they cater to different priorities and implementation philosophies. ### Empowering the Graph Thinkers in the Age of Generative AI *June 18, 2025* This white paper explores how [graph thinkers – those who naturally conceptualize information as networks](18/empowering-the-graph-thinkers-in-the-age-of-generative-ai.md) – are being empowered by advances in generative AI. Historically, these creative minds were constrained by technical limitations, requiring heavy implementation, dedicated development teams, or costly graph database infrastructure to realize their visions. The document demonstrates how Large Language Models and AI-assisted development platforms are democratizing the ability to build sophisticated applications without traditional programming, particularly enabling LLMs to serve as ephemeral graph databases. The research introduces MGraph-DB, a memory-first graph database co-developed by Cruz, and shows how it enables rapid experimentation with graph structures. Through case studies like the ThreatModCon 2025 keynote preparation and MyFeeds.ai project, the paper illustrates how graph thinkers can now leverage AI tools to dramatically compress development timelines without sacrificing quality. The document emphasizes that this democratization of innovation allows anyone with a graph-oriented mindset to turn ideas into tangible models and applications, fostering creativity and problem-solving across domains from business process mapping to enterprise knowledge management. ### No Code Development (NCD): A Paradigm Shift Beyond 'Vibe Coding' *June 18, 2025* This white paper proposes [No Code Development (NCD)](18/no-code-development--ndc--a-paradigm-shift-beyond-vibe-coding.md) as a more accurate and professional term for the emerging paradigm often referred to as "vibe coding." Drawing on Cruz's firsthand experience at ThreatModCon 2025, where he created multiple interactive visualizations and UI tools without writing manual code, the document argues that NCD better encapsulates this workflow's essence: a highly-iterative, feedback-driven process where developers focus on intent and orchestration while AI handles code generation. The paper examines why "vibe coding" falls short as a descriptor, potentially trivializing the skill involved and failing to communicate that real development work is occurring. Through the ThreatModCon case study, where Cruz was able to make adjustments to UIs up to 5 minutes before his keynote presentation, the document highlights the productivity gains of NCD. It also discusses the "air gap" between NCD and traditional engineering, exploring how context-switching into code disrupts creative flow. The research emphasizes the irony that seasoned software engineers are often the most effective NCD practitioners due to their domain knowledge and prompting skills. ### Empowering Workshops with Custom GPTs for GenAI Training *June 19, 2025* This white paper explores why [custom GPTs are powerful tools for workshops](19/empowering-workshops-with-custom-gpts-for-genai-training.md) and how they evolved into a commoditized tool in the AI landscape. OpenAI's introduction of GPTs in late 2023 allowed anyone to create tailored versions of ChatGPT for specific purposes, combining a large language model with user-provided instructions, optional domain knowledge, and tool integrations. Within two months of launch, users created over 3 million custom GPTs, demonstrating the demand for accessible AI customization. The document details key features that make GPTs workshop-ready: pre-loaded expert instructions, custom knowledge bases, integrated tools and skills, user-friendly creation interfaces, auto-generated icons and identities, and easy sharing capabilities. The paper provides a practical workshop plan including prerequisites, introductory demos, hands-on building sessions, iteration and tuning exercises, and use-case discussions. Real-world examples from executive training at financial services firms demonstrate how GPTs transform passive learning into active creation, with participants building functional AI assistants within 45-minute sessions. ### Using LLMs as Ephemeral Graph Databases *June 19, 2025* This white paper introduces the concept of using [Large Language Models as ephemeral graph databases](19/using-llms-as-ephemeral-graph-databases--empowering-the-graph-thinkers-in-the-age-of-generative-ai.md), harnessing an LLM's ability to understand and manipulate structured information without a persistent database. In this approach, the knowledge graph is created dynamically and exists only for the duration of an AI session, constructed at query-time to suit the user's context. The document provides a step-by-step tutorial modeling a cybersecurity risk management scenario, demonstrating how a non-programmer can create and query an ephemeral graph using only an LLM. The implementation walkthrough shows how to build a risk management graph starting from "User accounts can be compromised" and expanding to include causes (credentials leaked, no MFA), impacts (data exposure, compliance violations, financial losses), controls (password policies, monitoring), and stakeholders (system owners, executives). The paper demonstrates CRUD operations through natural language, showing how users can create nodes, establish relationships, update properties, and query the graph for insights. This approach provides graph thinking capabilities without requiring installation, configuration, or knowledge of query languages like Cypher or SPARQL. ### FIST Meets the Semantic Knowledge Graph *June 22, 2025* This white paper provides a comparative analysis of how the [FIST framework (Fast, Inexpensive, Simple, Tiny)](22/fist-meets-the-semantic-knowledge-graph-aligning-fast-inexpensive-simple-tiny-with-dinis-cruzs-g3-approach.md) aligns with Cruz's Semantic Knowledge Graph and G³ methodology. FIST, originally coined in the defense acquisition context to promote rapid, low-cost innovation, emphasizes streamlined processes, accelerated timelines, and restrained scope. The document demonstrates how these principles resonate strongly with modern knowledge graph engineering practices, even though FIST emerged in a military-tech context. The research maps each FIST principle to Cruz's approach: Speed and Agility achieved through automation and real-time LLM interaction, Cost Efficiency through open-source tools and serverless architecture, Simplicity through modular design and user-friendly interfaces, and Small Scale through tiny, focused knowledge graph modules. The paper includes concrete comparisons like the Marine Corps' Harvest Hawk project versus Cruz's Cyber Boardroom risk graph, and the Air Force's Condor Cluster supercomputer versus MyFeeds.ai personalized knowledge system. The analysis concludes that both FIST and the Semantic Knowledge Graph methodology prove that focusing on the smallest effective solution yields better outcomes. ### Project Plan: High Street GenAI Learning Hub *June 22, 2025* This project plan outlines the creation of a [GenAI Learning Hub on town high streets](22/project-plan-high-street-genai-learning-hub.md), a modern community space dedicated to making generative AI accessible to all citizens. Inspired by the concept that "Libraries are the only thing left on the high street that doesn't want either your soul or your wallet," the plan envisions a welcoming, café-like environment that demystifies AI technology and empowers citizens to use these tools positively. The hub would serve as a "third place" for creativity and education, anchoring high streets as community learning centers. The implementation plan details a phased approach starting with foundation and partnerships, moving through pilot operation, refinement and growth, to eventual scale-out and replication. The hub would offer café-style coziness with workstations, open-source AI software, educational content, maker hardware kits, and demonstration zones. Programming includes introductory workshops, themed short-courses, ongoing code/build clubs, guest talks, youth programs, and cross-generational learning events. The business model combines community support, earned income, and sponsorship, operating as a Community Interest Company to ensure the mission remains focused on community benefit. ### Data Tests for Neo4j: Bringing Automated Testing to Graph Databases *June 25, 2025* This white paper introduces [data tests for Neo4j](25/data-tests-for-neo4j-bringing-automated-testing-to-graph-databases.md), applying proven software testing principles to graph data itself. Much like unit tests catch code regressions, data tests catch graph regressions – unintended changes to the structure or content of your Neo4j database. The document outlines how to implement them using Python's PyTest and CI/CD pipelines, providing immediate feedback on any change's impact to ensure the data still meets expectations. The implementation details include defining expected conditions through brainstorming invariants and using past incidents, setting up ephemeral test environments via Docker or Testcontainers, loading synthetic or sampled data for testing, and writing tests using PyTest and the Neo4j driver. The paper provides concrete examples of data tests including schema compliance checks, uniqueness and duplicate detection, referential integrity validation, forbidden subgraph pattern detection, and expected graph metrics verification. Integration with CI/CD pipelines ensures that data tests run automatically on every push or pull request, creating a culture where data integrity is systematically ensured rather than hoped for. ### Ephemeral Neo4j Instances for On-Demand Graph Analytics *June 25, 2025* This white paper proposes an [ephemeral Neo4j architecture](25/ephemeral-neo4j-instances-for-on-demand-graph-analytics.md) for spinning up Neo4j instances on-demand in cloud environments, performing graph computations, then tearing them down. This approach combines the analytical power of graph databases with the cost-efficiency and elasticity of serverless computing, using existing open-source Neo4j as a commodity component deployed and disposed of as needed. The document focuses on AWS implementation using EC2 virtual machines and Fargate containers with Amazon S3 for storage. The four-stage workflow involves provisioning a new Neo4j instance, initializing and loading data from cloud storage, executing graph transformations and analytics, then exporting results and tearing down the instance completely. The paper details optimization strategies including pre-baked images for faster startup, bulk import for efficient data loading, and automated pipeline integration with CI/CD. Performance considerations address startup latency targets of under one minute, query performance optimization through proper instance sizing, and parallel task execution for scalability. The approach aligns with industry trends toward serverless and on-demand services, as evidenced by Neo4j's own Aura Graph Analytics Serverless offering. ### Using Ephemeral Neo4j Instances for a Cybersecurity Risk Graph Scenario *June 25, 2025* This tutorial demonstrates how to implement a [cybersecurity risk graph using ephemeral Neo4j instances](25/using-ephemeral-neo4j-instances-for-a-cybersecurity-risk-graph-scenario.md), bridging the conceptual LLM-as-graph approach into practical Neo4j implementation. The document provides step-by-step instructions for recreating the "User accounts can be compromised" risk scenario, showing how to spin up a Neo4j database on-demand, build the graph, perform analysis, and tear it down – providing the full power of a graph database with the flexibility of on-demand usage. The implementation walkthrough covers creating nodes for risks, personas, and events; establishing relationships to model impacts and causes; expanding the graph with downstream consequences and upstream preventive controls; adding assets and stakeholders to ground the abstract risk in reality; and simulating an incident to test the model. The tutorial includes complete Cypher queries for each step, demonstrating how to query the graph to answer critical questions like identifying at-risk systems, determining potential impacts, notifying appropriate stakeholders, and prioritizing remediation actions. The exercise validates the ephemeral approach for graph analytics, showing how complex graph modeling and querying can be performed without maintaining a permanent database server. --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/07/index.html (markdown twin: /2025/07/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/07/index.html* # July 2025 Published Materials > During July 2025, I published 12 comprehensive research documents with a strong focus on the practical applications of semantic knowledge graphs in cybersecurity and enterprise systems. The month's… --- During July 2025, I published 12 comprehensive research documents with a strong focus on the practical applications of semantic knowledge graphs in cybersecurity and enterprise systems. The month's publications demonstrate particular emphasis on cybersecurity innovations (7 papers covering SIEM reimagination, GDPR compliance, IAM evolution, and risk management), the evolution of AI from experimental LLMs to production-ready systems (2 papers on small models and AI-assisted development), and strategic approaches to modern challenges in news monetization and SaaS pricing models. My research this month showcases how ephemeral, serverless architectures combined with graph-based knowledge representation can fundamentally transform traditional enterprise security and compliance approaches. The publications reveal several interconnected themes in my work: the shift from continuous data ingestion to on-demand, context-aware processing (exemplified in the Ephemeral GenAI SIEM), the progression from large language models to deterministic code as AI matures, and the critical importance of finding "good enough" thresholds in risk management and product development. My exploration of GDPR compliance through Memory_FS and interactive Q&A graphs demonstrates practical implementations of theoretical concepts, while the proposals for graph-based cloud IAM and Project VulnAI show how these ideas scale to enterprise-level security challenges. Throughout July, I consistently explored how semantic knowledge graphs, sustainable AI practices, and serverless architectures converge to create more efficient, transparent, and economically viable solutions for modern organizations. # July 2025 Research Publications ## Overview | Date | Title | Focus Area | Key Concepts | |------|-------|------------|--------------| | 07/02 | [Ephemeral GenAI SIEM: A Serverless, Graph-Driven Approach to Security Event Management](02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md) | Cyber Security | Serverless architecture, semantic knowledge graphs, LLM integration, LETS pipeline | | 07/02 | [Using Memory_FS to Build a File-Based Representation of the GDPR Standard](02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md) | Cyber Security | Memory_FS framework, GDPR compliance, file-based storage, three-file pattern | | 07/03 | [LLM-Driven GDPR Compliance Q&A Graph – Technical Brief](03/llm-driven-gdpr-compliance-q-and-a-graph-technical-brief.md) | Cyber Security | Interactive chatbot, GDPR compliance, knowledge graphs, client-side implementation | | 07/04 | [FAQ - Evolving Semantic Graphs and Ontologies with LLMs and MGraph-DB](04/faq-evolving-semantic-graphs-and-ontologies-with-llms-and-mgraph-db.md) | Graphs | MGraph-DB, semantic graphs, ontology evolution, confidence scoring | | 07/04 | [From Free Scraping to Fair Compensation: Cloudflare's GenAI Crawler Charges and the Future of News Monetization](04/from-free-scraping-to-fair-compensation-cloudflares-genai-crawler-charges-and-the-future-of-news-monetization.md) | The Future of News | AI crawler fees, content monetization, micropayments, trust services | | 07/04 | [The Joy of Programming in the Age of AI-Assisted Development](04/the-joy-of-programming-in-the-age-of-ai-assisted-development.md) | Development and GenAI | Vibe coding, flow state, citizen developers, no-code development | | 07/04 | [Usage-Based Billable Entities: Aligning SaaS Pricing with Customer Usage](04/usage-based-billable-entities-aligning-saas-pricing-with-customer-usage.md) | Projects | Consumption-based pricing, value metrics, SaaS billing, usage tracking | | 07/06 | [Finding the "Good Enough" Threshold: Optimizing Risk, Creativity, and Product Decisions](06/finding-the-good-enough-threshold-optimizing-risk-creativity-and-product-decisions.md) | Cyber Security | Risk appetite, inefficiency delta, Wardley Maps, EVTP model | | 07/06 | [From Large Language Models to Small Models and Code: The Evolution of AI Solutions](06/from-large-language-models-to-small-models-and-code-the-evolution-of-ai-solutions.md) | Development and GenAI | LLMs, SLMs, deterministic code, AI commoditization, edge computing | | 07/06 | [Graph-Based Cloud IAM in the GenAI Agentic World](06/graph-based-cloud-iam-in-the-genai-agentic-world.md) | Cyber Security | Cloud IAM, least privilege, ephemeral permissions, semantic knowledge graphs | | 07/06 | [Semantic Knowledge Graphs, G³, and Sustainable AI: Aligning Innovations with ESG Objectives](06/semantic-knowledge-graphs-g3-and-sustainable-ai-aligning-innovations-with-esg-objectives.md) | Cyber Security | ESG alignment, sustainable AI, G³ concept, carbon-efficient computing | | 07/27 | [Project VulnAI: AI-Powered Vulnerability Risk Management Platform](27/project-vulnai-ai-powered-vulnerability-risk-management-platform.md) | Projects | Risk-based vulnerability management, AI analysis, semantic graphs, serverless architecture | ## Detailed Summaries ### Ephemeral GenAI SIEM: A Serverless, Graph-Driven Approach to Security Event Management *July 2, 2025* This comprehensive white paper introduces the [Ephemeral GenAI SIEM](02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md), a revolutionary approach to Security Information and Event Management that addresses the critical shortcomings of traditional SIEM solutions. The paper highlights how conventional SIEMs struggle with scale, context, and cost, often ingesting massive volumes of log data while generating thousands of alerts with little context. The proposed solution leverages serverless computing, semantic knowledge graphs, and generative AI to fundamentally redefine how security data is collected, analyzed, and acted upon, with a focus on on-demand data collection rather than continuous ingestion. The architecture employs a deterministic LETS (Load, Extract, Transform, Save) pipeline pattern where every processing step writes its output to persistent storage, enabling full traceability and reproducibility. The system uses Large Language Models for tasks requiring human-like reasoning while maintaining explainability through structured output schemas and controlled sandboxing. By treating cloud object storage as the database and using MGraph-DB as the in-memory graph engine, the solution achieves high scalability and cost-efficiency while providing context-rich investigations that bridge the gap from data detection to actionable intelligence. ### Using Memory_FS to Build a File-Based Representation of the GDPR Standard *July 2, 2025* This technical white paper presents an innovative method for converting the General Data Protection Regulation text into a structured, file-based representation using the [Memory_FS framework](02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md). The approach parses the official GDPR document into a hierarchical set of files, with each regulatory element (down to individual paragraphs or bullet points) captured using Memory_FS's three-file pattern consisting of content, config, and metadata files. This structured representation provides a robust foundation for further processing and analysis of the complex 99-article regulation. The methodology demonstrates how to maintain fidelity through round-trip conversions between formats (including Markdown and PDF) while ensuring the legal document's structure is preserved. The paper provides detailed implementation guidance with code examples for ingesting the GDPR document into Memory_FS and subsequently exporting or utilizing the data. The resulting Memory_FS output can feed into graph databases like MGraph-DB to generate knowledge graphs of the law, with support for intermediate representations (ZIP or SQLite) for cloud deployment, showcasing the advantages of Memory_FS's pluggable storage backends. ### LLM-Driven GDPR Compliance Q&A Graph – Technical Brief *July 3, 2025* This technical brief outlines the development of an [interactive chatbot UI](03/llm-driven-gdpr-compliance-q-and-a-graph-technical-brief.md) that guides users through GDPR compliance questions while dynamically building a knowledge graph from their answers. Unlike static questionnaires, this approach uses Large Language Models to adapt questions to the user's context in real-time, creating a personalized Q&A flow that captures up to 10 questions about GDPR compliance. The system constructs a graph of nodes and edges representing the user's context and answers, with the LLM using this growing graph to inform subsequent questions. The implementation is entirely client-side with no server storage, leveraging the user's browser and local storage to maintain state while using an LLM API for intelligence. The workflow includes a confirmation step where the system shows users what information has been captured for verification, followed by generating a final guidance report with personalized GDPR recommendations. The brief provides detailed prompt schemas for each phase of interaction and discusses front-end implementation details including chat interface design, state management, and optional graph visualization components. ### FAQ - Evolving Semantic Graphs and Ontologies with LLMs and MGraph-DB *July 4, 2025* This FAQ-style white paper addresses key technical questions about Dinis Cruz's approach to building and [evolving semantic knowledge graphs](04/faq-evolving-semantic-graphs-and-ontologies-with-llms-and-mgraph-db.md) using Large Language Models and the MGraph-DB platform. The paper explores how this architecture differs from traditional Semantic Web frameworks like OWL ontologies, focusing on concept identification, confidence scoring, reasoning methods, and graph lifecycle management. It explains how concepts are handled more fluidly than in OWL, with human-friendly labels and contextual uniqueness rather than rigid IRIs, allowing for emergent alignment through iterative refinement. The document details how confidence and relevance scores are calculated and used throughout the system, including LLM-generated relevance ratings and algorithmic composite scores. It describes the storage and management of graphs using MGraph-DB's file-based, version-controlled approach, and introduces the concept of an "ontology of ontologies" (G³) that enables multiple domain ontologies to coexist and interlink. This federated approach to knowledge management allows different teams to maintain their own taxonomies while enabling interoperability, representing a shift from centralized to distributed knowledge organization. ### From Free Scraping to Fair Compensation: Cloudflare's GenAI Crawler Charges and the Future of News Monetization *July 4, 2025* This white paper examines [Cloudflare's groundbreaking initiative](04/from-free-scraping-to-fair-compensation-cloudflares-genai-crawler-charges-and-the-future-of-news-monetization.md) to block AI crawlers by default from websites on its network unless they have permission or provide compensation to content owners. The paper contextualizes this move within the broader crisis facing content publishers, who bear infrastructure costs from AI scraping while being cut out of the value chain as AI models can regurgitate aggregated knowledge without directing users back to original sites. Cloudflare's permission-based model forces AI companies to obtain consent and potentially strike licensing deals before ingesting content. The paper outlines a comprehensive vision for news monetization in the GenAI era that extends beyond crawler access fees to include structured content APIs, trust and verification services, and micro-payments from readers. It provides detailed technical implementation guidance for enabling fair access and payments, including traffic identification and control, authentication mechanisms, metering and logging usage, and billing integration. The analysis demonstrates how publishers can create multiple, complementary revenue streams while maintaining quality journalism, with Cloudflare's initiative representing just the first step toward a more balanced ecosystem where content creators are compensated when their work generates value. ### The Joy of Programming in the Age of AI-Assisted Development *July 4, 2025* This white paper explores how recent advances in generative AI are unlocking the [joy of programming](04/the-joy-of-programming-in-the-age-of-ai-assisted-development.md) for a new wave of developers through "vibe coding" or No Code Development (NCD). The paper examines how AI-powered platforms dramatically lower barriers to entry, enabling citizen developers across business functions to rapidly turn ideas into working applications. It analyzes the psychological aspects of programming, including the flow state and creative delight that has long motivated software engineers, and how AI tools now deliver that instant interactivity to non-coders through natural language interfaces. The paper presents compelling evidence that this democratization of programming will actually increase rather than decrease the need for professional developers, as they will be needed to provide robust architectures, governance, and maintainability for the explosion of citizen-created applications. It explores how organizations can harness this influx of new programmers while maintaining software quality, with professional developers evolving into mentors and architects who build platforms and guardrails for safe no-code development at scale. The analysis concludes that programming skills are becoming widespread, with the creative joy of coding transitioning from a specialist's privilege to a universal experience. ### Usage-Based Billable Entities: Aligning SaaS Pricing with Customer Usage *July 4, 2025* This comprehensive white paper examines the shift from traditional fixed subscription models to [usage-based pricing](04/usage-based-billable-entities-aligning-saas-pricing-with-customer-usage.md) in SaaS, where customers pay only for what they actually use based on discrete billable units. The paper addresses the fundamental misalignment in traditional subscription models where customers often pay for unused capacity while heavy users consume far more resources than their subscription fee covers. It presents usage-based billing as a solution that aligns revenue with value, delivering fairness to customers and sustainability to providers. The document provides detailed guidance on defining billable units or value metrics, from cloud infrastructure units and API calls to data volume and operational events. It outlines the technical requirements for building a usage-based billing system, including instrumentation and metering, usage attribution, billing engines, real-time visibility, and payment management. Through analysis of successful implementations at companies like AWS, Snowflake, Twilio, and Stripe, the paper demonstrates how usage-based models enable faster growth, higher net retention, and better alignment between provider success and customer value. ### Finding the "Good Enough" Threshold: Optimizing Risk, Creativity, and Product Decisions *July 6, 2025* This white paper explores the critical concept of the ["good enough" threshold](06/finding-the-good-enough-threshold-optimizing-risk-creativity-and-product-decisions.md) across three scenarios: cybersecurity risk management, independent music production, and product design. The paper introduces the concept of the "Inefficiency Delta" - the gap between optimal outcomes and over-engineered or over-cautious approaches that yield diminishing returns. Through detailed case studies, it demonstrates how overshooting or undershooting the optimal point can hurt outcomes in each domain. The analysis applies Wardley Maps' Explorers-Villagers-Town Planners (EVTP) model to show how success comes from operating in the middle zone between chaotic exploration and rigid execution. The paper provides actionable strategies for finding the optimal point, including clearly defining acceptance criteria, using time boxes and cost caps, leveraging early feedback, breaking big bets into small ones, and aligning incentives with optimal risk-taking. The conclusion emphasizes that perfectionism, over-engineering, and ultra-conservatism are all manifestations of failing to define "good enough" and pulling the trigger when it's met. ### From Large Language Models to Small Models and Code: The Evolution of AI Solutions *July 6, 2025* This white paper charts the [evolution from Large Language Models to Small Language Models and ultimately to deterministic code](06/from-large-language-models-to-small-models-and-code-the-evolution-of-ai-solutions.md), presenting this as the natural maturation of AI technology. The paper frames this evolution through Wardley Maps, showing how language processing capabilities progress from novel, chaotic solutions (LLMs) to stable, commoditized utilities (code/APIs). It explains how LLMs served as exploratory tools that demonstrated what's possible, but their expense, non-determinism, and black-box nature make them unsuitable for many production use cases. The document details the rise of Small Language Models (SLMs) as focused specialists that can match or surpass large models on specific tasks while offering lower costs, faster performance, on-device processing capabilities, and better customizability. It then explores the path to full determinism, where AI capabilities are promoted into hard-coded solutions or API calls once thoroughly understood. The analysis concludes that the future of AI is not one mega-model to rule them all, but rather a constellation of smaller models, micro-models, and conventional software working in concert, with organizations using the right tool for each job. ### Graph-Based Cloud IAM in the GenAI Agentic World *July 6, 2025* This white paper proposes a revolutionary [graph-based IAM and permission workflow](06/graph-based-cloud-iam-in-the-genai-agentic-world.md) for cloud providers to address the security challenges posed by generative AI agents that can dynamically invoke cloud APIs in unpredictable ways. The paper highlights how current cloud IAM systems lead to over-provisioned permissions, where applications run with far more privileges than actually needed, creating significant security risks that are amplified when AI agents have flexible decision spaces influenced by prompts or even malicious injections. The proposed solution models cloud permissions, resources, and API calls as a semantic knowledge graph that can precisely determine the exact privileges required per action and issue ephemeral, context-specific credentials. The architecture would enable just-in-time permissions minted per action, with AI agents receiving temporary credentials scoped to exactly the permissions needed for each operation. Implementation details include using semantic knowledge graphs like MGraph-DB, creating policy decision engines, and potentially establishing cloud-agnostic abstraction layers. This approach would dramatically reduce the blast radius of compromised components while enabling intelligent, context-aware cloud security workflows. ### Semantic Knowledge Graphs, G³, and Sustainable AI: Aligning Innovations with ESG Objectives *July 6, 2025* This comprehensive white paper explores how [semantic knowledge graphs and the G³ concept](06/semantic-knowledge-graphs-g3-and-sustainable-ai-aligning-innovations-with-esg-objectives.md) (Graphs of Graphs of Graphs) align with efforts to make IT and AI systems more sustainable and ESG-compliant. The paper examines how these practices intersect with carbon-efficient computing, ethical AI, stakeholder transparency, open-source collaboration, responsible data governance, and AI explainability. G³ advocates connecting multiple graphs and ontologies rather than forcing a single dominant hierarchy, enabling systems to remain flexible, adaptive, and inclusive of diverse viewpoints. The analysis demonstrates how semantic knowledge graphs contribute to environmental sustainability through efficiency gains and knowledge reuse, support ethical AI by avoiding one-dimensional bias and ensuring transparency in meaning-making, and enable stakeholder empowerment through open-source development and knowledge sharing. The paper shows how these approaches facilitate responsible data governance by making relationships and rules explicit in machine-readable form, and enhance AI explainability through provenance-enabled systems that can trace their reasoning. The conclusion positions Dinis Cruz's work as exemplifying a path where advanced technology and ESG ideals exist in harmony, with knowledge elevated to the same stature as data and code in engineering priorities. ### Project VulnAI: AI-Powered Vulnerability Risk Management Platform *July 27, 2025* This detailed project brief presents [VulnAI](27/project-vulnai-ai-powered-vulnerability-risk-management-platform.md), a next-generation SaaS platform for AI-driven vulnerability management that prioritizes risk context over raw vulnerability counts. The platform leverages semantic knowledge graphs, automated AI analysis, and a deterministic data pipeline to unify diverse security data into a coherent risk knowledge base. Unlike traditional vulnerability management tools that overwhelm teams with endless lists, VulnAI helps security teams and developers make smarter decisions by focusing remediation efforts where they matter most to the business. The platform features risk-centric prioritization that contextualizes every vulnerability with business impact and exploitability, AI-powered analysis with full traceability through a controlled LETS pipeline, and a semantic knowledge graph backbone that enables complex querying and mapping of technical issues to business concerns. The architecture employs ephemeral and serverless components for cost-efficiency and scalability, with an open-source core that fosters community contributions and transparency. The implementation plan outlines a phased approach from foundation through targeted MVP to production hardening, with a clear business case demonstrating value for both customers (through improved risk posture and efficiency) and the SaaS business (through a large addressable market and differentiated offering). --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/08/index.html (markdown twin: /2025/08/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/08/index.html* # August 2025 Published Materials > During August 2025, I published 2 significant research documents focused on software development methodologies and practices in the age of generative AI. Both publications, authored in collaboration… --- ## Overview During August 2025, I published 2 significant research documents focused on software development methodologies and practices in the age of generative AI. Both publications, authored in collaboration with AI assistants (ChatGPT Deep Research and Claude Opus 4.1), address critical aspects of modern development workflows: rapid application development with AI assistance and offline-first development strategies. These works represent a concentrated exploration of how development practices must evolve to leverage AI capabilities while maintaining engineering rigor and operational resilience. The month's research demonstrates a strong emphasis on practical methodologies that bridge the gap between traditional software engineering and AI-augmented development. The Iterative Flow Development (IFD) methodology introduces a revolutionary approach to maintaining developer flow state while achieving 10-20x productivity gains, while the Surrogate Dependencies framework addresses the critical need for offline development capabilities. Together, these publications present complementary strategies for creating more efficient, resilient, and developer-friendly software engineering practices that acknowledge both the promise and challenges of AI-assisted development. ## Publications Overview | Date | Title | Focus Area | Key Concepts | |------|-------|------------|--------------| | 08/22 | [Iterative Flow Development (IFD) Methodology: JavaScript Web Application Implementation](22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.md) | Development and GenAI | IFD, Vibe Coding, Air-Gapped Mode, Flow State, Version Independence | | 08/22 | [Surrogate Dependencies: Simulating Backends for Offline-First Development](22/surrogate-dependencies-simulating-backends-for-offline-first-development.md) | Development and GenAI | Surrogate Dependencies, Offline Development, API Simulation, TDD, Service Virtualization | ## Detailed Summaries ### Iterative Flow Development (IFD) Methodology: JavaScript Web Application Implementation *August 22, 2025* This comprehensive white paper introduces [Iterative Flow Development (IFD)](22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.md), a novel methodology that transforms how developers collaborate with Large Language Models (LLMs) to achieve unprecedented development speed without sacrificing software quality. The methodology centers on preserving developer flow state while leveraging AI for rapid code generation, introducing dual operational modes - "vibe coding" for non-technical users and "air-gapped" for professional developers - that bridge the gap between rapid prototyping and production-ready software. Through a detailed case study of a text analysis application built in a single day (13,000+ lines of production-ready code), the paper demonstrates how IFD can achieve 10-20x productivity gains while maintaining architectural integrity through principles like version independence, real-data-first development, and zero external dependencies. The implementation details reveal how IFD's two-tier versioning system enables continuous experimentation while ensuring production stability, with minor versions evolving incrementally and major versions representing clean, standalone releases. The methodology's emphasis on Web Components and native browser APIs eliminates framework dependencies, while its structured approach to LLM collaboration through templated prompts and consolidation workflows ensures that AI-generated code meets production standards. The paper provides extensive practical guidance including architectural patterns, team workflows, and adoption strategies, positioning IFD as a paradigm shift from team-centric to flow-centric development that makes both rapid innovation and production excellence achievable within the same methodology. ### Surrogate Dependencies: Simulating Backends for Offline-First Development *August 22, 2025* This technical white paper presents [Surrogate Dependencies](22/surrogate-dependencies-simulating-backends-for-offline-first-development.md), a solution for enabling applications to run in fully offline mode by simulating backend systems with prerecorded JSON data. Originally proposed to align security and development needs, surrogate dependencies act as stand-ins for live APIs, serving consistent responses from local data instead of making network calls, thereby addressing challenges like fragile dev/QA environments, limited testing capabilities, and the inability to work offline. The paper demonstrates how this approach differs from traditional mocks and stubs by using real data captured from actual API calls, making it more realistic and easier to scale across many endpoints while remaining simpler than full service virtualization platforms. The implementation strategy centers on centralizing network communication through an abstracted layer that can switch between live and surrogate modes, organizing surrogate data in a structure that mirrors API routes, and integrating this capability into CI/CD workflows. The paper explores how surrogate dependencies enhance test-driven development by enabling integration tests to run without live backends, and discusses their increasing relevance in LLM-based development environments where AI assistants can leverage surrogate data for better code generation and testing. Through detailed code examples and best practices, the document provides a complete blueprint for implementing surrogate dependencies in modern web applications, particularly those using FastAPI backends, emphasizing how this approach fosters better API contract understanding, faster debugging cycles, and more resilient development processes. --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /admin/index.html (markdown twin: /admin/index.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/admin/index.html* > How diniscruz.ai is built and released: hand-written HTML plus pages generated from the markdown migrated from docs.diniscruz.ai, a markdown twin at every URL, llms.txt, and a validate-then-deploy pipeline to GitHub Pages. --- # How this site is built This is a static site on GitHub Pages, built the same way as [sgit.ai](https://sgit.ai) and [open-source.sgit.ai](https://open-source.sgit.ai), and sharing their design language. There is no framework and no JavaScript bundle. Every page works without scripts, and every page has a markdown twin. ## Two kinds of page | Kind | Source of truth | What the build does | |---|---|---| | **Hand-written** home, about, building, admin | The `.html` file itself, edited by hand | Rewrites the nav and footer from `admin/build/site_config.py`, fills the `` blocks (such as the latest-writing cards), and derives the `.md` twin from the HTML | | **Generated** essays, research hubs, talks | The markdown under `content/`, migrated from docs.diniscruz.ai with its front matter intact | Renders each file to HTML **at the same path it had on docs.diniscruz.ai**, expands the old MkDocs macros (PDF, LinkedIn, video and slide embeds), and writes the cleaned markdown as its twin | ## The indexing surfaces - **[sitemap.xml](../sitemap.xml)** and **[robots.txt](../robots.txt)**. Every indexable page, with the publication date as `lastmod` for essays. - **[feed.xml](../feed.xml)**, the RSS feed of the writing. It is also published as `feed_rss_created.xml` and `feed_rss_updated.xml`, the names the old MkDocs site used. - **Structured data.** A `Person` record on the home and about pages, `BlogPosting` on every essay (author, date, licence, word count), and canonical and Open Graph tags everywhere. - **[llms.txt](../llms.txt)**, which is self-sufficient: who I am and every page with its description. **[llms-full.txt](../llms-full.txt)** holds every page in one file, for agents that cannot follow links. ## Search and AI features Google's guidance for its AI features (AI Overviews and AI Mode) is that no special optimisation is needed beyond the fundamentals, so this site does the fundamentals and checks them on every build: - **Crawlable, indexable text.** Every page is static HTML that works without JavaScript. Each has one `

`, a real title and description, and a canonical URL that appears in [sitemap.xml](../sitemap.xml). `validate.js` fails the release if any of that is missing. - **Snippet-eligible.**`max-snippet:-1, max-image-preview:large` on every page, so search results and AI answers can quote and preview it in full. - **Structured data that matches the visible page.**`Person`, `WebSite`, `BlogPosting`, `BreadcrumbList` (generated from the same data as the visible breadcrumb) and `CollectionPage`. - **Internal links.** Every essay links to its topic hub, its neighbours in time, and the five nearest pieces on the same topic. - **No duplicates in the index.** The markdown twins and `llms-full.txt` repeat the HTML for agents. [robots.txt](../robots.txt) keeps Googlebot and Bingbot off them, so each page is indexed once, as HTML. Other crawlers and agents can still fetch them. **llms.txt** follows the [llmstxt.org](https://llmstxt.org/) format. Google has said its search does not use llms.txt, so it is here for the other agents and LLM tools that do read it, not for ranking. ## Release process 1. Bump `admin/build/version.txt` and add a row to [the release history](versions.md). 2. `pip install -r admin/build/requirements.txt` (python-markdown and PyYAML, once). 3. `python3 admin/build/build.py` regenerates every derived file. 4. `node admin/build/validate.js` checks the version, internal links, canonicals, twins and licence lines. 5. `git commit -am "site vX.Y.Z: ..." && git push origin dev` Every push runs `.github/workflows/deploy-pages.yml`. It rebuilds the site and fails if anything committed is stale, then validates, then deploys to GitHub Pages. Pull requests run the checks only. ## Adding an essay Drop a markdown file at `content/YYYY/MM/DD/slug.md` with the same front matter the old site used (`title`, `authors`, `date`, and optionally `description`, `tags`, `pdf_file`, `linkedin` and `back_link`), then run the build. It appears in the writing index, the feed, the sitemap, llms.txt and, if it is among the newest, on the home page. [← About](../about/index.md)[Moving from docs.diniscruz.ai →](migration.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /admin/migration.html (markdown twin: /admin/migration.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/admin/migration.html* > Why the essays moved from docs.diniscruz.ai to diniscruz.ai, how every old URL maps to a new one at the same path, and the steps that retire the old host without losing search ranking. --- # Moving from docs.diniscruz.ai The research and essays used to live at **docs.diniscruz.ai**, an MkDocs Material site. It never ranked well. The content sat on a subdomain rather than the name people search for, the pages carried little structured data, and many had no description. So the content has moved here, to the domain that carries the name. ## Every URL keeps its path The old site used `use_directory_urls: false`, so every page was a plain `.html` file, and this site keeps exactly those paths. **Moving the host is the only change**: | Old | New | |---|---| | `docs.diniscruz.ai/2025/06/07/.html` | `diniscruz.ai/2025/06/07/.html` | | `docs.diniscruz.ai/research/.html` | `diniscruz.ai/research/.html` | | `docs.diniscruz.ai/resources/presentations.html` | `diniscruz.ai/resources/presentations.html` | | `docs.diniscruz.ai/about.html` | `diniscruz.ai/about/index.html` (a redirect page stays at `/about.html`) | | `docs.diniscruz.ai/feed_rss_created.xml` | Same path, plus `diniscruz.ai/feed.xml` | The PDFs, audio and slides stay where they were, on `files.diniscruz.ai`. The markdown sources are copied under `content/` in this repository, with front matter unchanged. ## What changed for search - Every essay now has a meta description. Where the front matter had none (most of them), the build takes the first real paragraph. - Structured data: `BlogPosting` with author, date and licence on every essay, and a `Person` record that ties the site to LinkedIn, GitHub and the sgit.ai author pages. - A sitemap with publication dates, an RSS feed, canonical URLs on the apex domain, and a single chronological [writing index](../writing/index.md) that links to every piece. - Lighter pages: one small stylesheet and no framework JavaScript. ## Retiring the old host 1. **Best option, if the DNS is on Cloudflare:** a bulk redirect rule `docs.diniscruz.ai/*` → `https://diniscruz.ai/${1}` with status 301. That is a real permanent redirect, and all ranking signals pass through. 2. **Otherwise (GitHub Pages only):** run `python3 admin/migration/make_redirect_stubs.py ` and publish its output as the docs.diniscruz.ai site. Every old URL becomes a page with a canonical link to its new URL, an instant meta refresh and a JS redirect, which Google treats as a permanent move. A `404.html` catch-all covers anything else. 3. In Google Search Console, verify `diniscruz.ai` and submit `https://diniscruz.ai/sitemap.xml`. With 301s in place, also file a change of address from the old property. 4. Update the links that point at the old host: LinkedIn profile and posts, the [sgit.ai](https://sgit.ai/about/index.html) and [open-source.sgit.ai](https://open-source.sgit.ai/about/index.html) author pages, and the GitHub profile. **Order matters:** switch docs.diniscruz.ai to redirects only after this site is live on diniscruz.ai and the new URLs return 200. [← How this site is built](index.md)[Release history →](versions.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /admin/versions.html (markdown twin: /admin/versions.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/admin/versions.html* > Every release of diniscruz.ai, newest first: what changed in the site, its content and its build, version by version. --- # Release history Every release of this site, newest first. The version badge in the nav comes from `admin/build/version.txt`, and the validator checks that it agrees everywhere. | Version | Date | What changed | |---|---|---| | v0.1.1 | 27 Sep 2026 | SEO pass. A share image and large-preview robots directives on every page. BreadcrumbList structured data matching the visible breadcrumb, CollectionPage with ItemList on the research hubs, and a richer BlogPosting (image, section, PDF). "More on this topic" links on every essay. The home page's canonical URL is now the domain root. Sitemap lastmod on every URL. robots.txt keeps search engines off the markdown twins. llms.txt follows the llmstxt.org list format and has topic entry points. The validator now checks the search essentials on every page. | | v0.1.0 | 27 Sep 2026 | First pass. Home, about and building pages in the sgit.ai design language. All essays, research hubs and talks migrated from docs.diniscruz.ai at their original paths. Writing index, RSS, sitemap, structured data, markdown twins, llms.txt and llms-full.txt. The build, the validator, the GitHub Pages workflow, and the redirect kit for the old host. | [← Moving from docs.diniscruz.ai](migration.md)[Home →](../index.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/cyber-security.html (markdown twin: /research/cyber-security.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/cyber-security.html* # Cyber-Security > Index of articles on threat modeling, modern AppSec, and security innovation. --- ## Threat Modeling & Semantic Graphs - [Supercharging AppSec Threat Modeling Services with GenAI and Semantic Graphs](../2025/06/09/supercharging-appsec-threat-modeling-services-with-genai-and-semantic-graphs.md) - [Linking Threat Models with Semantic Business Graphs](../2025/06/02/linking-threat-models-with-semantic-business-graphs.md) - [Threat Models as Mandatory Disclosures: A Vision for Security Transparency](../2025/05/29/threat-models-as-mandatory-disclosures__a-vision-for-security-transparency.md) - [Advancing Threat Modeling with Semantic Knowledge Graphs](../2025/05/29/advancing-threat-modeling-with-semantic-knowledge-graphs.md) - [Using Threat Modeling and Semantic Graphs to Secure the Digital Supply Chain](../2025/05/30/using-threat-modeling-and-semantic-graphs-to-secure-the-digital-supply-chain.md) - [Graphs of Graphs of Graphs (G3) in Threat Modeling](../2025/05/30/graphs-of-graphs-of-graphs-g3-in-threat-modeling.md) - [Scaling Supply Chain Security using Threat Modeling Semantic Knowledge Graphs and Maps](../2025/05/30/scaling-supply-chain-security-using-threat-modeling-semantic-knowledge-graphs-and-maps.md) - [Using Ephemeral Neo4j Instances for a Cybersecurity Risk Graph Scenario](../2025/06/25/using-ephemeral-neo4j-instances-for-a-cybersecurity-risk-graph-scenario.md) ## Security Architecture & SIEM - [Ephemeral GenAI SIEM: A Serverless, Graph-Driven Approach to Security Event Management](../2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md) - [Graph-Based Cloud IAM in the GenAI Agentic World](../2025/07/06/graph-based-cloud-iam-in-the-genai-agentic-world.md) - [Security Implications of the Model Context Protocol (MCP) and the Need for Robust Infrastructure](../2025/06/10/security-implications-of-the-model-context-protocol-mcp-and-the-need-for-robust-infrastructure.md) - [Project VulnAI: AI-Powered Vulnerability Risk Management Platform](../2025/07/27/project-vulnai-ai-powered-vulnerability-risk-management-platform.md) - [Project Cybersage: AI-Powered Risk Contextualization & Security Reporting](../2025/04/10/project-cybersage__ai-powered-risk-contextualization_security-reporting.md) ## Compliance & Standards - [Using Memory_FS to Build a File-Based Representation of the GDPR Standard](../2025/07/02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md) - [LLM-Driven GDPR Compliance Q&A Graph – Technical Brief](../2025/07/03/llm-driven-gdpr-compliance-q-and-a-graph-technical-brief.md) - [Maturity Models vs. Traditional Standards in Application Security](../2025/04/02/maturity-modes-vs-traditional-standards-in-application-security.md) - [Semantic Knowledge Graphs, G³, and Sustainable AI: Aligning Innovations with ESG Objectives](../2025/07/06/semantic-knowledge-graphs-g3-and-sustainable-ai-aligning-innovations-with-esg-objectives.md) ## Risk Management & Decision-Making - [Finding the "Good Enough" Threshold: Optimizing Risk, Creativity, and Product Decisions](../2025/07/06/finding-the-good-enough-threshold-optimizing-risk-creativity-and-product-decisions.md) - [FIST Meets the Semantic Knowledge Graph: Aligning Fast, Inexpensive, Simple, Tiny with Dinis Cruz's G³ Approach](../2025/06/22/fist-meets-the-semantic-knowledge-graph-aligning-fast-inexpensive-simple-tiny-with-dinis-cruzs-g3-approach.md) - [Fail Safe, Not Fail Big: Cyber-Security-Inspired Strategies to Prevent the Next Iberian Grid Crisis](../2025/04/29/fail-safe-not-fail-big__cyber-security-inspired-strategies-to-prevent-the-next-iberian-grid-crisis.md) - [Second Stories: From Three Mile Island to Cybersecurity](../2025/02/10/second-stories__from-three-mile-island-to-cybersecurity.md) ## Digital Rights & Ethics - [Personal Content Rights: Protecting Individuals in the Age of Deepfakes and AI Cloning](../2025/06/15/personal-content-rights-protecting-individuals-in-the-age-of-deepfakes-and-ai-cloning.md) - [OAuth Security Concerns and Implications for the Model Context Protocol (MCP)](../2025/05/18/oauth-security-concerns-and-implications-for-the-model-context-protocol.md) - [Security Debrief: OpenAI's ChatGPT Connector GitHub App](../2025/05/18/security-debrief__openai_chatgpt_connector_gitHub_app.md) ## OWASP & Community - [History and Analysis of OWASP In-Person Summits](../2025/06/07/history-and-analysis-of-owasp-in-person-summits.md) - [Semantic OWASP: Leveraging GenAI and Graphs to Customise and Scale Security Knowledge](../2025/04/02/semantic-owasp__leveraging-genai-and-graphs-to-customise-and-scale-security-knowledge.md) - [Enhancing Cybersecurity Event Networking with Semantic Knowledge Graphs](../2025/04/06/enhancing-cybersecurity-event-networking-with-semantic-knowledge-graphs.md) ## SAST & Code Analysis - [O2 Platform's MethodStreams (2010 Open Source SAST engine)](../2025/02/11/o2-platforms-methodstreams-2010-open-source-sast-engine.md) - [Semantic Knowledge Graphs for LLM-Driven Source Code Analysis](../2025/05/29/semantic-knowledge-graphs-for-llm-driven-source-code-analysis.md) ## Community Learning - [Project Plan: High Street GenAI Learning Hub](../2025/06/22/project-plan-high-street-genai-learning-hub.md) --- See more Cyber-Security research documents in these LinkedIn posts: [Part 1](https://www.linkedin.com/posts/diniscruz_this-post-contains-multiple-examples-of-the-activity-7294760699866599426-9Za5/) and [Part 2](https://www.linkedin.com/posts/diniscruz_this-is-part-ii-of-a-post-containing-multiple-activity-7312404309508345856-sbDS/) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/development-and-genai.html (markdown twin: /research/development-and-genai.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/development-and-genai.html* # Development and GenAI > Hub for resources on generative AI, modern development methodologies, and AI-assisted programming. --- ## Development Methodologies - [Time as a Calibrator of Credibility and Trust in Information Systems](../2025/10/02/time-as-a-calibrator-of-credibility-and-trust-in-information-systems.md) - [Iterative Flow Development (IFD) Methodology: JavaScript Web Application Implementation](../2025/08/22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.md) - [Surrogate Dependencies: Simulating Backends for Offline-First Development](../2025/08/22/surrogate-dependencies-simulating-backends-for-offline-first-development.md) - [No Code Development (NCD): A Paradigm Shift Beyond 'Vibe Coding'](../2025/06/18/no-code-development--ndc--a-paradigm-shift-beyond-vibe-coding.md) - [LETS (Load, Extract, Transform, Save): A Deterministic and Debuggable Data Pipeline Architecture](../2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md) - [Scaling Kubernetes with One-Node Clusters: A Paradigm Shift in Cloud-Native Orchestration](../2025/02/13/scaling-kubernetes-with-one-node-clusters-a-paradigm-shift-in-cloud-native-orchestration.md) ## AI-Assisted Development - [The Joy of Programming in the Age of AI-Assisted Development](../2025/07/04/the-joy-of-programming-in-the-age-of-ai-assisted-development.md) - [From Large Language Models to Small Models and Code: The Evolution of AI Solutions](../2025/07/06/from-large-language-models-to-small-models-and-code-the-evolution-of-ai-solutions.md) - [Explorers, Villagers, and Town Planners: Understanding the Generative AI Divide](../2025/06/10/explorers-villagers-and-town-planners-understanding-the-generative-ai-divide.md) - [GenAI Legacy Code Refactoring – Business Plan](../2025/06/07/genai-legacy-code-refactoring-business-plan.md) ## Testing & Quality - [The Hidden Cost of Ephemeral Testing and the Case for Automation](../2025/06/15/the-hidden-cost-of-ephemeral-testing-and-the-case-for-automation.md) - [Data Tests for Neo4j: Bringing Automated Testing to Graph Databases](../2025/06/25/data-tests-for-neo4j-bringing-automated-testing-to-graph-databases.md) ## Workshops & Training - [Empowering Workshops with Custom GPTs for GenAI Training](../2025/06/19/empowering-workshops-with-custom-gpts-for-genai-training.md) - [Workshop Plan: User-Driven Semantic Persona Graphs Powered by GenAI](../2025/06/16/workshop-plan-user-driven-semantic-persona-graphs-powered-by-genai.md) ## Cloud & Infrastructure - [Comparing the EU FED Cloud vs. an Open-Source Federated Cloud Proposal](../2025/06/18/comparing-the-eu-fed-cloud-vs-an-open-source-federated-cloud-proposal.md) - [Intent-Based Feedback Loops in Cloud Environments](../2025/04/01/intent-based-feedback-loops-in-cloud-environments.md) ## Strategic Perspectives - [Think Different, Again: Reimagining Apple's Role in the AI Era](../2025/04/06/think-different-again__reimagining-apple-role-in-the-ai-era.md) - [Briefing on Canada's New Minister of Artificial Intelligence and Digital Innovation (vs UK and PT)](../2025/05/18/briefing-on-canada-new-minister-of-artificial-intelligence-and-digital-innovation-vs-uk-and-pt.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/europe-and-learning.html (markdown twin: /research/europe-and-learning.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/europe-and-learning.html* # Europe and Learning > Overview page linking GenAI opportunities in Europe and learning resources. --- ## Europe GenAI Opportunity - [Scaling Europe's Regulatory Superpower: From Static Cybersecurity Standards to Semantic Graphs](../2025/03/31/scaling-europe-regulatory-superpower.md) - [An Open-Source Sovereign Cloud for an Open Europe: The Case for a Federated, AI-Enabled, and Multilingual Digital Infrastructure](../2025/02/24/an-open-source-sovereign-cloud-for-an-open-europe.md) - [Portuguese as a Programming Language in the AI Era](../2025/02/11/portuguese-as-a-programming-language-in-the-AI-Era.md) - [Deterministic GenAI Outputs with Provenance (OWASP EU AppSec Lisbon )](../2024/06/28/deterministic-genai-outputs-with-provenance__owasp-appsec-lisbon__gslides.md) - [Europe's Strategic Opportunity in GenAI: A Deep Dive into Six Defining Trends](../2025/04/01/europe-strategic-opportunity-in-gen-ai__a-deep-dive-into-six-defining-trends.md) ## Learning - [Using Presentations Instead of CVs in Hiring](../2025/06/08/using-presentations-instead-of-cvs-in-hiring.md) - [Generative AI and the Future of Learning](../2025/02/12/generative-ai-and-the-future-of-learning.md) - [Navigating the AI Revolution: A Student's Guide to Generative AI in Education](../2025/04/22/navigating-the-ai-revolution__a_university_students_guide_to_generative-ai-in-education.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/graphs.html (markdown twin: /research/graphs.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/graphs.html* # Knowledge Graphs > Guides on graph databases, ontologies, semantic knowledge graphs, and ephemeral architectures. --- ## Graph Creation and Management - [User-Driven Semantic Persona Graphs Powered by GenAI](../2025/06/14/user-driven-semantic-persona-graphs-powered-by-genai.md) - [Using LLMs as Ephemeral Graph Databases: Empowering the Graph Thinkers in the Age of Generative AI](../2025/06/19/using-llms-as-ephemeral-graph-databases--empowering-the-graph-thinkers-in-the-age-of-generative-ai.md) - [Empowering the Graph Thinkers in the Age of Generative AI](../2025/06/18/empowering-the-graph-thinkers-in-the-age-of-generative-ai.md) - [FAQ - Evolving Semantic Graphs and Ontologies with LLMs and MGraph-DB](../2025/07/04/faq-evolving-semantic-graphs-and-ontologies-with-llms-and-mgraph-db.md) ## Ephemeral Graph Architectures - [Ephemeral Neo4j Instances for On-Demand Graph Analytics](../2025/06/25/ephemeral-neo4j-instances-for-on-demand-graph-analytics.md) - [Using Ephemeral Neo4j Instances for a Cybersecurity Risk Graph Scenario](../2025/06/25/using-ephemeral-neo4j-instances-for-a-cybersecurity-risk-graph-scenario.md) - [Data Tests for Neo4j: Bringing Automated Testing to Graph Databases](../2025/06/25/data-tests-for-neo4j-bringing-automated-testing-to-graph-databases.md) ## Knowledge Management Theory - [Bridging Niklas Luhmann's Ideas with Semantic Knowledge Graphs and G³](../2025/06/18/bridging-niklas-luhmanns-ideas-with-semantic-knowledge-graphs-and-g3.md) - [FIST Meets the Semantic Knowledge Graph: Aligning Fast, Inexpensive, Simple, Tiny with Dinis Cruz's G³ Approach](../2025/06/22/fist-meets-the-semantic-knowledge-graph-aligning-fast-inexpensive-simple-tiny-with-dinis-cruzs-g3-approach.md) - [From Top-Down to Organic Evolving Graphs, Ontologies, and Taxonomies](../2025/03/29/from-top-down-to-organic-evolving-graphs-ontologies-and-taxonomies.md) ## Applied Graph Solutions - [Jira as a Graph Database – Proposal for Atlassian Executives](../2025/06/03/jira-as-a-graph-database—proposal-for-atlassian-executives.md) - [Semantic OWASP: Leveraging GenAI and Graphs to Customise and Scale Security Knowledge](../2025/04/02/semantic-owasp__leveraging-genai-and-graphs-to-customise-and-scale-security-knowledge.md) - [Enhancing Cybersecurity Event Networking with Semantic Knowledge Graphs](../2025/04/06/enhancing-cybersecurity-event-networking-with-semantic-knowledge-graphs.md) - [Graph-Powered Legal Knowledge: An Open, Distributed, and AI-Assisted Roadmap](../2025/04/22/graph-powered-legal-knowledge__an-open-distributed-and-ai-assisted-roadmap.md) ## Personalized Briefings & Case Studies - [Personalized Briefing: Semantic Knowledge Graphs – Intersection of Dinis Cruz & Kerstin Clessienne's Work](../2025/06/08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md) ## Strategic Partnerships - [Proposal for Neo4j Collaboration with Dinis Cruz](../2025/06/03/proposal-for-neo4j-collaboration-with-dinis-cruz.md) - [Proposal: Strategic AWS Partnership with Dinis Cruz's GenAI and Graph Innovations](../2025/06/03/proposal-strategic-aws-partnership-with-dinis-cruz-genai-and-graph-innovations.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/organizational-transformation.html (markdown twin: /research/organizational-transformation.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/organizational-transformation.html* # Organizational Transformation > How organizations can leverage AI, semantic graphs, and modern methodologies to transform operations. --- ## Hiring & Talent - [Using Presentations Instead of CVs in Hiring](../2025/06/08/using-presentations-instead-of-cvs-in-hiring.md) ## Meeting Efficiency - [Project Agenda: GenAI-Powered Transformation of Meetings and Documentation](../2025/04/10/project-agenda__gen-ai-powered-transformation-of-meetings-and-documentation.md) ## Workflow Optimization - [Finding the "Good Enough" Threshold: Optimizing Risk, Creativity, and Product Decisions](../2025/07/06/finding-the-good-enough-threshold-optimizing-risk-creativity-and-product-decisions.md) - [The Hidden Cost of Ephemeral Testing and the Case for Automation](../2025/06/15/the-hidden-cost-of-ephemeral-testing-and-the-case-for-automation.md) ## Cultural Change - [Explorers, Villagers, and Town Planners: Understanding the Generative AI Divide](../2025/06/10/explorers-villagers-and-town-planners-understanding-the-generative-ai-divide.md) - [The Joy of Programming in the Age of AI-Assisted Development](../2025/07/04/the-joy-of-programming-in-the-age-of-ai-assisted-development.md) ## Knowledge Management - [Bridging Niklas Luhmann's Ideas with Semantic Knowledge Graphs and G³](../2025/06/18/bridging-niklas-luhmanns-ideas-with-semantic-knowledge-graphs-and-g3.md) - [From Top-Down to Organic Evolving Graphs, Ontologies, and Taxonomies](../2025/03/29/from-top-down-to-organic-evolving-graphs-ontologies-and-taxonomies.md) ## Training & Development - [Empowering Workshops with Custom GPTs for GenAI Training](../2025/06/19/empowering-workshops-with-custom-gpts-for-genai-training.md) - [Workshop Plan: User-Driven Semantic Persona Graphs Powered by GenAI](../2025/06/16/workshop-plan-user-driven-semantic-persona-graphs-powered-by-genai.md) ## Business Model Innovation - [Usage-Based Billable Entities: Aligning SaaS Pricing with Customer Usage](../2025/07/04/usage-based-billable-entities-aligning-saas-pricing-with-customer-usage.md) - [GenAI Legacy Code Refactoring – Business Plan](../2025/06/07/genai-legacy-code-refactoring-business-plan.md) ## Community Building - [History and Analysis of OWASP In-Person Summits](../2025/06/07/history-and-analysis-of-owasp-in-person-summits.md) - [Enhancing Cybersecurity Event Networking with Semantic Knowledge Graphs](../2025/04/06/enhancing-cybersecurity-event-networking-with-semantic-knowledge-graphs.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/projects.html (markdown twin: /research/projects.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/projects.html* # Projects and Innovation Lab > Collection of research projects, business ideas, and strategic partnerships exploring generative AI. --- ## Major Projects ### Security & Risk Management - [Project Voice2SIEM: Turning Customer Support Audio Into Real-Time Security Events](../2025/10/02/project-voice2siem-turning-customer-support-audio-into-real-time-security-events.md) - [Next-Generation API Security Platform: Semantic Graphs, GenAI Testing & Ephemeral Environments for 2025](../2025/09/01/next-generation-api-security-platform-semantic-graphs-genai-testing-ephemeral-environments-2025.md) - [Dinis Cruz's Research on API Security (2009-2025)](../2025/09/01/dinis-cruz-research-on-api-security-2009-2025.md) - [Project VulnAI: AI-Powered Vulnerability Risk Management Platform](../2025/07/27/project-vulnai-ai-powered-vulnerability-risk-management-platform.md) - [Project Cybersage: AI-Powered Risk Contextualization & Security Reporting](../2025/04/10/project-cybersage__ai-powered-risk-contextualization_security-reporting.md) - [Project SupplyShield: GenAI-Driven Supply Chain Risk Management and Compliance](../2025/02/15/project-supplyshield__genai-driven-supply-chain-risk-management-and-compliance.md) ### Web & Content Solutions - [Technical Briefing: Web Content Filtering Project](../2025/06/13/technical-briefing-web-content-filtering-project.md) - [Follow-Up Technical Vision: Optimizations, Deployment, and Security](../2025/06/13/follow-up-technical-vision-optimizations-deployment-and-security.md) - [Technical Debrief: Evolution from Electron‑Based to Python‑Native Web Content Capture App](../2025/05/22/technical-debrief__evolution-from-electron-based-to-python-native-web-content-capture-app.md) - [Project: Electron-Based Web Content Capture App (with Playwright & Python)](../2025/05/21/project__electron-based-web-content-capture-app-with-playwright-and-python.md) - [Project: Web Content Capture Extension with Pyodide and Serverless Backend](../2025/05/18/project__web-content-capture-extension-with-pyodide-and-serverless-backend.md) ### Business Transformation - [Project GenBnB: Enhancing Airbnb Host Workflows with GenAI](../2025/05/04/project-genbnb__enhancing-airbnb-host-workflows-with-gen-ai.md) - [Project InsightFlow: GenAI-Powered Transformation of Regulatory and News Feeds](../2025/05/03/project-insightflow__genai-powered-transformation-of-regulatory-and-news-feeds.md) - [Project Agenda: GenAI-Powered Transformation of Meetings and Documentation](../2025/04/10/project-agenda__gen-ai-powered-transformation-of-meetings-and-documentation.md) ### Data Integration & Synchronization - [Project Lumos: Serverless JIRA-to-GraphDB XYZ Connector](../2025/02/13/project-lumos__serverless-jira-to-graphdb-xyz-connector.md) - [Project JSync: JIRA Exporter and Synchronization System](../2025/02/09/project-jsync__jira-exporter-and-synchronization-system.md) - [Project StartLLM: Technical Proposal for 5x GenAI Projects](../2025/03/01/project-startllm__technical-proposal-for-5x-genai-projects.md) - [Project StartLLM: UC-01: Pentesting Insights Acceleration (PIA)](../2025/03/01/project-startllm__uc-01__pentesting-insight-acceleration.md) ## Strategic Partnerships - [Proposal: Strategic AWS Partnership with Dinis Cruz's GenAI and Graph Innovations](../2025/06/03/proposal-strategic-aws-partnership-with-dinis-cruz-genai-and-graph-innovations.md) - [Proposal for Neo4j Collaboration with Dinis Cruz](../2025/06/03/proposal-for-neo4j-collaboration-with-dinis-cruz.md) - [Jira as a Graph Database – Proposal for Atlassian Executives](../2025/06/03/jira-as-a-graph-database—proposal-for-atlassian-executives.md) - [Ephemeral Neo4j Instances for On-Demand Graph Analytics](../2025/06/25/ephemeral-neo4j-instances-for-on-demand-graph-analytics.md) ## Business Ideas & Plans - [GenLegalAdvise Project Plan](../2025/10/03/genlegaladvise-project-plan.md) - [GenAI Legacy Code Refactoring – Business Plan](../2025/06/07/genai-legacy-code-refactoring-business-plan.md) - [Project Plan: High Street GenAI Learning Hub](../2025/06/22/project-plan-high-street-genai-learning-hub.md) - [LinkedIn Vault: Professional Data Preservation Service](../2025/03/03/linkedin-vault__professional-data-preservation-servic.md) - [Scaling a Solo Cybersecurity Consulting Practice: Business Plan Research](../2025/04/07/scaling-a-solo-cybersecurity-consulting-practice__business-plan-research.md) ## Research & Technical Briefs ### LLMs & Compliance - [LLM-Driven GDPR Compliance Q&A Graph – Technical Brief](../2025/07/03/llm-driven-gdpr-compliance-q-and-a-graph-technical-brief.md) ### Pricing & Business Models - [Usage-Based Billable Entities: Aligning SaaS Pricing with Customer Usage](../2025/07/04/usage-based-billable-entities-aligning-saas-pricing-with-customer-usage.md) ### Market Research - [Research - AI-Powered Customer Service Solutions for Multi-Property Airbnb Hosts](../2025/05/04/research__ai-powered-customer-service-solutions-for-multi-property-airbnb-hosts.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/research-document.html (markdown twin: /research/research-document.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/research-document.html* # Research Document Catalog > Every research proposal and technical brief by Dinis Cruz in one table, by month, with a one-line summary and tags for each. --- This page summarizes recent research proposals and technical briefs. Articles are grouped by month with the most recent entries listed first. ## May 2025 | Date | Title | Description | Tags | | --- | --- | --- | --- | | 2025-05-18 | [Project: Web Content Capture Extension with Pyodide and Serverless Backend](../2025/05/18/project__web-content-capture-extension-with-pyodide-and-serverless-backend.md) | Prototype browser extension leveraging Pyodide and a serverless backend to capture webpage content, including a debrief on why the initial approach failed. | web-extension, pyodide, serverless, content-capture, prototype | | 2025-05-04 | [Project GenBnB: Enhancing Airbnb Host Workflows with GenAI](../2025/05/04/project-genbnb__enhancing-airbnb-host-workflows-with-gen-ai.md) | Explores using generative AI to automate property listings and customer interactions so Airbnb hosts can focus on high-value tasks. | airbnb, generative-ai, automation, property-management, customer-interaction | | 2025-05-03 | [Project InsightFlow: GenAI-Powered Transformation of Regulatory and News Feeds](../2025/05/03/project-insightflow__genai-powered-transformation-of-regulatory-and-news-feeds.md) | Describes a semantic knowledge graph pipeline that ingests regulations and news, using generative AI to deliver personalized client newsletters. | regulations, news, semantic-knowledge-graph, genai, personalization | ## April 2025 | Date | Title | Description | Tags | | --- | --- | --- | --- | | 2025-04-10 | [Project Cybersage: AI-Powered Risk Contextualization & Security Reporting](../2025/04/10/project-cybersage__ai-powered-risk-contextualization_security-reporting.md) | Proposal for an AI platform that contextualizes vulnerability data and generates clear, risk-based security reports for technical and executive audiences. | risk-contextualization, vulnerability-reporting, AI, security, automation | | 2025-04-10 | [Project Agenda: GenAI-Powered Transformation of Meetings and Documentation](../2025/04/10/project-agenda__gen-ai-powered-transformation-of-meetings-and-documentation.md) | Introduces a GenAI-driven workflow that prepares meetings, generates personalized briefs and tracks action items to reduce wasted time. | meetings, documentation, genai, action-items, productivity | ## March 2025 | Date | Title | Description | Tags | | --- | --- | --- | --- | | 2025-03-01 | [Project StartLLM: Technical Proposal for 5x GenAI Projects](../2025/03/01/project-startllm__technical-proposal-for-5x-genai-projects.md) | Strategic proposal outlining five focused generative AI projects designed to deliver quick wins and accelerate organizational adoption. | genai, strategy, adoption, quick-wins, proposals | | 2025-03-01 | [Project StartLLM: UC-01: Pentesting Insights Acceleration (PIA)](../2025/03/01/project-startllm__uc-01__pentesting-insight-acceleration.md) | Explores how generative AI can streamline pentesting workflows by summarizing findings and highlighting priorities for faster remediation. | pentesting, genai, workflow, summarization, security | ## February 2025 | Date | Title | Description | Tags | | --- | --- | --- | --- | | 2025-02-15 | [Project SupplyShield: GenAI-Driven Supply Chain Risk Management and Compliance](../2025/02/15/project-supplyshield__genai-driven-supply-chain-risk-management-and-compliance.md) | Introduces an AI-powered third-party risk platform that uses generative models and knowledge graphs to deliver continuous supply chain compliance. | genai, supply-chain, risk-management, compliance, knowledge-graph | | 2025-02-13 | [Project Lumos: Serverless JIRA-to-GraphDB XYZ Connector](../2025/02/13/project-lumos__serverless-jira-to-graphdb-xyz-connector.md) | Working Backwards plan for an open-source connector that streams JIRA data into GraphDB XYZ using a fully serverless architecture. | JIRA, GraphDB, serverless, data-streaming, open-source | | 2025-02-09 | [Project JSync: JIRA Exporter and Synchronization System](../2025/02/09/project-jsync__jira-exporter-and-synchronization-system.md) | Serverless pipeline that captures JIRA issue changes in real time and stores them in S3 and GitHub, exposing a FastAPI interface for querying the data. | serverless, JIRA, data-synchronization, AWS, devops | --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /research/the-future-of-news.html (markdown twin: /research/the-future-of-news.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/research/the-future-of-news.html* # The Future of news > Articles on news monetization, fact provenance, and trust. --- ## Research - [From Free Scraping to Fair Compensation: Cloudflare’s GenAI Crawler Charges and the Future of News Monetization](../2025/07/04/from-free-scraping-to-fair-compensation-cloudflares-genai-crawler-charges-and-the-future-of-news-monetization.md) - [Monetising Trust and Knowledge: How News Providers can leverage Personalised Semantic Graphs](../2025/02/02/monetising-trust-and-knowledge-for-news-providers.md) - [Journalists' Challenges with Digital Content Provenance and Trust](../2025/03/24/journalists-challenges-with-digital-content-provenance-and-trust.md) - [The Future of News: Building Trust Through Fact Provenance](../2025/02/05/the-future-of-news-building-trust-through-fact-provenance.md) - [The Future of News Monetization: Embracing Micro and Nano Payments](../2025/04/02/the-future-of-news-monetization__embracing-micro-and-nano-payments.md) - [Strengthening Trust in News: Implementing Identity Graphs for Authors and Sources](../2025/04/21/strengthening-trust-in-news__implementing-identity-graphs-for-authors-and-sources.md) ## Personalised briefings - [Personalised Briefing for Dan Raywood on the Future of News](../2025/06/06/personalised-briefing-for-dan-raywood-on-the-future-of-news.md) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /resources/presentations.html (markdown twin: /resources/presentations.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/resources/presentations.html* # Presentations > Talks and presentations by Dinis Cruz on GenAI, application security and semantic knowledge graphs, from OWASP AppSec EU Lisbon 2024 to The Grafter meetup. --- - [Using GenAI to graph and map your company’s data @ The Grafter meetup](../2025/05/07/using-genai-to-graph-and-map-your-companys-data__the-grafter__may_2025.md) (2025/05/07) - [Semantic OWASP: Leveraging GenAI and Graphs to Customise and Scale Security Knowledge](../2025/04/23/semantic_owasp__leveraging_genai_and_garphs_to_customise_and_scale_security_knowledge.md) (2025/04/23) - [My Journey Building a GenAI Startup: The Power of MVPs and CI Pipelines - PART 2](../2025/02/26/my-journey-building_a_genai_startup__the-power-of-mvps-and-ci-pipelines__part-2.md) (2025/02/26) - [My Journey Building a GenAI Startup: The Power of MVPs and CI Pipelines - PART 1](../2025/01/29/my-journey-building_a_genai_startup__the-power-of-mvps-and-ci-pipelines__part-1.md) (2025/01/29) - [Deterministic GenAI Outputs with Provenance (OWASP EU AppSec Lisbon )](../2024/06/28/deterministic-genai-outputs-with-provenance__owasp-appsec-lisbon__gslides.md) (2024/06/28) - [It’s 2024 and, with GenAI, we can finally make AppSec work](../2024/02/22/its-2024-and-with-genai-we-can-finally-make-appsec-work.md) (2024/02/22) --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/10/02/project-voice2siem-turning-customer-support-audio-into-real-time-security-events.html (markdown twin: /2025/10/02/project-voice2siem-turning-customer-support-audio-into-real-time-security-events.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/10/02/project-voice2siem-turning-customer-support-audio-into-real-time-security-events.html* # Project Voice2SIEM: Turning Customer Support Audio Into Real-Time Security Events *By Dinis Cruz and ChatGPT Deep Research · 2025-10-02* > Phone-based social engineering and vishing (voice phishing) attacks are on the rise, targeting customer support and help desk agents. Attackers impersonate customers or executives to manipulate… --- [PDF](https://files.diniscruz.ai/github/pdf/2025/10/02/project-voice2siem-turning-customer-support-audio-into-real-time-security-events.pdf) ## Problem and Architecture Overview Phone-based **social engineering and vishing (voice phishing)** attacks are on the rise, targeting customer support and help desk agents. Attackers impersonate customers or executives to manipulate agents into divulging sensitive information or performing unauthorized actions (password resets, fund transfers, etc.). With AI-driven voice cloning, fraudsters can even mimic a victim's voice to fool biometric checks or **bypass security questions by social engineering the agent**[[1]](https://www.tsys.com/insights/2024/10/22/new-voice-fraud-cloning-techniques-expose-a-vulnerability-of-call-centers#:~:text=Call%20center%20fraud%20is%20when,into%20giving%20them%20customer%20information). The business impact is severe -- breached accounts, financial loss, compliance violations -- and **detecting these threats during a call is extremely challenging**. Human agents may miss subtle cues of manipulation, especially if the caller is persuasive or the attack blends into normal support workflows. The technical challenge is to **analyze voice conversations on-the-fly for signs of phishing or manipulation**. This means turning an audio stream into structured data (transcripts, sentiment, detected intents) that security systems can evaluate. A straightforward approach is to treat each support call as an input to a pipeline of detection stages. The pipeline would transcribe the audio, analyze linguistic content and vocal tone for *urgency or stress*, extract semantic information (e.g. what requests are being made, what account or device is discussed), and then apply logic to decide if the call is suspicious. If a threat is likely, the system should generate a **security event** (e.g. alert in the Security Information and Event Management system, SIEM) so the incident can be investigated or action can be taken (like warning the agent or supervisor). **Real-time vs "fast enough" detection:** Importantly, this pipeline does **not** need to operate in true real-time (i.e. within milliseconds). Unlike a network firewall, we don't necessarily have to stop the conversation mid-sentence. The goal is to be *fast enough* -- ideally analyzing the call **within seconds to a minute** -- so that if a high-risk social engineering attempt is detected, we can intervene *before the attacker's goal is achieved*. For example, if the caller is trying to convince the agent to reset a password or reveal an OTP (one-time passcode), detecting the threat just in time could prompt a supervisor to intervene or require additional verification. In essence, as long as the system flags the call **before a fraudulent transaction completes**, it's effectively real-time for security purposes. ### Architectural Workflow At a high level, **Project Voice2SIEM** proposes a pipeline with multiple stages, each transforming the raw audio into more refined signals (see **Figure 1**). By the end, the system has enough understanding of the conversation to decide if it warrants a security alert. The stages include: 1. **Audio Input Capture:** The customer support call (which may be telephone audio or VoIP) is recorded or streamed into the system. This could be a real-time audio feed from the telephony system or a recording saved after the call. 2. **Speech-to-Text Transcription:** The raw audio is converted to text using Automatic Speech Recognition (ASR). This produces a transcript of the call with timestamps and speaker identification (who said what). Accurate transcription is crucial since all further analysis relies on the textual content. 3. **Urgency, Tone, and Emotion Detection:** Beyond the words spoken, the system analyzes *how* things are said. It looks at indicators of stress or emotion in the audio (e.g. elevated volume or pitch, rapid speech suggesting urgency, or pauses suggesting hesitation). In the transcript, this can be supplemented by sentiment analysis -- e.g. detecting if the **customer is angry, anxious, or insistent**[[2]](https://aws.amazon.com/transcribe/call-analytics/#:~:text=With%20Amazon%20Transcribe%20Call%20Analytics%2C,the%20issue%20was%20addressed%2C%20and). A manipulative caller might sound either unusually pushy or feign distress; these vocal features provide early clues. 4. **Semantic Extraction of Content:** In parallel with tone analysis, the transcript is parsed for **topics, intents, and requests** being made. For example, the system extracts what the caller is asking for ("reset my password", "update my email", "check a charge on my account"), any mention of sensitive information (account numbers, one-time passcodes), and entities like names or organizations. Natural Language Processing (NLP) techniques identify key entities and the overall intent of the call. The system might also detect conversation acts (caller provided authentication info, agent verified identity, caller requested an exception to policy, etc.). 5. **Conversation Graph Generation:** All the extracted data -- semantic facts, entities, the dialog sequence, emotional indicators -- are aggregated into a **graph representation** of the conversation. In this graph, nodes might represent the participants (customer, agent), the **utterances** (individual statements or questions), and important entities (like an account or ticket number). Edges can illustrate the flow of the conversation ("agent asks for verification after customer request") or relationships ("caller identity provided matches account name"). The graph format makes it easier to see the overall structure of the interaction and to apply pattern-matching for known attack scenarios. 6. **Scoring and Decision for SIEM Escalation:** Finally, a decision engine evaluates the graph and all collected signals to produce a **threat score**. This could be a simple rules engine or a machine-learning model trained on past call data. It considers factors like: *Did the caller exhibit high stress or urgency while requesting a sensitive action?* *Were there anomalies in authentication (e.g. multiple failed attempts, or the caller bypassing questions)?* *Does the sequence of events match known fraud playbooks (such as the "CEO impersonation" scam)?* If the score exceeds a threshold or certain red flags are present, the system generates a **security event** that is sent to the SIEM. This event would contain the pertinent details (timestamp, call ID, transcript highlights, reason for alert) so security analysts or automated responders can take action. *Figure 1: Conceptual Voice2SIEM pipeline converting a support call into a security alert. The process begins with audio input from a customer-agent conversation. The audio is transcribed to text, then analyzed for emotional tone (e.g. urgency, anger) and semantic content (topics, intents, requests). These insights feed into a* *conversation graph* *representing the flow of the call. Finally, a scoring logic decides if the graph indicates a likely social engineering attack, triggering a SIEM event.* This architecture acknowledges that **no single signal is conclusive**. A caller might be angry for legitimate reasons, or mention sensitive info as part of normal troubleshooting. It's the **combination of signals and the context** that reveals a social engineering attempt. For example, an attacker might start very friendly (low anger, positive sentiment) then suddenly create urgency ("I need this done *immediately* or I'll get you fired!") -- a stark sentiment change coupled with a policy exception request. The conversation graph for such a call would show an abnormal sequence (from small talk to threats) that differs from genuine customer calls. By breaking down the audio and analyzing these aspects step by step, the system can catch what a human agent alone might miss. ## Solution Using Commodity Cloud Services (AWS Reference Implementation) Designing and implementing the above pipeline from scratch can be complex. Fortunately, modern cloud providers offer building blocks that can be assembled to create this voice-to-SIEM analysis system. To demonstrate, we outline a reference solution using **Amazon Web Services (AWS)** serverless components (note that similar services exist on Azure and Google Cloud, so the approach is portable). The emphasis is on using **managed, pay-per-use services** to achieve this inexpensively at scale. We don't need to maintain servers or deep ML expertise -- we can leverage cloud AI APIs that are readily available[[3]](https://github.com/aws-samples/amazon-transcribe-live-call-analytics#:~:text=Amazon%20machine%20learning%20services%20like,You%20figure%20that). **Key AWS components in the pipeline:** - **Audio Ingestion (S3 + Lambda):** A call recording (for post-call analysis) or a live audio stream is fed into the pipeline. In AWS, one simple method is to use Amazon S3 as an ingestion point: the call system saves the audio file (e.g. MP3/WAV) to an S3 bucket at the end of the call. This event triggers an AWS Lambda function (via S3 event notification) to start processing. For live calls, AWS Kinesis Streams or Amazon Chime SDK can stream audio in real-time, but for our "fast enough" approach a short post-call delay is acceptable. The Lambda retrieves the audio from S3 and initiates transcription. - **Speech-to-Text with Amazon Transcribe:** The Lambda function uses **Amazon Transcribe** to convert the audio to text. Amazon Transcribe can operate in batch mode (transcribing a file from S3 asynchronously) or streaming mode (transcribing a live stream in real-time). In our reference design, the Lambda could call the Transcribe API to start a transcription job on the audio file. The result will be a transcript file (JSON or TXT) -- possibly stored back in S3 or returned to the Lambda after completion. Transcribe can provide word-by-word timestamps and distinguish between speakers (important for identifying "who said what" in the dialogue). - **Tone and Sentiment Analysis with Comprehend:** Once the transcript is ready (within seconds for an average call), another Lambda step (or the same Lambda if orchestrated sequentially) analyzes the text. **Amazon Comprehend**, a natural language AI service, can detect sentiment (positive, negative, neutral, mixed) of text and extract key phrases. Comprehend's sentiment analysis helps determine if the caller was angry, frustrated, or urgent during the call. This can be done at the overall call level and even per sentence to see if sentiment shifted over time. Additionally, Comprehend can perform entity recognition -- identifying names, dates, organizations mentioned -- which could flag if the caller mentioned things like a specific bank, a password, an account number, etc. These become pieces of metadata attached to the call record. - **Intent and Keyword Extraction:** While Comprehend covers basic NLP, AWS offers other tools for deeper insight. Amazon Transcribe itself has a feature called **Call Analytics** that can directly identify **call categories and issues**, like spotting if certain phrases (e.g. "not happy", "speak to manager") occurred[[4]](https://aws.amazon.com/transcribe/call-analytics/#:~:text=customer%20and%20agent%20sentiment%2C%20call,names%2C%20addresses%2C%20and%20credit%20card). Alternatively, one could use **Amazon Lex**, a conversational AI service, to parse the transcript or even actively listen during the call to identify intents. For example, Lex could be configured with intents such as "PasswordResetIntent", "VerifyIdentityIntent", "AccountUnlockIntent" etc., and the transcript (or live audio via Lex integration) would reveal if the caller's requests match any of these. The combination of Comprehend and Lex can thus provide a structured view of *what the caller was trying to achieve*. AWS's own blog notes that using services like Transcribe with Comprehend and Lex enables capturing both **insights and intents from conversations**[[3]](https://github.com/aws-samples/amazon-transcribe-live-call-analytics#:~:text=Amazon%20machine%20learning%20services%20like,You%20figure%20that). - **Orchestration and Data Flow:** All the above steps can be orchestrated with AWS Lambda functions passing data through Amazon S3 or in-memory. A simple approach is a **pipeline of Lambdas** triggered in sequence: audio file lands in S3 -> triggers Transcription Lambda -> writes transcript to S3 -> triggers Analysis Lambda -> which calls Comprehend/Lex -> outputs findings to a results database. AWS Step Functions (a serverless workflow service) could also manage this multi-step pipeline with error handling and retries built-in. Each service in the pipeline is pay-per-use: you pay only for the seconds of Lambda execution, the seconds of transcription, and the Comprehend API calls used, making it cost-efficient. - **Graph and Pattern Analysis:** In an AWS-only implementation, we might not explicitly build a graph database for the conversation (since that introduces a stateful component). However, we can simulate the graph analysis through structured data and search indices. For example, after analysis we could construct a JSON object that represents the conversation structure (speakers, sequence of intents, sentiment timeline, entities mentioned). This JSON could be indexed in **Amazon OpenSearch** (the AWS-hosted Elasticsearch service) to enable complex queries and visualization. OpenSearch can serve as a lightweight SIEM database where all call transcripts and alert scores are stored. Analysts could search this index for patterns (like all calls where a password reset was requested and caller sentiment was angry). If needed, Amazon Neptune (a managed graph DB) could be used to store the conversation graph and run graph queries -- but that adds complexity, so our reference keeps it simple with JSON data and OpenSearch. - **Scoring and SIEM Alerting:** The final step is deciding if a particular call is malicious. This logic can run in a Lambda function once all analysis data is available. The Lambda might use a set of rules (e.g., IF caller_sentiment = "angry" AND requested_action = "password_reset" AND auth_failed = true THEN high_risk) to compute a risk score. More advanced, it could use a machine learning model (perhaps SageMaker or even a Comprehend custom classifier trained on examples of fraudulent vs. normal calls). But clear, explainable rules are a good starting point. If the call is deemed suspicious, the system creates a **security event record**. In AWS, this could be an entry sent to **Amazon EventBridge**, which can route the event to various targets: an SNS notification to Security Engineers, a ticket in an incident management system, or even automatically calling an AWS Lambda to disable the customer's account until verified. Alternatively, the Lambda could index an "alert" document into the OpenSearch index with a field like `alert=true` so it shows up in SIEM dashboards. The key is that a tangible alert or log is generated and integrated with whatever SIEM or logging solution the company uses (could be Splunk, Elastic, Datadog, etc., via webhooks or connectors). - **Modularity and Cloud Portability:** All components here are loosely coupled. Audio files and transcripts reside in S3 (or could be any object store), triggers are event-driven, and each analysis piece is a replaceable module. For instance, if one wanted to switch to Google Cloud, you could use Cloud Storage instead of S3, Google's Speech-to-Text instead of Transcribe, and Cloud NLP for sentiment/intent instead of Comprehend/Lex. The pipeline concept remains the same. Because it's serverless, scaling to thousands of calls is simply a matter of AWS handling more Lambda invocations and more parallel transcribe jobs -- no infrastructure bottlenecks. Cost-wise, using these services means you pay per call-minute processed, which for sporadic or moderate call volumes is extremely cost-effective compared to hiring a team of human monitors. **Data Pipeline Example (AWS):** Below is a summary of how data flows in this serverless design: 1. **Ingestion:** Agent software or telephony records the call audio and uploads `call123.mp3` to S3. 2. **Transcription Trigger:** S3 event kicks off Lambda "TranscribeCall". It calls Amazon Transcribe to transcribe the audio file (language can be auto-detected if needed). Transcribe outputs `call123-transcript.json` to an S3 `transcripts/` folder. 3. **Analysis Trigger:** The upload of the transcript JSON triggers Lambda "AnalyzeCall". This function loads the transcript (could also get it from the Transcribe API result) and calls Amazon Comprehend for sentiment and entities. It may also invoke Amazon Lex (or a custom intent classifier) with the transcript to get recognized intents (e.g. `intent: ResetPassword`). The function compiles an analysis result JSON with fields like `sentiment_trend`, `key_phrases`, `detected_intent`, `caller_tone`, `auth_attempts`, etc. It saves this to S3 or sends it directly to the next step. 4. **Scoring & Alert:** A final Lambda "ScoreCall" (triggered by the presence of analysis results) evaluates all inputs. It might say: *sentiment went from neutral to highly negative when agent asked security question*, *caller requested high-risk action*, *caller provided account info after failing initial verification*. These factors are tallied into a risk score (say 85/100). If above threshold (e.g. 80), the Lambda sends an event to EventBridge with details (`alert: true, risk:85, call_id:123, reason:"high urgency password reset"`). EventBridge forwards it to the security notification topic and also logs it into OpenSearch. If the score is low, the call record might just be logged in OpenSearch with `alert: false, risk:10` for learning/tracking. 5. **SIEM Integration:** In our case, Amazon OpenSearch acts as a simple SIEM data store where all calls and alerts reside. Security teams can query and visualize this (e.g. see trends, or get an alert feed). If the organization has an existing SIEM (Splunk, QRadar, etc.), the EventBridge rule could instead call a webhook or use a connector Lambda to forward the event to that system in real-time. Through this AWS solution, we achieve a functional **voice-to-SIEM pipeline using off-the-shelf services**. We leveraged **Transcribe for speech-to-text**, **Comprehend for NLP insights**, and simple logic in Lambda for the decision -- demonstrating that even complex-sounding capabilities like emotion detection or intent recognition are accessible via APIs[[3]](https://github.com/aws-samples/amazon-transcribe-live-call-analytics#:~:text=Amazon%20machine%20learning%20services%20like,You%20figure%20that). All data (audio, text, results) is centralized in S3/OpenSearch which provides an audit trail. Moreover, the **serverless architecture** means we can handle bursts of calls or scale down to zero when no calls are happening, paying only for actual usage. This cloud reference implementation proves the concept: organizations can start detecting social engineering in calls *today*, without waiting for a specialized vendor product, by composing existing cloud services. ## Open Source and Semantic Graph-Based Implementation While cloud services are convenient, some organizations prefer open-source, self-hosted solutions -- for flexibility, **transparency, and avoiding vendor lock-in**. Furthermore, to push the envelope, we can design a system that not only detects threats but does so in a **fully explainable, graph-driven manner**, aligning with cutting-edge semantic analysis techniques. In this section, we present an open-source blueprint using Dinis Cruz's stack of tools and frameworks, which are geared towards building **type-safe, graph-centric, and auditable AI systems**. The goal is to show how the Voice2SIEM pipeline can be constructed with open technologies, yielding a high degree of control and introspection into how decisions are made. The open-source solution will use the following key components: - **Cache Service (FastAPI-based)** -- for structured storage of audio, transcriptions, metadata, and event outputs. - **MGraph-DB (Memory Graph Database)** -- for building and querying the semantic graph of each conversation. - **Type_Safe Schema System** -- to define all data models (call records, transcript segments, analysis results, alerts) with strict types, ensuring consistency and traceability across the pipeline. - **MyFeeds.ai LETS Pipeline** -- an architectural pattern (Load-Extract-Transform-Save) that guides the data flow in deterministic, debuggable steps. - **Persona Modeling and Scoring** -- a layer that infers the "persona" or likely intent of the caller (legitimate customer vs. potential fraudster) and scores threat likelihood based on how the conversation aligns with known patterns. Let's discuss each in the context of Voice2SIEM: **Cache Service for Data Ingestion and Storage:** The Cache Service is a lightweight, fast key-value store accessible via API (built with FastAPI and OSBot-FastAPI extensions). We use it as the glue between pipeline stages. For example, when a call audio is received, it can be stored as a binary blob in the Cache (under some unique key like `call/123/audio`). When the transcription step runs (as a microservice or job), it pulls the audio from Cache, produces a transcript object, and saves that back to the Cache (e.g. under `call/123/transcript`). Similarly, analysis stages store their outputs (tone analysis, intent extraction results, etc.) in the Cache with versioning. The Cache Service essentially provides a **central state repository** so that each stage of the pipeline is decoupled (they communicate by reading/writing data via the Cache API). Because it's backed by a fast datastore and supports JSON natively, it's ideal for persisting the conversation data at each step. This also means every intermediate artifact is saved -- enabling **replay, auditing, and debugging**. If an alert is raised on a call, we have the full chain of data (audio -> transcript -> graph -> score) stored for forensic analysis or model improvement. **Transcription and Analysis (Open-Source Tools):** For transcription, one could use an open-source ASR engine. A leading choice is **OpenAI Whisper** (which has a high-accuracy model that can run on-premise GPUs or even on CPU for smaller models). There are also others like Kaldi or Vosk. In our blueprint, we'll assume using Whisper for speech-to-text, integrated into the pipeline as a service (perhaps a container that the Cache Service can call, or an offline batch process that writes results to Cache). For text analysis (sentiment, intent), we can leverage open-source NLP libraries or models: for sentiment, models from HuggingFace (e.g. a RoBERTa sentiment classifier) can be used; for keyword/entity extraction, spaCy or transformers can identify entities; for intent, one might train a simple classifier or use an LLM with prompts. The **persona modeling** can even utilize a **large language model** in a controlled way: for instance, use an LLM to parse the transcript and fill in a structured **"CallAnalysis" JSON schema** with fields like `suspected_attack_vector`, `caller_persona`, `important_entities`. By defining the schema and validating it (with Type_Safe classes), we ensure the LLM's output is structured and can be parsed deterministically (a technique proven in Dinis's MyFeeds.ai project, where LLMs populate predefined JSON schemas)[[5]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L8-L16). Each of these analysis steps writes its JSON results to the Cache. This approach means even if we use advanced AI (LLM) for analysis, we **capture its reasoning in data form**, rather than a black-box judgment. **MGraph-DB for Semantic Graph Construction:** Once the transcript and initial analyses are done, we consolidate the information into a **semantic graph** using MGraph-DB. MGraph-DB is an in-memory graph database optimized for JSON and Python usage[[6]](https://pypi.org/project/mgraph-db/#:~:text=MGraph,suited%20for). We create nodes for key elements of the call: e.g. a node for the *Caller*, a node for the *Agent*, nodes for each *Utterance* (with properties like timestamp, sentiment, speaker), nodes for important *Entities* (like *Account* or *PIN Code* if they were mentioned), and perhaps nodes for *Intents* or *Requests* (like *ResetPasswordAction*). Then we add edges to connect these: *Caller* --(speaks)--> *Utterance1*; *Utterance1* --(intent)--> *ResetPasswordAction*; *Utterance2* --(sentiment)--> *Angry*; *Caller* --(provides)--> *AccountNumberEntity*, etc. The exact schema of the graph can evolve, but the idea is to create a rich representation that can be traversed to answer questions like "Did the caller provide credentials?" or "How did the agent respond after the caller got angry?". MGraph-DB, being type-safe and in-memory, allows us to quickly build and query this graph within a Python service. Its **type-safe nature** ensures we only create valid node/edge types as defined in our schema (catching mistakes early)[[7]](https://pypi.org/project/mgraph-db/#:~:text=)[[8]](https://github.com/owasp-sbot/OSBot-Fast-API/blob/f3bc57390c988d6cce2d9bd5e16d9618beb00868/docs/dev/briefs/v0.26.1__developing-fastapi-service-clients.md#L41-L48). We can also serialize this graph to JSON for storage or debugging, since MGraph-DB supports JSON persistence. **Type_Safe Schemas and Data Models:** All data entities in this pipeline are defined as classes using the **Type_Safe** system (from OSBot). For example, we might define `class Transcript(Type_Safe): ...` with fields for call_id, full_text, segments, etc., or `class Utterance(Type_Safe): speaker, text, timestamp, sentiment`. By using Type_Safe, we get runtime-checked, self-documenting data structures that can seamlessly convert to/from JSON[[8]](https://github.com/owasp-sbot/OSBot-Fast-API/blob/f3bc57390c988d6cce2d9bd5e16d9618beb00868/docs/dev/briefs/v0.26.1__developing-fastapi-service-clients.md#L41-L48). This means when we pass data between services (or even within the graph DB), we do so in a structured manner. It also aids transparency: each piece of data can be logged or inspected with confidence in its format. The Type_Safe schema definitions essentially act as the **contract** for each pipeline stage's input/output. Moreover, this system helps with auditability: for instance, a `SecurityAlert` class might require certain fields (e.g. reason, score, timestamp, evidence graph reference), ensuring no alert is created without sufficient data. Sharing these schema classes between the pipeline components (the Cache service, the analysis code, the graph builder, etc.) guarantees consistency -- similar to how the client-server model in OSBot shares schemas to have a single source of truth[[8]](https://github.com/owasp-sbot/OSBot-Fast-API/blob/f3bc57390c988d6cce2d9bd5e16d9618beb00868/docs/dev/briefs/v0.26.1__developing-fastapi-service-clients.md#L41-L48). **LETS Pipeline Orchestration:** We adopt the **LETS (Load-Extract-Transform-Save)** architecture to manage the pipeline execution in clear steps[[9]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L18-L26)[[10]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L32-L35). Here's how it maps to Voice2SIEM: - **Load:** Bring in the raw data (audio file) and save it unmodified. In practice, when a call audio arrives, we "load" it by storing it in the Cache (this is analogous to saving the raw RSS feed in MyFeeds.ai's pipeline[[11]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L40-L45)). This ensures the exact original audio is preserved. - **Extract:** Derive initial structured data from the raw input. This would be the transcription step (audio -> text) and perhaps parsing the text into structured dialogues. The transcript (with speaker turns) is saved as a JSON in the Cache. We might also extract other low-level info, like a timeline of who spoke when, or a list of detected keywords. The key is this stage is about structuring the raw audio into data we can work with (similar to how MyFeeds extracted article JSON from raw RSS)[[12]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L26-L34). - **Transform:** Perform higher-level transformations to enrich the data with semantics and insights. In our case, several sub-stages of Transform happen: sentiment analysis, intent detection, building the conversation graph, and scoring the risk. Each of these can be its own Transform step that takes the output of Extract or a previous transform and produces a new artifact. For example, "Transform 1" might take the Transcript and produce a **SentimentTimeline** object (list of utterances with sentiment tags). "Transform 2" might take Transcript + SentimentTimeline and produce the **ConversationGraph** (using MGraph-DB). "Transform 3" might take the ConversationGraph and produce a **ThreatHypothesis** (which encapsulates the potential attack vectors identified, e.g. "caller impersonating CEO scenario"). Each of these transforms saves its output to the Cache (and could be an API call or microservice in the implementation). By chaining multiple fine-grained transformations, we make the system easier to debug and extend[[5]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L8-L16). Crucially, we **persist after each transform**, so intermediate data is always available for inspection. If the final decision seems wrong, we can go back to see, for example, what the conversation graph looked like or what the sentiment analysis found, to pinpoint which stage misinterpreted the data. - **Save:** In LETS, every stage saves its output, but finally we also **save the final results to their destination**. In this context, the ultimate "save" is to log the security event (if any) and related data in a permanent store. The Cache service could serve in this capacity (it might append the alert to an "alerts" collection in its storage). Or we might output it to a SIEM system or even just a JSON file repository. The Save step here emphasizes versioning and traceability: we label the final outputs with version IDs and timestamps. If six months later we want to audit why a certain call was flagged, we can retrieve the exact data versions that led to it. This strong provenance tracking is a core benefit of the LETS approach[[13]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L12-L16) -- every decision can be traced back through the pipeline with records of each intermediate state. Using LETS in this pipeline means the solution is **fully reproducible and debuggable**. If an error is found in our analysis logic, we can re-run that step on the saved inputs to test a fix. It also means we can gradually improve each stage (swap out a model, fine-tune a rule) and have confidence how it impacts the end result, since we can replay past calls through the pipeline offline if needed. **Persona Modeling and Threat Scoring:** One of the powerful ideas we introduce is modeling the personas involved in the call -- especially the caller, who could be genuine or malicious. Over time, the system can build profiles of legitimate customer behavior versus known attacker tactics. For instance, a legitimate customer might answer verification questions slowly but correctly, exhibit frustration *after* repeated problems, but will comply with security checks. In contrast, an attacker persona might **either** act overly authoritative ("I'm in a huge hurry, I'm the CTO, just reset my password now") or overly distressed ("I've been locked out and this is an emergency, please help, I can't answer all these questions!"). By creating a library of these **persona archetypes** (which could be informed by real fraud cases), the system can compare a live call's features to the personas. The **persona modeling system** (inspired by MyFeeds.ai's approach of creating persona interest graphs[[12]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L26-L34)) could represent each archetype as a set of expected behaviors or markers. For example, an *"Impersonator Executive" persona* might have markers like: high authority tone, expresses urgency, drops names of company executives, attempts to bypass normal process. A *"Distraught User" persona* might: sound panicked, mention personal crises, push the agent to bend rules out of sympathy. These can be encoded in a data structure (even as a graph or just a profile object). During the Transform stage, we take the conversation data and try to **match it against these persona profiles**. This could be done with rules or possibly an ML model that classifies the call into one of several personas. If a strong match to a malicious persona is found, that heavily influences the threat score. It's not a binary thing -- we might see partial matches to multiple personas -- so the scoring system can weigh various factors. For example: *caller matched 80% of "Impersonator Executive" traits + caller requested password reset (a high-risk action) + call came from an unusual phone number (metadata anomaly)* -- combine these to yield a, say, 95/100 risk score, clearly over the threshold. **Graph-Based Detection of Patterns:** Because we have the conversation as a graph in MGraph-DB, we can also directly apply graph algorithms or queries to spot suspicious patterns. For example, we could query for a subgraph where: a *Caller node* is connected to an *Account node* via a *provided credential* edge, but the *Agent node* is connected to a *Verification step* that failed. This pattern (caller provided some info, failed verification, but is still pushing) could be indicative of fraud. Another pattern: the time between *Caller's request* and *Agent's action* is very short because the caller pressured the agent (could be calculated via timestamps in the graph). Or a sequence pattern: *Caller asks innocuous question -> builds rapport -> then asks for a sensitive action.* In the graph, that might appear as three utterances where the first has neutral intent, second is small talk, third is high-risk intent; the presence of that sequence can be automatically flagged by traversing the graph or by converting the conversation into a sequence of intent labels and running it through a sequence pattern matcher. We can even search across calls: if the same caller (or same voice print, if we did voice analysis) appears in multiple incidents, the graph of their interactions across calls could expose a fraud campaign. **Examples of Detectable Social Engineering Signs:** To illustrate, here are a few scenarios and how our system would catch them: - *Sequence Pattern:* An attacker often follows a script. For instance, "Friendly introduction" -> "Problem statement" -> "Urgent request with flattery/threat". A normal call might not have such a scripted progression. Our semantic extraction would label each segment (greeting, verification, request, etc.), and the conversation graph would show an abnormal transition from a casual chat to a critical demand. If our rules know this pattern (perhaps from past examples), the system flags it. For example, *"caller suddenly transitioned from calm to urgent while requesting a policy exception"* could be a rule derived from sequence analysis. - *Stress Indicators:* Suppose the caller's voice analysis shows **elevated stress or anger whenever security protocols are mentioned** (like the agent asking to verify identity causes the caller to raise their voice or heart rate if that could be measured). The sentiment timeline might show spikes of negative sentiment aligned with those moments. This is a red flag -- a genuine user might be annoyed at verification but wouldn't typically become aggressive; an attacker often does when impeded. The system would note *high emotional variance correlated with security steps* as a risk indicator. - *Anomalous Metadata:* Beyond content, contextual metadata can be telling. If the call came in at 3 AM local time from an overseas IP (for VoIP) or the phone number isn't one normally associated with the customer's account, those could be captured in the data model. Perhaps the account's profile says typical call time is daytime and this is highly out of pattern. Or the caller claims to be in one city but the telephony data suggests otherwise. Our pipeline can ingest such metadata at Load time (if available) and attach it to the graph (e.g. a node for "CallOrigin" with attributes). Any anomaly here (especially in combination with suspicious dialog) increases the score. After all these analyses, the open-source system arrives at a **ThreatLikelihood score** for the call, along with an explanation of why. Thanks to the structured approach, this explanation can be very specific: e.g. *"Alert: 95% likely social engineering. Detected persona: Impersonating Executive. Evidence: Caller insisted on urgent action (ResetPassword) with high anger (sentiment -0.85) after failing verification twice; call origin UK London, but user account holder is in USA; sequence matched known fraud pattern #3."* This kind of rich explanation is what a graph-based and type-safe pipeline can provide. It's **transparent and auditable**, unlike a monolithic "AI black box" solution. Every piece of that explanation links to an artifact in our system (transcript, sentiment value, graph pattern, metadata record). The security team can drill down to confirm each fact (because the data is in the Cache/graph) -- building trust in the system's verdict. In fact, the intermediate outputs themselves can be used for auditor training or improving processes (maybe the company discovers certain verification steps often cause false alarms, and they adjust policy). Finally, the open-source pipeline would emit the event just like the cloud one -- perhaps writing an entry to the Cache's event store or sending a message to a SIEM connector. The major difference is that with open components, the organization can **own the solution end-to-end**: data stays on their servers, models can be customized, and the logic can be adapted to their unique needs or expanded (for example, integrating a voice biometric check if available, or linking to a database of known fraud caller IDs). All core components (Cache Service, MGraph-DB, OSBot Type_Safe, etc.) are open-source (Apache 2.0 or similar licenses) developed by Dinis Cruz and community, meaning there's no license cost and one can contribute improvements. MGraph-DB's design for high-performance in-memory operation ensures even complex graph queries or large calls can be handled efficiently[[6]](https://pypi.org/project/mgraph-db/#:~:text=MGraph,suited%20for). The Type_Safe framework makes sure our system's APIs and data are robust and error-checked at development time, reducing runtime surprises[[8]](https://github.com/owasp-sbot/OSBot-Fast-API/blob/f3bc57390c988d6cce2d9bd5e16d9618beb00868/docs/dev/briefs/v0.26.1__developing-fastapi-service-clients.md#L41-L48). ## Conclusion and Call to Action **Project Voice2SIEM** demonstrates that turning customer support conversations into actionable security intelligence is not only possible -- it's achievable today with open technology and a bit of integration work. By combining voice-to-text, AI-driven analysis, and graph-based correlations, we can shine a light on what has traditionally been a blind spot in security monitoring (the content of phone calls). Importantly, the approach we presented emphasizes **trust, transparency, and auditability**. Every decision the system makes is backed by data artifacts and clear rules, so security teams can trust the alerts and verify why an alert was raised by tracing through the stored intermediate states[[13]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L12-L16). This is in stark contrast to opaque "AI magic" solutions; here, we're effectively opening the black box and making it a glass box. We also made a point to ensure the system respects the practical constraints of real call centers: it doesn't disrupt the live call flow, it works in near real-time, and it produces alerts quickly enough to matter. Whether implemented with cloud serverless components or a fully open-source stack, the blueprint is modular and flexible. Companies can start small -- maybe begin by just transcribing calls and doing sentiment analysis as a pilot -- and gradually build up the full pipeline as confidence grows. Because each piece is replaceable, improvements in AI (say a better speech recognition model or a new NLP technique) can be plugged in without redesigning the whole system. The authors (Dinis Cruz and the ChatGPT Deep Research team) are releasing this as an open blueprint with **no commercial agenda** -- our intent is purely to advance the state of the art in security monitoring and inspire others to build upon it. All the mentioned open-source tools (Cache Service, MGraph-DB, OSBot Type_Safe, etc.) are available in open repositories, and we invite readers to explore them, contribute, or adapt them to their own projects. We believe this kind of system would be incredibly valuable if deployed broadly: imagine a world where every help desk call or IT support call is quietly analyzed for potential fraud, providing a safety net for human agents who might be socially engineered. Many high-profile breaches could have been mitigated or even prevented if such technology were in place to catch the tell-tale signs of a con artist on the phone. **Call to Action:** If you find this project intriguing, consider contributing to its development or trying it out in your environment. You could start by using AWS's AI services to get quick wins on call analysis, or if you're more experimentally minded, deploy the open-source components and run some recorded calls through it to see what insights surface. Share your findings, build custom persona profiles that fit your industry, and help refine the detection logic. Since this is an open effort, improvements by one can benefit many. Ultimately, securing the "human layer" of support interactions is a shared challenge -- let's collaboratively turn the tide against voice-based social engineering. **Voice2SIEM can be a community-driven shield**, and we welcome you to join us in building it. Together, we can make customer support channels safer through transparency, open tech, and a healthy dose of innovation. [[1]](https://www.tsys.com/insights/2024/10/22/new-voice-fraud-cloning-techniques-expose-a-vulnerability-of-call-centers#:~:text=Call%20center%20fraud%20is%20when,into%20giving%20them%20customer%20information) New voice fraud cloning techniques expose a vulnerability of call centers | TSYS [[2]](https://aws.amazon.com/transcribe/call-analytics/#:~:text=With%20Amazon%20Transcribe%20Call%20Analytics%2C,the%20issue%20was%20addressed%2C%20and) [[4]](https://aws.amazon.com/transcribe/call-analytics/#:~:text=customer%20and%20agent%20sentiment%2C%20call,names%2C%20addresses%2C%20and%20credit%20card) Amazon Transcribe Call Analytics | Transcripts & Insights | AWS [[3]](https://github.com/aws-samples/amazon-transcribe-live-call-analytics#:~:text=Amazon%20machine%20learning%20services%20like,You%20figure%20that) GitHub - aws-samples/amazon-transcribe-live-call-analytics: Amazon Transcribe Live Call Analytics (LCA) Sample Solution [[5]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L8-L16) [[9]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L18-L26) [[10]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L32-L35) [[11]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L40-L45) [[12]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L26-L34) [[13]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/6e0e278d80d0ea5cdb125f998590a3b2b248021d/docs/2025/05/27/lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md#L12-L16) lets__load-extract-transform-save__a-deterministic-and-debuggable-data-pipeline_architecture.md [[6]](https://pypi.org/project/mgraph-db/#:~:text=MGraph,suited%20for) [[7]](https://pypi.org/project/mgraph-db/#:~:text=) mgraph-db · PyPI [[8]](https://github.com/owasp-sbot/OSBot-Fast-API/blob/f3bc57390c988d6cce2d9bd5e16d9618beb00868/docs/dev/briefs/v0.26.1__developing-fastapi-service-clients.md#L41-L48) v0.26.1__developing-fastapi-service-clients.md --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/10/02/time-as-a-calibrator-of-credibility-and-trust-in-information-systems.html (markdown twin: /2025/10/02/time-as-a-calibrator-of-credibility-and-trust-in-information-systems.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/10/02/time-as-a-calibrator-of-credibility-and-trust-in-information-systems.html* # Time as a Calibrator of Credibility and Trust in Information Systems *By Dinis Cruz and ChatGPT Deep Research · 2025-10-02* > In an era of information overload and rampant misinformation, time emerges as a critical factor in determining what information we trust. Traditional approaches to credibility tend to evaluate a… --- [PDF](https://files.diniscruz.ai/github/pdf/2025/10/02/time-as-a-calibrator-of-credibility-and-trust-in-information-systems.pdf) ## Introduction In an era of information overload and rampant misinformation, **time** emerges as a critical factor in determining what information we trust. Traditional approaches to credibility tend to evaluate a statement at the moment it is made -- a snapshot judgment of truth or falsehood. Dinis Cruz envisions a more dynamic paradigm: treating **every statement as an evolving entity** whose credibility is calibrated by the passage of time and the evidence that accumulates (or erodes) around it. In this vision, facts, opinions, hypotheses, and data points are not static declarations but living components of a knowledge ecosystem, each with attributes that can be **objectively extracted, tracked, and updated over time**. Trust is not a binary label stamped at publication; it is an emergent property that grows or diminishes as statements are corroborated, disproven, or refined by subsequent information. This whitepaper articulates that vision for an audience of AI researchers, cybersecurity professionals, and journalists. We explore the philosophical underpinnings and technical concepts of using time as the **ultimate arbiter of credibility** in information systems. Rather than prescribing a particular implementation, we discuss the frameworks and principles -- from **information taxonomies** to **semantic knowledge graphs** -- that can enable such a system. Two real-world initiatives led by Dinis Cruz, **MyFeeds.ai** and **The Cyber Boardroom**, will serve as running examples. These projects demonstrate how concepts like the **LETS data pipeline**, **LLM-driven extraction**, and **persona-based modeling** can be applied to build information systems where provenance, transparency, and temporal evidence tracking are first-class features. The goal is to show how, by encoding the temporal evolution of knowledge, we can fundamentally enhance credibility assessment. Over time, truth finds a way of asserting itself -- and our systems should be designed to capture that, providing a **clearer signal amid the noise**. In the following sections, we introduce a classification of information types, outline how AI (especially large language models) can extract and monitor these over time, and delve into the architecture of systems that put this into practice. We also highlight the importance of open-source, transparent infrastructure in earning trust. This paper is authored in a factual, professional tone, reflecting the style of prior whitepapers by Dinis Cruz, and is intended for publication on platforms like LinkedIn and Dinis's personal site. ## The Temporal Dimension of Trust Time has a unique role in the ecology of information: it is the **calibrator of credibility**. A statement made today might carry uncertainty; given weeks, months, or years, that statement could be bolstered by confirming data or undermined by contradictions. Consider a breaking news report that includes an eyewitness claim -- initially an unverified piece of information. Over subsequent days, investigations might provide evidence that either validates the claim as a fact or exposes it as false. The **credibility of the original claim is thus a function of time and evidence**. In science, a hypothesis must endure rigorous testing over time before it is accepted as proven; in journalism, initial reports are updated as more sources speak out or more documents come to light. Human societies have long used the test of time to judge ideas -- *"truth prevails in the end"* is a common refrain -- yet many modern information systems lack any notion of this temporal validation. In current social media and even some news feeds, **context collapse** is common: a two-year-old claim might circulate without indication that it was later debunked, or a prediction might still be treated as credible despite subsequent evidence against it. Dinis Cruz's vision addresses this gap by making the *timeline* of each piece of information an integral part of how machines assess trustworthiness. Any datum -- whether a factual assertion, an expert opinion, a hypothesis, or a raw statistic -- should carry with it a **history**: when it was first stated, what corroborations or refutations have appeared, and how its status has changed. By **explicitly tracking the evolution of information**, systems can present users with not just a claim, but the *current state of that claim's credibility*. In practical terms, this means building data architectures where **statements are objects with properties and links** that evolve. A claim might start with a low credibility rating, essentially a hypothesis awaiting verification. If multiple reputable sources later confirm it, the system elevates its status (and can even reclassify it from "hypothesis" to "fact"). Conversely, if credible evidence disproves it, the statement can be tagged as "disproven" or its credibility score lowered accordingly. Importantly, time-based trust doesn't imply that older information is automatically more credible -- rather, it means that **older information has simply had more opportunity to be tested**. A long-standing claim that has been continually supported by evidence earns a kind of trust that a fresh, untested claim cannot yet possess. By the same token, an old claim that has never been tested or that has languished without scrutiny might actually be less credible than a newer claim backed by immediate solid evidence. Thus, time is the calibrator in conjunction with evidence: it is the framework within which evidence accumulates. From a philosophical standpoint, this approach echoes the scientific method and investigative journalism -- continual gathering of proof and revisiting of prior assumptions. It acknowledges a reality often lost in AI systems: **knowledge is provisional** and our confidence in it should be proportional to the journey it has undergone. What Cruz proposes is essentially to bake that epistemological humility into information systems. Trust becomes a **dynamic metric**. Users (be they researchers, security analysts, or readers) could see not only *what* is known at a point in time, but *how* it came to be known and how that knowledge has changed. This temporal awareness is especially pertinent in cybersecurity and AI, where new vulnerabilities or discoveries can overturn "facts" quickly, and in journalism, where narratives evolve as stories develop. By treating time as a first-class dimension, we enable a richer form of credibility -- one that can be visualized as a timeline or knowledge graph rather than a static label. The sections below discuss how to systematically classify information and harness large language models and knowledge graphs to implement this vision. ## Classifying Information: Facts, Opinions, Hypotheses, and Data A foundational step in building time-aware credibility systems is establishing a clear **taxonomy of information types**. Not all statements are created equal -- a verifiable fact is different from a personal opinion, which is in turn different from a hypothesis or a raw data point. By categorizing statements, an system can apply the appropriate handling and validation logic to each. Dinis Cruz emphasizes four primary categories: **fact**, **opinion**, **hypothesis**, and **data**. Below we define each and explain their roles: - **Fact:** A statement about reality presented as objective truth, ideally supported by evidence. Facts are assertions that, in principle, can be verified or falsified. For example, *"The company's servers were breached on July 15, 2025"* is a factual claim. Facts initially may come with a confidence level (especially if just reported), and over time they can be corroborated by further evidence or challenged by contradictions. In our system, a fact would be linked to its sources (documents, interviews, sensors, etc.) and marked with its verification status. If later reports confirm it -- say, an official investigation verifies the breach -- the fact's credibility strengthens. If evidence emerges that the date was wrong or the breach never occurred, the statement's status would be downgraded (e.g., labeled as erroneous or retracted). Essentially, facts are **statements awaiting or bearing verification**, and their truthfulness is a function of evidence). Maintaining **traceability of facts to their evidence sources** is critical to this process. By tracking provenance, we ensure that every factual claim can point to *why* we believe it (or did at one time), aligning with Cruz's focus on transparency and provenance to minimize error and "hallucination" in AI. - **Opinion:** A statement of personal belief, interpretation, or judgment, which by nature is subjective. Opinions can be expert assessments (*"In my view, this security threat is being exaggerated"*), editorial commentary, or individual preferences. They cannot be "proven" true or false in the same way facts can, but they can be more or less persuasive or widely accepted. In an information credibility system, opinions are handled differently: rather than verifying them, the system might track *who* holds the opinion, their expertise, and how that opinion may shift over time or differ from other viewpoints. Opinions often add context or insight around facts (e.g., an analyst's opinion on why a breach happened). Classifying a statement as opinion ensures it isn't conflated with factual reporting. Large language models can be trained to detect opinionated language or phrases indicating subjectivity (e.g., "I think," "it is likely that," or tonal indicators), helping to label these correctly. Over time, one could even see how opinions trend -- for instance, initially a lone opinion might later become consensus (many others echo it) or remain controversial (sharply divided opinions). **Persona-based modeling** (discussed later) becomes crucial here, as understanding the source's identity and bias is key: the credibility of an opinion often depends on who expresses it. A statement like *"Our systems are secure enough"* carries different weight coming from a company's CEO versus an independent security researcher. Thus, opinions in the system carry metadata about their source and context, enabling users to factor in biases and perspective. - **Hypothesis:** A conjecture or tentative explanation that requires validation. Hypotheses often appear in investigative contexts (security analysts theorizing about an attack vector, scientists proposing a link between variables, journalists suspecting a cover-up). An example might be, *"The breach might have been an inside job,"* stated before any proof is available. Hypotheses are essentially questions framed as statements -- they signal *uncertainty and an invitation for further evidence*. In Cruz's approach, hypotheses are explicitly tagged as such and occupy an important place in the knowledge graph: they are nodes that expect evolution. As time passes, a hypothesis can be supported by facts (which might graduate it to accepted theory) or refuted by facts (leading to its rejection). One can imagine the system automatically updating a hypothesis's status when certain conditions are met -- e.g., if forensic data later show external IP addresses, the hypothesis of an inside job gets a lower credence or a "disproven" mark. Notably, hypotheses often drive the collection of new data. In the **Interactive Report Assistant** (one of Cruz's projects for AI-guided reporting), the AI explicitly captures hypotheses and even outstanding questions during a consultation, building a *knowledge base* that distinguishes between confirmed facts and speculative point. This structured capture means that the resulting report can label which findings are definitive and which are possible issues to investigate. By treating hypotheses as first-class citizens, an information system encourages a scientific mindset: everything is up for re-evaluation as new evidence comes in. Over a long term, tracking hypotheses can reveal how knowledge advances -- for example, a hypothesis from five years ago in medical research might now be a well-established fact, or might have been debunked, and that journey should be traceable. - **Data (Data Point):** A raw observation or measurement -- often numeric or categorical -- presented without interpretation. Data points are the building blocks of facts. For instance, *"Server logs show 5,000 failed login attempts between 1-2 AM"* is data. By itself, data may not be meaningful until placed in context (is 5,000 high? does it indicate an attack?). Data can also be statistical results, experimental readings, or quotes from sources. In our taxonomy, data points are captured and stored as evidence nuggets. They often feed into facts (supporting evidence for a factual statement) or can lead to new hypotheses (*given this unusual metric, could something be wrong?*). Ensuring data integrity and tracking its source is a key part of provenance. Data is often time-stamped, and its credibility might hinge on how it was collected (e.g., a properly calibrated instrument vs. an anecdotal report). In a dynamic credibility system, raw data might be reinterpreted over time: for example, an anomaly in data might later be explained by a calibration error (thus the data point would be flagged as faulty), or multiple independent datasets might confirm the same trend (boosting confidence in those numbers). Automation through AI can assist in extracting data points from text (via techniques like OCR for numbers in documents, or pattern recognition in text for "X% increase" statements) and storing them in the knowledge graph along with units and context. Over time, linking data points to the claims they support or refute is crucial. A single data point might be an outlier, but a time series or repeated measurement can solidify a fact. Hence, the **temporal tracking of data** itself -- noting how a metric changes -- can calibrate trust (for instance, a sudden spike in a security metric might at first seem like an attack, but if it drops the next day, the interpretation changes). By classifying statements into these categories, an information system gains clarity on how to treat each piece. **Facts** and **data** demand verification and are the basis of objective truth-seeking; **opinions** require context and source awareness; **hypotheses** call for monitoring and future resolution. This taxonomy is not just theoretical -- it is being operationalized in tools. The **Interactive Report Assistant** for example, uses a conversational AI to extract and differentiate facts, assumptions/hypotheses, questions, and evidence in real time while an expert conducts an assessment. It builds a structured representation (almost a mini knowledge graph) where each piece of captured information is labeled appropriately -- facts here, open questions there, etc. -- and even keeps track of *which statements have been confirmed by the user and which are tentative*. This ensures that when the final report is generated, every claim is either backed by confirmed input or clearly marked as an open issue, with the provenance of each fact traceable to the conversation snippet or document it came from. The ability to extract such a taxonomy from raw text is greatly enhanced by **Large Language Models (LLMs)**, which we discuss next. But even before the AI gets involved, having a human-understandable classification sets the stage for how information will flow through the system. It aligns with the principle that *different types of knowledge have different lifecycles*, and by recognizing that, the system can calibrate trust accordingly. A fact might have a lifecycle of verification steps; a hypothesis has a lifecycle of testing and either confirmation or abandonment; an opinion might shift with perspective or remain constant with its author; data points can accumulate into trends. Time affects each of these in unique ways, and a robust taxonomy is the first tool to manage that complexity. ## Extracting and Evolving Knowledge with LLMs Modern **Large Language Models** have demonstrated remarkable ability to read and interpret unstructured text. In the context of time-calibrated credibility, LLMs serve as the engine that **extracts structured knowledge** from raw information and helps update it as new inputs arrive. The idea is to leverage AI to do what humans do when researching: read documents (news articles, reports, transcripts), identify key claims and evidence, classify them (is this a fact? an opinion? who said it?), and flag their relationships to other information (does this support a prior claim? contradict it? raise a new question?). LLMs can perform or assist in all these tasks, making it feasible to maintain a rich, constantly updating knowledge base. **Initial Extraction:** When a new piece of content comes in -- say, a news article or an incident report -- an LLM can parse it and pull out the statements of interest. This involves natural language processing steps like entity recognition (finding the who/what/where), relation extraction (how entities relate, e.g. X attacked Y at time Z), and classification (is this sentence an assertion of fact or speculation?). For example, given a cybersecurity blog post about a newly discovered vulnerability, an LLM might extract: *Fact:* "A vulnerability CVE-2025-1234 was found in AcmeCorp's software;" *Opinion:* "The researcher believes it could be exploited widely;" *Data:* "70% of tested servers were affected;" *Hypothesis:* "It might be related to an earlier bug in a shared library." Each of these would be added as nodes or entries in the system, linked to the source document and time-stamped. Off-the-shelf LLMs (like GPT-based models or domain-tuned variants) can be prompted or fine-tuned to perform this multi-faceted extraction. Indeed, Dinis Cruz's projects use LLMs for content understanding tasks. In **MyFeeds.ai**, for instance, after raw articles are ingested, an LLM analyzes each article to extract **key entities, topics, and relationships**, effectively summarizing the article in a structured form. This creates a semantic representation (a mini knowledge graph) of each piece of content. Similarly, the **Semantic Content Filter** project uses an LLM (or a distilled model) to read web pages on the fly and generate a **semantic profile** of the page, identifying main topics and even classifying content by type (e.g. "this section is unverified information"). These practical uses underscore how LLMs can discern and label the components of information needed for our taxonomy. It's worth noting that while LLMs are powerful, they can also make errors or "hallucinate" facts. Therefore, Cruz's approach couples AI extraction with human or systematic verification steps and **provenance tracking**. The LLM might propose that "Statement X is a fact supported by Source Y," but the system will store that linkage so it can be checked. In the Report Assistant, every fact the AI captures is immediately shown to the user for validation or correction, creating a feedback loop that refines the AI's output in real time. This human-in-the-loop design ensures the knowledge base being built is accurate and trusted by its users. The emphasis on **traceability of facts to their origins** cannot be overstated -- as noted earlier, the Report Assistant was designed to maintain provenance for each extracted fact, and more broadly, all of Cruz's startups place importance on **provenance and explainability** of AI outputs. By keeping the source links, the system allows any claim to be audited: a user can ask *"How do we know this?"* and the system can point to the supporting evidence. This directly contributes to credibility, as trust in the system's information is reinforced when users see that nothing is conjured from thin air -- every assertion ties back to an input source or confirmed user input. **Temporal Updates:** Once the initial extraction has populated the knowledge base, the LLM's job is not done -- it now helps in **evolving that knowledge as new information arrives**. This is where time as a calibrator comes in. As the system ingests new content over days and weeks, it needs to reconcile it with existing knowledge. If a new source provides additional evidence for a fact, the system (with AI assistance) should link that evidence to the fact, possibly increasing a confidence score. If new content directly contradicts a previously stored claim, this conflict must be noted and ideally resolved (perhaps flagged for human review or majority-rules logic). LLMs can assist by analyzing the new information in context: for example, when a follow-up article comes, the model could recognize *"this statement from today's article refers to the same event as that statement from last week's article, but with a different detail"*. It might flag, *"Previously it was claimed 10,000 records were leaked; now this source says 5,000 -- discrepancy noted."* The system can then mark that fact as contested until further clarity. One way to implement this is through a **semantic knowledge graph** where each node (representing an entity or claim) accumulates links to sources and evidence over time. LLMs can help by merging nodes that appear to refer to the same real-world fact (e.g., "CVE-2025-1234" in various phrasings) and by creating relational links (e.g., "CVE-2025-1234 is confirmed by Source X on Date Y"). The **graph approach** is central to Cruz's solutions: MyFeeds.ai, the content filter, and the report assistant all convert unstructured text into a structured graph form. By doing so, it becomes much easier to **track the state of a particular node** (like a claim or an entity) as evidence is attached to it. The graph is a living data structure, and because it's stored in a database (or even as JSON files via the MemoryFS/GraphFS approach[17][18]), the system can update it incrementally. For instance, MyFeeds aggregates news and represents each article's knowledge as a graph; if another article comes in about the same vulnerability, the system can connect those graphs, enriching the picture of that vulnerability across sources[18]. Over time, if one of those articles turns out to be erroneous and is retracted, that could also be encoded in the graph (perhaps a property on the edge from that source saying "retracted" with a date). **Reclassification:** A particularly interesting aspect of using AI over time is the possibility of *reclassifying information as its nature changes*. A statement initially extracted as a *hypothesis* could later be reclassified as *fact* if evidence is found. The system might initially label *"It's likely an insider attack"* as a hypothesis. Later, if an investigation report definitively says "It was an insider," the system, via an update routine, could change that node from hypothesis to fact and annotate it with "confirmed by [source] on [date]." Likewise, opinions can shift to facts in some cases (for example, an expert's prediction that "X will happen" becomes a fact once X does happen). LLMs can be employed periodically or triggered by events to review the knowledge base: essentially asking the AI, *"Given this new document, do any existing nodes in the graph need to be updated or have their status changed?"* This is a complex task -- it requires maintaining consistent identifiers for the same real-world claims and having business logic about what constitutes sufficient evidence. Not all of this can or should be fully automated -- human oversight is valuable. However, the AI can do the heavy lifting of reading and comparing new text with stored assertions. For example, the system might store a hypothesis node: *"Cause of outage: misconfiguration (hypothesis)"*. When a postmortem report arrives a week later, the LLM might detect sentences like "The outage was caused by a misconfiguration of the firewall." It can then signal that this hypothesis is now confirmed by that report, prompting the system to mark it as a fact (and maybe move the original hypothesis node to an archive or link it as "now proven"). The system might also notify users or maintainers of this change -- effectively providing a **news feed of knowledge status updates** ("Hypothesis H has been confirmed as Fact, by source Z"). In this way, time and AI work together to calibrate what the system deems true. The **LETS pipeline** (Load, Extract, Transform, Save), which we will discuss more later, provides a structured way to carry out these regular updates[15][19]. Each time new data is loaded, the extraction and transformation steps include reconciling with existing saved knowledge. By breaking the process into discrete steps, it's easier to monitor and tweak how updates occur -- a design choice that favors transparency and control[20]. Dinis Cruz has highlighted that a structured pipeline improves *"transparency and tweakability -- critical for debugging AI decisions and maintaining provenance"*[20]. Indeed, when AI is used to adjust our knowledge base, we must be able to audit those adjustments. The provenance of every change (which AI suggestion or which source triggered it) should be logged, so that trust is maintained in the system's ongoing evolution. In summary, LLMs act as both the **miners of knowledge** (extracting structured claims from raw text) and the **maintenance crew** (continually comparing new information against the old and adjusting the structure). They work in concert with human experts and curated rules to ensure the knowledge base doesn't drift into error. By automating extraction and updates, we can keep pace with the torrent of information in domains like cybersecurity or global news, where no single person could manually track all the threads. The result is an always-current, evidence-weighted map of what is known and unknown. In the next section, we delve deeper into that map -- the semantic knowledge graph and the pipeline architecture that supports it, including how it handles provenance, source bias, and evidence trails over time. ## Semantic Knowledge Graphs and the LETS Pipeline Central to operationalizing a time-calibrated trust system is a robust **information architecture** that can store content, context, and connections. Dinis Cruz's approach relies on **semantic knowledge graphs** as the backbone for representing information, and a well-defined processing pipeline (called **LETS: Load, Extract, Transform, Save**) to manage how data flows from raw inputs to structured knowledge[15][17]. These components work together to ensure that every statement is captured with its provenance, that relationships (like support or contradiction) between statements are explicitly modeled, and that updates can be applied systematically as time goes on. ### Building a Knowledge Graph of Evidence A *semantic knowledge graph* is a graph data structure where nodes typically represent entities or concepts (people, organizations, events, claims, etc.) and edges represent relationships or interactions between them (e.g., "reported_by," "confirmed_by," "contradicts," "part_of"). By representing information in a graph, we gain two advantages for our purposes: **interconnectivity** and **explainability**. Interconnectivity means any given piece of information doesn't live in isolation -- it's linked to sources, related facts, and broader contexts. Explainability comes from the graph's ability to show why something was suggested or how a conclusion was reached (following the edges back to evidence). In Cruz's vision, whenever the system ingests an information source, it doesn't just store a blob of text. Instead, through LLM extraction, it **populates the graph** with structured data. For example, an article might produce nodes for the event it describes, the key actors involved, and the claims made, all connected by labeled edges indicating relationships (like "Actor A -> involved_in -> Event X"). If a claim is made, an edge from that claim node to the article's node could be labeled "asserted_in [source name/date]." If later another source confirms the same claim, another edge from the claim to that source is added, perhaps labeled "corroborated_by [other source]." Over time, the claim node accumulates a web of supporting or refuting links. This graphical approach directly supports **time-based credibility**: one can literally count or weight the number of confirmations vs refutations attached to a claim node, and factor in the credibility of each source node attached (we will address source credibility soon). The **semantic graph enables reasoning** -- for instance, the system can traverse the graph to find all evidence for a statement or to detect if two statements are in conflict (if they cannot both be true logically). MyFeeds.ai provides a clear example of semantic graph construction. It takes in articles and *"analyzes the article to extract key entities, topics, and relationships, forming a semantic representation (a graph) of the article's content"*[11][12]. So an article about a new ransomware might result in graph nodes like *Ransomware X*, *Healthcare Industry*, *Country Y*, *Data Breach*, with relationships encoding facts such as "Ransomware X caused a data breach in Healthcare sector in Country Y"[12]. This graph is then stored via a specialized interface called GraphFS (graph filesystem)[21]. The graph acts as a **machine-interpretable summary** of the article[22]. Why is this useful? Because now MyFeeds can do content matching at the level of concepts rather than keywords -- and it can explain recommendations by pointing to the graph overlap[23][24]. For instance, if a user's profile graph (more on persona graphs later) has a node "zero-trust architecture" and an article's graph has a related node "zero-trust security model," the system can match them even if the exact words differ, and then tell the user *"Recommended because it discusses zero-trust, which is one of your interests."* This transparency in *why* content was shown builds trust in the system's curation[24]. The broader point is that the graph is the structure that makes such explainability possible. Similarly, in our credibility context, a graph would allow the system to explain *why* a statement is considered credible or not: e.g., *"This claim is considered credible because it has been independently reported by three sources (Nodes A, B, C) and no contradictory evidence is in the graph."* If contradictory evidence exists, that too can be shown: *"However, one source (Node D) disputes this claim, hence it is marked contested."* This is far more informative than a simple true/false tag and provides nuance that professionals and analysts need. ### The LETS Pipeline -- Ensuring Structure and Provenance The **LETS (Load, Extract, Transform, Save) pipeline** is a disciplined workflow for data processing that Cruz employs to handle content ingestion and analysis[15][17]. It is inspired by classic ETL (Extract, Transform, Load) but adapted for the needs of AI-driven knowledge extraction. The phases are: 1. **Load:** Gather raw data from sources. In practice, this could mean fetching RSS feeds, API data, web pages, PDFs, or any input. For example, MyFeeds periodically pulls in new articles from dozens of cybersecurity news feeds[25][26]. In the context of our credibility system, Load might happen continuously -- new information is always coming (news articles, social media posts, internal reports) -- so this stage deals with connecting to those sources and retrieving the content, possibly in real-time or in batches. 2. **Extract:** Parse and clean the raw data to get it into a usable form. This often involves text extraction (removing HTML boilerplate, splitting into sections or sentences), metadata extraction (capturing author, publication date, etc.), and initial filtering. In MyFeeds, after loading, each item is parsed and its text and metadata are saved in a uniform storage (MemoryFS) and then converted into a JSON structure for further processing[27][28]. The idea is to have a structured representation of the content that can be fed into AI models. In our scenario, extraction also includes pulling out those facts/opinions/hypotheses as discussed earlier -- essentially preparing the inputs for the semantic analysis. By the end of Extract, we have the raw content distilled into something like: text segments, identified entities, and preliminary classifications. 3. **Transform:** This is where the heavy semantic lifting occurs -- using LLMs or algorithms to generate the knowledge graph entries and any additional annotations. The transform step takes the extracted content and applies the intelligence: it might call an LLM to generate the semantic graph of the content (as MyFeeds does)[12], or to classify each statement into our taxonomy, or to apply policies. In the Content Filter pipeline, for example, the transform step includes the LLM analyzing the page content and creating a semantic profile (a graph or metadata profile) of the page, including tags like "contains explicit language" or "phishing login form detected"[29][30]. In the credibility use-case, transform would link statements to prior knowledge (e.g., if it identifies a claim that's already known, it will connect it) and perhaps score them. If a source is known to be biased or unreliable, the transform step could note that -- e.g., *tag the incoming claims with a lower initial trust weight due to source reputation*. The transform phase is also where any complex logic like deduplicating information or reconciling conflicts can be executed. Essentially, transform turns input data into structured knowledge updates. 4. **Save:** Finally, the results of transform -- the updated knowledge graph nodes/edges, the cleaned content, the metadata -- are saved to persistent storage. This could be a graph database, a relational database, a set of JSON files, or a combination. Cruz's projects use abstractions like MemoryFS and GraphFS to simplify storing both file-like data and graph data uniformly[17][18]. Save is crucial for persistence (so that the next pipeline run knows the current state of the world) and for traceability. Saving doesn't just mean storing the final graph, but also logging the pipeline operations, recording timestamps, and possibly keeping previous versions (for auditing how something changed over time). By enforcing the LETS structure, the system gains **transparency and modularity**. Each step can be monitored and improved independently. For example, if some false information is getting through, one can check: did we fail to extract correctly, or was our transform rule not catching a contradiction? Because the pipeline is stepwise, debugging is easier than in a monolithic process. As noted in the multi-startup strategy document, *"This structured pipeline improves transparency and tweakability -- critical for debugging AI decisions and maintaining provenance."*[20] The controlled processing also means you can implement checks like saving intermediate outputs for review. In an environment where trust is paramount, this is important. We don't want a black-box AI magically altering our knowledge base; we want a clear record of how each piece of data was handled, what the AI suggested, and what was ultimately stored. Critically, **provenance** is woven throughout the pipeline. From the moment of Load, we tag every piece of information with its origin (source name, URL, timestamp, etc.). During Extract, if we break content into sentences, each sentence knows which document it came from. In Transform, when producing a knowledge graph node for a claim, we attach references to the source sentence or document node (the graph edge "asserted_in [source]" as mentioned). In Save, we ensure these references persist in the database. This way, any node or edge in the knowledge graph can be traced back to *who said it, where, and when*. Provenance is the bedrock of a credible system -- users and auditors can verify things for themselves. Dinis's emphasis on provenance in his designs (e.g., the Report Assistant explicitly stores which conversation snippet led to each fact[9]) aligns with best practices in both journalism and science, where citations and references are mandatory for trust. ### Content Provenance, Source Bias, and Evidence Trails With the graph and pipeline in place, the system can tackle the nuances of **source credibility and bias** as well as maintain **evidence trails over time**. Every source (be it a media outlet, a social media account, an internal document, or a sensor feed) can be represented as an entity in the graph with attributes (e.g., source type, known bias, reliability score). External data can be used to seed these attributes -- for instance, a news source might be annotated with its political leaning if known, or a scientific journal with its impact factor, etc. More dynamically, the system can compute a **trustworthiness score** for sources based on how their information pans out over time. If Source A's claims are frequently confirmed by others and rarely retracted, Source A's credibility rating could increase. If Source B has numerous claims later proven false or exaggerated, its rating would fall. These ratings can feed into how new information is treated: an unverified claim from a historically trustworthy source might be given more initial credence (or flagged as high priority to fact-check), whereas one from a dubious source might be tagged with a warning or held until corroboration appears. For example, the content filter could be configured to insert a banner like "**Warning: This news is from a source with a history of unreliability**." This kind of feature would be invaluable for readers and analysts trying to quickly judge new information. Analyzing **source bias** goes beyond just reliability scores. It also means understanding perspective: a source might be consistently skewing facts in a certain direction (e.g., minimizing certain risks or always pushing a particular narrative). An AI system can be trained to detect tonal or framing patterns that indicate bias. Over time, if one compares how different outlets report the same event, one might annotate "Outlet X tends to emphasize cybersecurity threats, while Outlet Y downplays them." This context can be included in the knowledge graph (perhaps as a "bias profile" node linked to the source). Then, when that source publishes something, the transform step can note, for example, *"Source X described this as 'massive' breach, but their bias profile suggests they often use hyperbole."* The system might then either normalize the language or flag to the user that *"Source X tends to exaggerate, consider verifying details from another source."* In MyFeeds, while the current focus is on personalization and relevance, the underlying tech of knowledge graphs and LLM analysis could well be extended to capture sentiment and bias in content -- especially since it's already extracting semantic meaning[12]. Indeed, the open nature of the platform means adding a "bias detection" transform is feasible. **Content provenance** in this context ensures that when we talk about a claim, we can always list its lineage. For example: *Claim:* "ZeroTrustCorp suffered a data breach leaking customer records." Provenance might tell us: first appeared in Source A (a tweet by a security researcher at time T1), then was reported by TechNews (article at time T2 citing Source A), then confirmed in an official press release (at time T3). The knowledge graph would have nodes for each source with edges like *Claim -> reported_in -> TechNews article (T2)*, *Claim -> reported_in -> press release (T3)*, and *TechNews article -> cites -> researcher tweet (T1)*. Traversing the graph, an analyst can see the chain of reporting -- essentially a **transparency timeline**. This helps in assessing credibility: by the time of the press release confirmation, the claim is pretty solid. Early on at T1, it was just a researcher's allegation (maybe credible if the researcher is known, but still single-source). If someone sees the claim on TechNews, the system can show: "source of this claim is a tweet by X." This alerts the user that the article might be based on a single external source and encourages caution until further confirmation. Maintaining such provenance links addresses one of the challenges in today's information landscape: repeated information can create an illusion of multiple sources when in fact everyone is quoting the same original source. The knowledge graph can cut through that by literally linking all those mentions to the common origin, preventing false amplification of confidence. Finally, **evidentiary support over time** is recorded as trails in the system. Each claim or hypothesis accumulates an **evidence trail** -- a list of supporting and opposing pieces of evidence along with timestamps. This could be visualized to users as a timeline ("Here is the history of this claim"). Such a trail might look like: - *Day 0:* Statement first made (by Source X). - *Day 1:* No verification yet (status: unverified). - *Day 2:* Another outlet Y reports the same (status: two reports, limited verification). - *Day 5:* Official source confirms (status: confirmed fact). - *Day 10:* An analysis questions a detail (status: mostly confirmed, minor contention). - *Day 20:* Further analysis resolves the contention (status: confirmed, consensus reached). This running log is essentially extracted from the evolving knowledge graph. It provides a narrative of how the truth unfolded, which is invaluable for deep analysis, audits, or simply understanding context. For example, journalists verifying a story can use this to ensure they've seen all angles. Cybersecurity professionals investigating an incident can track how early assumptions changed as forensic data came in. AI researchers could feed these trails into models to study how information solidifies or decays, perhaps improving the models' own calibration of uncertainty. MyFeeds.ai and the other systems already embody parts of this vision. MyFeeds doesn't just throw articles at users; it curates them based on a structured understanding and is very mindful of explaining *why* something is surfaced[24]. The Report Assistant ensures every generated line in a report ties back to input evidence[9], which is essentially leaving an evidence trail in the final product for the reader or auditor. The Semantic Content Filter even can insert annotations into content, like highlighting suspicious or unverified parts of a web page[14][31] -- this is a real-time manifestation of evidentiary labeling, warning users at the moment of consumption. All these contribute to an ecosystem where trust is continuously assessed and communicated. In summary, the combination of **knowledge graphs** and the **LETS pipeline** yields an infrastructure where information is systematically ingested, analyzed, and stored with rich relationships and provenance. This sets the stage for building applications that leverage time-based credibility. Next, we will discuss how **persona-based modeling** integrates with this to tailor the system's outputs to different users and use cases, and then we'll dive into the specific examples of MyFeeds.ai and The Cyber Boardroom to illustrate these principles in action. ## Persona-Based Modeling and Contextual Trust An innovative aspect of Dinis Cruz's approach is the use of **persona-based modeling** -- designing AI systems that understand and simulate the perspectives of different users or stakeholders[32][33]. In the context of credibility and information analysis, persona modeling plays a dual role. First, it helps tailor what information is presented and how it is presented, in order to maximize relevance and comprehension (which in turn affects how information is trusted by the end-user). Second, it allows the system to factor in biases and preferences associated with those personas, which is crucial when evaluating credibility. Essentially, the truth may be objective, but **the reception of information is subjective** -- who the user is can influence what they consider credible or what needs extra explanation. By modeling personas, we can adapt the system's behavior accordingly without compromising the underlying factual integrity. **User Persona Profiles:** MyFeeds.ai provides a clear example of persona-driven content delivery. The system builds a **profile graph for each user or persona** representing topics of interest, role, industry, and even content preferences[34][35]. A CISO at a bank might have a persona graph weighted towards regulatory compliance, financial sector threats, and high-level summaries, whereas a software engineer might have a persona focused on technical vulnerabilities, open-source tool news, and detailed analysis. These persona graphs are themselves part of the knowledge graph ecosystem -- they are essentially filters or lenses through which the global information graph is viewed. MyFeeds uses the intersection of the content's semantic graph and the user's persona graph to decide which articles to recommend[23]. This means the system doesn't just fling "top news" at everyone; it picks what matters to *you*, and it can even justify that choice ("we recommended this because it matches your interest in X")[24]. This increases trust in two ways: the user sees that the system understands their needs (making it feel like a credible assistant rather than a random feed), and the user can verify that the recommendations aren't arbitrary or biased -- they're tied to the user's own stated preferences. In a world of AI, giving users that transparency and control loop fosters trust in the system itself. **Persona Simulation for Communication:** The Cyber Boardroom takes persona modeling a step further by using it to simulate how different stakeholders perceive information. One of its distinguishing features is the ability to model perspectives of roles like CFO, CEO, CTO, etc., and even simulate a boardroom Q&A where the AI adopts those personas[32][36]. For instance, a security leader can input a report and ask the system to respond as a skeptical CFO might -- *"If presented this way, the CFO might worry about X or misunderstand Y"*[36]. This is immensely valuable for **trust-building in communication**. It forces the technical communicator to address likely concerns and clarify points before the real meeting, ensuring that when the information is delivered, it resonates and is convincing. Here, time is involved in a different sense: through iterative practice and refinement, trust is built up between the technical and non-technical sides. Over multiple board meetings, if the board consistently gets clear, persona-tailored answers, their trust in the security team's information increases. The Cyber Boardroom's persona simulation can be seen as a training ground for that trust -- the AI helps the human deliver information in the most credible way possible for each audience[37]. It essentially encodes empathy and perspective-taking into the information delivery process, which is a profound application of AI beyond just data crunching. By modeling "what if I am a CFO hearing this?", the system uncovers potential credibility gaps (e.g., jargon that a CFO might not trust or understand, or missing business context that a CEO would need). Filling those gaps ahead of time means the eventual communication is more **trustworthy** to its recipients. **Incorporating Source and Persona Bias:** Persona modeling also offers a structured way to deal with bias. We discussed source bias above -- now consider that users (personas) have their own biases too. A journalist might inherently distrust information from an anonymous source unless verified, whereas a corporate PR officer might be more skeptical of negative news until fully confirmed. By encoding such tendencies into persona profiles, the system can adapt its behavior. For example, a "Journalist" persona might cause the system to highlight "unverified" labels more prominently and provide easy access to source documents, aligning with a journalist's training to double-check facts. A "Casual Reader" persona might prefer the system to automatically filter out anything not yet corroborated (to avoid spreading rumors). The persona preferences could include thresholds for what confidence level is needed to show a claim as a fact. Another scenario: consider political or cultural bias -- if a user leans a certain way, they might initially distrust sources from the other side. While the ultimate goal is to present objective truth, a persona-aware system can *acknowledge* these biases in how information is presented. It might say, for instance, *"This claim comes from a source you don't usually follow, but note that three independent sources from across the spectrum have confirmed it"* -- essentially nudging the user to trust information that is well-evidenced even if it comes from an unlikely place. The key is not to reinforce bias, but to transparently work with it to gain the user's trust in the facts. Over time, an evidence-driven system might even help broaden a user's trust network by showing how reliability can come from many quarters if tracked objectively. **Persona in AI Reasoning:** Persona modeling isn't just for the user interface; it can be part of the AI's reasoning under the hood. When evaluating a claim, the system might apply different heuristics based on context: for a cybersecurity professional user, the system might emphasize technical evidence (logs, indicators of compromise), whereas for a board-level summary, it might emphasize authoritative statements (regulatory findings, law enforcement confirmations). The underlying data is the same, but the weight given can shift to match what that persona finds credible. Dinis's framework for defining personas includes factors like role, expertise, cultural context, and goals[38]. For example, a persona with low technical expertise but high need for certainty (like a board member) might trigger the system to only present facts that are confirmed and to avoid technical jargon, focusing instead on analogies and business impact[39][40]. On the other hand, a highly technical persona might be shown provisional findings with the caveat that they are unconfirmed, because that persona can handle uncertainty and may want the heads-up. By **integrating persona profiles into the LLM's prompts or the graph query process**, the system can effectively shape its outputs to maintain credibility in the eyes of the beholder. This is reminiscent of how a human analyst would brief different audiences: it's not about changing the facts, but about how you frame them and what you choose to highlight or contextualize. **Feedback Loop and Persona Evolution:** Personas themselves are not static -- they can evolve as the system learns from user feedback. MyFeeds.ai, for instance, updates a user's persona graph based on which articles the user reads or finds useful[41][42]. If the user consistently ignores news about a certain topic, the system might down-weight that topic in their profile; if they always click items about a new subject, it might add that subject to the interest graph. In terms of trust calibration, if the user frequently gives feedback like "this source is not credible" or "I don't believe this", the system could incorporate that into the persona bias (though ideally it would also aim to show why something is credible if evidence supports it -- feedback could also highlight places where the system needs to provide more evidence to convince the user). This adaptability means the credibility system becomes personalized over time -- not in the sense of filtering truth (we must avoid just creating echo chambers), but in tailoring *how* it engages the user to build trust in verified information. To ground these ideas, let's incorporate an example scenario: Suppose the system is analyzing a news article that makes a bold claim about a cyber-attack, sourced from an anonymous intelligence report. A **skeptical persona** (say an experienced analyst) might be given that information with a caution: *"Preliminary claim from an unverified source, treat with caution until more info emerges."* The system might even suggest questions that this persona would likely ask (because it knows the analyst persona values certain evidence): *"No malware hashes or IOCs provided -- you might want to see those for proof."* In contrast, a **general executive persona** receiving information on the same incident might get a different treatment: *"Early reports suggest a cyber-attack; confirmation pending -- we will update you when official statements arrive."* The exec doesn't need technical details (and might distrust them if confusing), but does need the assurance that the information is being validated. Both personas ultimately get the same outcome (if later an official report confirms the attack, both will be informed of that fact), but the journey to that point -- and how the uncertainty is communicated -- differs to maintain credibility for each audience. In summary, persona-based modeling is about **contextualizing credibility**. It acknowledges that trust is partly in the eye of the beholder and that a one-size-fits-all approach to presenting information can fall short. By designing AI systems that simulate and adapt to personas, Dinis Cruz's projects ensure that the sophisticated analyses (knowledge graphs, evidence tracking, etc.) actually translate into insights that different users trust and find useful. The Cyber Boardroom and MyFeeds.ai both exemplify this: one by translating tech to business language to earn trust at the board level[43][40], the other by curating feeds that feel almost eerily relevant to each user[44]. With persona modeling covered, we now move to highlight these real-world systems explicitly, drawing out how they implement the principles discussed and how they are paving the way for a new standard in information credibility. ## Real-World Examples: MyFeeds.ai and The Cyber Boardroom To illustrate the principles discussed, we turn to two of Dinis Cruz's ongoing projects: **MyFeeds.ai** and **The Cyber Boardroom**. Each addresses a different problem space (personalized information feeds and executive cybersecurity communication, respectively), yet both are built on the core ideas of semantic analysis, provenance, and adaptive presentation of information. They serve as concrete prototypes of how time-calibrated credibility and trust can be encoded in practical systems. ### MyFeeds.ai -- Personalized, Provenance-Rich Intelligence Feeds **MyFeeds.ai** is designed to combat information overload for professionals in cybersecurity, tech, and related domains[45]. The problem: there is an endless flood of news and alerts, but busy experts have limited time to sift signal from noise. Traditional news aggregators lack fine personalization and often miss context, while raw feeds and keyword alerts return heaps of irrelevant results[46]. MyFeeds tackles this by delivering **highly personalized news feeds** that are tailored to each user's role and interests, accompanied by concise summaries[47][48]. Under the hood, it exemplifies many of the concepts we've covered: - **Semantic Knowledge Graph Curation:** Instead of naive keyword matching, MyFeeds uses a semantic pipeline. It ingests many sources (RSS feeds, APIs) frequently, ensuring timely updates[25]. Each article is parsed and then transformed by an LLM into a semantic graph of its content[11][12]. So, all content becomes structured data linked by meaning. For example, an article on a new vulnerability might have graph nodes for the vulnerability identifier, affected software, potential impact, etc. This structured approach allows MyFeeds to match content to users at the concept level: it knows what the article is *about*, not just what words appear. This dramatically improves relevance -- a user interested in "supply chain attacks" will be shown an article about a compromised NPM package, even if the article doesn't use the phrase "supply chain," because the graph understands the relationship (NPM package hack is a type of software supply chain issue)[49]. - **LETS Pipeline and Memory Graph Storage:** MyFeeds leverages the LETS pipeline to handle its data flow[15]. **Load:** it fetches new content periodically. **Extract:** it parses articles and stores raw text and metadata in a uniform way using MemoryFS (an in-memory filesystem abstraction)[50]. **Transform:** it generates the semantic graph (using what Cruz built as GraphFS for storing graph data uniformly)[21]. **Save:** it stores both the raw content and graph representation. This disciplined pipeline means MyFeeds can process information systematically and at scale -- dozens of feeds, thousands of articles, continuously. The use of serverless functions and lean infrastructure means it only consumes significant resources when processing new content (keeping costs low and scalability high)[51][52]. This is critical for a system intended to monitor information streams around the clock. - **Content Provenance and Explainability:** MyFeeds emphasizes provenance in its recommendations[15]. Each piece of content retains a link back to its source, and when the system generates a summarized newsletter or feed for a user, it can explain *why each item is included*. For instance, a daily brief email might list 5 headlines with 2-sentence summaries; next to each, MyFeeds can indicate the source (e.g., "via Krebs on Security, reported 2 hours ago") and a rationale ("Chosen for you because it relates to [Cloud Security] in your profile")[24]. Users are not left guessing why something showed up -- the system is transparent about its reasoning. If a user wants to drill down, they could see the semantic connections (e.g., the article discusses AWS breaches and the user's profile has interest in cloud breaches). This fosters trust: professionals can rely on MyFeeds because it's not a black box, it's an assistant that shows its work. Moreover, provenance tracking means if there are conflicting reports on a story, MyFeeds could potentially show both and attribute them correctly, helping the user be aware of disagreements or evolving stories (a future enhancement might be to explicitly highlight when a story in yesterday's brief has an update or correction today). - **Adaptive Persona Feeds:** MyFeeds builds a graph for the user's persona (interests, role, etc.)[34] and continuously refines it based on feedback[41]. Suppose a user is an *investor* focusing on cybersecurity startups; their feed might prioritize funding news, major breaches with business impact, and tech trend analysis. If they start frequently reading AI-related security articles, MyFeeds will learn and expand their profile to include that. The system can also produce multiple persona feeds from the same content pool. As noted in Cruz's briefing, one CISO user asked for multiple feeds: one for themselves, one simplified for their non-technical executives, and one for their technical team[53]. MyFeeds delivered this by creating distinct persona profiles for each audience type and repackaging the same source content appropriately[53]. This demonstrates how powerful the combination of semantic graphs and persona modeling is -- the content can be filtered and reframed without manual effort, and each audience gets the information in the form they trust and understand. The CISO's bosses got a high-level brief (trust through clarity and relevance, no jargon), while the technical staff got a detailed feed (trust through completeness and technical accuracy). - **Evolving Content and Alerts:** Because MyFeeds runs continuously, it can catch the evolution of stories. If a Monday brief included "Company X breach reported, cause unknown," and by Tuesday there's an update "Cause identified as phishing," the Tuesday brief can include that development, possibly even referencing that it's an update to yesterday's news. This temporal linking is exactly what time-calibrated credibility is about. MyFeeds could in principle tag the Monday item as a hypothesis ("cause unknown") and then automatically follow up when the cause is confirmed, adjusting the status of that story from speculative to factual. While not explicitly stated, the underlying tech is ready for that kind of feature. In fact, Cruz originally built MyFeeds to support the Cyber Boardroom -- he needed a steady stream of content to discuss in board meetings without using sensitive data[54]. This means the feed had to be topical and up-to-date to simulate real-world scenarios. The MVP of MyFeeds was publicly demonstrated (with example newsletters like a "CEO Cybersecurity Brief" and an "Investor Tech Digest") and garnered positive feedback for its uncanny relevance[44]. This validation shows that focusing on semantic relevance and user context produces a qualitatively better information experience than generic feeds. - **Open-Source and Integration:** Importantly, MyFeeds (like the other projects) is built with an open-source core and a serverless-friendly architecture[55][56]. This means organizations could run their own instance or extend it. For instance, an enterprise might plug in internal data sources (threat intel feeds, internal incident reports) into MyFeeds to create a hybrid feed combining external and internal intel, all analyzed under the same graph framework. The serverless approach (deploying as cloud functions, etc.) means it's cost-efficient and scales with usage[52]. An open-source foundation also fosters trust: users (especially cybersecurity pros) can inspect how MyFeeds processes data and be confident there are no hidden agendas or leaks[57]. In the security community, open tools are often preferred for this reason -- they can be vetted[58]. MyFeeds's design reflects this ethos by focusing on interoperability (MemoryFS/GraphFS making it easy to connect systems) and transparency. In essence, MyFeeds.ai demonstrates how to deliver the *right* information at the *right* time in the *right* way, which is the crux of building trust. It doesn't overwhelm, it doesn't hide its logic, and it evolves as the world and the user's needs evolve. It shows that by using semantic understanding and tracking provenance, an automated system can actually earn the user's confidence in a domain where trust is paramount. ### The Cyber Boardroom -- Trust at the Nexus of Tech and Business **The Cyber Boardroom** addresses a very different challenge: bridging the communication gap between cybersecurity experts and business executives (such as corporate boards)[59]. Here the trust we're concerned with is the trust business leaders have in the information and advice coming from their security teams (and vice versa). Often, miscommunication or lack of context leads to misaligned expectations and skeptical board members who are not sure whether to take cybersecurity recommendations seriously. The Cyber Boardroom uses GenAI (generative AI) as an **intelligent translator and facilitator** in this relationship[60]. It exemplifies time-calibrated credibility in a more human-centric sense: through iterative dialog and persona simulation, it ensures that over time, the messages delivered to boards are consistently understandable, relevant, and backed by appropriate evidence -- all of which build credibility and trust in the eyes of leadership. Key features and principles of The Cyber Boardroom include: - **Bidirectional Translation:** The system works both ways -- it helps technical teams explain things to the board in plain business terms, and helps boards ask the right questions or get clarifications to relay back to technical teams[39][61]. For example, a CISO can input a detailed risk assessment or a technical report. The AI then produces a polished executive summary emphasizing business impact and critical points, stripped of jargon[62]. It might use analogies and focus on strategic priorities so that board members immediately grasp the significance without wading through technicalities[39]. Conversely, if a board member types a question like "Why do we need to invest $X more in cybersecurity next quarter?", the AI can interpret that against the technical data and either formulate an answer or translate it into actionable queries for the security team[61]. This translation is not done blindly -- it leverages a knowledge base of cybersecurity concepts mapped to business outcomes that the Boardroom maintains[40]. Essentially, the system "understands" common cybersecurity topics and how they relate to things a board cares about (financial impact, legal risk, operational continuity, etc.)[40]. Over time, as it's used in an organization, it can even incorporate specifics of that organization's context (like past incidents, industry regulations, the company's risk appetite) to make the translations more tailored. - **Persona Simulation and Training:** As mentioned earlier, a standout feature is the persona simulator for different stakeholders[63]. The user (say a CISO preparing for a board meeting) can run a mock presentation through the AI and get simulated responses from various personas: *"As a CFO, I'm concerned about the cost implications here,"* or *"As an outside director with legal background, I need clarity on regulatory exposure."*[64]. This allows the security leader to iteratively refine their message. It's essentially a **sandbox to preemptively address skepticism**. By the time the real meeting happens, many of the rough edges have been smoothed. This builds trust in two ways: the board gets a clearer, well-thought-out presentation (so they trust the presenter more), and the security leader feels more confident and in tune with the board's perspective (so they trust the board to understand, creating a virtuous cycle of open communication). Over repeated use, this can significantly improve the relationship; as the whitepaper notes, *"Over time, this can significantly improve mutual understanding and trust between tech and business leadership."*[65]. The time element here is the iterative practice and learning: the persona feedback loop effectively compresses what might take years of trial-and-error presentations into a much shorter learning curve. - **Knowledge Graph and Memory:** Under the hood, The Cyber Boardroom likely uses similar tech to MyFeeds for knowledge management. It maintains a knowledge base of cybersecurity concepts and maps them to business outcomes[40]. This sounds like a knowledge graph where nodes could be things like "ransomware attack" linked to "business downtime" and "revenue loss" as outcomes, etc. It factors in attributes like the stakeholder's role, concerns, and even personal style if known[40]. All this context forms a persona profile that the LLM uses to shape its output[40]. In essence, the system encodes things like: *"If audience is CFO, prioritize cost/benefit language; if General Counsel, highlight legal risk; if CTO, you can include some technical detail,"* and so forth[66][67]. This knowledge base would need to be built and refined over time -- presumably the system can learn from each interaction, noting what follow-up questions were asked by the real board and incorporating that into future simulations. Also, because the platform is used as a central hub, it can accumulate an institutional memory (with appropriate security) -- e.g., it could recall that *"Last quarter the board was particularly concerned about supply chain attacks"* and ensure that is addressed upfront in the next briefing if relevant. This temporal memory aspect again fosters trust: the board sees continuity and attentiveness to their past concerns, reinforcing that the tech team is responsive and thorough. - **Integration with Live Content:** The Cyber Boardroom doesn't operate in isolation; it can pull in live data like news or threat intel to enrich board discussions. In fact, one of the MVP features was the ability to ingest RSS news about cyber incidents and produce a tailored briefing for a given persona[68]. This was developed in tandem with MyFeeds.ai[69], showing synergy between the projects. For example, if a major cybersecurity incident is in the news on the day of a board meeting, the CISO can query the Athena bot (the Boardroom's AI assistant persona, as per demos) about that incident and get a quick rundown plus any relevant context on how it might affect their company, all in board-friendly terms. The system thus keeps the discussion timely and grounded in reality. From a credibility standpoint, this means board members aren't left in the dark about something they heard on the news -- the AI proactively brings it into the conversation with analysis, which increases the board's trust that the security team is on top of current events. It also showcases transparency: *"Yes, we're aware of breach X that's in headlines; here's what it means for us."* - **Outcome: Better Decisions and Trust Building:** The ultimate measure of The Cyber Boardroom's success is whether it *revolutionizes board-level decision-making* in cybersecurity, as intended. Early feedback from CISOs who tried the persona simulation has been that it's "especially insightful"[70] -- it surfaces concerns they hadn't thought of and helps them see their communication from the outside perspective[64][65]. This introspection leads to more polished communication. A polished, clear, evidence-backed presentation to the board results in fewer misunderstandings and more informed questions. Over successive quarters, the board can see a track record: "the security team consistently communicates well, provides data to back up claims, and addresses our concerns." That consistency is how **credibility is built over time**. It's analogous to how a news outlet builds trust by being accurate again and again. Here, the security program builds trust with leadership by communicating effectively again and again, aided by AI. Meanwhile, the security team gains trust in the board too -- seeing that when they articulate risks in business terms, the board responds constructively (e.g., approving budgets for critical defenses). This mutual trust can ultimately lead to better cybersecurity posture, as initiatives are understood and supported at the highest level. - **Open-Source and Security:** Just like MyFeeds, The Cyber Boardroom is built on an open-source core and can be deployed flexibly (cloud or on-prem) via a serverless model[71][52]. This is crucial because board discussions often involve sensitive data. Companies might be wary of putting that into a black-box SaaS. By having an open architecture, The Cyber Boardroom can be inspected for security and even hosted in a controlled environment. The open model also invites contributions from the community -- for example, new personas could be added (imagine a template persona for "audit committee chair" or "non-executive director with finance background"), or the knowledge base could be expanded with more scenarios. As noted, open-source establishes credibility with enterprise customers who can verify the integrity of the code[57]. For a tool meant to be in the boardroom, credibility of the tool itself is important -- it must be beyond reproach in terms of confidentiality and accuracy. Adopting open, transparent development helps in that regard, aligning with Cruz's overarching strategy of leveraging open-source as a trust and innovation catalyst[57][72]. In sum, The Cyber Boardroom showcases how time and iteration, combined with AI, can **calibrate trust in a human relationship context**. It's not marking a statement true or false over time, but rather refining a message over time so that it lands truthfully and effectively with an audience. It demonstrates that the same principles of evidence, context, and adaptation apply whether we are verifying a fact or persuading a person: provide the right context, check understanding, adjust, and do this repeatedly to build confidence. It complements MyFeeds by focusing on how insights are communicated and acted upon at the strategic level, closing the loop: data turns into intelligence (MyFeeds), which turns into decisions and actions (via Boardroom) -- all under a philosophy of **transparency, provenance, and continuous learning**. ## Open Infrastructure for Transparent, Scalable Analysis Underpinning the philosophy and examples above is a commitment to **open-source infrastructure and scalable architecture**. Dinis Cruz's vision is not only about *what* should be done (track credibility over time) but also *how* it should be enabled. The **credibility** of an information system is tied not just to the data it presents but to the trust users have in the system itself. By using open-source, transparent methods and modern cloud-native design (like serverless computing), these projects ensure that the system's operations are **auditable, adaptable, and capable of growing** with the data. **Open-Source as a Foundation:** All four of Cruz's synergistic startups (including MyFeeds.ai and The Cyber Boardroom) share an open-source core[73]. This is a deliberate strategy. Open-source software allows anyone to inspect the code for biases, errors, or security issues. For AI systems dealing with knowledge and truth, this is particularly important. Users -- especially in cybersecurity and research -- are rightly skeptical of black boxes. An open approach means the logic behind claim classification, evidence scoring, or feed curation can be scrutinized and validated. It establishes a baseline of **trust through transparency**[57]. Enterprises can vet the tools before adopting them, and independent contributors can suggest improvements or catch problems. Moreover, open-source encourages a community of practice: for example, researchers might contribute new modules for fact-checking or journalists might extend the taxonomy for new types of media. This collective innovation accelerates the development of robust credibility systems[74]. Cruz's experience as an open-source advocate (e.g., his creation of the OWASP O2 Platform for security testing) feeds into this approach -- he has seen how community-driven projects can shape industry standards[[75]](https://www.threatmodcon.com/speaker/dinis-cruz#:~:text=Technology%20Officer%20,source%20innovation)[[76]](https://www.threatmodcon.com/speaker/dinis-cruz#:~:text=application%20security%2C%20former%20OWASP%20board,source%20innovation). In these new ventures, being open means not reinventing the wheel for each project. Indeed, the startups reuse key building blocks -- a library for semantic graph handling created in one is reused in others[77][78] -- which speeds up development and keeps the design consistent. **Serverless and Lean Architecture:** Scalability is crucial because evaluating credibility over time can become data-intensive. There may be thousands of sources, millions of statements, and constant updates. The use of a **serverless deployment pipeline** means the system can scale out when needed (e.g., processing a burst of news during a major incident) and scale to zero when idle[79][52]. This ensures cost-effectiveness -- a key consideration for startups and also for any organization deploying such a system. They only pay for the compute they actually use, making it feasible to monitor vast amounts of information without a massive always-on infrastructure. A unified CI/CD and packaging approach across these projects allows them to run in various environments easily -- whether as cloud functions, containers, or on-prem appliances[79]. This flexibility means the tools can be brought to the data (important if certain data can't leave a corporate environment for privacy reasons). **Minimal fixed costs and high elasticity** also mean that as the user base grows or as more data streams are added, the system can accommodate that without a ground-up redesign[52]. It's essentially future-proofing the platform to handle the "firehose" of information we expect in the coming years. **Shared Components and Interoperability:** The startups were conceived to complement each other, sharing technology and passing data between them where useful[18][80]. This interoperability is a strength of an open, modular design. For example, the knowledge graphs are stored in standardized formats across the systems[81][18]. This means an insight discovered in MyFeeds (like a trending new threat topic) could be fed into The Cyber Boardroom's knowledge base to alert CISOs and boards about it[18]. Or the Report Assistant's structured output of a risk assessment could be used to generate an executive summary for the board, or to feed into MyFeeds as an internal news item. By **weaving these tools together**, Cruz envisions an ecosystem where data flows securely to where it's needed, and every piece of analysis reinforces others[80][82]. This reduces duplication of effort (each tool doesn't need to rediscover the same facts) and enhances consistency (a fact confirmed in one context is automatically updated everywhere). From an infrastructure perspective, this is facilitated by using common storage abstractions (MemoryFS and GraphFS to represent data uniformly as files or graphs)[17], and by keeping everything open so integration is straightforward (no proprietary formats or locked APIs). **Transparency in AI Workings:** Another reason open-source is vital is the need to **audit AI decisions**. When an AI model suggests that "Claim X is likely true" or filters out a piece of content as misinformation, stakeholders will want to know why. By having an open system, one can examine the rules or model outputs that led to that decision. For example, if a claim was flagged as false, was it because the AI found a contradicting source? Did it perhaps misinterpret something? Transparency allows developers and even end-users to ask these questions. In a closed system, one might suspect biases or errors but have no way to confirm; in an open system, one can look under the hood. The LETS pipeline structure supports this by design, since it logs each step's output and keeps the transformations modular[20]. It would not be hard, for instance, to output an intermediate file that shows "extracted claims and their initial confidence scores before and after reconciliation" for a given news article. Such traceability builds confidence that the system isn't arbitrarily labeling things as true or false -- it's following documented procedures that can be verified or contested as needed. **Security and Trust:** In cybersecurity applications, trust in the tool is paramount. By open-sourcing and building on well-tested components, Cruz's startups aim to be **secure by design** and earn the trust of security professionals (a notoriously tough crowd). The idea is that by the time these tools are being used in a critical environment (like a board meeting or processing a company's internal knowledge), they've been vetted by many eyes and perhaps formally verified or certified. Community vetting can catch vulnerabilities or logic flaws early. Additionally, from a user trust standpoint, being open-source aligns with the values of many in the target audience: journalists favor transparency, researchers value open data and methods, and security experts trust open scrutiny over closed promises[57]. We see this in the widespread adoption of open-source tools in security (like Wireshark, Metasploit, etc.) largely because people can ensure there are no malicious backdoors. Similarly, an open-source credibility engine can be trusted not to have hidden biases introduced for commercial or political reasons -- any such attempt could be discovered in the code or training data. **Community and Ecosystem:** Finally, open infrastructure fosters an ecosystem. Others can build atop these tools -- for example, someone might create a specialized plug-in for MyFeeds to handle a new domain (say medical news or financial markets) using the same pipeline. Or they might extend The Cyber Boardroom to other types of boardroom topics (risk in general, not just cyber). This means Dinis's core idea -- time-calibrated credibility -- can spread and adapt beyond his immediate implementations. It encourages **industry-wide adoption** of standards for tracking provenance and evidence. If multiple tools output knowledge graphs with similar schema for claims and evidence, these could interoperate or be aggregated. Imagine an open standard for representing "credibility of a claim" with timestamps, source references, and status -- much like RSS became a standard for syndicating feeds, a standard for credibility data could enable a whole new class of applications. By basing everything on open principles from the start, these projects are well-positioned to contribute to and benefit from such developments. In summary, the **infrastructure choices** reflect the same values as the system's logic: transparency, trust, and adaptability. Just as we want each statement's trustworthiness to be traceable and updatable, we want the system itself to be transparent and improvable. By leveraging open-source and serverless architecture, Dinis Cruz's projects not only address the technical challenges of building credibility systems but also the social and ethical ones -- ensuring the systems themselves merit the trust we place in them. ## Conclusion In a world awash with information and misinformation, **time** may be the most underutilized tool we have for discerning the truth. Dinis Cruz's vision, as articulated in this paper, is to explicitly harness the temporal dimension as a calibrator of credibility in our information systems. By treating each statement as an entity with its own life story -- from inception through evolution under the scrutiny of evidence -- we can transform how trust is built in the digital age. This approach moves us beyond static true/false judgments into a dynamic model where assertions are born as hypotheses or opinions, mature (or wither) as facts through corroboration or refutation, and are continually recontextualized as new data emerges. We introduced a taxonomy of information types (facts, opinions, hypotheses, data) as the foundation for this framework, recognizing that different kinds of statements demand different handling and validation. We saw how large language models and AI can serve as powerful allies in extracting these elements and populating **semantic knowledge graphs** that serve as living maps of knowledge. With systems like MyFeeds.ai and The Cyber Boardroom, we explored how these concepts are not merely theoretical -- they are being implemented in real products that address pressing needs: from personalized intelligence feeds that **keep professionals informed with context and provenance**[15], to boardroom assistants that **bridge the gap between technical truth and business trust**[32][65]. These examples underscore the practicality and versatility of the vision. They show, for instance, that an AI-curated news brief can gain a user's trust by explaining its recommendations and highlighting source credibility[24], or that an executive can trust a cybersecurity briefing because it's been honed through persona-driven simulations to preempt their concerns[36]. A recurring theme is that **transparency and traceability** are inseparable from credibility. A system that tracks the lineage of every claim, that can point and say "this is what we know and here's how we know it," inherently engenders more trust than one that cannot. By using time and evidence as the yardstick, we also inject a healthy dose of humility and resilience into our AI: humility in acknowledging uncertainty and separating what is known from what is conjectured, and resilience in being able to update and correct course as reality unfolds. In essence, the systems we build must themselves learn and adapt over time, much like the humans who operate them. We also emphasized the **importance of open, scalable infrastructure** in realizing this vision. The choice to build on open-source principles and serverless architectures is not just an implementation detail; it is a statement of values. It says that **truth-seeking should be a collaborative, transparent endeavor**, and that the tools for it should be accessible and trustworthy. By aligning engineering choices with the end goal of trust, Cruz's approach ensures the platform on which we measure credibility is itself credible. This alignment of content and platform -- having open data pipelines (LETS), common schemas, and community vetting -- means that the credibility system can be trusted not to distort or conceal. It also means it can scale to meet the challenge: as the volume of information explodes, the combination of cloud scalability and crowd-sourced improvement positions these systems to keep up with the deluge, extracting signal from noise. For AI researchers, this whitepaper offers a blueprint of how AI can move beyond isolated predictions and into the realm of **knowledge maintenance over time** -- a sort of longitudinal AI that remembers and revises. For cybersecurity professionals, it outlines tools that can enhance situational awareness and strategic communication, ensuring that security insights are both accurate and effective in driving decisions. For journalists and truth-seekers, it presents hope that technology can be harnessed to bolster fact-checking, provide context, and uphold the integrity of information in the public sphere. Looking ahead, one can imagine the principles outlined here being applied widely: social media platforms that tag posts with the current credibility status of their claims (and update them as facts emerge), scientific literature databases that track the replication and validation status of published results over time, or public knowledge bases (like Wikipedia or Wikidata) enhanced with temporal evidence graphs that show how our understanding of a topic has evolved. The concept of a "credibility timeline" could become a standard feature in information consumption, much like timestamps or view counts are today. In conclusion, by making **time the calibrator of credibility**, we align our information systems more closely with the reality of how knowledge works in the real world. Truth is a process -- one of inquiry, verification, and sometimes revision. Dinis Cruz's vision is to embed that process into the fabric of our digital knowledge tools. It is a vision of *information integrity through temporal context*, one that holds great promise for improving trust in the age of AI. As these ideas are implemented and refined in projects like MyFeeds.ai and The Cyber Boardroom, they lay the groundwork for a new paradigm in which *credibility is not just asserted, but demonstrated and earned over time*. By combining philosophical rigor with technical innovation -- from taxonomy and knowledge graphs to persona models and open infrastructure -- we can build systems that not only handle information, but genuinely *understand* and *honor* the journey each piece of information takes. In doing so, we equip ourselves and our societies with better defenses against falsehood and better tools to navigate an ever more complex information landscape. The ultimate calibrator, time, will tell how successful this approach will be, but the framework laid out here provides a clear and compelling path forward. **Sources:** The concepts and examples discussed in this paper are drawn from Dinis Cruz's work and voice memos, as well as related documentation and demonstrations of the mentioned platforms. Key references include the MyFeeds.ai architecture for semantic feed curation the Cyber Boardroom's persona-based communication approach, the Interactive Report Assistant's method of capturing facts and hypotheses with provenance, and Dinis Cruz's overall advocacy for building trust in AI through **provenance, transparency, and human-centered design**. These sources illustrate the marriage of philosophy and practice, showing the real-world momentum behind time-calibrated credibility systems. --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/10/03/genlegaladvise-project-plan.html (markdown twin: /2025/10/03/genlegaladvise-project-plan.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/10/03/genlegaladvise-project-plan.html* # GenLegalAdvise Project Plan *By Dinis Cruz and ChatGPT Deep Research and Claude Opus 4.1 · 2025-10-02* > Small businesses, freelancers, and independent consultants regularly encounter legal documents -- consulting agreements, NDAs, service terms, EULAs, data sharing policies -- that are lengthy, dense… --- > > **(NOTE: Needs final review)** > ## Problem Statement & User Pain Points Small businesses, freelancers, and independent consultants regularly encounter legal documents -- consulting agreements, NDAs, service terms, EULAs, data sharing policies -- that are lengthy, dense, and full of legal jargon. These individuals often **skim or skip** the fine print due to time and cost constraints, risking exposure to unfavorable terms. Key pain points include: - **Lack of Accessible Review:** Hiring a lawyer for every contract is expensive and slow. As a result, non-experts sign documents without understanding hidden risks (like unlimited liability or onerous IP clauses). - **Information Overload:** Legal text is verbose and complex. Users struggle to identify the *few critical clauses* (e.g. indemnities, liability caps, non-competes) buried in dozens of paragraphs. Important obligations or rights can be missed due to sheer volume. - **Inefficient Negotiation:** Even when red flags are spotted, users aren't sure how to **propose changes**. Crafting a polite yet firm negotiation email or redlined document is daunting without legal training. Many simply accept one-sided terms rather than negotiate improvements. - **Fear of Missing Something:** There's anxiety around "what did I overlook?" in a contract. Without a structured review, freelancers worry they might have missed a clause that could hurt them later (like automatic renewals or stringent confidentiality terms). In summary, the target users face a gap between the **importance of thorough legal review** and the **practical difficulty** of doing it. GenLegalAdvise aims to bridge this gap by providing an AI-driven, fast, and user-friendly way to understand and negotiate common legal documents. ## User Personas and Typical Workflow GenLegalAdvise is designed with several personas in mind, each with specific needs and workflows. Below are key user personas and how they would interact with the platform: - **Freelance Consultant (Solo Contractor):** Often signs NDAs or consulting agreements with clients. They would upload a contract and receive a plain-language summary of obligations, payment terms, and liabilities. The AI would highlight any *unusual or one-sided clauses* (e.g. an indemnity favoring the client only) and suggest redlines. The freelancer can then use the tool to draft an email back to the client requesting reasonable changes (like adding a mutual indemnity). - **Startup Founder / Small Business Owner:** Reviews software service agreements (SaaS terms, EULAs) from vendors. They paste lengthy terms of service into GenLegalAdvise. The platform quickly extracts key points like data ownership, service level commitments, and termination rights. It flags risky terms (e.g. excessive service provider liability disclaimers) and provides a summary they can share with their team. The founder might interactively tweak the AI's suggestions -- for instance, adjusting the tone of a negotiation email to a more formal style before sending it to the vendor. - **Independent Developer or Designer:** Signs boilerplate contracts with agencies or clients (e.g. work-for-hire agreements). They use GenLegalAdvise to ensure they are not inadvertently giving away IP rights to their pre-existing tools or violating any non-compete clause. The tool would identify any IP ownership clauses and advise if a carve-out is needed for prior work. The developer could then have the AI propose a contract clause that protects their existing IP, ready to be inserted as a redline suggestion. - **Tech-Savvy Lawyer or Legal Consultant:** While not the primary target, some legal professionals could use GenLegalAdvise as a **time-saver for first-pass reviews**. For a stack of similar NDAs, a lawyer could run them through the platform to get an initial issue list, then focus their expertise on the nuanced points. The workflow here might involve the lawyer feeding the AI custom prompts (or using a special "lawyer mode") to check compliance with certain laws or to ensure consistency with a client's standard contract preferences. This persona ensures the platform **complements** legal professionals by handling the grunt work, allowing them to add high-level judgment. **Typical Workflow:** Regardless of persona, the interaction with GenLegalAdvise would generally follow these steps: 1. **Document Ingestion:** The user uploads a legal document (or pastes the text). The system supports large documents (multiple pages) and automatically handles various formats (plain text, Word, PDF). 2. **AI Analysis (Parallel Models):** Upon ingestion, GenLegalAdvise uses multiple GenAI models in parallel to analyze the text. For example, GPT-4 and Anthropic Claude are both tasked with reviewing the contract's content simultaneously. This **cross-review** approach ensures thoroughness -- each model might catch things the other misses, and their outputs can be compared for consistency. (In practice, the system might prompt each model differently: one focused on summarizing terms, another on spotting red flags.) 3. **Key Extraction:** The platform consolidates the AI outputs to extract structured data about the contract. This includes identifying the parties, key dates/duration, payment terms, termination conditions, confidentiality obligations, liability clauses, indemnities, IP ownership, and any unusual obligations or rights. The output at this stage is a set of **"insights"**: e.g. *Clause 5 imposes unlimited liability on the Consultant*, *Clause 7 gives all IP to the Client, with no license back*. Each insight is linked to the source clause in the text for traceability. 4. **Risk & Red Flag Analysis:** GenLegalAdvise then evaluates which extracted items are potential risks or asymmetries. It highlights *red flags* (e.g. an indemnity that only goes one way, a very low liability cap for the other party, or broad non-compete language). These are presented in a concise list, with severity or importance indicated. For instance: "**Red Flag:** No liability cap for the consultant -- the consultant could face unlimited damages. **Location:** Clause 10." The system provides justification for why this is risky in layman's terms. 5. **Summary & Recommendations:** The user is presented with a **user-friendly summary** of the entire document. This summary describes the purpose of the contract and the main points in plain English (e.g. "This is a 12-month consulting agreement where you will develop software for Client X. You retain no IP rights to the work product. Either party can terminate with 30 days' notice. There is a confidentiality obligation lasting 2 years after termination..."). Alongside the summary, the platform lists **practical suggestions** for negotiation or redlines. Each suggestion is tied to a specific clause or issue -- for example, "Consider adding a cap on liability (e.g. limited to the contract value) in Clause 10" or "You may want to exclude indirect damages from the indemnity in Clause 8." These suggestions are generated by the AI based on best practices and the user's perspective (freelancer vs client). 6. **Interactive Refinement:** The user can then interact with the system to refine outputs. There might be a chat interface where the user asks follow-up questions ("Why is Clause 7 risky?") or requests specific drafts ("Generate an email to the client asking to clarify the IP clause"). The GenAI models, informed by the earlier analysis, will produce the requested content. Users can prompt the AI to adjust tone or detail level (for instance, "make this email sound more formal" or "simplify the summary for someone without legal knowledge"). This iterative loop helps users get to a final result they are comfortable with. 7. **Output & Export:** Finally, the platform allows users to export the results. This could include downloading a **redlined document** (where the AI's suggested clause changes are inserted as tracked changes in a Word document), copying the negotiation email text, and saving the plain-language summary and risk report as a PDF. The user leaves with a clearer understanding of their document and concrete next steps for negotiation. Throughout this workflow, the **turnaround time** is minutes, not days -- delivering fast advice in a structured format. The user remains in control, deciding which AI suggestions to adopt. Importantly, GenLegalAdvise maintains a record of the analysis (stored securely) so that the user can revisit the results later or re-run the analysis if the document is updated. ## Core Features and Solution Approach GenLegalAdvise's solution combines natural language understanding with software engineering best practices to deliver a powerful, structured review of legal documents. The core features and how they function are outlined below: - **Document Ingestion & Parsing:** The platform can ingest large legal texts (multi-page PDFs, Word docs, or raw text). Upon upload, the document is parsed into a structured format. This includes detecting section headings, clause numbering, and formatting, so that context (like "Section 10: Liability") is preserved. The content is then stored in an internal **MemoryFS** (in-memory file system) representation, treating the contract text as data that can be uniformly accessed and versioned[1]. This design (inspired by Dinis Cruz's MemoryFS) means the raw document and all its parts can be easily referenced or converted into other forms (like graph nodes or JSON) without losing place in the original. - **Parallel GenAI Analysis (GPT-4, Claude, etc.):** GenLegalAdvise employs an orchestration layer that queries multiple Large Language Models in parallel. Each model is given the **same document data** but possibly different prompts focusing on various perspectives. For example, GPT-4 might be prompted: *"Identify obligations and responsibilities of each party, and note any clause that seems heavily one-sided."* Meanwhile, Claude could be prompted: *"Summarize this contract in bullet points and list any clauses that a freelancer should be cautious about."* Running models in parallel speeds up analysis and provides a form of cross-validation -- the system can compare outputs to see if both models flag the same clauses as risky. Any discrepancies (say GPT-4 flags something Claude didn't) can be highlighted for the user or even fed into a resolution step (where the system asks a model to reconcile differences). This multi-LLM approach reduces the chance of a single model's oversight or hallucination dominating the result. - **LETS Pipeline for Processing:** The platform's backend follows a **LETS pipeline** -- Load, Extract, Transform, Save -- to structure the analysis process[2]. **Load:** The document is loaded into the system (and cached for reuse). **Extract:** The GenAI models extract key data points and clauses from the text (this includes pulling out structured elements like names of parties, dates, as well as semantic info like "this clause is about indemnification"). **Transform:** The raw AI outputs are transformed into structured insights and recommendations. For instance, if the AI identifies *"Clause 5 has unlimited liability"*, the system transforms that into a structured entry: `{ clause:5, issue:"Unlimited Liability", severity:"High", recommendation:"Add liability cap" }`. Additionally, a semantic representation of the contract is created here -- mapping relationships such as which obligations pertain to which party, which clauses relate to payment, confidentiality, IP, etc. Using Dinis's GraphFS approach, the system can represent these relationships in a **graph structure** where nodes might be clauses, obligations, or entities (parties), and edges describe relationships (e.g. *"Clause 8" ENFORCES "Confidentiality Obligation" ON "Consultant"*)[3][4]. **Save:** All results (the structured data, graphs, and even raw AI outputs) are saved via a caching layer for persistence. By saving both raw and processed forms of data, the pipeline improves transparency and makes it easier to trace how an AI observation became a final recommendation[2]. This structured approach also aids in debugging and refining the AI prompts over time, since each intermediate step is recorded rather than being a black box. - **Key Risk Identification & Highlighting:** One of the standout features is an automated "risk review." Using the data from extraction, GenLegalAdvise applies a set of rules and heuristics (developed in collaboration with legal experts) to flag high-risk or uncommon clauses. For example, it checks if liability is mutual and capped; if indemnities are one-sided or overly broad; if IP assignment is present without a license back for pre-existing IP; if there's an arbitration clause or governing law that might be unfavorable, etc. The GenAI contributes here by providing context -- it might explain, *"Clause 12 requires arbitration in the vendor's country, which could be costly for you as a freelancer"*. Each risk is displayed with an explanation of *why* it matters, in plain language. The system prioritizes these findings so the user sees the most critical issues first. For instance, truly severe red flags like "Unlimited liability for you" or "You grant all future IP rights" would be marked with a high-severity indicator. Milder issues (like "short 5-day payment timeline, which is unusual") might be lower priority. - **Redline Suggestions Generator:** For each identified issue, GenLegalAdvise can generate a proposed solution in the form of a contract redline or alternative clause wording. This feature uses the AI models to craft text that could replace or append to the problematic clause. It does so in a **practical, negotiation-friendly tone**. For example, if a non-compete clause is too broad, the suggestion might be: *"Limit the non-compete to 6 months and to the specific client industry, instead of an open-ended restriction."* The system could present this as a snippet of text or as an edit that the user can directly copy into the contract. By providing a concrete suggestion, the tool goes beyond issue-spotting and helps the user move towards resolution. These suggestions are informed by legal best practices (and, if the tool is used by lawyers, could be further refined by them). It's made clear, however, that these are starting points -- the user should confirm they fit their situation. - **User-Friendly Summaries:** In addition to the granular analysis, GenLegalAdvise outputs a high-level summary of the document that *anyone* can understand. This summary is akin to an executive brief -- it states the purpose of the agreement, the roles of each party, and the major points (deliverables, timeline, payment, confidentiality, IP ownership, termination). Legal jargon is either avoided or explained. For instance, instead of quoting a clause that says "The Consultant shall indemnify and hold harmless the Client...", the summary would say *"You agree to cover any losses the client suffers related to your work (indemnity), essentially meaning if something goes wrong because of your work, you might have to pay for it."* This level of clarity helps users truly grasp what they're agreeing to. The summary also notes any especially unusual terms by prefixing them (e.g. "**Unusual Term:** This NDA does not have a time limit, meaning your confidentiality obligation never expires."). - **Interactive Q&A and Editing Assistance:** The platform includes an interactive chat or Q&A interface (leveraging the LLMs) where users can ask follow-up questions about the contract. They might ask, *"What does clause 7 mean in simple terms?"* or *"Is there any clause that talks about data privacy?"*, and the system will answer based on the document's content. Because the system has a semantic understanding of the contract (via the knowledge graph and extracted data), it can pinpoint the relevant section and explain it or provide additional context. Users can also request the AI to **draft communications or edits**. A key use-case is drafting negotiation emails: e.g., *"Write an email to the client requesting to add a liability cap of 12 months of fees to the contract."* The AI will use the specifics of the contract (like referencing Clause 10 if that's the liability clause) to create a polite, professional email. The user can refine this by instructing, for example, *"shorten this email and make it more friendly"*, and the AI will adjust accordingly. This interactive loop continues until the user is satisfied. At all times, the underlying contract data and AI analysis remain available, so the AI's answers and edits remain grounded in the actual document (reducing hallucination risks). - **Semantic Graph & Obligation Mapping (Advanced Feature):** As an optional advanced feature (likely in a later version of the product), GenLegalAdvise can produce a **semantic knowledge graph** visualization of the contract. This graph shows the relationships between key elements: obligations, parties, and clauses. For example, one might see a node for "Consultant" and nodes for obligations like "Non-Disclosure" or "Liability", and edges that link them to the specific clause where that obligation appears. Different types of obligations could be color-coded (all payment-related nodes in green, confidentiality in blue, liability in red, etc.), giving a visual map of the contract's structure. This layered summary allows a power-user (or a lawyer) to quickly inspect how responsibilities and rights are distributed. It also aids in ensuring completeness -- e.g., one glance could show if all obligations are one-sided on the consultant. The graph is interactive: clicking a node might highlight the clause text associated with it. Under the hood, this is enabled by the aforementioned GraphFS integration: contract data is stored in a graph form, making such visualizations possible directly from the data structure[3]. The semantic graph feature emphasizes explainability and traceability, aligning with the system's goal of making legal documents more transparent. All these features are built to work in harmony. The **multi-model AI analysis** ensures comprehensive coverage, the pipeline and knowledge graph ensure structured and transparent data handling, and the user-facing tools (summaries, redlines, Q&A) provide multiple ways for the user to digest and act on the information. The end result is a platform that doesn't just dump an AI-generated blob of text, but rather delivers a *structured, annotated, and actionable* review of a legal document. ## Feature Roadmap (MVP to Future Enhancements) The development of GenLegalAdvise will be staged to deliver immediate value with a Minimum Viable Product (MVP) and then gradually add advanced capabilities (like the semantic graph) as the product matures. Below is the proposed feature roadmap: **MVP Features (Initial Release):** - *Document Upload & Basic Parsing:* Ability to upload or paste contracts (text and common file formats). Basic parsing to detect sections and split text for the AI to handle long documents. - *Single-Model Analysis:* Initially, to simplify, the MVP might use a single strong LLM (e.g., GPT-4) to perform the analysis. It will extract key points and flag obvious red flags using prompt-based logic. - *Key Clause Extraction:* Identification of parties, payment terms, dates, termination clause, liability clause, confidentiality clause, and any explicit IP ownership terms. These will be listed for the user, possibly in a tabular format for clarity. - *Basic Risk Highlighting:* A first pass at highlighting risky clauses, using a predefined list of patterns (e.g., phrases like "hold harmless" might trigger an indemnity flag, "in perpetuity" might trigger a duration concern flag). The model's output will supplement this by explaining the context of those clauses. - *Plain Language Summary:* An automatically generated summary of the contract's purpose and main terms, in 4-8 bullet points of simple language. - *Simple Suggestions:* For each flagged issue, a brief suggestion will be provided. In MVP this might be templated or simple text (not a full legal clause rewrite, but guidance like "You might want to negotiate this term"). - *Interactive Q&A (basic):* The user can ask one or two follow-up questions in a chat interface about the contract and get answers. This will be limited in scope to ensure reliability (for example, focusing only on clarifying meaning of clauses). - *User Interface:* A clean web-based UI where users can perform the above actions. The interface will likely have three main panels: one showing the original text (with highlights on problematic clauses), one showing the summary/risks, and one chat area for Q&A and suggestions. - *Serverless Backend & Caching:* The MVP will run on a serverless function backend (AWS Lambda or similar) to handle analysis requests, using the caching service to store results. This ensures even the MVP is cost-efficient and scalable from day one. Integration with the MGraph-AI Cache Service will be done at this stage for storing documents and AI outputs. **Post-MVP / Future Enhancements:** - *Multi-Model Parallel Analysis:* Introduce the full parallel model setup (GPT-4 + Claude, or others like Cohere or open-source LLMs). Develop a "consensus" mechanism to merge findings from multiple models and present a unified result, or highlight differences as needed. - *Redline Automation:* Move from just suggestions to actual redline generation. This could involve the AI producing marked-up text (e.g., in Word's Track Changes format or a markdown diff) that the user can download and send back. This feature might require careful formatting handling and could be offered for the most common document types first (like Word's DOCX). - *Expanded Document Types:* Support for more types of legal documents such as **privacy policies, data processing agreements (DPAs), employment offers or option agreements**, etc., which freelancers and small businesses also encounter. Each new type might involve developing specialized semantic knowledge graphs and prompt strategies to know what specific issues to look for (e.g., data processing agreements might involve GDPR-related nodes and relationships in the knowledge graph). - *Knowledge Base of Clauses:* Build a repository of common clauses and fallback terms. The system could recognize a clause from its knowledge base (e.g., a very typical NDA non-disclosure clause) and simply inform the user "this is a standard clause". Conversely, if it's a rare or aggressive clause, it would note that. Over time, this knowledge base (populated via open-source contributions or public domain legal texts) can make the AI's advice more grounded. - *Learning from Feedback:* Implement a feedback loop where users (especially lawyers) can mark the AI's outputs as helpful or not, and suggest corrections. For example, if the AI misses a risk or gives a bad suggestion, the user could flag that. These insights would be used to refine the prompts, enhance the semantic knowledge graphs, and adjust the system's rules, continually improving the platform. A community forum or GitHub repository could collect "recipes" for better prompts, improved semantic graph schemas, or new risk checks, given the open-source nature. - *Semantic Graph Visualization:* Introduce the semantic graph feature described earlier. This would likely start as an "beta" feature for power users. It might include a graph viewer in the UI where users can toggle on an interactive graph of the contract. Users could filter the graph to see, say, all payment-related obligations and navigate from there. Achieving this will leverage the GraphFS data already being collected; the challenge will be an intuitive visualization, possibly using existing open-source graph visualization libraries. - *Advanced Q&A and Agentic Assistance:* Evolve the chat interface into a more **agentic assistant** that can perform tasks. For instance, beyond Q&A, the user might say "Compare this consulting agreement to my last one" -- and the system (with user's permission and data) could fetch a previous contract from storage and produce a comparison. Another example: "Summarize the differences between this NDA and the template from X organization." This involves the system performing multi-document analysis, which would be a later-stage feature requiring careful design and likely more computing power. - *Collaboration and Multi-User Support:* As small businesses may have teams, eventually allow multiple users to collaborate on a document review. For example, a startup founder could share the GenLegalAdvise analysis of a contract with their co-founder or even their external lawyer via a secure link. That collaborator could add comments or feedback. Think of it as Google Docs-style commenting but on top of the AI analysis outputs. This drives the platform more into a productivity tool space. - *Mobile App or Integration:* Develop a mobile-friendly version or app for quick checks on the go. Alternatively, integrate GenLegalAdvise into platforms where these documents are encountered -- for instance, a plugin for email (to analyze an attachment contract directly in Gmail/Outlook) or integration with electronic signature platforms like DocuSign to "Review with GenLegalAdvise" before signing. Such integrations, however, would require API stability and likely come once the core system is robust. - *Regulatory Compliance Checks:* In the future, for certain document types, the tool could incorporate compliance checks (e.g., if a user is in the EU, does a contract have a GDPR clause; or checking if an employment contract complies with local labor law basics). This would require jurisdiction-specific data and possibly partnerships with legal experts, and may be offered as premium add-ons or templates rather than core features. Each of these future enhancements will be guided by user feedback and available resources. Thanks to the open-source strategy, community contributions may accelerate some of these features. For instance, an open-source contributor might build an experimental UI for the graph visualization or contribute prompt tuning for a new type of contract. The roadmap remains flexible, but grounded in the core mission: **make legal document review fast, accessible, and thorough for those without easy access to legal counsel**. ## Technology Stack and Integration with Open-Source Components GenLegalAdvise will be built on a modern, **type-safe** and modular tech stack that emphasizes reliability and leverages Dinis Cruz's existing open-source components for rapid development. Here is an overview of the planned technology stack and how each component fits into the system: - **Language & Framework:** The core platform will be developed in **Python**, taking advantage of its rich ecosystem for AI and web frameworks. Specifically, we will use **FastAPI** for the web API layer (which serves both the web frontend and any future API consumers). FastAPI is chosen for its performance, ease of use, and integration with Python type hints (enabling type-safe request/response models). The platform will be built on the **osbot-fast-api-serverless** framework, which makes it simple to create new FastAPI services that run in AWS Lambda via a fully tested CI pipeline[5]. This proven foundation has been successfully used across multiple services and provides a robust deployment pattern. - **Type_Safe Data Models:** We will use Dinis's `Type_Safe` classes, which provide the foundation for everything with excellent runtime type safety (not just in the IDE or on initialization like Pydantic). This means every piece of data -- a parsed clause, an AI-extracted issue, a recommendation -- will be validated against a schema at runtime. By using this proven type-safe design, we reduce runtime errors and ensure the system's components speak a common, expected data format[6]. For example, we will have a class `ContractIssue` with fields like `clause_number: int`, `issue_type: str`, `severity: str`, `recommendation: str`. All functions that handle contract issues will use this, preventing inconsistency. This approach provides confidence that our AI outputs (which can be unpredictable) are checked and normalized before use, with validation happening continuously during execution rather than just at startup. - **AI Models & Orchestration:** The GenAI models (GPT-4, Claude, etc.) will be accessed via their APIs (OpenAI, Anthropic). The platform will leverage **OSBot-LLMs** (available at https://llms.dev.mgraph.ai), which provides excellent multi-model support, type-safe JSON responses, and integrated caching and archiving capabilities. This proven system will manage prompts and combine multi-model outputs efficiently. For cost efficiency, the system will use models judiciously: e.g., use GPT-4 only for the longest or most critical analysis parts and GPT-3.5 or other cheaper models for simpler tasks (like drafting an email from a known summary). We will constantly monitor token usage and utilize the caching layer to avoid duplicate calls (if the same document was analyzed recently, etc.). - **MemoryFS for Storage Abstraction:** The platform will use the **MemoryFS** abstraction extensively for handling file operations and in-memory data management. MemoryFS provides a unified interface to handle files whether in memory or on disk/cloud, which is ideal for a serverless environment where local disk might be transient. For instance, when a user uploads a contract, it will be stored via MemoryFS -- abstracting whether it stays in memory or is persisted to S3 -- and given a unique content-addressed ID. This design allows easy passing of the document data between components, as everything can treat it like a file system operation (open, read, write) without worrying about the underlying storage details[8]. It also means down the line we could swap the storage backends (e.g., to a local disk, a different cloud) with minimal changes, thanks to the abstraction. - **GraphFS for Semantic Links:** All semantic data (the knowledge graph of the contract's clauses and concepts) will be managed with **GraphFS**. GraphFS is Dinis's concept for treating graph data through a filesystem-like interface. Essentially, it will let us create and traverse relationships (edges, nodes) as if navigating directories or files[9]. In GenLegalAdvise, when the AI identifies a relationship like "Clause 5 -> obligation -> Consultant", we can store that as something like a path `/contracts/{doc_id}/Clause5/obligation/Consultant` in GraphFS (hypothetically). This unified representation means we don't necessarily need a separate graph database; instead, our existing storage (S3 via MemoryFS) can hold these structures, and we can query them with GraphFS utilities. It's a very *developer-friendly* way to integrate knowledge graphs, leveraging file paths and JSON files to represent nodes/edges. Additionally, by storing these graphs in a standardized format (like JSON-LD or another common graph JSON), we ensure compatibility if we later integrate a dedicated graph database or need to export the data[4]. - **Semantic Knowledge Graph Construction:** Building on MemoryFS/GraphFS, the system will create a semantic knowledge graph for each document. This involves using the AI to extract entities (like party names, product names, jurisdiction mentions), obligations (e.g. nondisclosure, payment, liability), and their inter-relations. We will use an **ontology or schema** for legal documents to normalize this (for example, define categories like "Payment Term" -> relates to -> "Party" or "Duration" -> relates to -> "Obligation"). This structured data not only powers the advanced graph visualization, but even in the background it helps with reasoning. By converting the unstructured text into a structured form, we essentially give the AI (and the user) a second way to query the contract[3]. For instance, if the user asks, "Who has obligations in this contract?", we can answer by traversing the graph (which might show obligations of Consultant vs obligations of Client). This graph approach is core to the system's design philosophy and is shared with Dinis's other projects -- demonstrating the reuse of a successful pattern across domains[10]. - **Serverless Deployment (AWS Lambda):** The entire backend will be architected to run on **serverless infrastructure**, specifically AWS Lambda for the compute and AWS S3 for storage. Using the `osbot-fast-api-serverless` framework, our FastAPI app can be packaged as a Lambda function, giving us scalability and low cost overhead. Each analysis request can spin up in a Lambda, call the AI models, perhaps store results, and terminate -- we only pay for the compute time actually used. This is a proven approach in Dinis's other startups and keeps fixed costs minimal[11]. It also inherently scales: if 100 documents are uploaded at once, AWS will run as many Lambdas in parallel as needed (within account limits) to handle the load. Cold start times are mitigated by using techniques from Dinis's projects (like keeping the Lambda package lean and using provisioned concurrency if necessary for rapid response). - **MGraph-AI Cache Service Integration:** We will integrate GenLegalAdvise with the MGraph-AI Cache Service (available at https://cache.dev.mgraph.ai) for intelligent caching of content and AI responses. This cache service provides **content-addressable storage with multiple strategies** (direct, temporal, versioned, etc.) on top of S3, with a MemoryFS layer[12][13]. In practical terms, when a user uploads a document, the service can save it with a hash; if the *same* document (or even the same paragraph) is uploaded later, we recognize it via the hash and could skip re-processing it fully, pulling the cached results instead. Similarly, after AI models produce an output, we can cache those results keyed by the combination of input text + prompt. This means if our system or another service asks a similar question on the same text, we retrieve the answer instantly. The cache service's support for **semantic file storage** (introduced in v0.5.30) will be useful to store files under readable paths (e.g., `/contracts/{user}/{contract_name}/analysis.json`) which is great for debugging and manual inspection[14]. The cache's type-safe API responses (it returns JSON with metadata) further align with our type-safe design. By building on this service, we avoid reinventing storage logic and gain a robust, battle-tested caching layer that fits our serverless, AWS-based architecture out of the box. - **Integration with LETS Pipelines:** As noted, the system's processing flow follows the LETS (Load-Extract-Transform-Save) methodology. In implementation, we might literally incorporate code or libraries from Dinis's previous pipelines. For example, if there's an open-source library or template for a LETS pipeline (maybe something like an orchestrator that enforces these steps), we'll adopt it. This ensures each step's output is logged and available for the next, making the process transparent. It also helps with **provenance tracking** -- a concept Dinis emphasizes. We will tag every piece of data derived from the document with trace information (e.g., "this suggestion was generated from clause 7 via prompt X at time Y") and save that. If an issue arises, we can trace back how the AI arrived at a certain output. This level of detail is crucial in legal contexts to build trust that the AI isn't making things up without basis[2]. - **Frontend Technology:** On the front-end, a modern JavaScript framework like **React** (possibly with TypeScript for type safety) will be used to create a smooth user experience. It will communicate with the backend via REST API (or GraphQL if we decide on that). The front-end will handle uploading files, displaying the analysis results (including nice rendering of the original document text with highlights, which could be done with a library for rendering PDFs or using HTML if the doc is converted), and providing the chat interface. We might also incorporate a graph visualization library (like D3.js or vis.js) for the semantic graph feature in the future. Given the focus on structured output, the UI design will likely involve tables and accordions (for clauses and details) and an intuitive way to switch between the summary view and detailed view. The technology stack is chosen to be **open-source friendly and modular**. By using and extending open components (FastAPI, MemoryFS, etc.), we ensure that GenLegalAdvise can be built in a lean way without heavy proprietary software. Moreover, this stack positions the project to accept contributions: Python and JavaScript are widely known, and the architecture (serverless functions + S3 storage) is accessible to replicate for development. Security is also a consideration: handling legal documents means we'll enforce encryption (S3 buckets will be encrypted, data in transit via HTTPS) and we may allow self-hosting for those who are extra cautious (since it's open source, an organization could deploy their own instance). In summary, the stack is a blend of **AI capabilities, semantic data processing, and cloud-native infrastructure**, aligned with Dinis Cruz's architecture philosophy of type-safe design, serverless deployment, and semantic knowledge representation. ## Infrastructure and DevOps Considerations Building GenLegalAdvise on a solid infrastructure foundation is critical for reliability, scalability, and cost-effectiveness. We will use a cloud-native, serverless infrastructure with automation pipelines, ensuring the platform can scale to many users without large fixed costs. Key aspects of the infrastructure plan include: - **Serverless Architecture on AWS:** GenLegalAdvise will primarily run on AWS using a serverless approach. AWS Lambda will host the backend API and processing tasks. Each key function (document ingestion, AI analysis coordination, result compilation) can be a separate Lambda function or a set of functions behind a single API Gateway. This design means we **incur cost only per execution** and can scale automatically. As more users upload documents, AWS will allocate more Lambda instances to handle the load. We avoid maintaining servers or paying for idle capacity, keeping the operation cost extremely low when usage is low[11]. AWS API Gateway will provide the HTTPS endpoints, and AWS S3 will serve as the durable storage (for documents, cached results, etc.). We will also use AWS CloudFront (a CDN) if needed to serve static assets or to cache results geographically for performance. - **CI/CD Pipeline and Deployments:** We will implement a Continuous Integration/Continuous Deployment pipeline, likely using GitHub Actions (given the open-source nature) to test and deploy the code. Dinis's unified CI/CD approach means we can package the application (backend and possibly front-end) and deploy to AWS quickly[15]. Infrastructure-as-code tools like AWS SAM or Terraform will be used to define our cloud resources (Lambda, API Gateway, S3 buckets, IAM roles) so that the environment is reproducible. Every commit to main could trigger automated tests (including perhaps running some sample document analyses through a stubbed AI model for determinism) and then deploy to a development environment. We might maintain separate stages: "dev", "staging", "prod" with corresponding AWS setups, which allows testing new features on a staging environment with limited users before full release. - **Caching Layer (MGraph-AI Cache Service):** We will deploy the MGraph-AI Cache Service (https://cache.dev.mgraph.ai) as part of our infrastructure. This service is serverless (running on AWS Lambda + S3) and can be integrated directly into our architecture. The cache will handle storing the content and results. The cache service provides multiple **caching strategies** out-of-the-box (direct, temporal, versioned, semantic)[16][17]. For GenLegalAdvise, we will use: a **direct cache** for content (store by hash of document text), a **temporal cache** for keeping history of analyses (so a user can see previous versions or analysis runs), and a **semantic file cache** for organizing outputs by user or project (e.g., all files for User123's ContractABC under one folder path). The cache service also helps in multi-user scenarios with its namespace feature (each user or team could be a namespace, isolating their data)[18]. Deploying this cache service in our AWS environment ensures low-latency access (since Lambdas and S3 in the same region communicate quickly) and secure storage. - **Cold Start and Performance Optimizations:** Lambda cold starts can be an issue, especially for a Python app that might include heavy libraries (like AI SDKs). To mitigate this, we will use techniques such as keeping the deployment package slim (excluding unnecessary libraries), possibly using Lambda layers for large libraries (so they're cached by AWS), and considering provisioned concurrency for critical parts (maybe keep 1 instance warm during business hours). The **osbot-fast-api-serverless** framework we're using already has cold start optimizations built-in[19]. For instance, it delays heavy imports and uses lightweight stub servers. We will also monitor response times; if needed, we could adjust memory allocation (more memory in Lambda can mean faster CPU performance) to ensure the AI calls and processing happen swiftly. The aim is for the user not to wait more than a few seconds for initial results on a moderate-size contract (perhaps longer for very large documents or during peak loads). - **Cost Management:** Cost-effectiveness is a key design tenet. Using serverless ensures we only pay for what is used, as noted. We'll also implement caching to avoid repeat AI calls, which are the most expensive part (GPT-4 calls have a cost in USD per 1K tokens). By caching results, if the same clause or same contract needs analysis, we won't spend on AI again unnecessarily. We'll likely also implement rate limiting or user-specific quotas to control abuse (especially on a free tier). Monitoring tools (like AWS CloudWatch and custom dashboards) will track usage and costs. If we detect certain features are expensive (e.g., the graph generation might call the AI a lot), we can tune those (maybe make them optional or batch the calls). Thanks to open-source and serverless, fixed costs like software licenses or idle server time are essentially zero[11], making the platform financially sustainable even with many free users, as long as we manage the AI call costs smartly. - **Security & Privacy:** From an infrastructure standpoint, we handle sensitive documents, so security is paramount. All data at rest in S3 will be encrypted (AWS SSE). Data in transit will be encrypted via HTTPS. We will enforce strict IAM roles such that Lambdas only have access to the specific S3 paths (namespaces) they need. For example, a Lambda handling user X's request can only read/write in the cache namespace for user X. API endpoints will use secure tokens or keys for authentication when we have user accounts. In the open-source spirit, individuals or companies who self-host can integrate with their identity systems or run it isolated in their VPC. We will also consider data retention policies -- small users might not want their documents stored indefinitely. Our cache could, for instance, use a temporal strategy to auto-expire data after a user-defined period (unless they save it). This ties into our use of **MemoryFS temporal** capabilities if needed[20]. - **DevOps and Monitoring:** We will incorporate logging at various levels (each step of LETS pipeline logs events, each AI model call logs input size and outcome length, etc.). CloudWatch Logs will capture these, and we can build alarms for anomalies (e.g., sudden spike in errors or huge cost usage). For DevOps, since the project is open-source, we might involve the community in code reviews via GitHub. Using Infrastructure-as-Code means contributors can spin up a dev environment (on their own AWS account) relatively easily by following our documentation, which fosters outside contributions. We'll also document how to run the system locally (perhaps in a Docker container that emulates the AWS services locally for testing). - **Integration with Dinis's Ecosystem:** Because this project shares DNA with others, there are synergy opportunities at the infrastructure level. For example, if Dinis's other startups (like MyFeeds.ai or Cyber Boardroom) are running in the same overarching environment, they could potentially share the cache service or the graph database. This isn't a given, but we'll design with interoperability in mind. The standardized graph and file storage means if, say, a Cyber Boardroom instance wanted to pull in a "legal risks summary" from GenLegalAdvise for a board report, it could read the data from our S3 (with permission) in a known format. This kind of cross-venture integration is facilitated by using consistent open formats and APIs across projects[21]. - **Scalability Testing:** We will perform load tests to ensure the system can handle realistic scenarios. A likely usage pattern is small bursts of activity (e.g., a user uploads a contract, then maybe another, then goes idle). We'll simulate concurrent uploads by multiple users to see that our Lambda concurrency scales and that our cache (S3) doesn't become a bottleneck. AWS is quite elastic, but we need to be mindful of any limits (like Lambda concurrency limits or API Gateway throughput). Early testing might be done with smaller models (or mocking the AI calls) to focus on system throughput. When including real AI calls, we'll check how many can run in parallel without hitting rate limits of the AI API providers. If necessary, we might queue some requests or degrade gracefully (for instance, if 50 people upload at once and we are limited by AI API, we might process a few at a time and inform others of a short wait). Overall, the infrastructure is designed to be **robust, low-cost, and scalable from the get-go**. It uses serverless principles to avoid the pitfalls of big up-front investments in servers or operations. Instead, we lean on cloud providers for auto-scaling and on open-source DevOps practices for rapid iteration. This approach not only saves cost but also aligns with a lean startup model -- we can handle a growing user base without a significant rewrite or migration. It's also attractive to potential collaborators or investors, as it shows we can grow efficiently and securely. ## Synergies with Legal Professionals While GenLegalAdvise is aimed at empowering non-lawyers, an important principle of the project is to **complement, not replace, professional lawyers**. The platform is designed with input from legal professionals and in a way that can integrate into a lawyer's workflow, rather than working against it. Here's how GenLegalAdvise synergizes with legal professionals: - **Augmenting Lawyers' Efficiency:** For lawyers (especially those serving small businesses or startups), GenLegalAdvise can handle the *initial triage* of a document. Instead of a lawyer spending an hour reading a 10-page contract to spot standard issues, they could use the tool to get a quick rundown. The lawyer can then focus their time on the nuanced aspects and on advising the client about the implications and negotiation strategies. This means lawyers can serve more clients faster, or focus on higher-value analysis, improving their productivity. In this sense, GenLegalAdvise is like an assistant that does the heavy lifting of reading and summarizing, under the lawyer's supervision. - **Incorporating Legal Expertise:** We plan to involve legal professionals in the development and refinement of the platform. Their expertise is crucial in defining what constitutes a "red flag" or what a good suggestion is. For example, an experienced contract attorney will know that an indemnity clause might be acceptable in one context and not in another -- these nuances can be built into the AI's prompt engineering or rule-based checks through collaboration. As an open-source project, we invite lawyers to contribute by reviewing the output quality and suggesting improvements. Perhaps a lawyer might contribute a better prompt for summarizing limitation of liability clauses, or provide a list of top 10 issues to always check in an NDA which we incorporate into the logic. - **Customizable Playbooks:** Law firms or individual lawyers could extend GenLegalAdvise with their own "playbooks." A playbook could be a set of custom rules or model prompts reflecting a firm's philosophy or a client's standards. For instance, a freelance lawyer working with startups might add a rule: "Flag if equity compensation is mentioned in a contractor agreement" because that's something they particularly care about. The platform could allow loading such custom checks (perhaps as simple configuration files or Python plugins) so that the AI's analysis aligns with what a human lawyer would do for that client. This makes GenLegalAdvise a flexible aide for different legal practices rather than a one-size-fits-all black box. - **Review and Approval Workflow:** GenLegalAdvise could include a mode where a lawyer can "review" the AI's output before it goes to the end-client. Imagine an independent consultant uses the tool and gets a summary and suggestions; they could then share this with their lawyer. The lawyer might use a special interface to see each flagged issue and either approve it, modify the advice, or add additional notes. The end result would be a lawyer-approved version of the analysis. This kind of workflow would increase trust for end users (knowing a human verified it) and keep lawyers in the loop, using the AI as a preparatory step. It opens possibilities for **lawyer-AI collaboration**, perhaps even as a service: some lawyers might advertise that they use such AI tools to provide quicker turnaround and pass some savings to the client. - **Not a Lawyer, and We Know It (Disclaimers):** The platform will make clear that it is not a substitute for professional legal advice, especially for complex or high-stakes agreements. By being open about its role, GenLegalAdvise sets the stage for cooperation with lawyers. Lawyers can feel more comfortable that clients won't just take the AI output and ignore them. In fact, the tool might often advise users to *consult a lawyer* when it encounters something particularly unusual or outside its confidence zone. For example, if a contract has a very domain-specific clause (like a patent license grant), the AI might flag it and say "This involves specialized legal considerations; consider getting a professional opinion." This kind of self-awareness, guided by programming, will be built in to hand off to humans when appropriate. - **Education and Training:** GenLegalAdvise's detailed explanations and semantic graphs could be used by junior lawyers or law students as a learning aid. By seeing how an AI breaks down a contract and identifies issues, they can learn to do the same. We might collaborate with legal clinics or education programs to pilot the tool as a teaching assistant. The more the tool is used by people with legal knowledge, the better it can become (through feedback), and those users also benefit by cross-checking their own analyses against the AI (a kind of double-check). - **Legal Partner APIs and Integration:** Law firms or legal-tech companies could integrate GenLegalAdvise via an API (as mentioned in the business model). For instance, a contract management software used by a law firm could call GenLegalAdvise API to pre-analyze uploaded contracts and populate fields in their system (like filling a risk checklist automatically). This extends the lawyer's capabilities. We foresee that some lawyers might build specialized services on top of GenLegalAdvise (given it's open source, they can even run their tailored version), such as a niche version for, say, real estate leases or healthcare contracts, with more domain-specific checks. GenLegalAdvise, by being open and extensible, becomes a foundation that legal professionals can build upon to deliver faster or more consistent service. - **Maintaining Quality and Trust:** By inviting legal professionals into the process, we ensure the tool's advice stays **grounded and trustworthy**. Lawyers think in terms of worst-case scenarios and precise language -- this mentality can help guide the AI to avoid hallucinations or over-generalizations. For example, a lawyer contributor may insist the summary always mentions governing law if it's present, since that's important -- we can then adjust the model prompts to always extract governing law clauses. These kinds of refinements, driven by professional standards, will make the output more reliable. Over time, if the legal community sees the tool as beneficial rather than adversarial, they are more likely to contribute domain knowledge, which in turn improves GenLegalAdvise for all users. In essence, GenLegalAdvise is positioned as a **co-pilot for legal reviews**. It does the tedious drafting and reading work at machine speed, but leaves the final judgement and complex reasoning to humans. By aligning with the interests of legal professionals and proving useful to them, the project can tap into a wealth of knowledge and also avoid the resistance that comes when technology tries to displace professionals. Instead, we aim to **empower lawyers** to deliver their services more effectively, and **empower non-lawyers** to know when and what to ask lawyers by giving them preliminary insights. This collaborative approach will drive adoption and continual improvement in the legal review ecosystem. ## Open-Source Strategy and Licensing GenLegalAdvise will be an **open-source project** from day one, aligning with the core philosophy that transparency and collaboration lead to better software -- a principle strongly advocated by Dinis Cruz. The open-source strategy is not just a licensing decision, but a fundamental approach to building trust, accelerating innovation, and creating community-driven momentum. Key points of this strategy include: - **Licensing Model:** We plan to release GenLegalAdvise under a permissive open-source license, likely the **MIT License** or Apache 2.0. This will allow individuals, startups, or law firms to freely use, modify, and even commercialize the core platform (with attribution), which encourages wide adoption. We want as few barriers as possible for adoption -- if a freelancer wants to self-host GenLegalAdvise on their laptop, they can; if a legaltech startup wants to integrate it into their product, they can do so as well. We believe that the value we provide (and monetize, see Business Model) will be in hosted services and additional layers, not in restricting the core IP. An open license also makes it easier for other developers to contribute without legal entanglements. - **Community Contributions and Collaboration:** By open-sourcing the project, we invite a global community of developers and legal experts to contribute. This could be in the form of code, semantic graph schemas, knowledge base entries, or simply by filing issues and feature requests. We will maintain a public repository under established GitHub organizations like **OWASP-SBot** and **The-Cyber-Boardroom**, where related open-source components and libraries have been successfully developed. The repository will contain the full source code, documentation, example contracts, semantic knowledge graph templates, and perhaps a library of sample AI prompts for various legal analyses. The project can leverage existing Python libraries from these organizations, including the MemoryFS, GraphFS, and Type_Safe implementations. We'll encourage an environment where even users who aren't coders can contribute by sharing feedback or by helping to curate a list of "known problematic clauses" that we can encode in the knowledge graph. Open sourcing invites scrutiny and improvement: others can inspect the code for bugs or security issues (crucial for a tool that handles sensitive documents) and propose fixes -- this peer review improves quality[22]. - **Transparency and Trust:** Legal advice is an area where trust is critical. Users need to trust that the platform isn't misusing their data and that the advice is impartial and based on facts. By having the code open, anyone can audit how we handle documents (ensuring, for instance, that we're not silently sending documents to third parties beyond the stated AI models) and how the logic works. Enterprise users (like a company legal department) would be more willing to adopt or integrate an open-source tool because they can vet it for compliance. Also, open source means we can integrate more easily with other open legal data initiatives, and we can quickly adapt to any changes (for example, if OpenAI updates their API, the community might even contribute the fix). In the cybersecurity domain, Dinis has noted that open-source tools are preferred because they can be vetted[23] -- the same likely holds in legaltech for slightly different reasons (vetting for correctness and privacy). - **Open Data and Knowledge:** In addition to code, some outputs or knowledge bases can be open-sourced. For example, if we accumulate a list of common clauses and what they mean, that could be published as an open dataset or incorporated into the documentation. We might maintain a wiki of legal terms and model prompts (e.g., "How to prompt GenLegalAdvise to check for X in a contract"). The semantic knowledge graphs and relationship schemas developed for legal document analysis could be shared as templates for others to extend. The idea is to create an ecosystem where people building anything related to legal document analysis see GenLegalAdvise as a reference and starting point, particularly for semantic graph-based approaches. - **Avoiding "Closed-Source Temptation":** We will refrain from keeping any core feature proprietary. Some companies open-source a "lite" version and keep an "enterprise" version closed. Our strategy is different: **the full core capability** (document analysis, risk identification, suggestions, etc.) will be open. We might develop some commercial add-ons or services (see Business Model), but those will be more about convenience (hosted service, or human lawyer network) rather than core functionality. This clarity ensures the open-source project remains truly useful and not crippled. It also means the community won't feel like they're contributing to something only to have advanced features walled off -- instead, any improvement goes back to the commons. - **Ecosystem and Shared Innovation:** Dinis's multi-startup strategy emphasizes that open source allows cross-pollination of tech across projects[24][25]. GenLegalAdvise, by being open, might benefit from innovations in those other projects and vice versa. For example, if Cyber Boardroom (one of Dinis's projects) develops a new way to present AI findings in a report, we could adapt that for our summaries. If MyFeeds.ai improves its semantic parsing pipeline, we might incorporate that improvement to better parse contracts. Conversely, advancements we make in legal document analysis (like better long-context handling or a new graph schema for obligations) could feed back into those projects. This synergy only works smoothly if projects are open-source and modular. We will actively document and share any such breakthroughs, contributing to a "GenAI for documents" knowledge base that others can leverage. This aligns with the vision of building GenAI-powered solutions that reinforce each other as accelerators[26]. - **Credit and Community Building:** We will credit contributors and co-authors (as we did at the top of this document). The project will acknowledge Dinis Cruz's leadership and the community's input. Perhaps we will organize this under the banner of an open initiative (maybe an "Open Legal Tech" collective or as part of existing ones like OpenJS or OWASP if relevant). We might present the project in open-source forums or legal innovation conferences to gather interest. This not only helps improve the tool but could also attract early adopters who are power users giving us valuable feedback. - **No Vendor Lock-in:** Users (especially enterprise) fear being dependent on a tool and then having it yanked away or changed. By being open-source, GenLegalAdvise assures them that the core will always be available. Even if our startup pivoted or stopped, the code remains for others to pick up. This assurance can be a selling point, especially for something as sensitive as legal document processing -- a company might hesitate to adopt a closed SaaS for contract review due to risk of the service shutting down, but if it's open-source, they know they could self-host if needed. In this way, open-source strategy directly supports the business by making customers more comfortable using the tool. - **License Choice and Contributions:** With a permissive license, even commercial entities might contribute back improvements because it benefits them to have a robust central codebase. If someone builds an extension for a unique use-case, they might contribute it upstream rather than maintaining a fork, to benefit from future updates. We will encourage this by being very receptive to pull requests and by designing the system to be **extensible** (with plugin-like architecture for custom rules, etc., to make contributions easier without needing to alter core logic drastically). In summary, open source is not just a tagline for GenLegalAdvise; it's how we plan to achieve a **semantic legal review platform that people can trust and build upon quickly**. It accelerates development (more eyes, more ideas), accelerates adoption (transparency builds trust), and even opens up monetization avenues that don't rely on locking down IP. This approach follows the path of successful open-source based companies (for example, how HashiCorp or Red Hat provided services around open tools) -- we aim to provide value on top of the open core, knowing that the widespread use of the open core is our best marketing. ## Business Model and Sustainability While GenLegalAdvise is an open-source project, a sustainable business model will ensure its longevity and continuous improvement. We envision a **hybrid model** that balances free community use with paid offerings for those who need more advanced features, support, or convenience. Below are the key components of the business and sustainability strategy: - **Transparent Consumption-Based Pricing:** GenLegalAdvise operates on a pay-per-use model where customers only pay for what they consume. Every action has a clear, published price: document analysis (per page or per document), storage of analysis results in semantic knowledge graphs (per GB per month), and retention of supporting evidence chains. Pricing is completely transparent with a public pricing calculator showing exact costs before any operation. For example: analyzing a 10-page contract might cost $X, storing the resulting knowledge graph for 30 days might cost $Y, and maintaining the full evidence chain with all AI reasoning steps might cost $Z. The platform adds a reasonable markup to cover infrastructure and development costs, but all base costs (AI API calls, storage, compute) are visible to users. This transparency builds trust and allows users to control their spending precisely. - **Usage-Based Feature Tiers:** While not a subscription, users can choose different processing levels at different price points. Basic analysis might use a single AI model and provide essential risk identification. Premium processing (at a higher per-document cost) could include multi-model analysis for cross-validation, semantic graph visualization, or deeper relationship mapping. Users decide per-document which level of analysis they need. For instance: basic contract review at $X per page, comprehensive multi-model analysis at $2X per page, full semantic graph with obligation mapping at $3X per page. All options are transparently priced and users can mix and match based on document importance. Volume discounts could apply automatically (e.g., 10% off after 100 pages in a month) but are clearly stated upfront. The system might offer "analysis credits" that users can pre-purchase at a discount for budgeting purposes, but these are optional and never expire. - **Enterprise or Self-Hosted Model:** For larger organizations (or firms) that have strict data requirements, we offer transparent enterprise pricing. Since the product is open-source, they could self-host it; our business can provide value via **support contracts, integration services, or custom feature development** with clear, published rates. For instance, a law firm might want to run GenLegalAdvise on their private cloud and integrate it with their document management system -- we would provide a transparent quote for setup and customization. Enterprise clients receive volume-based pricing tiers that are clearly defined (e.g., >1000 pages/month gets 20% discount, >5000 pages/month gets 30% discount). They can also opt for enhanced data retention in semantic knowledge graphs with transparent storage pricing per GB. All enterprise pricing is public and calculator-based, ensuring no hidden costs. We might also offer SLAs (service-level agreements) for uptime and priority support channels at clearly stated prices. - **Legal Partner Network (Marketplace):** A transparent marketplace connecting users to legal professionals with consumption-based pricing. After AI analysis, users can request human lawyer review at clearly published rates (e.g., $X per page reviewed, $Y per hour of consultation). All lawyer rates are transparently displayed before engagement. The platform takes a published percentage (e.g., 15% facilitation fee) that users can see. Lawyers set their own rates which are publicly visible, creating market competition. Users only pay for actual review time or specific deliverables, never flat fees or retainers. The system tracks and displays time spent, pages reviewed, and running costs in real-time. Since our platform has already done the basic analysis and created semantic knowledge graphs, lawyers can review contracts more efficiently, reflected in their pricing. All transactions are itemized with clear breakdowns of lawyer fees, platform fees, and any data storage costs for the lawyer's notes or modifications to the semantic graphs. - **API and Integration Licensing:** We offer a transparent pay-per-call API for other software providers who want to integrate GenLegalAdvise capabilities. Every API endpoint has a published price per call (e.g., $X per document analysis, $Y per semantic graph query). Volume tiers are automatic and transparent (e.g., calls 1-1000 at full price, 1001-10000 at 10% discount, etc.). Electronic signature platforms or contract lifecycle management (CLM) software can integrate our API with full visibility into costs. Real-time usage dashboards show current consumption, costs, and projections. API customers can set spending limits and receive alerts at configurable thresholds. All pricing changes are announced 30 days in advance. The pricing model includes separate transparent charges for data retention (keeping analysis results available via API) and semantic graph storage. - **Value-Added Cloud Services:** While the open-source project can be self-run, we anticipate many users will use our hosted version for convenience. We can have value-added services in the cloud version that might not be as straightforward in self-host (though not impossible). For instance, maintaining updated AI models -- we integrate new model versions as they come (GPT-5, etc.), and optimize our semantic knowledge graphs and prompting strategies for legal text analysis. Our hosted service would always have the latest improvements to the semantic graph schemas and relationship mappings. We might also aggregate anonymized patterns from many contract analyses (if users permit) to enhance our knowledge graphs and improve the AI's understanding of legal relationships. Those improvements continuously roll into the service. Essentially, paying us means you get the most powerful, up-to-date instance of GenLegalAdvise without having to manage it. This is similar to how open-source database companies offer hosted databases with tuning and maintenance included. - **Cost Management and Transparency:** Complete pricing transparency is core to our model. Users see exactly what each operation costs before executing it: AI API costs (with our markup clearly shown), storage costs for semantic knowledge graphs, and data retention fees. A real-time dashboard shows: current charges, cost breakdown by component (AI calls, storage, processing), and projected costs for pending operations. Users can set spending limits and receive alerts. We publish our markup percentage (e.g., "25% markup on AI API costs for operational sustainability"). Historical pricing is maintained publicly, and any changes are announced 30 days in advance. The platform provides cost optimization suggestions (e.g., "Using basic analysis instead of multi-model would save 60% on this document"). By showing exact costs and our markup separately, users understand they're paying for convenience, infrastructure, and continuous development, not hidden fees. - **Grants or Sponsorships:** Given the open-source and public-good nature (access to justice, helping small entities with legal documents), we could seek grants or sponsorships. For instance, an organization focused on access to legal services might fund development of certain features (like a special module for nonprofit use). Or cloud providers might give credits to support the infrastructure in early stages. These are not recurring revenue, but they can help bootstrap the project and are worth pursuing. - **Cross-Selling with Related Products:** If we consider Dinis's portfolio, there might be opportunities to cross-sell services. For example, if a company uses MyFeeds.ai (the personalized news feed), and they are concerned about regulatory news, we might offer them GenLegalAdvise to review their compliance documents or vendor contracts. Not directly obvious, but if a foot is in the door with one tool, the trust can extend to another. Similarly, if Cyber Boardroom targets boards for cybersecurity discussions, those same boards might care about legal documents related to cybersecurity (like policies, contracts) -- GenLegalAdvise could assist their legal team, and could be packaged as an add-on in a deal. This synergy means we should keep branding somewhat consistent and highlight how these tools complement each other (all being open, semantic-driven, etc.). - **Monetizing Knowledge (Independently):** Over time, GenLegalAdvise might accumulate extremely valuable aggregate insights -- e.g., statistics like "80% of freelancers agree to unlimited liability in our dataset" or "average payment terms have moved from 30 days to 45 days in the past year." These kinds of insights (anonymized and aggregated) could be packaged into reports or subscriptions for interested parties (perhaps insurance companies, policy makers, or media). It's a bit speculative, but the point is that the data and patterns from usage have value beyond individual transactions. We would, of course, be careful and only do this in ways that respect privacy and align with our users' interest (for instance, releasing an annual "State of Freelance Contracts Report" could actually be great PR and indirectly monetize by attracting more users). - **Ensuring Sustainability:** The ultimate goal is that revenue from the above streams supports continuous development, hosting costs, and provides a profit margin to fund growth (marketing, support staff, etc.). By having multiple streams -- subscriptions, enterprise services, API licensing, and partnerships -- we diversify our income (not reliant on just one). This also makes us more resilient: if, say, one model of AI becomes too expensive or a competitor offers something similar for free, we still have other value-adds and customer bases to rely on. Community contributions lower our R&D costs in a sense, since volunteers can build features, but we'll likely maintain a core team (funded by the business) to guide the project and handle the heavy lifting. The business model is crafted to align with the open-source ethos: **transparent consumption-based pricing where users only pay for what they use**. Every cost is visible -- from AI API calls to semantic graph storage to evidence chain retention. By showing our markup transparently and charging only for actual usage, we build trust and allow users to control their costs precisely. This approach ensures GenLegalAdvise remains accessible to occasional users (who pay only when needed) while scaling naturally for heavy users (who benefit from volume discounts). The transparency of pricing, combined with the open-source nature of the core platform, means users always understand what they're paying for: the convenience of a hosted service, continuous improvements to semantic knowledge graphs, and the infrastructure to deliver reliable legal document analysis. ## Conclusion GenLegalAdvise is an ambitious project at the intersection of legal tech and AI, inspired by the vision of making expert knowledge accessible through open-source, semantic-driven platforms. By focusing on a real pain point -- the challenge of understanding and negotiating everyday legal contracts -- it has the potential to empower individuals and small businesses worldwide. The approach outlined above leverages the latest in GenAI (using multiple LLMs, knowledge graphs) and proven software architecture patterns (serverless, type-safe design, caching, pipelines) to ensure the solution is not only smart, but also robust, cost-efficient, and scalable from the start. Importantly, this project plan emphasizes **publishing the idea and design openly**. In the spirit of open-source innovation, we encourage entrepreneurs, developers, and legal experts to take these ideas and build upon them. This comprehensive plan provides a blueprint -- from user needs to technical stack and business considerations -- that can jump-start development efforts. However, at this time, the authors (Dinis Cruz and the ChatGPT Deep Research collaborator) are sharing this concept as a contribution to the community, rather than embarking on building it as a proprietary venture. We believe in seeding good ideas and allowing the open-source and startup ecosystem to carry them forward. In conclusion, GenLegalAdvise could be a transformative tool that demystifies legal documents using AI and collaboration. Whether it's adopted in parts or as a whole, we hope this plan spurs new solutions for accessible legal advice. By openly sharing the strategy and technical foundation, we aim to catalyze innovation in legal tech, much like open-source has done in cybersecurity and other fields -- an embodiment of "engineering and AI as accelerators for business and societal good."[26] With the community's interest and effort, GenLegalAdvise or a similar project can become a reality. We're excited to see how others will take this plan, improve it, and implement it to bring about faster, clearer, and fairer legal document review for everyone. --- ## References [1] [2] [3] [4] [9] [10] [11] [15] [21] [22] [23] [24] [25] [26] Dinis Cruz's Multi-Startup Strategy_ Open-Source Innovation Across Four Synergistic Ventures.pdf file://file_00000000dfc8620ab30a653d3cb44f61 [5] [6] [8] [12] [13] [14] [16] [17] [18] [19] [20] v0.5.30__cache-service__llm-brief.md file://file_00000000e37462469bc7ee47e7b1bd30 [7] DinisCruz (Dinis Cruz) · GitHub https://github.com/DinisCruz --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/09/01/dinis-cruz-research-on-api-security-2009-2025.html (markdown twin: /2025/09/01/dinis-cruz-research-on-api-security-2009-2025.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/09/01/dinis-cruz-research-on-api-security-2009-2025.html* # Dinis Cruz's Research on API Security (2009-2025) *By Dinis Cruz and ChatGTP Deep Research · 2025-09-01* > API security has been a persistent theme in Dinis Cruz's work, spanning early insights in 2009--2010 through to innovative ideas in 2025. His contributions center on how to test and analyze web… --- [PDF](https://files.diniscruz.ai/github/pdf/2025/09/01/dinis-cruz-research-on-api-security-2009-2025.pdf) ## Introduction API security has been a persistent theme in Dinis Cruz's work, spanning early insights in 2009--2010 through to innovative ideas in 2025. His contributions center on how to **test and analyze web services/APIs securely**, often by blending static analysis, dynamic testing, and developer-centric tooling. A recurring motif is that *every API and web service is inherently vulnerable at some level* -- thus requiring proactive design, rigorous testing, and automation to mitigate risks[\[1\]](https://diniscruz.blogspot.com/2010/01/#:~:text=When%20doing%20source%20code%20analysis%2C,vulnerability%20pattern)[\[2\]](https://diniscruz.blogspot.com/2010/01/#:~:text=From%20a%20secure%20design%20point,not%20possible%20to%20exploit%20them). Cruz's writings emphasize creating security tests early, making security "invisible" to developers (through automation), and leveraging advanced techniques (like AST manipulation and knowledge graphs) to understand and secure APIs. This debrief summarizes Cruz's key publications on API security (with publication dates) and distills their core ideas, providing AppSec professionals and next-gen AI-driven API security teams a cohesive view of his work. *(Authored by Dinis Cruz and ChatGPT Deep Research.)* ## Key Publications and Concepts on API Security The table below lists Cruz's main writings on API security and testing, with a brief description and their key ideas: | **Publication (Date)** | **Summary & Key Ideas** | |------------------------|-------------------------| | **"Every API is (at some level) vulnerable" (Blog – Jan 9, 2010)** | Introduces the premise that any API/function with capability likely harbors vulnerabilities. Advocates mapping out "layers" of vulnerable functions and tracing if/when they become externally accessible (i.e. exploitable). Stresses that secure API design should *wrap internal flaws* so they can't be abused externally (e.g. .NET's Code Access Security treating partial-trust boundaries as the attack surface). Published 2010-01-09. | | **"ESTAPI idea" (Blog – Jun 2011)** | Proposes the **ESTAPI** (Enterprise Security Testing API) concept. Envisions a standardized way to expose or automate security tests for web apps/services. *Key idea:* make security testing a first-class function via an API, so that dynamic tests and security checks can be uniformly applied. *(Published 2011-06; part of "ESTAPI" series of 5 posts.)* | | **"Fixing/Encoding .NET code in real time (Response.Write)" (O2 Blog – Nov 7, 2011)** | Demonstrates runtime patching of vulnerable code using OWASP O2 Platform. Hooks the `.NET Response.Write` to auto-encode output, preventing XSS on the fly. Emphasizes making security **"invisible/automatic" for developers** by fixing issues in dev environments transparently. Key idea: security frameworks/tools should actively prevent vulns (and allow testing such fixes) during development. | | **"If you're not blowing up the database, you're not testing" (Blog – Apr 2012)** | Argues that effective web service testing must be aggressive. If tests never break anything (e.g. never cause DB errors or crashes), they are too shallow. Recommends testers simulate malicious and extreme conditions – even at risk of "blowing up" components – to reveal security weaknesses. *(Published 2012-04; underscores the need for thorough, adversarial testing mindset.)* | | **"Journey into testing WebServices in TeamMentor" (Blog – Apr 2012)** | A narrative of initial attempts at testing the **TeamMentor** product's web services. Documents setting up a testing environment and early findings. Likely covers using Python scripts and SOAP/REST clients to enumerate API calls. *Key ideas:* the importance of understanding an API's functionality before testing and dealing with real-world complexities (authentication, state, etc.) during security testing. *(Published 2012-04.)* | | **"First, you create tests for WebServices..." (Blog – Apr 2012)** | Advocates a *test-first approach* for web service security: before or as you build a web API, create a suite of security tests. By writing tests that target each endpoint (including negative cases), developers can catch vulnerabilities early. Emphasizes that security requirements should be codified as tests from the start – similar to Test-Driven Development but for security. *(Published 2012-04.)* | | **"Roadmap for testing WebServices" (Blog – May 2012)** | Outlines a structured plan for web service testing. Proposes steps such as: **1)** Enumerate and understand all API endpoints; **2)** Write unit and integration tests for each (covering auth, input validation, business logic); **3)** Use automation (scripts or tools) to fuzz inputs and simulate abuse; **4)** Incorporate tests into CI. Likely highlights needed tooling and roles (QA, AppSec) for each step. *(Published 2012-05.)* | | **"What is the formula for WebServices security testing?" (Blog – May 2012)** | Explores whether a "formula" or repeatable methodology can be applied to test any web service. Suggests combining **static analysis** (to find hidden endpoints or dangerous calls) with **dynamic testing** (to actually try attacks). The "formula" involves mapping trust boundaries, identifying inputs/outputs of each service, then systematically verifying authentication, authorization, input handling, and error handling for each. *(Published 2012-05.)* | | **"Big security challenges with creating an API" (Blog – Jun 2012)** | Reflects on pitfalls in designing secure APIs. Topics include: managing authentication/authorization complexity, avoiding data over-exposure, versioning and backward compatibility introducing security debt, and the difficulty of providing flexibility to developers without opening attack surface. Uses examples (likely from building TeamMentor's API) to illustrate how design decisions can inadvertently create vulns. *(Published 2012-06.)* | | **TeamMentor Web Services Testing Series (A. Doraiswamy & Cruz – Apr 2012)** | A four-part blog series (April 28, 2012) by Arvind Doraiswamy with Dinis Cruz's help, detailing the process of testing **TeamMentor**'s SOAP/REST web services. It covers setting up a Python test harness (using *Suds* for SOAP), writing scripts for each API method, and logging/visualizing results. The series uncovered auth issues and logic flaws, and produced an **authorization mapping spreadsheet** for TeamMentor's API. Key lesson: a systematic, script-driven approach can reveal subtle authorization gaps and drive the creation of proper access-control matrices. | | **"Creating an API to create WebServices (on the fly)" (Blog – May 2013)** | Describes an innovative **meta-API** that can dynamically generate new web service endpoints. Cruz explored how an application could accept high-level definitions and spawn actual web APIs at runtime. The experiment (circa May 2013) demonstrated both the power and security implications of on-the-fly API generation. *Key ideas:* automation in API creation could help testing (by generating test endpoints or mocks), but it must be carefully secured to avoid abuse (since an API that creates APIs could be misused if not locked down). | | **"O2 Platform's MethodStreams (2010): Open Source SAST Engine" (Tech Blog – Feb 11, 2025)** | A retrospective on the **MethodStreams** feature of OWASP O2 Platform (originally developed 2010). MethodStreams automatically **traverses a web service's call tree** to aggregate all its relevant code into a single view. By extracting the AST of each method in the execution path, it produces a consolidated file with only the code that implements that API endpoint. This greatly streamlines code review and security analysis for large apps: e.g., a web method that spanned 20+ files and 50k lines could be reduced to a 3k-line snippet of just what matters. The document (2025) highlights how MethodStreams exemplifies using static analysis to aid **targeted API review**, and it encourages modern tools to adopt similar programmatic AST manipulation for security. | | **"Semantic Knowledge Graphs for LLM-driven Source Code Analysis" (Article – May 29, 2025)** | Explores the use of **semantic knowledge graphs** in combination with Large Language Models to analyze source code (including APIs). Cruz notes that capturing code facts (functions, data flows, auth checks, etc.) in a graph can help an LLM reason about security properties of an API. This approach could automate detection of missing authorization or data exposure by having the LLM traverse a knowledge graph of the application's API structure. *(Published 2025-05-29; indicates the forward-looking integration of AI/LLM into API security testing.)* | | **"AST (Abstract Syntax Tree)" (Medium blog – Jul 2017)** | Discusses the power of ASTs for code understanding and security. Shares examples where AST parsing enabled writing tests to ensure every exposed web service method calls an authorization routine. Introduces **code transformations** like MethodStreams and automated refactoring to fix vulnerabilities. Emphasizes that developers can leverage AST tooling to enforce security requirements (e.g., "no API method without input validation" can be checked via AST-based unit tests). This ties back to his earlier work by showing how *automation and code parsing* can systematically improve API security. | | **SecDevOps Risk Workflow (Leanpub e-book – v0.1 draft, 2017)** | An early-stage book where Cruz frames how to embed security into DevOps pipelines. While not solely about APIs, it reinforces relevant ideas: *security unit tests as part of CI*, "security champions" in dev teams, and risk acceptance workflows. One notable point: encouraging high code coverage in tests so that one can see exactly what parts of an app (or API) are exercised by tests versus which are "dark" (unhit by any test). In the API context, this means ensuring no endpoint is left untested or with dead code that could become a vulnerability. Cruz also highlights making security decisions visible and automating as much as possible – principles that next-gen API security tools should incorporate. | *(Table: Dinis Cruz's publications on API security, with dates and key ideas.)* ## Evolution of Themes in Cruz's API Security Research ### Early Insights: "Every API is Vulnerable" (2010--2011) Cruz's early writings set the tone by asserting that **any API function can be a weakness** if it processes untrusted input or performs sensitive actions. In *"Every API is (at some level) vulnerable"* (Jan 2010), he notes that when doing code reviews, we often debate whether a given API is *inherently* vulnerable -- and his conclusion is usually "yes, if it does something powerful"[\[1\]](https://diniscruz.blogspot.com/2010/01/#:~:text=When%20doing%20source%20code%20analysis%2C,vulnerability%20pattern). For example, a data-layer method that executes an SQL query from a string is *"vulnerable by design to SQL injection"* -- the real question is whether an attacker can reach it[\[15\]](https://diniscruz.blogspot.com/2010/01/#:~:text=A%20good%20example%20is%20a,put%20a%20payload%20on%20it). This led to the practice of **mapping internal APIs and tracing call chains** to see if untrusted users can invoke them. Cruz recommended identifying layers of vulnerabilities and *"mapping them upwards"* through consuming functions until you either hit an entry point or determine the issue isn't exploitable[\[4\]](https://diniscruz.blogspot.com/2010/01/#:~:text=I%20think%20that%20one%20of,or%20not%29%20the%20problems). This approach laid a foundation for later ideas like MethodStreams (which automates gathering those call chains). In 2011, Cruz proposed the **ESTAPI (Enterprise Security Testing API)** concept. At a high level, ESTAPI was about providing a uniform interface or service in applications for security testing. While details were still exploratory, the goal was to make it easier to *invoke security checks or inject test scenarios* via an API. This reflects Cruz's philosophy of baking testing capabilities into systems, so that testers (or even malicious hackers) don't have to "go around" the application to test it -- instead, the system would openly present a testing surface. Although ESTAPI did not become a standard, the idea prefigured today's trends of self-test endpoints or integrated scanning hooks in web apps. Also around this time, Cruz was deeply involved in OWASP's O2 Platform tool. One notable blog from Nov 2011 describes using O2 to **patch a live application against an XSS vulnerability** by auto-encoding outputs. In *"Fixing/Encoding .NET code in real time"* (2011) he shows a proof-of-concept where, in a development environment, a dangerous call (`Response.Write` of unencoded data) is intercepted and fixed on the fly. He argues this kind of automation could make security **"invisible" to developers by happening automatically**[\[5\]](https://o2platform.wordpress.com/category/fixing-code/#:~:text=Fixing%2FEncoding%20,Write). The broader implication for API security is that frameworks and platforms should help developers by transparently handling common security tasks (like output encoding or input validation), thereby reducing the chance that an API is deployed with a trivial vulnerability. This early focus on automation and developer experience carries into his later SecDevOps work. ### 2012: Developing a Methodology for Testing Web Services In 2012, as web APIs and SOAP/REST services were becoming core to applications, Cruz's blog output on the topic peaked. He published a series of posts that effectively build a **methodology for API security testing**. A key principle was: *start by writing thorough tests*. In *"First, you create tests for WebServices"* (Apr 2012), he suggests that before trying to secure or break an API, one must have a baseline of **functional tests** that cover each endpoint and use-case. This ensures the tester understands expected behavior and can detect when a security test causes a deviation or failure. It's an approach analogous to test-driven development -- here used for security, to the point of recommending that if you're developing a web service, write its security tests alongside its features. Cruz's *"Journey into testing WebServices in TeamMentor"* and related 2012 posts put this into practice. TeamMentor was an application with its own web service API, and along with colleague **Arvind Doraiswamy**, Cruz helped create a Python-based test harness to exercise it. They leveraged SOAP clients (the **Suds** library in Python) to enumerate all web service methods, call them with various inputs, and check responses. Through this journey, they produced artifacts like a spreadsheet mapping each web service method to the user roles allowed to invoke it[\[6\]](https://security.stackexchange.com/questions/14598/is-there-a-spreadsheet-template-for-mapping-webservices-authorization-rules#:~:text=the%20O2%20Platform,think%20came%20out%20quite%20well) -- essentially an **authorization matrix**. This was in response to real issues: for example, they discovered cases where certain API calls lacked proper authorization checks, which such a matrix would quickly highlight. This work underscored the value of systematically *mapping "who can do what" in an API* and using that as both a guide for testing and a deliverable to the dev team. (In fact, Cruz posted on StackExchange about creating these authorization mapping spreadsheets[\[16\]](https://security.stackexchange.com/questions/14598/is-there-a-spreadsheet-template-for-mapping-webservices-authorization-rules#:~:text=I%27m%20looking%20for%20a%20spreadsheet%2Ftemplate,be%20mapped%2C%20visualized%20and%20analyzed) to encourage others to do the same for their APIs.) Another notable mantra from 2012 was encapsulated in the provocatively titled *"If you're not blowing up the database, you're not testing"*. Here, Cruz stresses the importance of **aggressive test scenarios**. In practice, when testing a web service, this means doing things like sending extremely large payloads, deliberately malformed data, or bizarre sequences of requests -- essentially trying to make the API or its backend components fail. If none of your test cases ever trigger a crash, heavy load, or an error in the logs (e.g., *blowing up the database* with a crazy query), then your testing may be too superficial. While this approach can cause downtime in a production environment (hence should be confined to dev/test setups), the philosophy is that only by pushing systems to their limits do you expose security-critical weaknesses (such as SQL injection flaws that could dump an entire database, or logic bugs that corrupt data). Cruz's point was that **robust APIs should be able to handle abusive or unexpected inputs gracefully** -- and if they can't, it's better for a tester to discover that than an attacker. By May 2012, in *"What is the formula for WebServices testing?"* and the *"Roadmap for testing WebServices"*, Cruz attempted to generalize these lessons. He broke down API security testing into repeatable steps -- from **reconnaissance** (understand the API surface: endpoints, parameters, technologies) to **test design** (cover authentication, authZ, input validation, business logic, error handling) and **execution** (automation tools, fuzzers, etc.). An interesting aspect he highlighted was combining **static analysis with dynamic testing**. For instance, static analysis might find a hidden API endpoint or developer-backdoor function not documented; testers could then include that in their dynamic test plan. Or static analysis could reveal that a web method calls a database query with no parameterization -- flagging it as a SQL injection candidate to verify via dynamic test. This blended approach was ahead of its time, foreshadowing the modern IAST (Interactive Application Security Testing) tools that instrument apps to get the best of both worlds. Cruz's writing from this period is essentially a call to be *thorough and systematic*: use every tool available (code review, unit tests, integration tests, fuzzing, runtime monitoring) to poke at your APIs from all angles. In *"Big security challenges with creating an API"* (Jun 2012), Cruz took a step back to the design perspective. Here he enumerated common hard problems when developing APIs securely. For example, **authentication & authorization**: providing a smooth developer experience (DX) often conflicts with enforcing strict security. If an API is too permissive by default, it's a risk; if it's too locked-down, developers might hack around it. **Excessive data exposure** was another challenge -- many APIs tend to return more data than necessary (e.g., returning full user records where only names are needed), which can accidentally leak sensitive info. Cruz also discussed **state management** and how stateful vs stateless designs carry different security considerations (like replay attacks or race conditions in state changes). The takeaway was that designing an API is not just an engineering problem but a security one: choices made early (URL structure, payload format, use of HTTP verbs, etc.) all can introduce or mitigate classes of vulnerabilities. This blog served as a checklist for API designers to think like an attacker *from the design phase onward*. ### Advanced Tooling: MethodStreams and Dynamic API Generation (2013) By 2013, Dinis Cruz's focus expanded to more advanced ideas, leveraging his tooling background. In *"Creating an API to create WebServices"* (May 2013), he explored a meta-level concept: What if an application exposed an API that let you spin up *new API endpoints* on demand? This sounds esoteric, but it had practical motivation. For testing purposes, such a capability could let one create test endpoints (perhaps to echo inputs or perform specific tasks) without redeploying the server -- useful for hooking into an application's internal functions dynamically. It also mirrored how cloud services and integration platforms were evolving -- programmatically defining endpoints and their logic. Cruz's experiment likely used scripting (maybe an O2 Platform script or reflection in .NET) to register new web service methods at runtime. The **security angle** is double-edged: on one hand, this could enable *on-the-fly security tests* or probes. On the other, if such a feature were present in production (even inadvertently, like a debug interface), an attacker could abuse it to create backdoor methods. Cruz acknowledged this risk -- it's a reminder that meta-programming features must be guarded. The broader point from this exploration was visionary: *APIs can be made extensible and dynamic, but doing so safely requires careful design*. This anticipated today's serverless and plugin-based architectures, where endpoints and functions come and go dynamically. Another significant development was the deeper discussion of **MethodStreams** in the context of the O2 Platform. Although MethodStreams had been built in 2010, Cruz documented it thoroughly in a February 2025 post, reflecting on its impact. The MethodStreams approach is worth reiterating: given a particular entry-point method (say a REST API handler or SOAP operation), O2's MethodStreams module will find *all* methods it calls, then all methods those call, and so on -- essentially the full call graph -- and then **merge all those code fragments into one linear stream** of code[\[8\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=Back%20in%202010%20I%20was,execution%20path%20of%20those%20methods)[\[17\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=1,all%20the%20objects%20from%204). The result is a single file or view showing exactly what that API does from start to finish, including data validations, business logic, and data access. In a large enterprise application, where understanding a single web request might otherwise demand opening dozens of files and classes, MethodStreams hugely simplifies the reviewer's task[\[10\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=during%20my%20security%20review). For security testing, this means you can more easily spot where a validation is missing or where an unsafe call is made, since the context is preserved in one place. Cruz pointed out that this capability *"made a massive difference when doing code reviews"*[\[18\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=tree%20of%20a%20particular%20method,difference%20when%20doing%20code%20reviews) -- it was like having an X-ray of the application's internals focused on one API endpoint. In today's terms, this aligns with **trace-based static analysis** or **code property graphs** that security researchers use; Cruz was implementing a form of that a decade earlier. His advocacy in 2025 is for toolmakers to revive and modernize this idea: use ASTs and parsing to help engineers and AI systems visually and programmatically understand the full scope of an API's behavior. It's an approach that can complement traditional testing by revealing hidden paths (e.g., a method might indirectly call a dangerous API deeper in the stack -- MethodStreams would surface that, guiding the tester where to focus). ### Recent Work: AI, Knowledge Graphs, and Integrating Security into DevOps (2017--2025) In the latter part of the decade, Cruz's research pivoted to how emerging technologies like machine learning and knowledge graphs could tackle longstanding security problems -- including API security. His 2017 Medium article on *"AST (Abstract Syntax Tree)"* not only recaps the MethodStreams and AST-driven testing concepts, but also demonstrates writing **automated tests from ASTs**. For example, he posits an organization might require that *"every exposed web service method must call an authorization check."* Enforcing this manually is hard; but Cruz shows you can parse the code, find all public web methods, and then programmatically verify that within their AST there is a call to the `Authorize()` function (for instance)[\[11\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=A%20good%20example%20is%20how,every%20exposed%20web%20services%20method%E2%80%9D)[\[12\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=But%2C%20when%20you%20have%20the,requirement%2C%20here%20is%20what%20happens). He even suggests generating documentation from these AST-based tests -- turning them into living standards. This approach empowers the next generation of **"secure by design"**: if developers can easily create such AST-driven unit tests, they get immediate feedback when they violate an architectural security rule (much like a unit test failing). It's an evolution of the "test-first" idea, augmented with modern parsing and automation. Moving into 2025, Cruz wrote about *"Semantic Knowledge Graphs for LLM-driven Source Code Analysis."* This cutting-edge piece suggests that to harness AI (Large Language Models) for code security, we should feed them richer, structured information -- namely **knowledge graphs** that encode facts about the code. In context of API security, imagine a graph where nodes are API endpoints, data models, permissions, internal functions, etc., with edges representing calls, data flows, or access requirements. An LLM with this graph could answer complex questions (e.g., "Which API endpoints can modify customer records and do they all require admin tokens?") without having to read raw code line by line. This resonates with Cruz's long-standing vision of *making security knowledge consumable by tools at the exact moment needed*[\[19\]](http://diniscruz.blogspot.com/2010/09/can-we-please-stop-saying-that-xss-is.html#:~:text=Packs%27%2C%20since%20from%20my%20point,one%20of%20theses%20Rule%20Packs)[\[20\]](http://diniscruz.blogspot.com/2010/09/can-we-please-stop-saying-that-xss-is.html#:~:text=Remember%20that%3A%20security%20knowledge%2C%20that,existent). Back in 2010, he lamented that security advice buried in documentation is "almost as good as non-existent" unless it's available in context[\[19\]](http://diniscruz.blogspot.com/2010/09/can-we-please-stop-saying-that-xss-is.html#:~:text=Packs%27%2C%20since%20from%20my%20point,one%20of%20theses%20Rule%20Packs) -- fast forward to 2025, and he's effectively proposing to codify that knowledge in graphs that AI agents can use during code analysis. It's a direct bridge from the static/dynamic analysis integration he championed into the AI era: the static facts (graph) combined with dynamic reasoning (LLM) might finally crack some hard problems of scale in API security analysis. Finally, Cruz's work on **SecDevOps** (Secure DevOps) provides a philosophical backdrop. He consistently argued that security testing (including for APIs) must be integrated into the development lifecycle -- not a one-off external audit. In his SecDevOps Risk Workflow book, he highlights practices like: treating security findings as normal bugs, using CI pipelines to run security tests, and giving developers self-service security tools. For API security, one practical implication is to include API security scans in CI (e.g., automated fuzzing or dependency checks every build). Another is to measure test coverage for API endpoints -- ensuring that any new API functionality is accompanied by corresponding security tests. Cruz mentions the idea of tracking how much of an application's code is exercised by tests, and warns against leaving large swaths untested especially in APIs exposed to the outside[\[21\]\[22\]](https://leanpub.com/secdevops/read#:~:text=It%20is%20very%20important%20especially,application%20that%20actually%20doesn%E2%80%99t%20exist). This quantification of coverage can drive risk discussions: a CISO could ask, "We have 100 APIs, but only 70% of them are covered by automated tests -- what about the remaining 30%?" Making such gaps visible is key to *managing API risk continuously*, rather than in the spurts of occasional pen-tests. In summary, across 15+ years, Dinis Cruz's contributions to API security revolve around **proactive testing, developer enablement, and smart automation**. From manual code review tips in 2010, to building custom tools in 2012, to leveraging ASTs and AI by 2025, his work consistently pushes towards making it easier to understand and secure the complex web of functionalities that modern APIs expose. ## Takeaways for Next-Generation API Security Solutions Cruz's research offers a wealth of guidance for those building the next generation of API security products -- including GenAI-driven platforms and traditional AppSec tools. Here are key takeaways and how they align with what CISOs, AppSec teams, and developers are seeking in 2025: - **1. Security Testing Must Be Automated and Integrated** -- One of Cruz's core messages is to embed security tests into the development process. Tools should make it trivial to create and run API security tests as part of CI/CD. For example, an ideal product might automatically generate baseline security test cases for each new API endpoint (much like Cruz manually wrote tests for each web service in 2012). This addresses CISO concerns by providing continuous assurance, and developers appreciate it when security checks run seamlessly in their pipeline rather than as last-minute firefights. - **2. *"Think how someone could abuse it"* -- threat modeling at design time** -- Cruz often starts with the mindset of an attacker or an abuse-case (famously noting: *"when you develop your awesome, flexible, dynamic thingamajiggy, think about how someone could abuse it"*[\[23\]](https://docs.huihoo.com/javaone/2014/CON2488-RESTing-on-Your-Laurels-Will-Get-You-Pwned.pdf#:~:text=Goals%20and%20Main%20Point%20u,could%20abuse%20what%20you%20developed)). Next-gen solutions should help teams model threats to their APIs early. This could mean providing templates for common API threat scenarios (like broken object access, injection, auth bypass) and integrating those into user stories or design reviews. A product that can take an API definition (e.g. OpenAPI spec) and automatically highlight risky elements (e.g. "This endpoint exposes user email addresses -- is that intended?") would embody this principle. CISOs value this because it reduces the chance of design-level flaws; developers value not having to manually brainstorm every possible abuse case thanks to assistive tools. - **3.** Every **API is vulnerable -- discover all and monitor them** -- An implication of "every API is vulnerable" is that organizations need visibility into all their APIs and an understanding of their exposure. Tools should emphasize API discovery (including shadow or undocumented APIs) and mapping of internal dependencies. Cruz's approach of mapping internal function call graphs and building authorization matrices can inspire features: for instance, an API security platform that automatically builds a **knowledge graph of the application's API endpoints, the data they touch, and the guards (auth, validation) around them**. This directly appeals to CISOs' needs for inventory and risk assessment, and helps AppSec teams pinpoint weak links. Modern "API security posture management" solutions echo this, and Cruz's work reinforces how crucial comprehensive mapping is. - **4. Combine Static and Dynamic Analysis for API code** -- As shown by MethodStreams and AST-based testing, static analysis can uncover deep issues and guide dynamic tests. Next-gen products should blend these techniques. For example, if static analysis finds that an API method does not validate an input, the tool could then prompt a dynamic fuzz test on that parameter to confirm exploitability. This two-pronged approach yields more reliable results (fewer false positives, more confirmed vulnerabilities) -- something CISOs want to see in reports. Developers benefit too, as they get concrete evidence of issues (a failing test or proof-of-concept) rather than theoretical warnings. GenAI can assist here by interpreting static analysis findings and suggesting or even executing relevant dynamic tests, much like an expert pen-tester would -- a concept very much in line with Cruz's vision of augmented analysis. - **5. Developer-Friendly Remediation and "Invisible" Security** -- A consistent theme in Cruz's writing is making life easier for developers to do the right thing. Whether it's auto-encoding outputs or providing ready-made test scripts, the idea is to reduce friction. Products targeting API security should therefore focus not just on finding problems, but also on *fixing* them or preventing them. This might include secure frameworks that automatically handle common vulnerabilities (for instance, a library that sanitizes inputs to DB queries to prevent SQL injection, much as Cruz's runtime fix did for XSS). Another example is providing **code-fix suggestions** when a vulnerability is found -- e.g., "This API endpoint is missing an authorization check; here is a patch to add one." Development teams are more likely to embrace security tools that save them time and integrate with their workflow (IDE plugins, pull request scanners, etc.). By striving for that "invisible/automatic" feel -- where security is embedded and doesn't slow down development -- vendors will meet the approval of both dev teams and security officers. - **6. Leverage AI and Knowledge Graphs for Scale** -- Modern applications might have hundreds of microservice APIs, which is challenging to analyze manually. Cruz's recent research into knowledge graphs and LLMs provides a blueprint for tackling this complexity. Next-gen solutions should utilize AI to correlate information across APIs and detect patterns humans might miss. For example, an AI-driven analysis might identify that *two different APIs together could be abused in a logic attack* (one API exposes a user's ID and another, separate API accepts that ID to delete an account -- together they create a vulnerability). A knowledge graph that links data flows and privileges can feed an AI to surface such issues. Importantly, AI should also help with **prioritization** -- something CISOs care deeply about. Among thousands of API endpoints, which pose the highest risk? Cruz's mindset of evidence-based risk (e.g., focusing on APIs that connect to critical data or have known vulnerability patterns) can be operationalized by AI analyzing the code and usage telemetry. The goal for new products should be to reduce noise and highlight *meaningful, context-rich security findings*, effectively applying the kind of deep insight that Cruz demonstrated manually, but at machine speed and cloud scale. - **7. Continuous Verification and Observability** -- Finally, Cruz's approach implies that security isn't a one-time event. Just as he advocates continuous testing and even runtime checks, new API security companies should think in terms of *ongoing verification*. This could mean runtime protection (RASP for APIs) that monitors for abnormal behavior and even blocks attacks. It also means providing dashboards where AppSec teams and product owners can see the "security health" of their APIs in real time -- test coverage, known vulnerabilities, recent attack attempts, etc. By continuously measuring and reporting (e.g., "95% of our APIs have up-to-date security tests and no critical vulns"), security becomes a quantifiable aspect of quality, which is exactly how Cruz frames it in SecDevOps discussions[\[24\]](https://leanpub.com/secdevops/read#:~:text=This%20is%20a%20book%20about,risks%20are%20accepted%20and%20understood)[\[25\]](https://leanpub.com/secdevops/read#:~:text=,DevOps%20workflows%2Ftools%20will%20be%20shown). Such capabilities align with what management (CISOs) want -- assurance over time -- and what developers need -- feedback loops to improve. In conclusion, Dinis Cruz's body of work on API security serves as a road map for building effective solutions in this space. The emphasis on **proactive testing, automation, developer enablement, and intelligent analysis** is now reflected in the market's expectations for API security tools. By learning from these contributions, next-generation products can better empower developers to create secure APIs from the start, give AppSec teams the deep visibility and automation they need, and ultimately provide CISOs the confidence that their organization's APIs are not the weakest link. The future of API security, as Cruz's research suggests, will belong to those who integrate security seamlessly into the DNA of software development and leverage smart technology to stay ahead of attackers. **Sources:** Dinis Cruz's personal blog posts and articles[\[1\]](https://diniscruz.blogspot.com/2010/01/#:~:text=When%20doing%20source%20code%20analysis%2C,vulnerability%20pattern)[\[7\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=One%20of%20the%20most%20powerful,added%20to%20the%20O2%20Platform)[\[23\]](https://docs.huihoo.com/javaone/2014/CON2488-RESTing-on-Your-Laurels-Will-Get-You-Pwned.pdf#:~:text=Goals%20and%20Main%20Point%20u,could%20abuse%20what%20you%20developed)[\[6\]](https://security.stackexchange.com/questions/14598/is-there-a-spreadsheet-template-for-mapping-webservices-authorization-rules#:~:text=the%20O2%20Platform,think%20came%20out%20quite%20well)[\[5\]](https://o2platform.wordpress.com/category/fixing-code/#:~:text=Fixing%2FEncoding%20,Write), as summarized and interpreted above. [\[1\]](https://diniscruz.blogspot.com/2010/01/#:~:text=When%20doing%20source%20code%20analysis%2C,vulnerability%20pattern) [\[2\]](https://diniscruz.blogspot.com/2010/01/#:~:text=From%20a%20secure%20design%20point,not%20possible%20to%20exploit%20them) [\[3\]](https://diniscruz.blogspot.com/2010/01/#:~:text=I%20quite%20like%20the%20picture,connect%20to%20the%20outside%20world) [\[4\]](https://diniscruz.blogspot.com/2010/01/#:~:text=I%20think%20that%20one%20of,or%20not%29%20the%20problems) [\[15\]](https://diniscruz.blogspot.com/2010/01/#:~:text=A%20good%20example%20is%20a,put%20a%20payload%20on%20it) Dinis Cruz Blog: January 2010 [\[5\]](https://o2platform.wordpress.com/category/fixing-code/#:~:text=Fixing%2FEncoding%20,Write) Fixing Code « OWASP O2 Platform Blog [\[6\]](https://security.stackexchange.com/questions/14598/is-there-a-spreadsheet-template-for-mapping-webservices-authorization-rules#:~:text=the%20O2%20Platform,think%20came%20out%20quite%20well) [\[16\]](https://security.stackexchange.com/questions/14598/is-there-a-spreadsheet-template-for-mapping-webservices-authorization-rules#:~:text=I%27m%20looking%20for%20a%20spreadsheet%2Ftemplate,be%20mapped%2C%20visualized%20and%20analyzed) web service - Is there a spreadsheet/template for Mapping WebServices Authorization Rules? - Information Security Stack Exchange [\[7\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=One%20of%20the%20most%20powerful,added%20to%20the%20O2%20Platform) [\[8\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=Back%20in%202010%20I%20was,execution%20path%20of%20those%20methods) [\[9\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=Since%20in%20the%20O2%20Platform,write%20a%20new%20module%20that) [\[10\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=during%20my%20security%20review) [\[11\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=A%20good%20example%20is%20how,every%20exposed%20web%20services%20method%E2%80%9D) [\[12\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=But%2C%20when%20you%20have%20the,requirement%2C%20here%20is%20what%20happens) [\[13\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=,this%20looks%20like%20in%20practice) [\[17\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=1,all%20the%20objects%20from%204) [\[18\]](https://medium.com/@dinis.cruz/ast-abstract-syntax-tree-538aa146c53b#:~:text=tree%20of%20a%20particular%20method,difference%20when%20doing%20code%20reviews) AST (Abstract Syntax Tree). AST (Abstract Syntax Tree) is a graph... \| by Dinis Cruz \| Medium [\[14\]](https://leanpub.com/secdevops/read#:~:text=Once%20you%20get%20high%20degree,and%20what%20doesn%E2%80%99t%20get%20covered) [\[21\]](https://leanpub.com/secdevops/read#:~:text=It%20is%20very%20important%20especially,application%20that%20actually%20doesn%E2%80%99t%20exist) [\[22\]](https://leanpub.com/secdevops/read#:~:text=It%20is%20very%20important%20especially,application%20that%20actually%20doesn%E2%80%99t%20exist) [\[24\]](https://leanpub.com/secdevops/read#:~:text=This%20is%20a%20book%20about,risks%20are%20accepted%20and%20understood) [\[25\]](https://leanpub.com/secdevops/read#:~:text=,DevOps%20workflows%2Ftools%20will%20be%20shown) Read SecDevOps Risk Workflow \| Leanpub [\[19\]](http://diniscruz.blogspot.com/2010/09/can-we-please-stop-saying-that-xss-is.html#:~:text=Packs%27%2C%20since%20from%20my%20point,one%20of%20theses%20Rule%20Packs) [\[20\]](http://diniscruz.blogspot.com/2010/09/can-we-please-stop-saying-that-xss-is.html#:~:text=Remember%20that%3A%20security%20knowledge%2C%20that,existent) Dinis Cruz Blog: Can we please stop saying that XSS is boring and easy to fix! [\[23\]](https://docs.huihoo.com/javaone/2014/CON2488-RESTing-on-Your-Laurels-Will-Get-You-Pwned.pdf#:~:text=Goals%20and%20Main%20Point%20u,could%20abuse%20what%20you%20developed) docs.huihoo.com --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/09/01/next-generation-api-security-platform-semantic-graphs-genai-testing-ephemeral-environments-2025.html (markdown twin: /2025/09/01/next-generation-api-security-platform-semantic-graphs-genai-testing-ephemeral-environments-2025.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/09/01/next-generation-api-security-platform-semantic-graphs-genai-testing-ephemeral-environments-2025.html* # Next-Generation API Security Platform: Semantic Graphs, GenAI Testing & Ephemeral Environments for 2025 *By Dinis Cruz and ChatGPT Deep Research · 2025-09-01* > Modern API security requires going beyond traditional scanning -- it must blend into the API testing lifecycle and leverage cutting-edge AI to map and probe complex behaviors. This white paper… --- [PDF](https://files.diniscruz.ai/github/pdf/2025/09/01/next-generation-api-security-platform-semantic-graphs-genai-testing-ephemeral-environments-2025.pdf) ## Introduction ## Executive Summary Modern API security requires going beyond traditional scanning -- it must **blend into the API testing lifecycle** and leverage cutting-edge AI to map and probe complex behaviors. This white paper proposes a next-generation API security solution (circa 2025) that harnesses **semantic knowledge graphs, GenAI/LLMs, and hybrid testing workflows** to secure APIs as part of overall quality. Key innovations include: - **Comprehensive API Mapping** -- Every endpoint, dependency, and call flow is modeled in a **semantic knowledge graph** using an open-source engine (MGraph-DB). This graph tracks what data each API touches and what guards (auth, validation) protect it\[1\], giving security teams and architects a **holistic view of the attack surface**. - **GenAI-Assisted Analysis & Simulation** -- Large Language Models (LLMs) act as intelligent assistants to understand and test the API. They parse code and logs to identify vulnerabilities (e.g. missing auth checks) by reasoning over the graph\[2\]. They also simulate complex abuse scenarios -- stitching together multi-step attacks that combine endpoints and business logic -- which helps uncover edge-case flaws humans often miss\[3\]. - **Hybrid White-Box/Black-Box Testing** -- The solution integrates **static code analysis and dynamic testing** in a seamless workflow. For instance, if static analysis finds an input not validated, it automatically launches a fuzz test on that parameter\[4\]. This two-pronged approach produces **concrete evidence** of vulnerabilities (reducing false positives) and even uses AI to suggest or execute those tests like an expert pentester\[4\]. - **Ephemeral, Isolated Test Environments** -- Built on a serverless, stateless architecture, each security test spins up **on-demand with no persistent infrastructure**[\[5\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L28). The system uses **Memory_FS** (an in-memory filesystem) to package application code and data in a sandbox, and can generate **surrogate dependencies** (simulated backend services or data) to safely exercise the API in isolation. This enables aggressive testing (including destructive or offline scenarios) without impacting real systems, ensuring APIs handle failure and abuse gracefully\[6\]. - **Developer-Centric and "Invisible" Security** -- The platform is designed to integrate with developer workflows and shorten the feedback loop. It encourages a test-driven security mindset where **abuse cases are tested early** (even during design)\[7\]. Developers get results as failing test cases, detailed call traces, and even **code-fix suggestions** (e.g. "add an auth check here")\[8\]. Many common issues are auto-mitigated by the framework (e.g. output encoding), making security as transparent as possible during development\[9\]. This developer-first approach ensures security is a natural extension of building and testing APIs, not a roadblock. - **Multi-Stakeholder Views & Reporting** -- Security findings are translated into **persona-specific insights**. The same knowledge graph can generate a technical report for engineers, compliance checklists for auditors, and executive summaries for leadership. Inspired by Dinis Cruz's **MyFeeds.ai** project, the system uses AI to turn raw metadata into narrative "stories" tailored to each audience[\[10\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/06/08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md#L38-L41). For example, a CISO or board member might see a GenAI-generated briefing on API risks in business terms -- an idea influenced by Cruz's work on The Cyber Boardroom, which uses GenAI to assist board-level decisions[\[11\]](https://www.cyber-leaderssummit.com/speakers/dinis-cruz#:~:text=Dinis%20Cruz%20is%20the%20Chief,technical%20expertise%20with%20commercial%20leadership). This ensures security data is actionable at every level, from dev team to boardroom. - **Open, Extensible & Community-Driven** -- The solution is open source at its core and **highly extensible**. Key components like the graph database (**MGraph-DB**) and AI logic are open, with standard data formats (JSON, Markdown) to avoid lock-in[\[12\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L27). A credit-based SaaS offering can provide on-demand scale (LLM calls, cloud resources) while allowing on-premise deployment for sensitive environments. The community is invited to contribute custom test modules and share findings -- in fact, newly discovered vulnerabilities can be published (responsibly) as proof of the system's efficacy, driving adoption and improvement. The business model centers on support, expertise, and usage-based credits rather than proprietary licensing, aligning incentives with user success. Overall, this paper presents a vision for API security as an **integrated, AI-powered workflow** -- one that maps and understands APIs deeply, anticipates abuse cases, and empowers both developers and security teams to build resilient, secure services from design to deployment. ## Introduction APIs are the backbone of modern applications, from microservice architectures to serverless backends. As we enter 2025, organizations face an explosion of API endpoints and interconnections -- and with that comes an expanded attack surface. High-profile breaches and logic flaws in APIs have shown that *even well-designed APIs can harbor hidden vulnerabilities*. Dinis Cruz famously asserted over a decade ago that **"Every API is (at some level) vulnerable"**\[13\] -- not because developers are careless, but because any sufficiently powerful function can be abused if accessed improperly. The challenge is two-fold: **knowing where those vulnerable functions are** and **ensuring attackers can't reach them**\[14\]. Traditional approaches struggle to keep up. Manual code reviews and point-in-time pen-tests catch some issues but don't scale to hundreds of microservices. Generic scanners often miss business logic flaws and produce noisy false positives that frustrate developers. Meanwhile, agile development and continuous deployment mean APIs change rapidly -- security must be continuous and automated to be effective. This paper argues that **API security needs to evolve into an AI-assisted, developer-centric discipline**. It should be treated as a natural subset of API quality assurance -- alongside functionality, performance, UX, and resilience -- but with a special focus on the *abuse cases* that others ignore. In practice, that means expanding the scope of testing to include malicious and extreme scenarios. Security testing isn't just about checking a few OWASP Top 10 issues; it's about **thinking like an attacker** and pushing the API to its limits. As Cruz quipped, *"when you develop your awesome, flexible, dynamic thingamajiggy, think about how someone could abuse it"*\[7\]. Our solution operationalizes this mindset by automatically scrutinizing APIs for how they could fail or be exploited under adverse conditions. At the same time, next-gen API security must mesh with modern infrastructure. This means being cloud-native (leveraging ephemeral, serverless execution), handling the complexity of microservices and third-party APIs, and using advanced tooling like knowledge graphs and machine learning to augment human analysis. The goal is ambitious: **to map, test, and secure every interaction in and around an API, without drowning teams in noise or slowing down development**. The following sections detail how such a system works, drawing on Dinis Cruz's extensive research and tooling in this space -- from *MethodStreams* for static code analysis to *MGraph-DB* for semantic graphs, and from *Memory_FS* for data storage to *MyFeeds.ai* for AI-driven content generation. By building on these ideas, we outline an architecture that is technically unique, highly adaptive, and poised to redefine API security testing in 2025. ## API Security as Part of the Quality Lifecycle A core principle of this approach is that **API security is an extension of API quality**. It sits at the intersection of functional testing, reliability engineering, and user experience -- ensuring not just that an API works, but that it *fails safely* and *cannot be misused*. In practical terms, this means every practice applied to make APIs reliable and user-friendly should have a security analog. For example, if developers write unit tests for expected inputs, we add tests for unexpected or hostile inputs; if they monitor uptime, we also monitor unusual patterns that might indicate abuse. Security doesn't live in a silo -- it piggybacks on the same DevOps pipelines and feedback loops used for feature development. However, security's unique contribution is to cover the **"abuse cases"** -- scenarios where a user does something *not by design*. Normal QA might not try sending a 10MB payload where 1KB is expected, or attempt to perform actions out of order, but a malicious actor will. Our system ensures those cases are tested. As Cruz advocated, *"if you're not blowing up the database, you're not testing"*\[6\] -- meaning that effective testing should include attempts to break things and observe how the system copes. By including such adversarial tests (in isolated environments), we uncover security-critical weaknesses in error handling, data validation, access control, and more. For instance, does the API gracefully handle a flood of requests or weird characters in input, or does it crash and expose debug info? Does it enforce business rules consistently, or can a sequence of calls bypass them? These are quality issues *and* security issues. To bake security into the lifecycle, the solution provides tooling at multiple stages: - **Design & Threat Modeling:** Even before code is written, the system can analyze API designs (like OpenAPI/Swagger definitions) and highlight risky elements or assumptions. It brings templates of common API threat scenarios (e.g. broken object level authorization, mass assignment, injection points) into the design discussion. This is essentially *shifting left* on threat modeling -- providing an automated "attacker's perspective" early on. As an example, if an OpenAPI spec shows an endpoint that returns all user details, the tool might flag: "This endpoint exposes email addresses -- is that intended?"\[15\]. This guides architects to consider abuse cases from the start. - **Development & CI:** As code is implemented, developers can write security unit tests in parallel with features (a practice Cruz championed back in 2012)\[16\]. Our platform assists by generating baseline security tests for each endpoint -- covering authentication, authorization, input validation, and expected error handling. These tests run in CI just like functional tests, failing the build if a regression opens a hole. The key is making this **frictionless**: the tests are auto-generated or easy to write, and running them is as fast as any unit test thanks to our lightweight, stateless test harness. - **Continuous Scanning & Runtime:** Once the API is live (in staging or production), the system continues to monitor and probe it in safe ways. It might instrument the application to watch for dangerous patterns (like SQL queries constructed from user input, akin to runtime taint tracking) or use passive monitoring of logs for suspicious activity. If the app has observability hooks, those feed into our knowledge graph as real-time validation of the assumptions we tested. In production, the focus is on **observability and quick feedback** rather than brute-force testing -- for example, detecting an abnormal sequence of API calls that could indicate a logic abuse attempt, and alerting on it (or even blocking it if a preventable exploit is detected). This continuous verification aligns with Cruz's view that security isn't one-off -- you need ongoing "security health" metrics and runtime defenses\[17\]. By embedding into each phase, API security moves from a checkbox at release time to a constant quality attribute. Developers and testers start to see security as just another aspect of making a "good API". Importantly, many security features double as user-experience improvements -- for instance, strict input validation not only stops attacks but also catches user mistakes, and rate limiting not only thwarts abuse but also ensures fair usage. By framing security in terms of **resilience and quality**, we get buy-in from engineering teams and avoid the old "us vs. them" dynamic. Our system reinforces this by providing value to developers (like catching bugs early, generating tests, and even fixing code) rather than simply reporting problems. ## Mapping the API Attack Surface with Knowledge Graphs At the heart of the solution is an **API Knowledge Graph** -- a living map of all the pieces that make up the API's surface area. This goes far beyond a list of endpoints. The knowledge graph encodes: **endpoints** (HTTP routes, parameters, etc.), **internal functions** and microservice calls (the implementation behind each endpoint), **data models** (what data is read or written), and **security controls** (auth checks, validation logic, permission rules). By connecting all these nodes, we gain a powerful holistic view: for any given endpoint, we can see its entire execution flow, what it impacts, and where its weak points might be. To construct this graph, we leverage both static analysis and dynamic inspection: - **Static Code Traversal (MethodStreams):** Borrowing from Dinis Cruz's OWASP O2 Platform, we use a technique similar to *MethodStreams* to automatically traverse each API endpoint's call tree\[18\]. Essentially, given an entry point (say a controller method in a web service), the system retrieves its implementation and then recursively pulls in any function it calls, and so on. The result is a consolidated view of the endpoint's logic -- often spanning dozens of functions across multiple files -- distilled into one structure. In the O2 MethodStreams example, a web method that originally spanned 50k lines across 20 files was reduced to a 3k-line snippet of just the relevant code\[18\]. We achieve a similar result but instead of a flat snippet, we feed this into the knowledge graph: each function becomes a node, and "calls" relationships link the flow. Data definitions (like database tables or objects) are also nodes, linked to the functions that use them. This static map reveals critical info, such as "this endpoint ultimately executes SQL queries via Function X" or "these 3 modules process input from that HTTP parameter". It's a *white-box blueprint* of the API's internals. - **Dynamic Endpoint Discovery:** In parallel, we enumerate the API from the outside (black-box style). For REST/HTTP, this could mean parsing OpenAPI specs or documentation, and also actively crawling the API (logging in with test accounts and exercising known endpoints to see responses). For GraphQL or gRPC APIs, we introspect their schemas. The goal is to not miss any exposed entry point -- including "hidden" ones that might not be documented. This discovery also populates the graph: we add nodes for each **entry point** and link them to the internal code graph from the static analysis. If an endpoint isn't reachable due to conditions (feature flags, user roles), we note those prerequisites as properties in the graph. - **Infrastructure & Context Mapping:** Beyond code, a deployment context matters for security. So our graph incorporates metadata about where each API runs (which cloud, URLs, ports), its dependencies (datastores, external APIs it calls), and trust boundaries. For example, a connection from the API to an external payment service is modeled, as is the fact that the payment service is outside our security boundary. We link config files or IaC (Infrastructure-as-Code) data that define these connections. The result is that our knowledge graph can answer questions like: *"Show all external calls made by endpoint /purchase -- are they secure?"* or *"Which internal APIs depend on the Customer database?"* This comprehensive mapping is crucial because modern attacks often chain multiple systems. As Cruz pointed out, organizations need **visibility into all their APIs and how they connect** -- including internal and "shadow" APIs -- to truly assess exposure\[1\]. A graph approach naturally excels at mapping these connections. We chose a graph model because of its flexibility and power in capturing relationships. Each node (endpoint, function, data object, etc.) can have properties (e.g. "requires auth token", "sensitive personal data") and edges can denote various relationships (calls, reads, is guarded by, etc.). This format is ideal for reasoning about security: one can traverse the graph to find, say, all paths from a public endpoint to a sensitive data store, or detect if any path lacks an authorization check node. In fact, by encoding guard conditions as nodes (like an "AuthZ check" function), we can ask the graph questions like: *"Does every path from public entry points to admin-only functions include an AuthZ check?"* If the answer is no, we've found a potential access control gap. Technologically, the platform uses **MGraph-DB**, an open-source, in-memory graph database engine, to manage this knowledge graph. MGraph-DB was designed by Cruz specifically for GenAI and ephemeral scenarios -- it operates as a **"memory-first" graph DB** that can serialize to JSON for persistence[\[19\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L7-L10). This means we can load and query the graph **blazingly fast in-memory**, but still export it to a file (or cloud storage) whenever we want to save state, without needing a separate database service. In our serverless architecture, this is key: an analysis function can spin up, instantiate the graph from JSON in object form, do all the traversals and queries needed, then write back any updates and terminate. No always-on database required. The JSON-based storage also makes the graph data easily sharable and inspectable -- security teams can even load it into Neo4j or visualization tools if they like. The use of a semantic graph and MGraph-DB is inspired by Cruz's broader research into using graphs for security context; it provides a foundation where **all security-relevant knowledge about the API is interconnected and queryable**\[20\]. Finally, once the initial graph is built, it doesn't stay static. Each subsequent phase of testing and analysis (described in later sections) will **enrich the graph**. When we run dynamic tests, the results (e.g. "received HTTP 500 error with stack trace from endpoint X") are added as new nodes or annotations. When the LLM identifies a potential logical link ("Data from endpoint A can feed into endpoint B"), that too is recorded. Over time, the knowledge graph becomes a living security model of the application -- one that the AI and humans can both learn from. It's the single source of truth that drives the rest of the system's intelligence. ## GenAI-Augmented API Understanding and Simulation A distinguishing feature of our solution is the deep integration of **Generative AI (LLMs)** in analyzing and testing the API. Rather than treating AI as a black box that magically finds issues, we use it in a controlled, assistive manner at specific steps in the workflow -- essentially as an amplifier for human insight and an automation of tedious reasoning tasks. The LLMs help in two broad areas: **understanding the system's behavior** (code, logs, data flows) and **simulating intelligent attacks or misuse**. **1. Reasoning Over Code and Graphs:** By converting code and runtime information into a semantic graph, we make it far more digestible for an AI. Instead of prompting an LLM with thousands of lines of source code (which could be hit-or-miss), we prompt it with a structured summary: nodes and relationships that represent the program's logic. For example, we might serialize a portion of the knowledge graph and ask the LLM a question like: *"Identify any API endpoints where user input flows into a database query without sanitation."* The knowledge graph provides the facts (function A takes parameter X, passes to function B, which builds an SQL string), and the LLM can interpret that pattern as a SQL injection risk even if the code uses custom patterns that rule-based scanners might miss. Cruz's research highlights this approach -- capturing code facts in a graph and then using LLMs to infer security properties (like missing authorization checks or data exposures)\[2\]. Essentially, the LLM traverses and interprets the graph as a seasoned code reviewer would, looking for lapses in the expected security controls. We also apply LLMs to make sense of **unstructured output** during testing. APIs often produce logs, error messages, or stack traces that are meant for developers, not testers. Our system takes those raw texts and asks the LLM to extract security-relevant info. For instance, if a fuzz test yields an error log, we prompt: *"Parse this stack trace to identify the module and line of code where the error occurred, and the type of error."* With structured output enforcement, the LLM returns JSON like `{ "exception": "NullPointerException", "module": "PaymentService", "line": 212, "message": "Account ID is null" }`. This is immediately added to the graph, linking that error to the `PaymentService` function node. In another case, if we see an authentication token in a log, we can have the LLM check whether it's properly formatted or if sensitive data (like a password) accidentally appears in logs. These tasks normally require a human to manually read through logs and connect dots; the LLM speeds that up dramatically. **2. Simulating Abuse and Logic Attacks:** One of the most powerful uses of GenAI here is **planning and simulating multi-step attacks**. Traditional testing tools struggle with business logic vulnerabilities -- issues that arise not from a single request, but from a *sequence* of interactions or a clever misuse of a feature. This is where an LLM, with its ability to "think" through scenarios, shines. We effectively give the LLM a sandbox in which to play the role of an attacker: it has access to the API documentation (or a description of the API capabilities from our graph) and it can propose strategies to achieve a malicious goal. For example, we might prompt: *"Given this e-commerce API, how might a low-privilege user obtain another user's order history?"* The LLM might generate an attack chain: first call endpoint A to enumerate user IDs, then call endpoint B with another user's ID to fetch their data (assuming no proper check). If that chain makes sense, our system can then execute it step by step in the test environment to see if it actually succeeds. This approach turns the AI into a sort of autonomous penetration tester, exploring the API's functionality with malicious intent. Importantly, every LLM-generated hypothesis is **validated** via testing -- we don't trust it until we see it working. But even a hypothesis that fails can be insightful: maybe the attempt was blocked by a validation, which is good to know (it reinforces that control's presence in our graph). We also record these AI-suggested scenarios as test cases, contributing to a growing library of logic tests for the API. A concrete scenario Cruz described is where two different APIs, when used together, create a security issue that neither has alone\[3\]. For instance, API #1 exposes a user's internal ID (perhaps via an account export feature), and API #2 allows deletion of accounts by ID. Individually, each might require auth and seem fine, but an attacker could use API #1 to fetch someone else's ID (if not properly access-controlled) and then feed it to API #2 to delete that other account. These are subtle issues spanning multiple endpoints and even multiple microservices. Our LLM is adept at spotting such possibilities by cross-referencing what each part of the graph does. It can notice: "Endpoint `/exportData` returns userId; Endpoint `/deleteAccount` takes userId and doesn't double-check ownership -- could these be miscombined?" This pattern recognition across the graph, combined with a bit of creative evil thinking, is something human threat modelers do in brainstorming sessions. Here we have the AI do it continuously, at scale. **3. Taint Tracking and Data Flow Analysis:** While not an LLM task per se, we incorporate classic **taint analysis** into the pipeline and then use AI to interpret the results. Taint analysis marks user-controlled data and follows it through the code to see if it reaches sensitive sinks (like file writes, exec calls, or database queries) without proper handling. Our static analysis engine performs this on the code graph. The outcome is essentially a subgraph highlighting "taint paths". Now, an LLM can take those paths and reason about their exploitability. For example, it might see a path: `QueryParam "q" -> searchBooks(query) -> string concatenation -> SQL execute`. The static tool flags it, and the LLM confirms, "This looks like SQL injection if an attacker crafts the `q` parameter." It may even suggest a payload (e.g. `q = "' OR 1=1--"`) to test it. Similarly, for XSS, if tainted data flows to an HTTP response, the AI can infer whether output encoding was present or not and come up with a proof-of-concept script injection. This human-like judgement on static analysis findings helps filter out which ones are serious versus which ones might be false positives or mitigated by some framework (which the static tool might not fully understand). Essentially, **AI provides context and intent to raw dataflow results**, focusing our dynamic testing on the most dangerous flows. **4. Ensuring Determinism and Safety in AI Usage:** We are careful to avoid the "black box AI" pitfall. All LLM interactions are sandboxed and constrained to produce **deterministic, repeatable outputs**[\[21\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L8-L15). We achieve this by using structured prompts and expecting answers in a strict format (often JSON or a specific list of steps). For example, when asking the LLM to propose an attack chain, we might say: "List the sequence of API calls and payloads as a JSON array." If the LLM's output doesn't parse or doesn't meet criteria, we know to discard or retry. This way, anyone can replay the exact same prompt on the same model and get the same result, which is crucial for trust. We also log every AI decision and tie it back into the knowledge graph as evidence (for transparency). If the AI suggests a vulnerability, the evidence might be "LLM reasoning chain X" plus the actual test that confirmed it. This level of provenance is important to convince stakeholders that we're not reporting hallucinated issues -- **every finding is backed by either code analysis or a real exploit attempt** (usually both). By weaving AI into our analysis, we get the best of both worlds: the thoroughness of automation and the cleverness of human-like insight. GenAI becomes a force multiplier for the security team, handling tasks like reading voluminous code, correlating distant pieces of info, and exploring creative attack strategies in minutes rather than days. It's like having a tireless junior security researcher who combs through everything and occasionally surfaces, "Hey, have you considered this weird edge case?" -- except it can do this across an entire fleet of APIs simultaneously. ## Hybrid Testing: Marrying White-Box and Black-Box Techniques Identifying potential issues via graphs and AI reasoning is only half the story -- we need to **validate and exploit** those issues to prove they're real and impactful. This is where our platform's **hybrid testing workflow** comes in, seamlessly blending static (white-box) analysis with dynamic (black-box) testing. The tight integration of the two, mediated by AI, is a key differentiator of our approach. In traditional security tools, static analysis (SAST) and dynamic analysis (DAST) are separate silos. SAST might warn "Function X might be vulnerable to SQL injection" while DAST might blindly fuzz parameters hoping to trigger an error. Our solution treats them as complementary phases of one process, using each's strengths to cover the other's weaknesses: - **Static to Dynamic Handoff:** When our static analysis (augmented by AI understanding) flags a potential problem, the system automatically generates a targeted dynamic test case for it. For instance, if the analysis of code/graph suggests that endpoint `/search` is vulnerable to SQL injection via parameter `q`, the platform constructs a request to `/search?q=' OR '1'='1` (or some similar payload). It knows what payload to try either from built-in libraries or from the LLM's suggestion (the LLM might have suggested a classic `' OR 1=1 --`). This request is executed against the test environment and the response is observed. If we get an error or strange behavior (like a dump of data or a performance slowdown), the system notes that the issue is confirmed. This removes ambiguity -- static analysis may have false positives, but a successful exploit attempt is undeniable. As Cruz noted, combining static and dynamic in this way yields **more reliable results -- fewer false alarms and more concrete proofs**\[22\]. Security teams love this because it means reports come with evidence ("here's the data I extracted using the flaw"), and developers appreciate that it's not just theoretical ("we've shown this crash can actually happen"). - **Dynamic to Static Feedback:** Conversely, when we do black-box testing and find something odd (say a 500 Internal Server Error), we feed that back into the static analysis phase. The knowledge graph is updated: e.g., an edge is added from endpoint node `/uploadFile` to an "Exception" node indicating a null pointer dereference was observed. Now, with that clue, we can dive into the code (perhaps via MethodStreams) specifically around `/uploadFile` and trace why that null pointer might occur -- which could reveal a missing check or an assumption. The static analysis can then generalize: "If this happened for this input, could it happen for other inputs or other APIs that use the same library?" Essentially, dynamic findings help focus and prioritize static exploration. - **Fuzzing and Behavioral Testing:** Beyond specific exploit payloads, our dynamic testing includes **intelligent fuzzing**. We generate a variety of inputs (random strings, very large numbers, special characters, JSON structures etc.) for each endpoint to see how it handles them. But unlike naive fuzzing, our fuzz is guided by the graph and AI context. For example, if an endpoint expects a date, we'll try some boundary dates and invalid formats. If the AI knows a field is likely an ID, it might try negative numbers, or numbers that look like SQL meta-characters. This **guided fuzzing** finds issues like improper input validation, denial-of-service vectors, or weird logic bugs. And when something is found (like an error), it links back to our graph so we know exactly which component choked on which input. - **Environment Control for Testing:** We also instrument the test environment to observe internal behavior during dynamic tests (when possible). If we can, we attach monitors or use debug modes that log SQL queries or internal exceptions. This gives deeper insight into what happened on the server side for each test. For instance, if a fuzz payload triggers a slow response, we might see in logs that a full table scan was done -- indicating a performance issue that could be exploited (and is also a security concern if it can be triggered repeatedly to cause denial of service). All these internal observations are again fed into the knowledge graph. The **end result** of the hybrid workflow is a virtuous cycle: static analysis finds a potential issue -\> dynamic test confirms and further illuminates it -\> both results enrich the knowledge graph -\> which then helps find more issues. By cycling like this, we gradually build a very complete picture of API weaknesses. This approach was foreshadowed by Cruz's early vision of an *Enterprise Security Testing API (ESTAPI)* -- where an application might even expose testing hooks for authorized tools\[23\]. While we don't require apps to implement ESTAPI, we carry the spirit of it: we treat the app as both **subject and partner** in testing. We're not purely attacking from outside (like a black box hacker) nor just inspecting from inside (like a static auditor); we're doing both cooperatively. Notably, **GenAI is an active participant in the hybrid workflow**. It helps translate static findings into test plans and conversely helps explain dynamic anomalies with static context. One can imagine it as the coordinator saying: "Static analysis says X might be an issue, let's try Y to confirm it. Oh, the app responded with Z, which means\... let me check the code around that -- yep, looks like a null pointer in module M." This is exactly what a skilled security engineer would do, but automated. As Cruz's research suggests, an AI assistant can act *"much like an expert pen-tester"* to interpret results and drive deeper testing\[24\]. The difference is we can run this at machine speed and across a huge number of endpoints concurrently. ## Ephemeral Test Environments and Surrogate Dependencies Safety and realism in testing are often at odds: you want to test aggressively (even destructively), but you don't want to bring down real systems or violate data integrity. Our solution resolves this by using **ephemeral, isolated test environments** that can faithfully mimic production without the same risks. This environment strategy has several facets: **1. Memory_FS and Ephemeral State:** We use **Memory_FS**, a type-safe in-memory filesystem framework, to set up each test run's environment. Memory_FS provides a unified interface to manage files and data in-memory, or swap in different storage backends if needed[\[25\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md#L22-L25). When we prepare to test an API, we essentially create an *instance* of the application in Memory_FS: this could include the API's code, configuration files, and even a subset of its database (if available). Because Memory_FS supports pluggable backends (in-memory for speed, or persisted like SQLite/ZIP for portability)[\[26\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md#L16-L24), we can easily snapshot the entire state of an application into a single archive. For example, we might take a nightly dump of the production database schema (sans real sensitive data), and a copy of the latest code -- package that into a Memory_FS instance, and that becomes the sandbox for testing. All writes the tests make (like adding records, or uploading files) stay within this memory image. This means we can run even destructive operations -- like the "blow up the database" tests -- freely, then simply discard the memory image when done. Using Memory_FS in this way gives us **fast spin-up and tear-down** of test environments. It's akin to containerizing the app, but even lighter weight since it's just file system virtualization with strong typing and versioning. Our pipeline might start a test by launching a function (or container) that loads the Memory_FS archive of the app, runs the app (perhaps as an in-memory server), executes tests, then destroys it. This aligns perfectly with a serverless philosophy: *compute and state exist only for the duration of the test*, then everything is saved (as artifacts) or terminated[\[5\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L28). No long-lived test servers are needed, and parallel tests don't conflict since each has its own instance in memory. **2. Surrogate Dependencies:** Many APIs depend on external services (payments, OAuth providers, third-party data APIs). Testing against live dependencies is risky (you might trigger real transactions) and unreliable (network calls, rate limits). We introduce **surrogate dependencies** -- essentially, stand-in services or data that mimic the real ones for testing purposes. For any given integration, the system can either **record** real interactions and replay them, or use AI to **simulate** the dependency's behavior. For example, if the API calls an address verification service, our surrogate might simply return a canned "success" response for any reasonable input during tests. We integrate these surrogates by intercepting outbound requests from the app (via configuration or a proxy) and routing them to our dummy implementations. This allows the API to think it's talking to Service X, when it's actually hitting a local function that returns deterministic responses. Surrogate dependencies enable *black-box + white-box hybrid* on an ecosystem level: we are effectively performing a controlled black-box test against external interfaces by using a white-box stub. They also allow testing of failure modes -- we can program a surrogate to time out or return errors to see how the API handles it, something we wouldn't dare do with a live payment API. In our knowledge graph, we mark these surrogate interactions clearly, so we remember what was real and what was simulated. The Memory_FS three-file pattern (content, config, metadata) is handy here: we store surrogate data (content), along with config (how to route calls to it) and metadata (original source info), all versioned. This approach was inspired by Cruz's *offline-first development* ideas, where having local backends empowers developers to work and test without live connectivity issues. **3. Isolation and Parallelization:** Each test scenario runs in isolation, meaning one test cannot accidentally affect another. Because of the stateless, serverless design, we can run many tests in parallel across multiple ephemeral instances. This drastically cuts down total testing time -- important when you might have hundreds of endpoints each with dozens of test cases. If a particular test causes a crash in its environment, only that ephemeral instance dies; our orchestrator just notes the crash (as a finding) and moves on, while other tests continue unaffected. This isolation is a boon for reliability of the security test harness: unlike a traditional long-running test server that might become unstable after many attacks, here we "reboot" the world from scratch for each fresh test or batch of tests. **4. Realism through Data Seeding:** A security test is only as good as the scenarios it can exercise. We seed the ephemeral environments with **representative data** to make tests realistic. For example, we generate or import some user accounts with various roles, some sample records (orders, files, etc.) so that business logic can be meaningfully executed. In many cases, anonymized or fake data can mirror the edge cases in production (like one user has maximum privileges, another has none; one order has a weird status, etc.). The LLM can assist by generating data that hits edge conditions -- e.g., creating an account with an admin role but no associated profile, if we want to see how the API handles partially missing data. All this data lives only in memory during the test, but it ensures that when we attempt an unauthorized action or a complex sequence, the app has context to react realistically (for instance, denying access to another user's record rather than just "record not found"). By combining Memory_FS, surrogates, and this ephemeral model, our solution achieves **safe yet thorough testing**. We can execute tests that no one would dare run directly against production (dropping tables, spamming transactions, etc.), and we can do so repeatedly without cleanup hassles. The stateless design (with JSON/ZIP snapshots as the only persistent artifact) also means this system can be offered as a **managed cloud service without storing customer secrets**: a client could upload a sanitized snapshot of their app to our system, we test it in our ephemeral sandboxes, and return results, without ever needing access to the live environment. It's effectively *serverless security testing*, analogous to how ephemeral SIEM approaches avoid storing all the data centrally[\[5\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L28). Moreover, this architecture is **cost-efficient and scalable**. No idle resources are consuming budget -- we only incur compute costs when actually running tests, and storage costs for lightweight JSON/ZIP files of state. This ties into the credit-based model: users spend credits when they run analyses, reflecting actual resource usage. A stateless model makes those costs predictable and contained. ## Developer Collaboration and Feedback Loops A critical success factor for any security solution is acceptance and use by developers. We design our platform to foster **collaboration with developers, not confrontation**. The philosophy is to create tight feedback loops where the system and the developers learn from each other. Several features enable this: **1. Integrated into Dev Workflows:** We provide plugins and integrations for the tools developers already use -- IDEs, code review systems, CI pipelines, and ticketing systems. For example, as a developer writes code, an IDE plugin can show the API knowledge graph view for the function they're editing, highlighting where input validation is expected or where that code is called by an open endpoint. If they introduce a risky change (like removing an auth check), the plugin can warn them immediately by querying the local graph model. When a pull request is opened, our solution can automatically run the relevant security tests for the changed endpoints and comment on the PR with results ("The security regression test for endpoint /dataExport failed -- missing authorization now")\[8\]. This makes security feedback immediate and contextual, similar to how unit tests or lint results are given today. **2. Developer-Readable Results:** Rather than dumping raw scanner output, the findings are presented as **code-enhanced, reproducible scenarios**. If a SQL injection is found, we provide a test case that triggers it and point to the exact line in code (via the graph) where the fix is needed. We might say: *"Vulnerability:* *SQL Injection in* `SearchService.java` *line 45.* *Proof:* *sending* `q=' OR '1'='1` *to* `/search` *returns all records.* *Fix:* *parameterize the query or use the SafeQuery API."* This format gives developers everything needed: what the issue is, how to reproduce it (which they can turn into a unit test), and guidance on how to fix. By making the issues **actionable**, we turn security into just another bug to be squashed. We also prioritize false positive elimination -- developers quickly lose trust if a tool cries wolf. Thanks to our hybrid approach, most findings are backed by a real exploit or at least a very convincing static trace. If something is uncertain (maybe a very edge theoretical case), we mark its confidence level and often automatically ask the developer: *"Could this scenario happen in your deployment? If not, mark as safe."* -- effectively treating them as the domain expert to confirm or reject the finding. **3. Auto-Remediation and Suggestions:** Whenever possible, the platform doesn't just point out a problem, but offers a solution. We leverage the code understanding and LLM to suggest **patches or mitigations** for simple issues. For example, if an endpoint is missing an authorization check, we might generate a code snippet that checks the user's role, tailored to the coding style and framework in use\[8\]. If an input isn't validated, we might propose a validation rule (e.g. "add `if (input.length > 100) return error`"). These suggestions appear in the report and in developer tools, and in some cases can even be applied automatically (with developer approval). This follows Cruz's idea of making security fixes "invisible" or automatic for developers whenever possible\[9\]. By reducing the workload to fix issues, we further integrate with the developer's goals (nobody likes spending cycles on security bug fixing, so we help speed it up). **4. Human-in-the-Loop Learning:** Perhaps the most unique aspect is how developer input **feeds back into the system's intelligence**. Each finding or hypothesis in the knowledge graph can be marked as confirmed, fixed, or a false positive. When a developer says "this is a false positive because X," that note is stored. Over time, the AI can learn from these dispositions, avoiding similar false alarms in the future. If a developer fixes an issue, our system sees the code change (through integration with version control) and can update the graph (e.g., now there's an auth check node where there wasn't before). This closes the loop in a way most tools don't -- the system continuously validates its model against reality and expert feedback[\[27\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/06/08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md#L36-L39). Just as the AI suggests tests and finds issues, it also listens when humans correct it or code changes render an issue moot. Additionally, as new features are developed, developers can leverage the system to **validate design ideas**. Suppose a team wants to add a new API method -- they can write an OpenAPI snippet or pseudocode and run it through our analysis preemptively. The AI can then comment on the design: "This new endpoint will allow bulk deletion of users -- ensure only admins can call it. Consider adding rate limiting." It's almost like a pair programming assistant focused on security. **5. Continuous Verification and Metrics:** We provide dashboards that are meaningful to dev teams: for instance, a live score of security test coverage ("95% of endpoints have security tests passing"), and trending graphs of vulnerability counts, time-to-fix metrics, etc. This turns security into a quantifiable quality metric that teams can improve sprint over sprint\[17\]. It also helps identify areas of code that are "dark" (untested)\[28\] so developers can add coverage there (because what isn't tested likely isn't secure). By visualizing these and allowing developers to drill down (backed by the graph data), we encourage a culture of measurable security improvement. In essence, this solution treats developers as first-class users, not just subjects of security attention. It speaks their language (code, tests, CI/CD) and **augments their capabilities**. The outcome is a positive feedback loop: developers write better code and tests (with AI help), which makes the security analysis smarter, which finds deeper issues, which developers then fix -- raising the security bar continually. This collaborative approach is crucial for long-term success; it ensures the tool is seen as a helpful coworker rather than an annoying auditor. ## Persona-Based Reporting and The Cyber Boardroom Perspective Different stakeholders care about API security from different angles. A developer wants the stack trace and line number; a security architect wants to see the system design implications; a compliance officer might need to know if data privacy rules are followed; an executive just wants to know the business risk. Our platform addresses this by producing **persona-based outputs** from the central knowledge graph, using techniques similar to those in Cruz's MyFeeds.ai and Cyber Boardroom initiatives. **1. Security Team Dashboard:** For AppSec engineers and architects, we provide a rich dashboard driven by the knowledge graph. This includes an interactive map of the API attack surface -- you can explore nodes and edges visually, filter to see, for example, all endpoints lacking certain defenses, or all data flows that involve personal data. The dashboard highlights high-risk endpoints (perhaps scored by the AI considering factors like accessible without auth + touches critical data). One can click on an endpoint and see all linked information: code, recent test results, known vulnerabilities, related incidents. Essentially, it's an **interactive threat model** of the application that stays up to date. This addresses the CISO concern of inventory and understanding exposure\[1\], giving security teams a continuous view of "what do we have and where are the weak points?". Additionally, the security view includes compliance mappings -- for instance, tagging parts of the graph with OWASP Top 10 categories, or GDPR data classifications. This way, a security lead can ask, "Show me all APIs that handle personal data and whether we have tested them for data leakage." The system can answer via the graph relationships and even generate a report on it. **2. Developer & DevOps View:** As mentioned, developers see issues in their tools, but there's also a macro view for engineering managers or DevOps leads. They can see metrics like "security technical debt" -- number of outstanding vulns by severity, areas of code with frequent issues, etc. They can drill into a service and see if security tests are part of its CI, when they last passed, etc. This view is about integrating with the dev planning process: for instance, showing that Team X has 5 open security findings of high severity, or that Module Y hasn't been security-tested since major changes. It helps prioritize work and also celebrate improvements (e.g., "module X had 10 vulns last quarter, now zero"). Because our system is open and data-backed, these metrics are credible and traceable to actual evidence. **3. Compliance and Risk Reports:** Many APIs carry regulatory obligations (privacy, financial transactions, etc.). For compliance officers or auditors, our platform can produce evidence reports that specific controls are in place and tested. For example, a GDPR-focused report might list all endpoints that handle personal data and confirm whether data access is properly authorized and if data is encrypted in transit/storage. Since our knowledge graph knows which data is sensitive and which endpoints touch it, we can automate a lot of the compliance checking. We even model certain compliance requirements as subgraphs (like a template of what needs to connect to what for a control to be in place). The system can then highlight if any expected link is missing (say, an "audit logging" node should be linked to all payment transactions -- if not, that's a compliance gap). This not only saves time preparing for audits but also ensures security and compliance are aligned -- often the same technical measures satisfy both. **4. Executive/Board-Level Summaries:** High-level stakeholders, like a CIO, CISO, or board members, aren't interested in technical detail but in overall **risk posture and progress**. Inspired by *The Cyber Boardroom* (Cruz's GenAI-powered board briefing tool), our solution can generate executive-ready summaries from the underlying data[\[11\]](https://www.cyber-leaderssummit.com/speakers/dinis-cruz#:~:text=Dinis%20Cruz%20is%20the%20Chief,technical%20expertise%20with%20commercial%20leadership). Using natural language generation, it can produce something like: *"In Q3, the API security platform analyzed 120 endpoints and identified 3 critical vulnerabilities, all of which were remediated within 5 days. The most significant was an authorization flaw in the Orders API, which could have exposed customer data; it was fixed and tests confirm the control. Current risk level of our API estate is* *LOW* *with no known critical issues. Ongoing improvements include better input validation across services and enhanced monitoring for abuse patterns. The security posture has improved compared to last quarter, with average time-to-fix down from 10 to 6 days."* Such a narrative is backed by real data (and we can include charts or KPIs) but distilled by AI to be concise and business-focused. The multi-stakeholder approach ensures that **each role gets the information they need, in the form they need it**. It's not one report fits all. The key to enabling this is the richness of the knowledge graph and the use of GenAI to tailor outputs. MyFeeds.ai demonstrated how combining structured knowledge with generative narrative can produce highly personalized content[\[10\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/06/08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md#L38-L41). We apply that here: the structured data is our test results and API model; the personas are dev, security, compliance, executive; and the generative component assembles the right facts into the right narrative for each. For collaboration, these outputs can be delivered in the channels stakeholders prefer -- developers get Jira tickets or GitHub issues, security teams get an interactive portal (and maybe Slack alerts for critical findings), executives might get a periodic PDF report or a presentation generated for the quarterly risk review. The system can even power an **interactive Q&A bot** for these personas: e.g., a board member could ask in plain language "Have we tested our APIs against the latest OWASP vulnerabilities?" and the bot (using the knowledge graph and LLM) could answer with a summary and confidence level, referencing the tests done. This not only multiplies the value of the platform beyond just "finding bugs" -- it turns it into a knowledge and communication tool about our API security. In large organizations, that helps break down silos: everyone from engineers to board members stays informed and aligned on the state of API security. It bridges the gap that often exists where technical details don't make it up to decision-makers, or strategic priorities don't trickle down to devs. By having one system that feeds tailored info upward and downward, we create a **Cyber Boardroom effect** inside the company: the understanding that API security is being managed with transparency and intelligence, and that it's a continuous team effort. ## Architecture and Delivery Model The solution's architecture is designed for **cloud-native deployment, scalability, and openness**. Key characteristics include: - **Serverless Execution:** Every component of the analysis can run in a serverless function or container, triggered on demand. Whether it's parsing code, running a test, or generating a report, nothing runs 24/7 unless needed. We use cloud object storage (or a secure file store) as the system of record for inputs and results, rather than a persistent server. This approach -- similar to the Ephemeral SIEM concept -- yields a highly scalable and cost-efficient system with *zero idle infrastructure*[\[5\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L28). A user can spin up hundreds of test executions in parallel if needed, and pay only for the compute they consume. It also means updates to the system are easy to deploy (functions can be updated independently) and there is no single point of long-lived failure. - **Stateless & Stateless (Data Persistence):** The combination of Memory_FS and MGraph-DB gives us a unique persistence model: everything is essentially stored as **files/objects** (code, test artifacts, graphs) that can be versioned and moved, rather than locked in a database. This stateless design is what allows easy **local deployment** or on-premises use -- a team could run the whole stack on their laptop or private cloud by pointing it to their file system instead of our cloud storage. The heavy use of simple formats like JSON, Markdown, ZIP means it's easy for users to inspect or export their data. We want users to *own their security data*; the value we provide is in the analysis orchestration and AI, not in hoarding the data. This aligns with the open philosophy that the **knowledge itself is the valuable output and remains portable**[\[12\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L27). - **Open Source Core and Extensibility:** All core components -- the graph engine, the analysis rules, the testing harness, and integration SDKs -- are open source. This invites a community of contributors (researchers, practitioners) to extend the platform with new rules or adapters. For example, if someone creates a better taint analysis module or a new set of tests for GraphQL-specific issues, it can be contributed and shared. The system is built with an extensible plugin architecture: new languages, API protocols, or vulnerability classes can be added by plugging in modules that produce or consume from the knowledge graph. This way, the platform can evolve with technology. If your organization uses a new RPC framework, you can write a module to parse its routes into the graph, rather than wait for a vendor update. **Customization and personalization** are also supported -- companies can add organization-specific rules (like "alert if endpoint doesn't have our custom audit logging call") easily via config or scripting. We consider this a key differentiator: security tools often force one-size-fits-all, but our framework lets power users tailor it to their policies and threat models. - **Credit-Based SaaS Model:** While the core is open, we anticipate offering a cloud service where heavy computations (especially LLM calls) and large-scale orchestrations are managed for the user. This service would operate on a **usage-based (credit) model** -- similar to how cloud providers charge for function invocations or how GPT API charges per token. Each security analysis or test or AI query might cost a certain number of credits, reflecting the compute resources used. This model ensures users only pay for what they use and can predict costs by the scope of their testing (which they control). It also encourages efficient testing -- e.g., focusing intensive analysis on critical parts of the app, while perhaps running lighter checks on less critical ones, as decided by the user. The credit system could also allow a **freemium approach**: basic usage (for open-source projects or small apps) might be free or very low-cost, while enterprise-scale use (lots of APIs, very frequent testing) requires purchasing additional credits or a subscription. - **Support and Services:** As an open platform, another part of the model is offering **professional support, training, and consulting** for organizations that need help integrating the solution or interpreting results. This might include helping a dev team set up their CI pipelines with the tool, or doing an initial security mapping of a legacy system using our platform and handing over the knowledge graph to the client. Because the product is inherently technical and touches many aspects (DevOps, SecOps, compliance), some companies will value expert guidance -- this becomes a revenue stream that complements the software usage. The key is that these services are optional; a savvy team can fully self-serve with the open tools, or they can opt for convenience and assurance via our managed service and expertise. - **Community and Knowledge Sharing:** We plan to foster a community where **vulnerabilities and test cases can be shared (responsibly)**. For example, if our system finds a zero-day vulnerability in a popular open-source API framework, after coordinated disclosure we can publish a case study showing how the tool found it, including anonymized graph snippets or test payloads\[3\]\[29\]. This serves as proof of efficacy and also as educational content for others. We could maintain a repository of known API vulnerability patterns (like a library of graph motifs that indicate certain flaws) which the community can contribute to. Over time, as more companies use the tool, they might contribute sanitized models of their APIs or custom tests (especially if they develop tests for very domain-specific logic) -- these could be added (with permission) to a collective knowledge base. In essence, every time the platform finds a new type of vulnerability, that knowledge can help *all users* by updating the AI models or rule sets. - **Privacy and Data Security:** Since security tooling will naturally handle sensitive code and perhaps data, we ensure strong isolation and encryption in the SaaS offering. Memory_FS archives and test results can be encrypted with customer-managed keys. If a user chooses local deployment, none of their code or data ever leaves their environment; even if they use our AI features, they can opt for on-prem LLMs or anonymized prompts. This flexibility is important to gain trust, especially in industries with strict IP or data control (e.g., finance, healthcare). Our open model actually helps here too -- customers can inspect exactly what the system is doing with their data since the code is open. In summary, the architecture is built to be **flexible, transparent, and scalable**. It embraces modern cloud principles (serverless, stateless microservices), which not only scale well but also reinforce a usage-based business model. The openness ensures that users are never hostage to a vendor -- they can always run the core themselves or extend it to new needs. Yet by offering a managed service on top, we combine the best of both worlds: community-driven innovation with the convenience of a polished SaaS for those who want it. ## Community Impact and Adoption Strategy To truly succeed, a security solution must not only be technically sound but also win hearts and minds. Our strategy for adoption hinges on building **credibility through results, fostering an open community, and demonstrating unique value** in a crowded space. **1. Demonstrating Efficacy with Real Vulnerabilities:** Nothing convinces like real-world results. Early on, we plan to use the platform on open source projects and publicly share the findings (responsibly). For example, we might pick a popular open API project, run our tool, and discover a subtle logic flaw or injection issue that has slipped past its maintainers. We would then coordinate disclosure, help fix it, and publish a detailed post (or even a white paper) showing how our approach found that issue -- essentially a **proof of concept of our methodology**. By publishing such case studies, we not only improve security of those projects but show the broader industry that this approach works. These success stories can drive interest from organizations: "If it found a new vulnerability in X project, imagine what it might find in our systems." It's akin to how some security companies publish research on new vulnerabilities -- we are turning our tool into a research engine and sharing the outputs. This also naturally contributes to the community's knowledge: each published finding enriches the collective understanding of API security. **2. Community Involvement and OWASP Alignment:** We will actively engage with communities like OWASP (Open Web Application Security Project) and other security forums. Given Dinis Cruz's own history with OWASP, we can align our project with OWASP's mission (perhaps even contribute it as an OWASP project). This lends credibility and invites a ready audience of AppSec professionals to participate. By being open source, students, researchers, and independent developers can use the tool on their projects and contribute improvements. We might run community contests or bug bounty-like challenges using our tool -- e.g., "secure this intentionally vulnerable API using our platform and win prizes," which both educates and promotes usage. **3. Differentiating in the Market:** The API security space is heating up, but our approach is deliberately distinct. We emphasize that **we are not just an "API scanner" -- we are a platform for API *understanding and collaborative testing***. The integration of LLM-driven knowledge graphs and the developer-centric philosophy set us apart. Where other tools might be black-box fuzzers or static analyzers with new packaging, we offer a genuinely new architecture (as described in this paper). Our marketing will focus on these differentiators: - *Developer-first:* It's built as much for devs as for security, whereas many tools cater only to security teams. - *AI + Graph intelligence:* We bring semantic context and AI reasoning -- buzzwords, yes, but we back them with concrete features (graphs, LLMs) that are not common in most security tools. - *Open and extensible:* While competitors might be closed SaaS, our open approach is attractive to enterprises wary of lock-in and to users who want transparency. - *Multi-faceted testing (functional, security, resilience):* We position API security testing as an extension of overall testing, which resonates with modern quality engineering trends. By highlighting these, we aim to be seen not just as a "security tool" but as an enabler of **better software engineering** overall. This broadens our appeal. **4. Gradual Onboarding via Credit Model:** The credit-based model allows new users to start small and see value quickly. An organization can, for instance, spend a few credits to map one API and run basic tests -- if they see interesting results, they'll be enticed to expand usage. The granular usage model lowers the barrier to trial (no need for a big commitment or installation -- just run it on one service). We will likely have a generous free tier for open source projects, which further spreads awareness (open source maintainers improve their projects for free, and they become references for us). **5. Ecosystem and Partners:** We would explore partnerships with cloud providers and DevOps platforms. For example, an integration with GitHub Actions could make it one-click for users to add our security tests to their CI. Cloud providers might showcase our tool as a way to secure serverless and API gateway setups. Since our approach is stateless and could run on any cloud, we can partner rather than compete with infrastructure providers -- potentially analyzing their customers' workloads with their blessing. We might also integrate with popular API management platforms or gateways, so that whenever someone defines a new API route, a hook triggers our analysis. **6. Continuous Improvement via AI Feedback:** As the user base grows, the AI models (if using a centralized service) will get smarter. We can train on the patterns of vulnerabilities found (without using private code, just the abstract patterns) to improve the AI's suggestions. For local deployments, we might allow users to opt-in to share anonymized meta-data to contribute to this learning. This means the tool's effectiveness will scale not just with computing power, but with **community wisdom**. In closing, the ambition of this approach is to elevate API security to a new plane -- one where **mapping and securing an API is as intuitive and automated as running a suite of unit tests**, amplified by the insights of AI and the connectivity of knowledge graphs. It's a vision where security is woven into the fabric of development and operations, driven by data and collaboration rather than fear or checkbox compliance. The technical uniqueness lies in how the pieces (graphs, LLMs, memory virtualization, etc.) come together, but the ultimate measure of success is the real-world impact: more secure APIs, fewer breaches, faster development of robust features, and a community that shares in advancing the state of the art. By adhering to the principles outlined by Dinis Cruz's research -- from "every API is vulnerable"\[13\] to "think how it can be abused"\[7\] to leveraging AI and graphs for scale\[3\] -- we believe this solution can significantly advance how organizations protect the critical APIs that power their business. The path forward is one of open collaboration, continuous innovation, and relentless focus on making security practical for those building the future. \[1\] \[2\] \[3\] \[4\] \[6\] \[7\] \[8\] \[9\] \[13\] \[14\] \[15\] \[16\] \[17\] \[18\] \[20\] \[22\] \[23\] \[24\] \[28\] \[29\] Debrief\_ Dinis Cruz's Research on API Security (2009--2025).docx [\[5\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L28) [\[12\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L20-L27) [\[19\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L7-L10) [\[21\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md#L8-L15) ephemeral-genai-siem-a-serverless-graph-driven-approach-to-security-event-management.md [\[10\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/06/08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md#L38-L41) [\[27\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/06/08/personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md#L36-L39) personalized-briefing-semantic-knowledge-graphs-dinis-cruz-kerstin-clessienne.md [\[11\]](https://www.cyber-leaderssummit.com/speakers/dinis-cruz#:~:text=Dinis%20Cruz%20is%20the%20Chief,technical%20expertise%20with%20commercial%20leadership) Dinis Cruz - Cyber Leaders\' Summit 2025 [\[25\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md#L22-L25) [\[26\]](https://github.com/DinisCruz/docs.diniscruz.ai/blob/ab78b74dfe5b45822f16ecf09e3e84391a97d948/docs/2025/07/02/using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md#L16-L24) using-memory_fs-to-build-a-file-based-representation-of-the-gdpr-standard.md --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/08/23/llm-workflows-stateflow-service-technical-brief.html (markdown twin: /2025/08/23/llm-workflows-stateflow-service-technical-brief.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/08/23/llm-workflows-stateflow-service-technical-brief.html* # LLM Workflows/Stateflow Service - Technical Brief *By Dinis Cruz and ChatGPT Deep Research · 2025-08-23* > The LLM Workflows/Stateflow Service is a proposed stateless web service for executing AI-driven workflows with well-defined, deterministic steps. It acts as a state machine for orchestrating Large… --- [PDF](https://files.diniscruz.ai/github/pdf/2025/08/23/llm-workflows-stateflow-service-technical-brief.pdf) ## Introduction The **LLM Workflows/Stateflow Service** is a proposed stateless web service for executing AI-driven workflows with well-defined, deterministic steps. It acts as a **state machine** for orchestrating Large Language Model (LLM) calls and other actions in a controlled sequence. Unlike "agentic" AI systems that freely decide each action, this service uses a **pre-defined flow (blueprint)** to ensure predictability and provenance. The goal is to harness the power of LLMs (for analysis or content generation) within strict boundaries: the LLM can perform complex subtasks, but it *never* dictates the workflow's overall path[\[1\]](https://arxiv.org/html/2508.02721v1#:~:text=Agent%20framework,as%20parsing%20an%20error%20log). This design follows the *"Blueprint First, Model Second"* principle, decoupling high-level logic from the probabilistic nature of LLMs[\[1\]](https://arxiv.org/html/2508.02721v1#:~:text=Agent%20framework,as%20parsing%20an%20error%20log). The result is an execution environment where every step is explicit and auditable, yielding reliable behavior even in complex multi-step AI tasks. **Key collaborators:** The creation of this service is a collaborative effort between human developers and AI assistants (LLMs). Developers will define the workflow blueprints, implement the engine, and integrate services, while advanced LLMs (such as ChatGPT) assist in brainstorming, research, and even code generation. For example, this very document is co-authored by Dinis Cruz and ChatGPT, reflecting the human-AI partnership in designing the system. By leveraging LLM support during development, the team can explore design options and documentation rapidly, then validate and refine them with human expertise. This collaboration ensures the final system design is both innovative and technically sound. ## Design Principles and Goals - **Deterministic Control Flow:** All workflows are executed according to a fixed, predefined sequence of steps (a *stateflow*). The LLM is not allowed to randomly alter the flow or insert unforeseen steps. This contrasts with fully agentic systems where the next action is decided by the model at runtime. By enforcing determinism, the service guarantees predictable and repeatable executions. Recent research underscores the importance of this approach: using an explicit blueprint and a deterministic engine yields verifiable and auditable processes, as opposed to unpredictable agent behaviors[\[1\]](https://arxiv.org/html/2508.02721v1#:~:text=Agent%20framework,as%20parsing%20an%20error%20log). In a similar vein, frameworks like LangGraph favor explicit state machines for multi-agent orchestration, resulting in workflows that are *"predictable and auditable"* -- easier to debug and trust in enterprise settings[\[2\]](https://www.zenml.io/blog/langgraph-vs-autogen#:~:text=LangGraph%20is%20ideal%20for%20building,to%20debug%20and%20guarantee%20behavior). - **Stateless Execution:** The service itself maintains no persistent state between steps. Each step in a workflow is executed independently, given its input and the overall flow definition, and it produces an output (and possibly a next-step decision). Any state needed (e.g. variables or results passed from one step to another) is included in the input/output data rather than stored in the server. This stateless design aligns well with serverless deployment and makes the system highly scalable. It also simplifies reasoning about execution, since each invocation of the service is pure (no hidden state from prior runs). - **Workflow as Data (Blueprint):** Workflows are defined declaratively as data structures (a **blueprint** or **flow definition**) rather than hard-coded logic. An expert (developer) or an automated design tool specifies the sequence of actions, conditional branches, and allowed operations in a machine-readable format. The workflow blueprint can be represented in JSON or a similar structured schema, akin to how one might describe an API contract in OpenAPI. This blueprint is essentially a directed graph or state machine: each node represents a step (action to perform) and edges define transitions to the next step. The service's engine interprets the blueprint to decide what to do at each step and where to go next. Crucially, the LLM is *only* invoked within specific steps for bounded tasks -- for example, to generate text or analyze data -- but **never to choose the next step in the state machine**[\[1\]](https://arxiv.org/html/2508.02721v1#:~:text=Agent%20framework,as%20parsing%20an%20error%20log). This ensures the flow's path remains under deterministic control of the blueprint. - **Provenance and Auditability:** Every action taken by the workflow (LLM call, API call, decision branch, etc.) is logged and traceable. Because the sequence is predetermined (except for conditional logic which is still explicitly defined), we can provide a complete audit trail of what happened and why. This is critical in enterprise or safety-critical contexts where understanding the decision process is required. The structured nature of the flows makes it easier to inspect and verify correctness before execution, and to analyze outcomes after execution. - **Limited LLM Scope (Safety):** The use of LLMs within the workflow is intentionally constrained. LLMs will be used for tasks like natural language processing, summarization, translation, or analysis -- **not for making control decisions**. By bounding the LLM's role, we avoid the unpredictability that comes if an LLM "agent" had free rein. Essentially, the LLM behaves as a powerful subroutine under supervision. This mitigates risks of the LLM drifting off-task or producing unintended actions. It also allows insertion of validation steps (for example, after an LLM produces an output, the next step could validate that output against rules or schemas before proceeding). - **Budgets and Resource Controls:** Each workflow (and even each actor or tool within the workflow) can be assigned a **budget** -- for example, a limit on the number of API calls, the number of tokens processed by an LLM, or a time limit. These budgets are integrated into the flow logic. If a component exceeds its budget, the workflow can take a predefined path (such as aborting or invoking a fallback step). By designing budget constraints into the blueprint, we ensure the system cannot accidentally run away in loops or incur unbounded costs. In our scenario, for instance, the Persona service or LLM might only have a certain token budget; if translating or answering a question would exceed that, the flow would stop or switch to an error state. This built-in guardrail is another measure to keep executions safe and predictable. - **Flexible Complexity:** The system is meant to handle simple to very complex workflows. A simple use-case might be a single-step call (e.g. \"call this LLM with prompt X and return answer\"). A complex use-case could involve multiple agents, branching logic, loops (controlled by explicit conditions), and tool integrations. The design should support both extremes. This implies the blueprint language must be expressive enough for conditions, branching (if/else), parallel execution (if needed), and looping (perhaps via recursive flows or explicit loop constructs) -- all while remaining human-comprehensible. The service should execute any such defined flow **reliably**, since each part is under strict governance of the blueprint. - **Integration with Knowledge Graphs (Future Scope):** Although not the first priority, the vision includes leveraging **semantic knowledge graphs** in conjunction with the flow engine. The idea is that the workflow steps and their context could be informed by a knowledge graph that contains relevant domain information or state. For example, a step might query a knowledge graph for facts that an LLM then uses in its prompt. Or the flow blueprint itself could be stored/versioned as a graph of nodes (steps) and edges (transitions), enabling visualization and analysis of the workflow structure. By combining flows with semantic knowledge representations, we can achieve more powerful reasoning while still keeping the execution controlled. This is an area of ongoing research and will be explored as the service evolves. ## Service Architecture and Components The **LLM Workflows/Stateflow Service** architecture is composed of several key components that together enable the definition and execution of these workflows: - **Workflow Definition (Blueprint):** At the core is the **workflow blueprint**, a complete specification of the states and transitions for a given process. This can be a JSON document or Python object (defined via dataclasses or Pydantic models) that lists all the steps in order, along with any branching logic. Each **Step** in the blueprint typically contains: - An **action** identifier -- e.g. `"call_LLM"` or `"call_service:persona.translate"` -- which tells the engine what to do. - **Parameters/input** for that action -- e.g. the prompt to send to the LLM, or the data to send to an API. Parameters might include references to outputs of previous steps (for example, step 3 can use the result produced by step 2). - **Result handling** info -- e.g. where to store the result of this step (in a context object or variable map that gets passed along). - **Transitions** -- rules for what the next step is. In the simplest case, each step names a single next step. In more complex cases, a step may have multiple possible next steps depending on its result (for example, a branching step could say: *if evaluation score \>= 8 go to step X, else go to step Y*). If a step is terminal, it can be marked as an **End** state. - Optional **metadata** like timeouts or retries for that step, if an action might fail and should be retried. The blueprint is analogous to a workflow script. The service includes a **Blueprint Interpreter** (or engine) which reads this definition and drives execution accordingly. Because the blueprint is declarative, we can inspect and validate it before running, and even visualize it as a flowchart. - **Execution Engine:** The engine is a stateless component that takes a workflow blueprint plus the current state (inputs and context) and executes one step of the workflow. It can be implemented as a FastAPI endpoint (for example, `/execute_step`) which accepts a JSON payload containing the workflow definition (or a reference to it), the current step to execute, and any necessary input data or context. The engine performs the action of that step and returns the result (and the identifier of the next step to execute, if any). Important responsibilities of the engine include: - **Action Dispatch:** Based on the step's action identifier, call the appropriate function or external service. For example, if the action is `call_LLM`, the engine will invoke the configured LLM API (such as an OpenAI or Anthropic model) with the given prompt and parameters. If the action is `persona.translate` or `persona.respond`, the engine calls the Persona service endpoint. If it's a generic HTTP API call, the engine will make that HTTP request (possibly the blueprint could contain the URL or a reference to a pre-registered service with credentials). - **Budget Checking:** Before executing an action, the engine checks the remaining budget for that actor/service. This could be implemented by maintaining a transient budget counter in the context that tracks tokens or calls. For instance, if the Persona service has 100 tokens budget and translating the message is estimated to use 30, the engine verifies budget \>= 30, deducts it, then proceeds. If budget would be exceeded, the engine can skip to a special termination or error step defined in the blueprint. - **Result Handling:** After performing the action, the engine takes the output and places it in the workflow's state (for example, storing it under a variable name). If the blueprint specifies post-processing (like simple transformations) or validations (e.g. ensure the result fits an expected format), those are done here as well. - **Next Step Logic:** The engine determines which state to execute next. In straightforward sequences, the blueprint might indicate a `next` step explicitly. If the current step was a branching/choice step or if the action outcome dictates the path, the engine evaluates the condition and picks the appropriate next step. For example, an evaluator step might branch to different follow-up steps depending on the score (high score leads to a "successful completion" branch, low score leads to an "escalation" branch). If there is no next step (the step was marked as an end), the workflow is complete and the engine returns the final result of the workflow. Notably, the engine is **agnostic to the overall workflow** -- it doesn't keep a running memory of all previous steps beyond what is carried in the input context. This makes it easy to scale or even to *pause and resume* workflows by simply not calling the next step immediately. It is the responsibility of the **Orchestrator** (described next) to drive the engine through the steps. - **Workflow Orchestrator:** Since the service executes one step at a time, we need an orchestrator to manage multi-step workflows from start to finish. This could be a separate client application, a calling service, or even a simple loop in a script that keeps invoking the engine until the workflow ends. For example, suppose we have a workflow with steps A -\> B -\> C. The orchestrator would: - Call `/execute_step` for step A with initial input. Receive result and next step = B. - Call `/execute_step` for step B with result of A. Receive result and next step = C. - Call `/execute_step` for step C with result of B. Receive result and next step = none (end). - Collect final result. The orchestrator could be implemented as part of the client using this service, or we could provide a utility in the service that takes a whole blueprint and automatically steps through it. However, having the orchestrator outside the core engine adds flexibility: workflows can be paused, inspected mid-run, or even modified between steps if needed (for advanced use cases). In a serverless deployment (e.g., AWS Lambda via OSBot-FastAPI-Serverless), the orchestrator might be a state machine service (like AWS Step Functions itself or Azure Durable Functions) that triggers the next lambda invocation. Alternatively, a simple loop in an API endpoint (if using a container or persistent service) could also orchestrate the steps synchronously. - **Integrated Services and Tools:** The workflow steps can involve various external services: - An **LLM Service**: e.g., an API call to an OpenAI or a local model. The flow can specify which model and prompt to use. The LLM service endpoint would need to be reachable by the workflow engine and secured (API keys, etc.). The service may have multiple models (GPT-4, GPT-3.5, etc.) and the blueprint could choose which one by an identifier. - The **Persona Service**: (at `persona.qa.mgraph.ai` as mentioned) which has two main capabilities: - *Persona Translation*: take an input message and translate or rephrase it for a target persona or in a target language (e.g., explain a technical incident in board-level terms, or convert English to Portuguese, etc.). - *Persona Response*: given a prompt or question, respond **as** the persona (generating an answer the persona would give). This service is itself likely powered by an LLM under the hood, but from our workflow's perspective it's a black-box API that we call with certain parameters. The workflow can include steps that call `persona.translate` or `persona.respond` with the appropriate persona and message. - **Evaluation Service**: an AI or rule-based system that evaluates a given piece of content. In the example scenario, this is used to rate how well a response was understood or how effective it was. This could be another LLM prompt (e.g., asking an evaluator model to score the answer on clarity), or a more deterministic function that checks for certain keywords, etc. We treat it as an external call (perhaps another endpoint like `evaluator.rate_response`). - **Other APIs/Tools**: The system could integrate with any HTTP API or internal function. For instance, steps could call a database, send an email, trigger a CI/CD pipeline, run a shell command (if allowed), etc. These would be included by defining new action types in the blueprint and teaching the engine how to execute them. The design is extensible; new action handlers can be added to the engine plugin-style. - **Budget Manager:** As mentioned, budgets for each actor/service must be enforced. A simple implementation is to include budget counters in the workflow context passed through steps. For example, the context could have `{ "budget": {"persona": 1000, "llm": 1500, "evaluator": 500} }` indicating remaining token or API call counts. Each time the engine is about to call one of these, it subtracts an estimated cost. If the cost is not known exactly (e.g., tokens to be consumed by an LLM call depends on output length), the engine can either make a conservative guess or get the info after the call (some LLM APIs return token usage). If a budget hits zero or goes negative, the engine can trigger a special transition (for example, go to a "Budget Exceeded" fail state). The blueprint can define what to do in that case (maybe notify the user, or just terminate). The **Budget Manager** could be just a piece of logic in the engine or a separate helper module. - **Logging & Monitoring:** Every request to the engine and every action invocation should be logged. This includes inputs, outputs, and any decision points. For debugging complex workflows, it is invaluable to see a trace of steps. We can integrate this with the OSBot-FastAPI's event system (which already provides request/response tracking and can store events)[\[3\]](https://pypi.org/project/osbot-fast-api/#:~:text=A%20Type,AWS%20Lambda%20integration%20through%20Mangum). Logs can be written to a database or a cloud logging service. We might also have a **monitoring dashboard** that shows active workflows, their progress, and any errors. For long-running workflows (if any), monitoring would show the current state and waiting transitions. Given the stateless nature, most workflows will run quickly or be orchestrated externally, but monitoring is still useful for analyzing results and performance. We will also capture metrics like time per step, tokens used, etc., to help optimize the flows and LLM usage. - **Security & Access Control:** As an OWASP Security Bot project, security is important. The service should authenticate and authorize requests (using API keys or tokens, as supported by OSBot-FastAPI middleware[\[4\]](https://pypi.org/project/osbot-fast-api/#:~:text=%2A%20Type,%E2%86%94%20BaseModel%20%E2%86%94%20Dataclass%20conversions)). Only authorized systems or users should be able to submit a workflow for execution or call the engine. Also, the actions allowed in a workflow might be gated: for example, we might restrict certain flows from calling arbitrary URLs unless explicitly allowed, to prevent misuse. Each integrated service (Persona, LLM, etc.) will have its own credentials (API keys), which the engine must handle securely (likely stored in environment or a vault, not hard-coded in the blueprint). The blueprint might reference a service by name and the engine knows which credentials to apply. In summary, the architecture separates the *declaration* of what to do (blueprint) from the *execution* of how it's done (engine and orchestrator), with clear interfaces to external AI services and strict control logic enveloping any LLM calls. This design maximizes reliability and clarity, ensuring that even as we automate complex tasks with AI, the process remains transparent and governable. ## Workflow Definition and Standards Defining the workflow blueprint in a robust way is a critical aspect of this service. We want a format that is **expressive, standard-compliant (when possible), and easy for developers to author and maintain.** Given that our implementation language is Python (with **OSBot-FastAPI** and **OSBot-FastAPI-Serverless** frameworks), we have a couple of choices for how to represent workflows: 1. **Python Type-Safe Classes (Code as Blueprint):** We can create Python classes (using OSBot's Type_Safe or Pydantic models) to represent the elements of a workflow -- e.g., a `Flow` class containing a list of `Step` objects, where each Step has fields like `id`, `action`, `parameters`, `transitions`, etc. Developers can then construct these classes in code or load them from a JSON. OSBot-FastAPI will automatically handle conversion between these classes and JSON schemas[\[3\]](https://pypi.org/project/osbot-fast-api/#:~:text=A%20Type,AWS%20Lambda%20integration%20through%20Mangum). This means we get strong typing and validation for free. For example, if a step is missing a required field or has an invalid next step reference, the model validation can catch it before execution. This approach treats the blueprint almost like writing a program (in Python), which is then serialized to JSON when needed. 2. **JSON/YAML Workflow Schema (Data as Blueprint):** Alternatively, we define a pure JSON (or YAML) schema for the workflow -- a text-based format that could be authored by hand or generated by tools. There are **existing standards and schemas** in the industry that we can draw inspiration from: 3. **BPMN 2.0:** A long-standing standard for business process modeling (usually visual diagrams stored in XML)[\[5\]](https://github.com/woodyhayday/FlowSpec#:~:text=%2A%20BPMN%202.0%3A%20An%20industry,both%20visual%20and%20code%20views). BPMN is very expressive (supports events, gateways, etc.), but it might be overly complex for our needs and not JSON-friendly by default. 4. **BPEL:** An older XML-based language for web service orchestration[\[5\]](https://github.com/woodyhayday/FlowSpec#:~:text=%2A%20BPMN%202.0%3A%20An%20industry,both%20visual%20and%20code%20views). Also quite heavy and tied to WS-\* services. 5. **AWS Step Functions (Amazon States Language):** A JSON-based state machine definition used in AWS Step Functions[\[6\]](https://github.com/woodyhayday/FlowSpec#:~:text=,workflow%20definition%20tool%20integrated%20with). This is a practical and widely used format. It represents workflows as a set of states (Task, Choice, Parallel, etc.) with a `StartAt` and explicit `Next` transitions or `Choice` branches. It's quite suitable for our concept since it inherently models step-by-step execution with choices. The downside is that it's AWS-specific in some of its integration details, but we could adopt the structure. For example, we'd have \"Task\" states for calling LLM or Persona, and \"Choice\" states for branching on evaluation results, etc., all expressed in JSON. 6. **Azure Logic Apps:** Similar to Step Functions, uses JSON (and often designed via a visual editor)[\[7\]](https://github.com/woodyhayday/FlowSpec#:~:text=,both%20visual%20and%20code%20views). Also could be a reference, though tied to Azure connectors. 7. **Workflow Description Languages (WDL/CWL):** These are specialized languages for scientific and data workflows[\[8\]](https://github.com/woodyhayday/FlowSpec#:~:text=,and%20scientific%20computing%2C%20emphasizing%20reproducibility). They emphasize reproducibility and usually model batch processing of data (like DNA sequencing pipelines). They might be too domain-specific for our interactive use-cases, but they show how to define steps, inputs, and outputs clearly. 8. **Argo Workflows / Tekton:** Kubernetes-native workflow engines using YAML[\[9\]](https://github.com/woodyhayday/FlowSpec#:~:text=,excel%20in%20containerized%20CI%2FCD%20pipelines). They often focus on CI/CD pipelines and container tasks. The concept of defining DAGs of steps is similar, though Argo's YAML could be more low-level (each step is basically a container spec). 9. **Custom DSLs:** Many teams end up creating their own lightweight JSON/YAML schemas for workflows[\[10\]](https://github.com/woodyhayday/FlowSpec#:~:text=,to%20their%20specific%20automation%20needs). This is likely the approach we will take: design a JSON schema tailored to our AI workflow needs, while borrowing ideas from the above standards for structure and best practices. For instance, we might incorporate Step Function's idea of states and transitions, but simplify it, or include a field for natural language description like BPMN does, etc. Given that we aim for a **type-safe** and easily maintainable solution, our plan is to define a **JSON-based schema** for the workflow and also represent it with Python classes for convenience. We can start from scratch or use a nascent standard like **FlowSpec**. *FlowSpec* is an open initiative to create a standardized JSON schema for AI automation workflows[\[11\]](https://github.com/woodyhayday/FlowSpec#:~:text=A%20lightweight%20standardized%20JSON%20schema,step%20workflows). It recognizes that many workflow tools share a flow-chart backbone, and it attempts to unify this in a portable way. In FlowSpec, a workflow is defined by a title, description, a list of steps, and transitions between steps[\[11\]](https://github.com/woodyhayday/FlowSpec#:~:text=A%20lightweight%20standardized%20JSON%20schema,step%20workflows). Each step has fields for what action to execute, its inputs, expected outputs, and what the next step(s) are depending on outcomes. It even allows global default transitions (like what to do on any failure)[\[12\]](https://github.com/woodyhayday/FlowSpec#:~:text=,)[\[13\]](https://github.com/woodyhayday/FlowSpec#:~:text=%2A%20BPMN%202.0%3A%20An%20industry,intensive%20pipelines%20and). FlowSpec also enumerates existing workflow standards (as we did above) to validate the approach of a common schema[\[13\]](https://github.com/woodyhayday/FlowSpec#:~:text=%2A%20BPMN%202.0%3A%20An%20industry,intensive%20pipelines%20and). After researching these options, **our recommendation** is to adopt a JSON state machine schema inspired by AWS Step Functions and FlowSpec. This gives us a known structure (states, next, choice, etc.) but we will customize it for our needs (for example, integrate the notion of budget and our specific action types). We will keep the schema human-readable and not too verbose. For instance, a simple flow with two steps might look like: { "workflowName": "Simple Q&A", "startAt": "AskQuestion", "states": { "AskQuestion": { "type": "Task", "action": "call_LLM", "parameters": { "prompt": "Answer the user's question: {user_question}" }, "resultVar": "answer", "next": "EvaluateAnswer" }, "EvaluateAnswer": { "type": "Task", "action": "evaluator.rate_response", "parameters": { "response": "{answer}" }, "resultVar": "score", "end": true } } } In this pseudo-JSON: - `startAt` specifies the entry step. - We have two states: one calls an LLM to get an answer, the next calls an evaluator to score that answer, then ends. - We use placeholders like `{user_question}` and `{answer}` to indicate passing data between steps (the engine would replace those at runtime with actual values from context). - This format is quite similar to Amazon States Language (each state has a `Type` and either a `Next` or `End`)[\[14\]](https://states-language.net/#:~:text=%7B%20,1%3A123456789012%3Afunction%3AHelloWorld%22%2C%20%22End%22%3A%20true)[\[15\]](https://states-language.net/#:~:text=When%20this%20state%20machine%20is,or%20a%20runtime%20error%20occurs), but with an `action` field that is our custom addition to specify what the Task does (since we are not tying directly into AWS Lambda ARNs as AWS does[\[16\]](https://states-language.net/#:~:text=,true%20%7D%20%7D)). We will formalize such a schema and provide a JSON Schema definition for it (so it can be validated). The use of Python with OSBot-FastAPI means we can also create corresponding classes. For example, a `TaskState` class and a `ChoiceState` class that inherit from a base `State` class, etc., enabling developers to construct workflows in Python fluidly. The **OSBot-Fast-API** toolkit will assist by ensuring these classes convert to Pydantic models easily, preserving the strong types[\[3\]](https://pypi.org/project/osbot-fast-api/#:~:text=A%20Type,AWS%20Lambda%20integration%20through%20Mangum). This approach satisfies our need for a clear contract for workflows, while leveraging existing best practices from industry standards[\[13\]](https://github.com/woodyhayday/FlowSpec#:~:text=%2A%20BPMN%202.0%3A%20An%20industry,intensive%20pipelines%20and)[\[17\]](https://github.com/woodyhayday/FlowSpec#:~:text=,to%20their%20specific%20automation%20needs). To summarize, the workflow definition will likely be expressed as JSON but with first-class support in Python. It will incorporate ideas from state machine standards (like having explicit states, transitions, start/end, etc.) and will be designed to be **readable, easy to modify, and rigorous**. By doing this, we make it easier for developers to create new workflows or adjust existing ones, and possibly even enable *LLM-assisted workflow authoring* in the future -- for instance, an LLM could take a high-level description and output a draft JSON workflow, which a developer then reviews and fine-tunes. The use of a standard schema also opens the door to visualization tools or workflow editors down the line. ## Example Workflow: Persona-Based Communication and Evaluation To illustrate how the stateflow service works, let\'s walk through a detailed example workflow. This scenario involves translating and conveying a critical message between two types of personas in an organization, and evaluating the communication's effectiveness. We will use the previously described actors: - **Actor A:** The originator of the message (could be a human user or an automated alert). In our scenario, the message is: *\"A ransomware attack has hit Division X, which will impact the P&L (profit and loss) for this quarter.\"* - **Actor B:** The Persona Service, which can assume different personas. We will use it in two modes: - *Translator mode:* to rephrase a message for a target persona's understanding. - *Responder mode:* to generate a reply as if coming from a persona. - **Actor C:** The Evaluator service, which will judge the quality of responses (e.g., does the response answer the question clearly, does the target persona understand the message, etc.). The organizational context is that **Board Members** care about financial terms like P&L but might not understand technical cybersecurity jargon, whereas **CISOs (Chief Information Security Officers)** understand ransomware but might not grasp business impact jargon. Our message contains both technical (ransomware) and financial (P&L) terms, so it's challenging for either persona to fully understand without translation. We will construct a workflow that explores different communication paths: **1. Direct Communication to CISO (No Translation):**\ Actor A sends the original message directly to a CISO persona (via Actor B's responder mode acting as a CISO). The flow steps might be: - *Step 1:* `persona.respond` as **CISO** with input = \"Ransomware attack on Division X will impact P&L this quarter.\"\ → (Actor B generates a response as it thinks a CISO would reply. This CISO likely understands the ransomware part but may be confused or less concerned about P&L specifics. The response might say something focusing on cybersecurity mitigation but not address financial impact fully.) - *Step 2:* `evaluator.rate_response` on the CISO's reply, with criteria like *completeness*, *clarity*, *appropriateness for the question*.\ → (Actor C returns a score or feedback. We expect this might be a mediocre score if the CISO persona missed the financial aspect.) - *Step 3:* End. (We record the score and perhaps the content of the CISO's answer.) Expected outcome: The CISO's answer might mention technical steps (e.g., *"We are investigating the ransomware attack on Division X and working to contain it."*) but not translate that into business terms. The evaluator might note that the board (who cares about P&L) would not get a full picture from this answer. The score could be low or moderate. **2. Direct Communication to Board Member (No Translation):**\ Actor A sends the same message directly to a Board Member persona (Actor B acting as a board member): - *Step 1:* `persona.respond` as **BoardMember** with input = \"Ransomware attack on Division X will impact P&L this quarter.\"\ → (Actor B generates a response a board member might give. The board member persona might latch onto the P&L impact but be unsure about the technical details, possibly responding with something like *"How severe is the ransomware attack and what are the projected losses?"*) - *Step 2:* `evaluator.rate_response` on the Board Member's reply.\ → (We expect the board member's answer might not be directly useful because the board persona might actually ask questions or express confusion about the ransomware aspect. The evaluator likely scores this low in terms of addressing the problem, since the board member persona didn't provide a solution or clear action.) - *End.* This path shows how a mismatched communication (technical message to non-technical persona) might fail. The board member didn't provide a satisfying answer because they themselves didn't fully understand the technical side. The evaluator would likely flag that the communication was ineffective. **3. Translated Communication to CISO:**\ Now we improve the communication. Actor A's original message will first be translated to the CISO's \"language\" (i.e., reframed in cybersecurity terms), then delivered to the CISO persona, and evaluated: - *Step 1:* `persona.translate` target=**CISO**, input = \"Ransomware attack on Division X will impact P&L\...\"\ → (Actor B returns a **translated message** that a CISO would immediately grasp. For instance, it might elaborate the technical threat and downplay financial jargon: *"Division X has been hit by ransomware, affecting operations; this could have a significant business impact this quarter."*) - *Step 2:* `persona.respond` as **CISO** with input = **translated message from Step 1**.\ → (Now, receiving a message phrased in his context, the CISO persona can respond more appropriately. The answer might be like: *"Understood. We have isolated the affected systems and are initiating incident response. We estimate recovery in 48 hours. Financial impact is being assessed in collaboration with finance."* This is a more complete answer covering both tech and acknowledging financial impact, because the question was framed in terms the CISO cares about.) - *Step 3:* `evaluator.rate_response` on the CISO's new reply.\ → (Actor C would likely give a higher score here, since the response is clear, addresses the issue, and bridges to business impact. The evaluator might note that the communication was effective for the target audience.) - *End.* We expect this translated workflow to yield a good outcome: the CISO persona understood the question after translation and responded in a way that likely satisfies a board or oversight evaluator. **4. Translated Communication to Board Member:**\ Similarly, translate the message for a Board Member, then get a response: - *Step 1:* `persona.translate` target=**BoardMember**, input = original message.\ → (This might produce something like: *"We estimate a hit to this quarter's profits due to a cyber incident (ransomware in Division X)."* Essentially explaining ransomware impact in terms a board cares about, possibly avoiding jargon.) - *Step 2:* `persona.respond` as **BoardMember** with input = translated message.\ → (Now the board persona, fully aware of the financial framing, might respond appropriately, e.g.: *"Understood. Ensure all necessary resources are allocated to IT to resolve this quickly. Let's prepare a statement for stakeholders about the financial impact."*) - *Step 3:* `evaluator.rate_response` on this reply.\ → (Likely another high score -- the board member persona's answer is on point when the question was phrased in their terms.) - *End.* This shows that with proper translation, even a non-technical persona can engage effectively. **5. Back-and-Forth Dialogue (CISO ⟷ Board, Mediated by Translations):**\ We can extend the scenario to simulate an interactive dialogue between the CISO and Board Member personas. The idea is to have multiple turns: - First, the CISO receives a translated question (as in #3) and responds as CISO. - Then take the CISO's response, translate it for the Board, get a Board persona reply. - Then translate that reply back to CISO's terms, get CISO's next response. - Continue this exchange for a few iterations or until a **budget limit** is reached (to prevent infinite loops). In the workflow blueprint, this could be represented by a loop or recursive transitions. For example: - Step 1: `persona.translate` to CISO (original message) -\> output ciso_msg. - Step 2: `persona.respond` as CISO (ciso_msg) -\> output ciso_reply. - Step 3: `persona.translate` to Board (ciso_reply) -\> output board_msg. - Step 4: `persona.respond` as Board (board_msg) -\> output board_reply. - Step 5: **Loop condition**: If `board_reply` or some context indicates conversation should continue AND budgets remain, go back to Step 1 (or a specific step) with `board_reply` now serving as the \"original message\" (Actor A's input) for the next round, targeting CISO again. - If loop ends (either a set number of rounds reached or budget exhausted), proceed to evaluation or finalization: - Step 6: `evaluator.rate_response` on the final response or on the overall dialogue quality. - End. This looping construct is explicitly controlled. The blueprint would contain a **Choice** or condition check after Step 4 to decide whether to loop or exit. The *budget* for each persona ensures that, say, we don't allow more than N exchanges or Y tokens. For instance, we might give each persona service 3 calls budget. Each `persona.respond` call uses 1. So at most 3 rounds of responses per persona can happen (which is 3 CISO replies and 3 Board replies, for a total of 3 cycles) before the budget prevents further calls. During this back-and-forth, each translation ensures both parties understand each other's messages in their own context. The **Evaluator** at the end might evaluate the overall success of the communication. Perhaps it looks at the final outcome: did they reach a mutual understanding or plan? We could even have the evaluator step after each reply, storing intermediate scores, but in practice it might suffice to evaluate at the end or only log the conversation. This complex example demonstrates the power of the workflow approach: - We can coordinate multiple AI calls (translations, persona responses, evaluations) in a sequence that achieves a larger goal (effective communication). - Because it's all in a defined flow, we avoid chaos: e.g., the Board and CISO personas will not talk over each other or go off on tangents; they only respond when prompted by the workflow. - If something fails (say one of the steps returns an error or empty response), we could have failure paths defined. For example, if `persona.respond` fails due to no available LLM, the blueprint could go to a step that sends a default apology message or logs the failure. - The budget prevents infinite loops or runaway costs, which is something ad-hoc agent loops might suffer from. In summary, this Persona Communication workflow shows a realistic use-case where **deterministic orchestration of LLM-powered services** adds significant value. It ensures that two different knowledge domains (technical vs business) can interact via AI intermediaries in a structured manner. The *stateflow service* makes it feasible to design such an interaction as a series of controlled steps, rather than leaving the entire conversation flow to an unpredictable AI agent. Each step's outcome is evaluated and can trigger specific next steps, which is exactly the kind of fine-grained control we need for enterprise applications. ## Additional Workflow Examples Beyond the persona translation scenario, the LLM Workflows service can support a wide range of other workflows. Here are a few **example use-cases** to demonstrate its versatility: - **Simple LLM Q&A Workflow:** *(\"Single-step answer\")* -- The user provides a query, and the workflow simply calls an LLM to get an answer and returns it. This is essentially a one-step workflow (plus maybe an evaluator or format step). While trivial, it shows how even a simple LLM invocation can be wrapped in the workflow for consistency (logging, budgeting, etc.). For instance, a question-answering bot could be just a workflow with one Task: `call_LLM` (with a certain prompt structure) and then End. - **Knowledge Base Retrieval and Answering:** *(\"Tool-augmented query\")* -- A workflow can integrate a search or database lookup before calling the LLM. Steps might be: (1) take a user question, (2) use a custom action `call_web_search` or `query_knowledge_graph` to retrieve relevant info, (3) feed the results into an `call_LLM` step that formulates an answer using those results, (4) maybe an evaluator step to check confidence or filter out any disallowed content. This deterministic sequence ensures the LLM's answer is grounded in retrieved data (addressing factuality), and each part is controlled (for example, if the search returns nothing, we could have a conditional branch to skip the LLM call and respond with "no data found"). - **Automated Code Assistant Workflow:** -- Consider a developer asking for code assistance. The workflow could involve multiple specialized steps: (1) `call_LLM` with a \"planner\" prompt that breaks down the request (e.g., "write a function to do X") into tasks, (2) loop through sub-tasks where for each task we call either a coding LLM to generate code or a testing tool to verify the code, (3) integrate results, (4) evaluate final code. For example, the first LLM might produce a pseudocode or list of steps, the workflow then calls a code-generation model to implement each step, then calls a compilation or test action to check it, if a test fails perhaps branch to a debugging LLM step, etc. Using a workflow ensures each step (planning, coding, testing, fixing) is done in order and under budget. This is much safer than an autonomous coding agent that might go into an infinite loop or execute code unsafely -- our workflow can explicitly restrict what happens (like only allow running tests in a sandbox, etc.). - **Content Moderation Pipeline:** -- An enterprise might use a workflow to filter and respond to user-generated content. For example: (1) `moderation_model` step (could be an LLM or a dedicated model) to classify a piece of text (is it hate speech, spam, etc.?), (2) a `choice` state that branches: if content is OK, proceed to next step, if not OK, go to a rejection message step or escalation, (3) maybe an `auto-response` step that uses an LLM to draft a polite reply or explanation, (4) an `approve` step where either automatically send the reply or require a human approval (this could be an integration point where the workflow pauses until a human intervenes -- possible by having the orchestrator not call the next step until a signal). This kind of workflow could automate moderation while keeping humans in the loop for tough cases, all defined by policy in the blueprint. - **Incident Response Workflow:** -- In a cybersecurity context (relevant to OWASP), imagine a workflow triggered by a security alert. Steps could be: (1) parse the alert details (maybe using regex or an LLM to summarize), (2) `choice` to categorize severity, (3) if severe, call a script or API to isolate affected systems, (4) call LLM to draft a notification email to the IT team or management, (5) log the incident to a database. Each of these is a deterministic step. The LLM is used in a constrained way (only to generate the email text), while decisions like \"if severe then isolate systems\" are hard-coded in the flow logic (not left to the AI). This ensures that important actions (like isolating systems) happen exactly when they should according to a predefined protocol, but we still benefit from AI in parsing and communicating information. - **Multi-Language Customer Support:** -- A customer writes in with a query in language X. The workflow: (1) detect language (maybe a small model or library call), (2) if not English, `translate` to English via an LLM or translation API, (3) use an LLM to draft an answer in English, referencing a support knowledge base if needed, (4) translate the answer back to the customer's language, (5) send the reply via an API. This workflow uses two LLM calls (one to answer, and possibly the same or another to translate) and ensures the final answer is in the customer's language. By orchestrating it, we can ensure translation happens both ways and include fallback steps (if translation fails, perhaps route to a human agent). It's a controlled agent that can autonomously handle many support tickets in multiple languages without ever deviating from the defined process. These examples scratch the surface. Essentially, any time we want an **LLM or AI-driven process with multiple steps** and we care about controlling those steps, this service can help. It provides the skeleton to plug in various AI and non-AI functions into a flowchart of actions. By keeping the workflows **declarative** and using this service, organizations can codify complex procedures that involve AI into a form that's *transparent, testable, and tunable*. Need to change the persona or the prompt? Just update the blueprint. Want to add a step to log to a new database? Add it to the blueprint. Because the execution is isolated per step, these modifications won't affect other steps' correctness. This modularity and clarity is much harder to achieve if one tries to hard-code logic intermingled with LLM prompts in a single blob. Our service enforces good separation of concerns. ## Implementation Considerations (Python & OSBot Framework) The service will be implemented in **Python 3.11+** (per OSBot-Fast-API requirements[\[18\]](https://pypi.org/project/osbot-fast-api/#:~:text=,18)) using the **OSBot-Fast-API** library and its serverless extension. Here we outline how we leverage these technologies and other implementation details: - **OSBot-Fast-API for Type Safety:** OSBot-Fast-API provides a strong type-safe layer on top of FastAPI[\[3\]](https://pypi.org/project/osbot-fast-api/#:~:text=A%20Type,AWS%20Lambda%20integration%20through%20Mangum). We will define the data models for our workflow blueprint (and related requests/responses) as classes using OSBot's `Type_Safe` or Pydantic. This ensures that when a workflow JSON is received, it's automatically validated against our schema. It also makes it easy to return structured responses. For instance, the `/execute_step` endpoint can be defined to accept a `WorkflowStepExecutionRequest` object (containing the blueprint or reference and current step data) and return a `WorkflowStepResult` object. OSBot-Fast-API will handle converting those to JSON for us, and we can be confident in the structure. Strong typing will catch mistakes early -- e.g., if a transition refers to a step that doesn't exist, we can detect that when loading the workflow. - **Serverless Deployment:** OSBot-Fast-API-Serverless enables deploying the FastAPI app to AWS Lambda easily. Our service, being stateless and lightweight in memory (each step execution is quick), is a good candidate for serverless. We could deploy the entire service as a Lambda behind an API Gateway. Each step execution call would be a separate invocation (which is fine given statelessness). The benefits: scaling automatically with load and zero server maintenance. We do need to be mindful of cold start times (Python Lambdas can have a few seconds cold start; using smaller models or warming mechanisms might be considered if needed). Also, token-based LLM calls can be slow (hundreds of milliseconds to seconds), but those are external API waits; the Lambda timeout should be set sufficiently high to allow an LLM call to complete (perhaps 30 seconds or more for large prompts). For now, we focus on functionality, and we have the flexibility to also run the service in a container or on a VM if needed (FastAPI is versatile). - **Integration of LLM APIs:** We will integrate with LLM providers via their Python SDKs or HTTP APIs. Likely, OpenAI's API (for GPT-4 or others) will be used initially (assuming we have keys). Calls to these will happen inside the engine's action dispatch. We must handle errors (network issues, rate limits) gracefully -- possibly by catching exceptions and either retrying (with backoff) or moving to a failure state. We should also use streaming responses only if necessary; otherwise synchronous calls returning the full output are simpler. For local LLMs or alternative providers, we can abstract the LLM call behind a common interface so that switching out is easy (for example, have a `LLMService` class with a method `generate(model, prompt)` that can call OpenAI, or HuggingFace pipeline, etc., based on configuration). - **Testing Workflows:** Because of the deterministic nature, we can unit test workflows by simulating the orchestrator. For a given blueprint, we can run the engine step by step and assert the final outcome or intermediate states. We can also create dummy action handlers for tests (e.g., instead of calling the real LLM API, use a stub that returns a fixed string or uses a local small model). This way, we can verify logic (like branching) without external dependencies. OSBot-Fast-API's testing utilities (like the built-in test server[\[4\]](https://pypi.org/project/osbot-fast-api/#:~:text=%2A%20Type,%E2%86%94%20BaseModel%20%E2%86%94%20Dataclass%20conversions)) will help in writing these tests. Each workflow example can have an automated test case that ensures it runs to completion and yields expected evaluator scores, etc., which is important for continuous integration. - **Performance and Caching:** Calling LLMs is the slowest part. We might implement caching at the step level. For example, if the same `persona.translate` is called with the exact same input frequently, we could cache the result to avoid redundant API calls. This could be done in-memory (for a single lambda invocation sequence) or even persisted (like a small Redis or DynamoDB cache keyed by input). However, caching needs to be designed carefully (e.g., LLM outputs might not be identical every time unless using deterministic prompts). Perhaps more straightforward is to avoid duplicate calls in the same workflow execution -- since a single workflow might reuse a result anyway via variables. - **Choosing Standards Libraries:** For state machine logic within Python, we might use or draw inspiration from libraries like `transitions` (a Python state machine library) or others, but given our custom requirements, we will likely implement the transition handling ourselves. The logic is not too complex given a proper data structure for the blueprint. - **Error Handling and Recovery:** If a step errors out (throws an exception, or returns a response indicating failure), the engine should catch that and either: - Move to a predefined error state if the blueprint defined one (like how AWS Step Functions has a `Catch` mechanism). - Or return an error back to the orchestrator. Since orchestrator is external, it\'s probably better to handle it within the workflow. We can allow steps to have a `on_error: ` field. This way, e.g., if an LLM call fails, we go to a specific step (maybe an apology message, or a cleanup). If no on_error is specified, the engine can return an error code and the whole workflow aborts. This aspect should be defined in our schema for completeness. - Logging the error is important regardless, for debugging. - **Collaboration and Iteration:** As development proceeds, the team (human developers) will refine the blueprint schema and engine logic, often with the assistance of LLMs for ideas or troubleshooting. This synergy will continue as we implement new features. For instance, if we want to introduce a new kind of step (say, a Parallel step to do two things concurrently), we might consult resources or have an LLM suggest how to implement thread pools or async patterns in FastAPI. However, all changes will be reviewed and tested by developers to ensure they meet the determinism and security criteria. - **Documentation and Accessibility:** We will document the schema and usage of the service thoroughly (potentially even auto-generating part of the docs from the schema, similar to how OpenAPI does for APIs). Since Dinis Cruz and the team are building this in the open (likely on GitHub), the documentation will credit the contributions of both the developers and the AI (ChatGPT) that helped along the way. This technical brief itself can serve as a living document to guide implementation, and as we integrate feedback and real-world testing, the design may be adjusted. The flexibility of our approach (thanks to Python and JSON) means we can iterate quickly. In conclusion, the **LLM Workflows/Stateflow Service** is a cutting-edge approach to making LLM-based systems more robust, transparent, and controllable. By blending established workflow orchestration concepts with the latest AI capabilities, and implementing it with modern Python frameworks, we aim to create a service that developers and AI systems can **collaboratively use and improve**. It will empower the creation of AI-driven applications that have the creativity of LLMs *and* the reliability of traditional software -- a combination that is increasingly essential in high-stakes applications[\[19\]](https://arxiv.org/html/2508.02721v1#:~:text=process,demonstrate%20that%20the%20Source%20Code)[\[20\]](https://arxiv.org/html/2508.02721v1#:~:text=Foundation%20Model%20is%20thus%20strategically,agents%20in%20handling%20OOM%20errors). With this foundation, we anticipate a new class of solutions where humans specify the *roadmap* (workflow) and AI fills in the *details* (content), all under a structure that ensures safety and effectiveness. [\[1\]](https://arxiv.org/html/2508.02721v1#:~:text=Agent%20framework,as%20parsing%20an%20error%20log) [\[19\]](https://arxiv.org/html/2508.02721v1#:~:text=process,demonstrate%20that%20the%20Source%20Code) [\[20\]](https://arxiv.org/html/2508.02721v1#:~:text=Foundation%20Model%20is%20thus%20strategically,agents%20in%20handling%20OOM%20errors) Blueprint First, Model Second: A Framework for Deterministic LLM Workflow [\[2\]](https://www.zenml.io/blog/langgraph-vs-autogen#:~:text=LangGraph%20is%20ideal%20for%20building,to%20debug%20and%20guarantee%20behavior) LangGraph vs AutoGen: How are These LLM Workflow Orchestration Platforms Different? - ZenML Blog [\[3\]](https://pypi.org/project/osbot-fast-api/#:~:text=A%20Type,AWS%20Lambda%20integration%20through%20Mangum) [\[4\]](https://pypi.org/project/osbot-fast-api/#:~:text=%2A%20Type,%E2%86%94%20BaseModel%20%E2%86%94%20Dataclass%20conversions) [\[18\]](https://pypi.org/project/osbot-fast-api/#:~:text=,18) osbot-fast-api · PyPI [\[5\]](https://github.com/woodyhayday/FlowSpec#:~:text=%2A%20BPMN%202.0%3A%20An%20industry,both%20visual%20and%20code%20views) [\[6\]](https://github.com/woodyhayday/FlowSpec#:~:text=,workflow%20definition%20tool%20integrated%20with) [\[7\]](https://github.com/woodyhayday/FlowSpec#:~:text=,both%20visual%20and%20code%20views) [\[8\]](https://github.com/woodyhayday/FlowSpec#:~:text=,and%20scientific%20computing%2C%20emphasizing%20reproducibility) [\[9\]](https://github.com/woodyhayday/FlowSpec#:~:text=,excel%20in%20containerized%20CI%2FCD%20pipelines) [\[10\]](https://github.com/woodyhayday/FlowSpec#:~:text=,to%20their%20specific%20automation%20needs) [\[11\]](https://github.com/woodyhayday/FlowSpec#:~:text=A%20lightweight%20standardized%20JSON%20schema,step%20workflows) [\[12\]](https://github.com/woodyhayday/FlowSpec#:~:text=,) [\[13\]](https://github.com/woodyhayday/FlowSpec#:~:text=%2A%20BPMN%202.0%3A%20An%20industry,intensive%20pipelines%20and) [\[17\]](https://github.com/woodyhayday/FlowSpec#:~:text=,to%20their%20specific%20automation%20needs) GitHub - woodyhayday/FlowSpec: FlowSpec: Automation Workflow Schema - A lightweight JSON schema for defining automations and multi-step workflows. Designed for AI Automation Workflows [\[14\]](https://states-language.net/#:~:text=%7B%20,1%3A123456789012%3Afunction%3AHelloWorld%22%2C%20%22End%22%3A%20true) [\[15\]](https://states-language.net/#:~:text=When%20this%20state%20machine%20is,or%20a%20runtime%20error%20occurs) [\[16\]](https://states-language.net/#:~:text=,true%20%7D%20%7D) Amazon States Language --- *This page is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).* ============================================================================== PAGE: /2025/08/22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.html (markdown twin: /2025/08/22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.md) ============================================================================== *[diniscruz.ai](/index.md) · site v0.1.1 · canonical: https://diniscruz.ai/2025/08/22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.html* # Iterative Flow Development (IFD) Methodology: JavaScript Web Application Implementation *By Dinis Cruz and ChatGPT Deep Research and Claude Opus 4.1 · 2025-08-22* > Generative AI (GenAI) is rapidly transforming software development, enabling new coding paradigms where human developers collaborate with Large Language Models (LLMs) as intelligent assistants. --- [PDF](https://files.diniscruz.ai/github/pdf/2025/08/22/iterative-flow-development-ifd-methodology__javascript-web-application-implementation.pdf) · [LinkedIn article](https://www.linkedin.com/pulse/introducing-iterative-flow-development-ifd-how-i-built-dinis-cruz-ciice/) ## Introduction Generative AI (GenAI) is rapidly transforming software development, enabling new coding paradigms where human developers collaborate with Large Language Models (LLMs) as intelligent assistants. However, harnessing LLMs effectively for **production-ready** software requires more than ad-hoc "AI coding" – it demands a disciplined methodology that preserves software engineering rigor while maximizing developer creativity and flow. **Iterative Flow Development (IFD)** is a new methodology developed through collaboration between human expertise and AI assistance that addresses this need. IFD builds upon prior concepts like "vibe coding" – the idea of coding guided by an AI partner in a creative flow state – but adds robust structure and quality control. The core goal is to maintain the developer's **flow state** and focus on user experience (UX) while leveraging the speed of AI code generation, ultimately delivering working software in extremely short cycles. This white paper provides a comprehensive overview of IFD's philosophy, architecture, and workflows, and demonstrates its advantages in productivity, cost, and software quality over traditional development methods. > **Note on Scope:** While this document primarily uses JavaScript web application examples to illustrate IFD concepts, the methodology itself is technology-agnostic. IFD principles can be applied to backend services (Python, Go, Node.js), mobile applications, data pipelines, ML models, or any software development context where rapid iteration with AI assistance is beneficial. The focus on JavaScript and Web Components here is chosen for its accessibility and the case study's specific implementation. ## Bridging Two Worlds: Vibe Coding and Professional Development A defining characteristic of IFD is its ability to operate in two distinct yet complementary modes, making it accessible to both non-technical innovators and professional developers. This dual-mode flexibility addresses a critical gap in modern AI-assisted development: how to harness the creative speed of "vibe coding" while maintaining the engineering rigor required for production software. In **vibe coding mode**, users with no programming experience can build functional applications by simply describing what they want in natural language. The AI directly creates and modifies code in the development environment, allowing rapid experimentation and immediate visual feedback. This democratizes software creation, enabling domain experts, designers, and business stakeholders to transform ideas directly into working prototypes without writing a single line of code. In **air-gapped mode**, professional developers maintain deliberate separation between the AI and their codebase. They use AI as a powerful code generation assistant but manually review, modify, and integrate all suggestions. This preserves code ownership, ensures security, and maintains architectural integrity while still benefiting from AI acceleration. Critically, IFD treats both modes as first-class approaches, not as "amateur" versus "professional" paths. Organizations can leverage vibe coding for rapid prototyping and requirement validation, then seamlessly transition to air-gapped development for production hardening. Business users might create versions v0.1 through v0.3 via vibe coding, exploring ideas and validating user experience, while developers consolidate these experiments into a production-ready v1.0 using air-gapped techniques. This methodology thus serves as a bridge between the democratization promise of AI-assisted development and the quality requirements of professional software engineering. It enables organizations to leverage the domain expertise of non-technical team members while ensuring that production systems meet professional standards for security, performance, and maintainability. ## Philosophy and Principles IFD's philosophy centers on keeping the developer in an optimal **flow state** – a state of uninterrupted focus and creativity – throughout the development process. In practice, this means minimizing context-switching and letting the developer concentrate on high-level design, UX, and business logic, while the LLM handles repetitive boilerplate coding tasks. This approach echoes the spirit of "vibe coding," where developers follow an intuitive, UX-driven coding *vibe* with AI assistance. IFD formalizes that intuition with guiding principles to ensure **engineering discipline** and reliability: • **Flow State Preservation:** The process is designed to avoid breaking the developer's concentration. The developer communicates desired features in natural language and the LLM generates code suggestions, allowing rapid iteration without jumping between disparate tools. By keeping the momentum on solving user-facing problems, IFD sustains creativity and momentum. Technical details (like syntax or boilerplate) are offloaded to the AI, and *never* allowed to block creative thinking. • **Dual-Mode Flexibility:** IFD uniquely supports two complementary workflows that serve different needs and skill levels. In **vibe coding mode**, non-technical users can build functional applications by describing what they want in natural language, with the AI directly creating and modifying code in the development environment – no coding expertise required. The user simply accepts or rejects changes, focusing purely on functionality and UX. In **air-gapped mode**, professional developers maintain deliberate separation between the AI and codebase, manually reviewing and integrating AI-generated code to ensure quality, security, and architectural coherence. This dual nature makes IFD a bridge between rapid business-driven prototyping and professional software engineering. Teams can even transition between modes as projects mature: starting with vibe coding for rapid exploration, then switching to air-gapped development for production hardening. • **UX-First Development:** The developer's flow is anchored in the end-user experience from the start. Before writing code, IFD advocates sketching the user journey and defining UX success criteria. Developers then describe the intended UX to the LLM (for example, "a chat interface with resizable textarea, live character count, and send-on-Enter"), letting the AI draft the initial UI implementation. This ensures that **user experience drives development** rather than technical infrastructure. The developer iteratively refines the AI's output to meet UX quality, e.g. adding input validation or focus handling that the LLM's first pass missed. This UX-centric, iterative design echoes the "vibe coding" focus on getting the *feel* right early, but with systematic refinement steps. • **Real-Data-First Approach:** IFD promotes using **real APIs and data from day one**, with no mocked data or stubbed services. This principle ensures the software is built and tested against real-world conditions, catching integration issues or API misunderstandings immediately. The backend (e.g. a FastAPI service) is created at project start and is called by the UI from the first version. Developers are encouraged to test API endpoints in isolation (e.g. with `curl` or http clients) before integrating them into the frontend, verifying that the real service behaves as expected. By getting **instant feedback from real systems**, the team avoids the common pitfall of code that works on fake data but breaks in production. This approach also enables *cache-aware* development – IFD components consider caching and performance from the beginning since they deal with real data volumes and latency. • **Version Independence**: Major versions in IFD are self-contained releases of the application that fully work on their own. The development follows a two-tier versioning approach: minor versions (e.g., v3.1, v3.2, v3.3...) evolve incrementally within a shared codebase, accumulating features and improvements, while major versions (v1.0, v2.0, v3.0, v4.0) are standalone extractions that contain no dependencies on previous versions.The transition from the last minor version to a major version (e.g., v3.6 to v4.0) is purely a **consolidation and packaging** step with absolutely no functional or logical changes. All end-to-end tests and integration tests should pass identically between these versions. This "publishing" step locks in all incremental changes without introducing new variables. The last minor version (e.g., v3.6) is what QA teams, business users, and product owners sign off for release – the major version (v4.0) is simply that same code, extracted and made standalone. • **Progressive Enhancement:** IFD embraces an **incremental build-up of features**. The first iteration (v0.1) intentionally implements only the core Minimum Viable Product (MVP) functionality, nothing more. Subsequent versions (v0.2, v0.3, etc.) each add a focused set of enhancements or new features on top of the previous conceptual foundation. Importantly, features must prove their value in these 0.x versions *before* they are consolidated – experimental features that don't pan out can be discarded in a later version without affecting the main line. Only once a feature has been validated through use (and possibly iterated on through multiple versions) is it merged into the production candidate v1.0. This guards against over-engineering or premature optimization. The methodology explicitly warns against common pitfalls like trying to build complex future features in v0.1 or adding too many features at once. By focusing each version on one major area (for example, v0.2 might polish UI, v0.3 improve data handling, v0.4 add monitoring, etc.), teams maintain clarity of purpose and high velocity. • **Zero External Dependencies:** A standout principle of IFD is avoiding heavy frameworks or libraries – instead, solutions are built with **native web platform capabilities** only. By using standard ES6+ JavaScript, Web Components, browser APIs, and modern HTML/CSS, the project eliminates external dependency overhead. This yields several benefits: no library version conflicts or upgrades to chase, smaller bundle sizes for performance, and easier debugging since stack traces point to your own code. It also prolongs the longevity of the codebase – there's no risk of a third-party framework becoming obsolete or changing licensing. All techniques rely on evergreen platform features (for example, using `querySelector` and DOM APIs in lieu of jQuery, or the Fetch API instead of Axios). While "zero dependencies" might not fit every scenario, IFD demonstrates that for many apps, the native web platform is powerful enough. This principle reinforces developer skills in web standards and keeps the architecture lean and maintainable. These six principles – Flow State, Dual-Mode Flexibility, UX-First, Version Independence, Real Data, Progressive Enhancement, and Zero Dependencies – form the bedrock of IFD. Underlying them is a respect for both creative development *and* sound engineering. IFD can be seen as a response to the free-form "vibe coding" mindset: it captures the good (preserving flow, fast iteration, creative freedom with AI) while avoiding the bad (lack of structure, fragile code, overlooked quality). By adhering to these principles, IFD aims to deliver **UX-first, high-quality software in record time** without sacrificing maintainability or confidence. ## Methodology: Iterative Development Flow with LLMs IFD defines a clear methodology for how developers work with LLMs and evolve the software through versions. The workflow can be visualized as a continuous loop between the developer's intent and the AI's code generation, with the human remaining the architect and quality guardian at each step. Importantly, IFD supports two distinct operational modes that cater to different users and project phases. ### Two Modes of Development #### Vibe Coding Mode In vibe coding mode, non-technical users or developers seeking maximum speed can build applications through natural language dialogue with the AI. The AI has direct access to the development environment and automatically creates or modifies code based on the user's descriptions. This mode follows this cycle: 1. The user describes desired functionality or improvements in natural language 2. The AI directly generates and integrates code into the current version 3. Changes appear immediately in the development environment 4. The user tests the functionality and provides feedback 5. The AI refines based on observed results and user input This mode enables business stakeholders, designers, and domain experts to create functional prototypes without coding knowledge. They simply "vibe" with the AI, accepting or rejecting changes, focusing entirely on whether the application does what they want. #### Air-Gapped Mode In air-gapped mode, professional developers maintain deliberate separation between the AI and the codebase. This mode provides greater control, security, and code quality assurance. The workflow follows: 1. The developer conceives a solution and describes it to the LLM 2. The LLM produces code suggestions or stubs 3. **The developer manually reviews and integrates** the AI-generated code (maintaining the "air gap") 4. The new code is run and tested against real backend/data 5. The developer refines the code through additional prompts or manual editing This air gap has several advantages: it forces clear requirement articulation, prevents blind trust in AI output, maintains developer ownership, and encourages batching changes into meaningful chunks. Rather than using an AI plugin that directly modifies code, IFD advocates keeping this deliberate separation: the developer interacts with the LLM (e.g. via a chat interface or API) and then manually transfers the generated code into the project. This forces the developer to clearly think through and articulate requirements (since you have to describe the needed code in prose), prevents blindly trusting the AI – the human must review and integrate the code, maintaining ownership – and encourages batching changes into meaningful chunks rather than constant one-line edits. In practice, this means the developer might work in an IDE like PyCharm or VSCode, and separately have the ChatGPT/Claude interface where they prompt for code, then copy results into their files. The slight friction of the air gap ironically **improves** efficiency by reducing churn and encouraging more thoughtful prompts. ### Mode Selection and Transition Teams typically choose modes based on: - **Project phase**: Vibe coding for initial exploration, air-gapped for production - **User expertise**: Non-coders use vibe mode, developers may prefer air-gapped - **Security requirements**: Critical systems demand air-gapped review - **Speed vs. control tradeoff**: Vibe for maximum velocity, air-gapped for maximum confidence Projects often transition between modes. A common pattern: - **v0.1-v0.3**: Business users rapidly prototype via vibe coding - **v0.4-v0.5**: Developers review and enhance in air-gapped mode - **v1.0**: Professional consolidation using air-gapped approach This flexibility allows IFD to serve as a bridge between business innovation and engineering rigor. ### Feature Iteration Process Regardless of mode, each **feature iteration** in IFD follows a similar conceptual loop focused on rapid feedback and continuous improvement. The key difference is whether the AI directly modifies code (vibe mode) or provides suggestions for manual integration (air-gapped mode). In both modes, iterations complete rapidly – often in minutes for small features. By structuring work into micro-iterations, IFD enables continuous feedback and prevents analysis-paralysis. The motto remains: *"iterate rapidly without overthinking"*. #### Entering the Flow To start an IFD project, preparation ensures productive flow regardless of chosen mode. The **pre-development checklist** includes having a clear project vision and problem definition, identifying target users and core features for the MVP (v0.1), and preparing the development environment. A simple FastAPI backend should be running with at least skeleton endpoints for core functionality (since the frontend will call real APIs from the outset). **For Vibe Coding Mode:** When operating in vibe coding mode, you'll need to set up the AI with direct access to your development environment, allowing it to create and modify code directly based on your natural language descriptions. Prime the AI with comprehensive project context and constraints at the start of each session, including your project goals, technical requirements, and any specific patterns to follow. Ensure that your real backend services and APIs are running and accessible, as the AI will be generating code that calls these endpoints from the outset. Most importantly, focus your communication on describing the desired outcomes and user experience rather than implementation details – let the AI handle the technical translation while you concentrate on what the application should do and how it should feel to users. **For Air-Gapped Mode:** In air-gapped mode, begin by preparing your development environment with your preferred IDE and version control system, ensuring you have a comfortable workspace for reviewing and integrating code. Set up a separate AI interface such as ChatGPT or Claude in a browser or dedicated application, maintaining deliberate separation between the AI and your codebase. Develop a practice of creating structured prompts that include clear technical requirements, context, constraints, and success criteria – treating prompt writing as a form of technical specification that will yield better AI outputs. Establish a consistent workflow for reviewing AI suggestions, integrating selected code into your project, testing immediately, and iterating based on results, ensuring you maintain full control and understanding of every piece of code that enters your codebase. Both modes require "priming" the LLM with project context – this might mean providing a summary of the project's goal and any relevant technical constraints at the start of the LLM session. For example, the developer might feed the LLM a brief like: *"I'm building a single-page text analysis app. Constraints: pure JavaScript (ES6), custom Web Components only, FastAPI backend at /api, no external libraries."* This context setting is crucial for effective AI assistance. The methodology recommends **structured LLM briefs** that outline the context, technical constraints, and specific task at hand, including success criteria for the feature. By providing this upfront clarity, the developer ensures the LLM's output aligns with the overall architecture and requirements. #### Version-by-Version Workflow Development in IFD proceeds through a **two-tier versioning system**: minor versions that evolve incrementally within a major version series, and major versions that represent standalone production releases. The typical progression follows this pattern: **Minor Versions (Incremental Development):** - **v0.1, v0.2, v0.3...** through **v0.n**: Incremental development within a shared codebase - **v1.1, v1.2, v1.3...** through **v1.n**: Post-release patches and features - **v2.1, v2.2, v2.3...** through **v2.n**: Next major feature set development **Major Versions (Standalone Releases):** - **v1.0**: First production release (consolidated from v0.n) - **v2.0**: Second major release (consolidated from v1.n) - **v3.0**: Third major release (consolidated from v2.n) Within a major version series, minor versions share a codebase and build incrementally upon each other. Each minor version adds features, fixes bugs, or improves existing functionality. The codebase evolves continuously, with each minor version being potentially shippable. Experiments and alternative implementations are managed through **Git branches** or **feature toggles**, not by creating separate version folders. When transitioning from the last minor version to a major version (e.g., v0.9 to v1.0, or v2.6 to v3.0), the process is purely administrative: 1. **Code Extraction:** The last minor version's code is copied to a new, standalone directory 2. **Dependency Cleanup:** Any references to previous versions are removed (though there shouldn't be any) 3. **Documentation Update:** Version numbers and release notes are updated 4. **Test Verification:** All existing tests pass without modification 5. **Stakeholder Sign-off:** QA, business users, and product owners approve the last minor version before it becomes the major release **No functional or logical changes occur between the last minor version and the major release.** If v0.9 is the last minor version before v1.0, then v1.0 is functionally identical to v0.9 – it's simply packaged as a clean, standalone release. This ensures that what stakeholders approve is exactly what gets released, with no last-minute surprises or integration issues. Each version sits in its own directory (e.g. `/versions/v0.1/`, `/versions/v0.2/`, etc.), containing all the code and assets for that iteration. Crucially, earlier version directories are never modified once created – new versions might copy code from them, but do not create interdependencies. This enforces the **"no shared code between versions"** rule. If, for example, a developer wants to reuse a component from v0.1 in v0.2, they copy the file forward into the v0.2 folder rather than importing it across versions. While this duplicates code, it prevents tangled dependencies and allows each version to evolve freely (or be discarded) without impacting others. Every version is expected to be **complete and functional on its own**, with no reliance on files in other version directories. This means each version folder might have its own `index.html`, its own set of components, styles, and utilities. IFD provides clear file organization guidelines for this; for example, a version folder may contain a structured sub-tree of components, services (for API clients), utils, and CSS, all self-contained. **In Vibe Coding Mode:** The following describes how version management and code generation work when using vibe coding mode for rapid prototyping: - **Automatic Version Management:** AI automatically creates new version directories when you request new iterations, handling all file organization without manual intervention - **Natural Language Implementation:** User describes features in plain language, and the AI implements them directly in the codebase without requiring coding knowledge - **Seamless Version Isolation:** Version isolation happens automatically behind the scenes, ensuring each iteration remains independent without user configuration - **Functionality-First Focus:** Focus remains purely on whether the application works as intended, with code structure and quality concerns deferred to later consolidation **In Air-Gapped Mode:** These practices ensure controlled, deliberate development when working in air-gapped mode for production-quality code: - **Manual Version Control:** Developer manually manages version directories, creating new folders and copying files forward with full visibility into the structure - **Deliberate Code Migration:** Code is deliberately copied and modified between versions, allowing selective incorporation of proven features and improvements - **Architectural Oversight:** Developer ensures clean separation of concerns and maintains architectural integrity throughout the version progression - **Dual Focus on Function and Quality:** Focus includes both delivering functionality and maintaining code quality standards from the start, balancing speed with sustainability #### Version 0.1: The Foundation The **v0.1** iteration is kept deliberately simple and focused. According to the IFD playbook, v0.1's purpose is to establish the core architecture and solve the primary use-case with minimal extras. A checklist for v0.1 ensures the basics are in place: project structure, one or two core components functioning, basic UI working, and an API call integrated end-to-end. Any tendency to over-engineer at this stage is discouraged – no complex state management, no premature optimization, and definitely no "nice-to-have" features that distract from the core problem. For example, if building a text analysis app, v0.1 might allow a user to input text and get a simple analysis result from the backend. Features like rich UI polish, caching, multi-view dashboards, etc., are left for later versions. This disciplined scoping of v0.1 ensures the team **proves the concept** quickly and establishes a working baseline. #### Subsequent Versions: Continuous Integration With a solid v0.1 in hand, subsequent minor versions (v0.2, v0.3, ...) each incrementally build upon the previous version within the same evolving codebase. Rather than creating isolated experiments in separate folders, the IFD methodology promotes **continuous integration** where each minor version represents the current state of the product with all accumulated improvements. The development pattern for minor versions follows this approach: - **v0.2 – UI Polish:** Improve the user interface based on v0.1, fixing initial bugs, refining layout and styling, adding responsiveness, etc. - **v0.3 – Data Enhancements:** Build upon v0.2 by introducing caching mechanisms, input validation, better handling of data outputs, etc. - **v0.4 – Monitoring & Logging:** Add to v0.3 with analytics, logging of user actions, and performance metrics for debugging - **v0.5 – Advanced Features:** Enhance v0.4 with more complex capabilities or integrations - **v0.9 – Pre-release Candidate:** The final minor version with all features integrated, tested, and ready for stakeholder approval Each minor version must be **potentially shippable** – it should work completely and could theoretically be deployed to production. This discipline ensures continuous quality and prevents the accumulation of half-finished features. **Managing Experiments and Variations:** When exploring different approaches (e.g., alternative UI designs or competing algorithms), teams should use: 1. **Feature Toggles:** Multiple implementations can coexist in the same codebase, controlled by configuration flags. This allows A/B testing and gradual rollout without code divergence. 2. **Git Branches:** Experimental features are developed in separate branches and only merged into the main minor version line when proven valuable. Failed experiments simply have their branches deleted, never polluting the main codebase. 3. **Progressive Enhancement:** Rather than replacing functionality between versions, new features are added alongside existing ones, with deprecated features removed only after their replacements are proven. The key principle is that by the time a minor version series is ready for major version consolidation (e.g., v0.12 becoming v1.0), all experiments have been resolved, all chosen features are integrated and working together, and the codebase represents a cohesive whole rather than a collection of competing alternatives. #### Version 1.0: Consolidation and Production Readiness #### Version 1.0: Publishing for Production After the minor version iterations reach a stable, feature-complete state, **v1.0** represents the formal production release. Critically, v1.0 is **not a development phase** – it's a publishing and packaging step that creates a standalone version from the last minor iteration. The transition from the last minor version (e.g., v0.9) to v1.0 involves: 1. **Stakeholder Sign-off:** QA teams, business users, and product owners review and approve v0.9 (or whatever the last minor version is). This is the version they test, validate, and approve for production release. 2. **Code Extraction:** The approved minor version's code is copied to a new v1.0 directory, creating a completely standalone codebase with no dependencies on any previous versions. 3. **Documentation and Metadata:** Version numbers are updated, release notes are finalized, and deployment configurations are set for production. 4. **Test Verification:** All existing end-to-end and integration tests are run against v1.0 to verify they pass identically to the last minor version. **No test changes should be needed** – if tests need modification, that's a red flag that functional changes have crept in. 5. **Final Packaging:** The v1.0 directory becomes the deployable artifact, containing everything needed for production deployment. **Absolutely no functional or logical changes occur during this consolidation.** The v1.0 code should be functionally identical to v0.9. Any bugs, features, or improvements discovered after sign-off are deferred to v1.1 (the first minor version of the next series). This approach has several critical benefits: - **What you test is what you deploy:** Stakeholders approve a working system (v0.9), not a theoretical merge - **Zero integration risk:** Since all features were already integrated in minor versions, there's no last-minute integration surprises - **Clear rollback path:** If issues arise, you can return to any previous major version - **Clean codebase:** Each major version is self-contained, making maintenance and understanding easier The quality criteria for v1.0 remain high, but these are **achieved during minor version development**, not added during consolidation: - **Clear Separation of Concerns:** Already established in minor versions - **Thorough Documentation:** Accumulated throughout minor version development - **Performance Benchmarks:** Met and verified in the final minor versions - **No Major Known Bugs:** Resolved during minor version iterations Essentially, v1.0 is the polished, standalone packaging of what was already proven to work in v0.9, containing only tested and integrated features, with all experimental code either incorporated or discarded during the minor version progression. ### LLM Collaboration and Prompting A cornerstone of the IFD workflow is effective use of the LLM as a **pair programmer**. Rather than writing boilerplate or routine code, the developer delegates those to the AI through well-crafted prompts. IFD documentation provides **LLM workflow templates** for common development tasks to streamline this communication. #### Prompt Templates and Patterns For example, when creating a new component, a template prompt provides structured sections to ensure comprehensive AI understanding: - **Component Purpose:** A clear statement of what the component does and why it exists in the application architecture - **Technical Requirements:** Specific constraints like "use ES6 class extending HTMLElement, no external dependencies, include its own CSS, and integrate with API endpoints X and Y" - **Desired Functionality:** An itemized list of features the component must support, from core behaviors to edge cases - **UI Requirements:** Visual and interaction specifications including layout, styling needs, and responsive behavior - **Events to Handle:** Both DOM events (clicks, input changes) and custom events the component should emit or listen for By supplying a structured prompt with sections (Context, Requirements, etc.), the developer ensures the LLM is aware of the important details. This often yields a surprisingly complete initial implementation from the AI – including not just the JavaScript class code, but also a stub of the CSS and how it should be used in HTML. Other templates include: - **Adding features to existing components**: Current component code is pasted in and the prompt describes what new feature to insert - **Debugging**: Provide the error and relevant code, ask the AI to fix it with logging and error handling - **Minor version development**: Incremental feature additions within the evolving codebase These templates encapsulate best practices in prompting so developers can systematically get the most out of the LLM. Essentially, IFD treats prompt-writing as a new form of development art – part of the engineer's skill set is to communicate with the AI clearly and precisely, much like writing a mini design spec, which the AI then turns into code. #### LLM Excellence in Major Version Consolidation LLMs demonstrate particular strength during the minor-to-major version transition (e.g., v3.6 to v4.0), making them ideal partners for the consolidation step. When provided with complete context – the last major version's code plus all subsequent minor versions – LLMs can perform highly reliable refactoring and optimization without the hallucination issues that often plague greenfield code generation. **Why LLMs Excel at Consolidation:** 1. **Complete Context Eliminates Guesswork:** With all the v3.x code available, the LLM has full visibility into every implementation detail, API contract, and component interaction. There's no need to "imagine" how something might work – it's all there in the provided code. 2. **Pattern Recognition Across Versions:** LLMs can identify duplicate code, similar patterns, and optimization opportunities across the entire minor version series, suggesting consolidations that human developers might miss. 3. **Consistent Refactoring:** The LLM can apply consistent code style, naming conventions, and architectural patterns across all components, eliminating the inconsistencies that naturally accumulate during rapid minor version development. 4. **Safe Optimization:** Since the functionality is already proven and working, the LLM can focus purely on code quality improvements: reducing redundancy, improving performance, enhancing readability, and standardizing patterns. **Test-Driven Consolidation Process:** The key to confident LLM-assisted consolidation is maintaining an **immutable test suite** that governs the transition: 1. **Freeze the Test Suite:** Before consolidation begins, lock all end-to-end and integration tests from v3.6. These tests become the unchangeable contract that v4.0 must fulfill. 2. **LLM Consolidation Prompt:** Provide the LLM with: - All code from the last major version (e.g., v3.0) - All code from subsequent minor versions (v3.1 through v3.6) - The frozen test suite as the acceptance criteria - Clear instructions that all tests must pass without modification 3. **Iterative Refinement:** The LLM consolidates the code, potentially through multiple iterations, merging duplicate functionality, standardizing interfaces, optimizing performance, and cleaning up technical debt. 4. **Test Verification:** After each LLM consolidation pass, run the frozen test suite. Any test failure means the consolidation introduced a functional change and must be corrected. 5. **Human Review:** While tests ensure functional equivalence, human review confirms that the consolidated code maintains architectural integrity and follows organizational standards. **Consolidation Prompt Template:** ``` Given the following code: - Last major version (v3.0): [complete codebase] - All minor versions (v3.1-v3.6): [complete codebases] Consolidate these into a clean v4.0 release that: 1. Maintains identical functionality (all existing tests must pass unchanged) 2. Removes code duplication across minor versions 3. Standardizes component patterns and interfaces 4. Optimizes performance where possible 5. Improves code documentation 6. Creates a standalone codebase with no external version dependencies The following test suite must pass without any modifications: [Include all e2e and integration tests] Generate the consolidated v4.0 code structure. ``` This approach transforms the major version consolidation from a risky integration exercise into a controlled optimization process. The LLM handles the mechanical work of merging and refactoring, while the frozen test suite ensures that no functionality is lost or altered. The result is a clean, optimized major version that is functionally identical to the last minor version but with superior code quality and maintainability. #### Maintaining Developer Control Despite heavy use of AI generation, IFD keeps the developer **firmly in charge** of architecture and critical decisions. The developer decides what components exist, how they interact, and when to override the LLM's suggestions. Often the AI's first output will be tweaked by the developer to match the desired UX or to fix small errors. For instance, in a chat interface feature, the AI might output a basic send-button handler; the developer then refines it to trim empty input, maintain focus, and optimistically update the UI for responsiveness. This human-guided refinement is crucial – it ensures the final product has the polish and correctness that pure AI generation might lack. Over time, as the LLM and developer iterate, the code converges to meet all requirements. The **flow state** is maintained because the developer is never stuck on rote coding; they're either describing the next feature to the AI, integrating results, or testing the live app. All of these are engaging tasks closely tied to the problem being solved, rather than fighting with configuration or waiting on builds. ### Testing and Quality Assurance in Flow Testing is woven into the IFD workflow in a very immediate, **real-time** manner. Since every version uses the real backend and data, every manual test exercise yields meaningful results. The methodology encourages developers to test features *as soon as they are implemented* in the browser, clicking through the UI or calling APIs, rather than writing extensive mock-based unit tests upfront. #### Testing in Production Conditions This is not to say automated testing is ignored, but the priority is given to **"testing in production conditions"** from the start. For example, if implementing a text analysis API, the developer would quickly deploy the FastAPI server locally and try actual analysis requests (via the UI or via direct HTTP calls) to see end-to-end behavior. Any errors (exceptions, incorrect responses) would surface immediately and can be addressed by adjusting either the frontend or backend on the spot. This tight loop catches integration issues (like mismatched data formats or CORS problems) early, before they become large debugging tasks. #### Visual and Interactive Testing IFD also advocates for **visual and interactive testing** during development. These patterns provide immediate feedback and validation during the development flow: - **Temporary UI Instrumentation:** Adding console logging to track user actions and state changes, providing a real-time audit trail of application behavior during testing - **Immediate Visual Feedback:** Implementing visible responses like button highlights, loading spinners, or color changes to confirm that user interactions are being registered and processed - **Real Sample Datasets:** Using production-like data samples to test UI components under realistic conditions, exposing layout issues, performance problems, or edge cases early The backend can provide a test data endpoint that returns sample inputs. The frontend component can then loop through these samples and assert that outputs contain expected elements (this is a form of lightweight integration test). Because it's using real data and the real processing logic, such tests increase confidence that the feature truly works, not just in a contrived unit test environment. #### Progressive Refactoring Across Versions When it comes to ensuring robustness, IFD relies heavily on **progressive refactoring** across versions. The approach follows three stages: 1. **"Make it work"** (early versions) – Focus on delivering functionality that meets requirements, even if the code is not perfect 2. **"Make it right"** (middle versions) – Refactor for clarity, maintainability, and edge-case handling 3. **"Make it fast"** (later versions) – Optimize performance-critical parts (adding caching, debouncing, etc.) This staged approach is exemplified by a simple feature's evolution: - Initial click handler may directly call `fetch` and dump results to the DOM (stage 1) - Later, it is rewritten to validate input, use async/await and proper error handling (stage 2) - Later still, it's optimized with caching and debouncing to handle rapid or repeated inputs efficiently (stage 3) By spreading out these improvements over versions, IFD ensures that at each stage the codebase is working and delivering value, and only then invests in polishing it. This reduces wasted effort on premature optimizations and lets real usage inform where refactoring is needed. #### Production Readiness Standards **Production readiness** is not an afterthought in IFD – each version is meant to be potentially shippable, and especially v1.0 is held to strict standards. The methodology defines clear criteria for **code quality, architecture, performance, and maintainability** that should be met: **Code Quality:** - **Consistent Error Handling:** Standardized try-catch patterns and error boundaries throughout the application, with meaningful error messages and graceful fallbacks for user-facing failures - **Memory Management:** Proper cleanup of event listeners, timers, and observers in component lifecycle methods, preventing memory leaks that degrade performance in long-running single-page applications **Architecture:** - **Separation of Concerns:** Each component or service handles a single responsibility, with business logic, presentation, and data access clearly separated into appropriate modules - **Event-Driven Communication:** Components communicate through custom events rather than direct method calls, maintaining loose coupling and enabling independent development - **Stateless Design:** Components favor functional patterns and avoid internal state where possible, making them more predictable and easier to test **Performance:** - **Lazy Loading Strategy:** Heavy components load only when needed using dynamic imports, reducing initial bundle size and improving time-to-interactive - **Action Debouncing:** Rapid user actions like search typing or scroll events are debounced to prevent excessive API calls or expensive computations - **Strategic Caching:** Frequently accessed data and expensive operation results are cached with appropriate invalidation strategies, balancing freshness with performance **Maintainability:** - **Intuitive File Organization:** Consistent directory structure with clear naming conventions that make finding and adding code straightforward for any developer - **Interface Documentation:** Component APIs, expected props, emitted events, and extension points are clearly documented, enabling safe modifications and extensions All of these are supported by code patterns in the IFD architecture guide (e.g., examples of implementing debounce in a search component, or caching API responses in-memory). The idea is that by the time the team has iterated to v1.0, they have baked in a professional level of quality. Any quick-and-dirty aspects from early versions should either have been refactored or left out of the consolidation. The result is a codebase that, despite being produced rapidly with AI help, meets conventional standards for readability and reliability. This is a critical point – IFD does not trade quality for speed, it attempts to deliver both by focusing on **flow** and smart use of AI for grunt work, while the human developers enforce quality through continuous testing and final consolidation. The flexibility to work in either vibe coding or air-gapped mode ensures that teams can adapt the methodology to their specific needs while maintaining the core benefits of rapid, iterative development with LLM assistance. ## Technical Architecture in IFD The architecture prescribed by IFD is deliberately simple and **web-native**, to maximize development agility and long-term maintainability. At its core is a **100% native frontend** built with standard Web Components (custom HTML elements) and modern JavaScript, paired with a lightweight **FastAPI backend** for serving data. This section outlines the key architectural patterns and how they differ from or improve upon traditional frameworks. ### Web Components and Encapsulation IFD projects utilize the Web Components standard (i.e. classes extending `HTMLElement`) as the unit of front-end modularity. Each major UI element or logical piece of the app is implemented as a custom element, encapsulating its own structure, style, and behavior. An example component structure from the IFD guide looks like: a class with a constructor (setting up initial state), a `connectedCallback` to render HTML and initialize events when the element is added to the DOM, and a `disconnectedCallback` to clean up when removed. Within the component's `render()` method, it generates its inner HTML (often injecting a template string) and caches references to important sub-elements for later updates. Event listeners are set up either in `connectedCallback` or a dedicated method, and a `cleanup()` method ensures any event handlers or timers are removed in `disconnectedCallback`. All styling for the component lives in a corresponding CSS file or `