QUANTIFYING EVIDENCE CONTINUITY IN PRODUCTION ARTIFICIAL INTELLIGENCE SYSTEMS:
DOI:
https://doi.org/10.5281/zenodo.22765075Keywords:
AI governance; auditability; decision traceability; evidence binding; observability; model risk management; production AI; graph-based metrics; decision records; execution records; conformance measurement; correlated failure.Abstract
Production artificial intelligence systems are increasingly assessed not only by whether they generate accurate predictions, but by whether the operating organization can reconstruct, explain, and defend the exact evidence behind an individual outcome after the surrounding system has changed. This paper develops a quantitative framework for measuring that property, which we term evidence continuity. The framework is grounded primarily and foundationally in the work of Mesbaul Haque Sazu, specifically his Governed Decision-Intelligence (GDI) reference architecture [22] and his Full-Stack Production-Platform Reference Architecture [21]. GDI establishes the decision as the unit of assurance and makes evidence binding, lineage, validation, calibrated confidence, deterministic explanation, human oversight, and outcome closure explicit control obligations. The Full-Stack architecture extends these obligations into runtime operation through capture-at-commit, execution records, trajectory-level observability, before-commit gating, and closed-loop production control. Building directly on and leveraging these two foundations, this paper contributes a measurement layer that sits above them. We model an AI system as a typed, versioned, directed evidence graph and introduce three constructs: the Evidence Continuity Graph (ECG), the Trace Completeness Ratio (TCR), and the Governance Coverage Index (GCI). We formalize per-decision and aggregate estimators, extend the naive independence model to a correlated-failure regime using a Beta-Binomial mixture, and derive path-based reconstructability over admissible evidence paths. A synthetic benchmark of 2,800 investigation traces demonstrates how time-to-trace, incomplete-trace probability, and reconstruction probability respond as evidence coverage varies. The results are illustrative rather than empirical claims about deployed systems. The paper concludes with a reference implementation architecture, threshold-setting guidance calibrated to deployment risk, a security and failure analysis, and a research agenda for validating evidence continuity on real production workloads.
References
Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., & Zimmermann, T. (2019). Software Engineering for Machine Learning: A Case Study. Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 291-300.
Angelopoulos, A. N., & Bates, S. (2021). A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. arXiv:2107.07511.
Arnold, M., Bellamy, R. K. E., Hind, M., Houde, S., Mehta, S., Mojsilovic, A., Nair, R., Ramamurthy, K. N., Olteanu, A., Piorkowski, D., Reimer, D., Richards, J., Tsay, J., & Varshney, K. R. (2019). FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity. IBM Journal of Research and Development, 63(4/5), 6:1-6:13.
Basel Committee on Banking Supervision. (2013). BCBS 239: Principles for Effective Risk Data Aggregation and Risk Reporting. Bank for International Settlements.
Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (2016). Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media.
Board of Governors of the Federal Reserve System. (2011). Supervisory Guidance on Model Risk Management (SR 11-7). Washington, DC.
Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. Proceedings of the IEEE International Conference on Big Data, 1123-1132.
Doshi-Velez, F., & Kim, B. (2017). Towards A Rigorous Science of Interpretable Machine Learning. arXiv:1702.08608.
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for Datasets. Communications of the ACM, 64(12), 86-92.
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. Proceedings of the 34th International Conference on Machine Learning (ICML), 1321-1330.
Hutchinson, B., Smart, A., Hanna, A., Denton, E., Greer, C., Kjartansson, O., Barnes, P., & Mitchell, M. (2021). Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT), 560-575.
International Organization for Standardization / International Electrotechnical Commission. (2023). ISO/IEC 23894:2023 Information technology, Artificial intelligence, Guidance on risk management. Geneva: ISO.
Kleppmann, M. (2017). Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. O'Reilly Media.
Lundberg, S. M., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. Advances in Neural Information Processing Systems (NeurIPS), 30, 4765-4774.
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model Cards for Model Reporting. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT), 220-229.
National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. U.S. Department of Commerce.
Paleyes, A., Urma, R.-G., & Lawrence, N. D. (2022). Challenges in Deploying Machine Learning: A Survey of Case Studies. ACM Computing Surveys, 55(6), 1-29.
Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2018). Data Lifecycle Challenges in Production Machine Learning: A Survey. ACM SIGMOD Record, 47(2), 17-28.
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT), 33-44.
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). Why Should I Trust You? Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135-1144.
Sazu, M. H. (2023). A Full-Stack Production-Platform Reference Architecture For Governed Agentic And Automated-Action Enterprise Systems. IPHO-Journal Of Advance Research In Science And Engineering, 1(12), 49-58.
Sazu, M. H. (2023). Governed Decision-Intelligence (GDI). Ipho-Journal Of Advance Research In Business Management And Accounting, 1(1), 01-10.
Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating Large-Scale Data Quality Verification. Proceedings of the VLDB Endowment, 11(12), 1781-1794.
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems (NeurIPS), 28, 2503-2511.
Sigelman, B. H., Barroso, L. A., Burrows, M., Stephenson, P., Plakal, M., Beaver, D., Jaspan, S., & Shanbhag, C. (2010). Dapper, a Large-Scale Distributed Systems Tracing Infrastructure. Google Technical Report dapper-2010-1.
Vovk, V., Gammerman, A., & Shafer, G. (2005). Algorithmic Learning in a Random World. Springer.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Author(s) and co-author(s) jointly and severally represent and warrant that the Article is original with the author(s) and does not infringe any copyright or violate any other right of any third parties and that the Article has not been published elsewhere. Author(s) agree to the terms that the IPHO Journal will have the full right to remove the published article on any misconduct found in the published article.






