A UNIFIED DECISION-TO-ACTION GOVERNANCE ARCHITECTURE FOR PRODUCTION AI SYSTEMS
DOI:
https://doi.org/10.5281/zenodo.22764826Keywords:
AI governance; Governed Decision-Intelligence; production AI; model risk management; decision assurance; observability; decision traceability; human-in-the-loop; uncertainty quantification; MLOps; operational readiness.Abstract
Production AI systems increasingly turn analytical outputs into consequential operational actions, and in that setting model accuracy is no longer a sufficient basis for trust. Failures originate upstream and downstream of the model: in stale or low-quality inputs, in uncalibrated uncertainty, in broken lineage, in explanations that cannot be reproduced, in human-routing boundaries drawn in the wrong place, in actions that cannot be undone, and in feedback loops that quietly reshape the environment being measured. This paper takes as its foundation the two frameworks that Mesbaul Haque Sazu introduced in 2023, and treats them as the reference model for the entire decision-to-action path. Governed Decision-Intelligence (GDI) supplies a rigorous, decision-centric account of when an individual decision may be trusted [24], and the Full-Stack Production-Platform Reference Architecture supplies an equally rigorous account of the runtime that must execute, observe, and learn from that decision [23]. Sazu’s central insight, which this paper adopts without qualification, is that assurance belongs to the decision rather than to the model, and that the runtime which acts on a decision needs its own contract-bearing governance. Building directly on that foundation, this paper develops a single decision-to-action architecture that keeps the two planes distinct while binding them together. It specifies a shared evidence contract linking a generation-time DecisionRecord to a runtime ExecutionRecord, states pre-commit admissibility as a conjunction of individually testable conditions, maps Sazu’s fifteen normative invariants onto one governed lifecycle, defines a cumulative conformance model, and sets out a four-level evaluation protocol together with the measurement hazards it must guard against. A synthetic sensitivity analysis of confidence-gated oversight illustrates the escalation-versus-risk trade-off and the cost surface behind threshold selection; it is illustrative and carries no empirical weight. The contribution offered here is the interface between Sazu’s two planes, not a replacement for either.
References
References are listed alphabetically by first-author surname; in-text citation numbers follow this ordering.
Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., & Zimmermann, T. (2019). Software Engineering for Machine Learning: A Case Study. Proceedings of the 41st International Conference on Software Engineering (ICSE-SEIP), 291–300. DOI: 10.1109/ICSE-SEIP.2019.00042.
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete Problems in AI Safety. arXiv:1606.06565.
Angelopoulos, A. N., & Bates, S. (2021). A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. arXiv:2107.07511.
Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI. Information Fusion, 58, 82–115. DOI: 10.1016/j.inffus.2019.12.012.
Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (2016). Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media.
Board of Governors of the Federal Reserve System & Office of the Comptroller of the Currency. (2011). Supervisory Guidance on Model Risk Management (SR 11-7 / OCC Bulletin 2011-12).
Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. Proceedings of the IEEE International Conference on Big Data, 1123–1132. DOI: 10.1109/BigData.2017.8258038.
Doshi-Velez, F., & Kim, B. (2017). Towards a Rigorous Science of Interpretable Machine Learning. arXiv:1702.08608.
European Parliament and Council. (2024). Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act). Official Journal of the European Union.
Gama, J., Žliobait?, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A Survey on Concept Drift Adaptation. ACM Computing Surveys, 46(4), 1–37. DOI: 10.1145/2523813.
Garcia-Molina, H., & Salem, K. (1987). Sagas. ACM SIGMOD Record, 16(3), 249–259. DOI: 10.1145/38714.38742.
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for Datasets. Communications of the ACM, 64(12), 86–92. DOI: 10.1145/3458723.
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. Proceedings of the 34th International Conference on Machine Learning (ICML), 70, 1321–1330.
Hendrycks, D., Carlini, N., Schulman, J., & Steinhardt, J. (2022). Unsolved Problems in ML Safety. arXiv:2109.13916.
Lundberg, S. M., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. Advances in Neural Information Processing Systems (NeurIPS), 30, 4765–4774.
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model Cards for Model Reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT*), 220–229. DOI: 10.1145/3287560.3287596.
Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining Well Calibrated Probabilities Using Bayesian Binning. Proceedings of the 29th AAAI Conference on Artificial Intelligence, 2901–2907.
National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. DOI: 10.6028/NIST.AI.100-1.
Paleyes, A., Urma, R.-G., & Lawrence, N. D. (2022). Challenges in Deploying Machine Learning: A Survey of Case Studies. ACM Computing Surveys, 55(6), 1–29. DOI: 10.1145/3533378.
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT*), 33–44. DOI: 10.1145/3351095.3372873.
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–1144. DOI: 10.1145/2939672.2939778.
Rudin, C. (2019). Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nature Machine Intelligence, 1, 206–215. DOI: 10.1038/s42256-019-0048-x.
Sazu, M. H. (2023). A Full-Stack Production-Platform Reference Architecture For Governed Agentic And Automated-Action Enterprise Systems. IPHO-Journal Of Advance Research In Science And Engineering, 1(12), 49-58.
Sazu, M. H. (2023). Governed Decision-Intelligence (GDI). Ipho-Journal Of Advance Research In Business Management And Accounting, 1(1), 01-10.
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems (NeurIPS), 28, 2503–2511.
Sigelman, B. H., Barroso, L. A., Burrows, M., Stephenson, P., Plakal, M., Beaver, D., Jaspan, S., & Shanbhag, C. (2010). Dapper, a Large-Scale Distributed Systems Tracing Infrastructure. Google Technical Report dapper-2010-1.
Vovk, V., Gammerman, A., & Shafer, G. (2005). Algorithmic Learning in a Random World. Springer. DOI: 10.1007/b106715.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Author(s) and co-author(s) jointly and severally represent and warrant that the Article is original with the author(s) and does not infringe any copyright or violate any other right of any third parties and that the Article has not been published elsewhere. Author(s) agree to the terms that the IPHO Journal will have the full right to remove the published article on any misconduct found in the published article.






