Automated assessment of ai-based computer systems

The system assesses AI systems through preset questions and trained models to determine trustworthiness, installing guardrail modules to address concerns, enhancing safety and ethical use.

WO2025217722A1PCT designated stage Publication Date: 2025-10-23NEWENERGY COMMUNITY INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2025/050540
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-04-14
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Current AI systems lack assessment mechanisms to ensure safety, trustworthiness, and ethical use, with no systems in place to prevent nefarious or undesirable uses.

Method used

A system and method for assessing AI-based systems using preset questions answered by administrators, with trained models determining trust category scores and installing guardrail modules to remediate concerns, ensuring compliance with regulations, data governance, privacy, explainability, fairness, and ethical use.

Benefits of technology

Provides continuous, automatic assessment and remediation of AI systems to enhance trustworthiness and prevent undesirable uses, ensuring ethical operation and compliance with regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025050540_23102025_PF_FP_ABST
    Figure CA2025050540_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for assessing AI-based computer systems. Preset questions are provided to administrators of an AI-based computer system being assessed. The answers to these preset questions are received and are then assessed by multiple trained trust category models, with each model being trained to assess the answers relative to its trust category. Each answer is assessed relative to each of the various trust categories and relative to the relevance of the question to the each of the trust categories. A rating is assessed for each answer and, for each of the trust categories, an overall score is determined. A remediation module also receives the answers and at least one answer may cause the installation of a guardrail module in the AI-based computer system. Each guardrail module is designed to remediate or address a trust concern revealed by the answer that caused the installation of the module.
Need to check novelty before this filing date? Find Prior Art

Description

AUTOMATED ASSESSMENT OF AI-BASED COMPUTER SYSTEMSTECHNICAL FIELD

[0001] The present invention relates to artificial intelligence (Al). More specifically, the present invention relates to systems and methods for assessing Al-based systems for trustworthiness as the systems relate to specific categories / topics.BACKGROUND

[0002] The rise in the use of and interest in artificial intelligence based systems has led to a need in ensuring that such systems are safe and trustworthy. Such Al systems, while useful tools, may be used for less than useful or even nefarious purposes.

[0003] Currently, Al systems are unfettered in that only the benevolence of their designers and their users are what prevents such nefarious or less than ideal uses. There are no systems that gather data about Al systems, assess such systems, or operate to prevent undesirable uses. Similarly, no systems or methods assess such Al systems for trustworthiness or whether a specific Al system use is ethical or not.

[0004] There is therefore a need for such assessment systems and a need for such systems to be automatic and continuously evolving.SUMMARY

[0005] The present invention provides systems and methods for assessing Al-based computer systems. Preset questions are provided to administrators of an Al-based computer system being assessed. The answers to these preset questions are received and are then assessed by multiple trained trust category models, with each model being trained to assess the answers relative to its trust category. Each answer is assessed relative to each of the various trust categories and relative to the relevance of the question to the each of the trust categories. A rating is assessed for each answer and, for each of the trust categories, an overall score is determined. A remediation module also receives the answers and at least one answer may cause the installation of a guardrail module in the Al-based computer system. Each guardrail module isdesigned to remediate or address a trust concern revealed by the answer that caused the installation of the module.

[0006] In a first aspect, the present invention provides a system for assessing an Al-based computer system, the system comprising:- a database of preset questions, said preset questions being sent to administrators of said Al based computer system and received answers corresponding to said preset questions being received from said administrators;- plurality of trained category models, each corresponding to one of a set of trust categories, each trained category model being trained to:- receive said received answers and said preset questions,- for each received answer, to provide a score based on said received answer and based on a preset relevance of a corresponding question to a specific one of said trust categories, and- produce a trust category score based on scores from said received answers;- a trained remediation model to determine remediation measures to implement on said Al-based computer system based on said trust category scores and on received answers;- a database of guardrail modules, said guardrail modules being for installation on Al-based computer system as remediation measures, each guardrail module being installed on said Al-based computer system to prevent or remediate a trust concern; wherein said trust concern is based on at least one of: a received answer or a trust category score.

[0007] In a second aspect, the present invention provides a method for assessing an Al -based computer system, the method comprising:a) providing a set of preset questions to administrators of said Al-based computer system, said set of preset questions relating to a plurality of topics regarding said Al-based computer system; b) receiving answers from said administrators to said set of preset questions; c) determining a plurality of trust category scores, each of said plurality of trust category scores corresponding to one of a set of trust categories and each one trust category score being based on said answers and said questions as said answers and questions relate to a particular one of said set of trust categories; d) determining one or more remediation measures to implement based on at least one of: said answers and one or more of said plurality of trust category scores, said remediation measures being measures to be applied to said Al-based computer system to thereby prevent said Al-based computer system from being used in a particular way or to prevent said Al -based computer system from functioning in a particular way; e) causing an installation of one or more guardrail modules in said Al-based computer system, said one or more guardrail modules being implementations of at least one of said remediation measures determined in step d).

[0008] In a further aspect, the present invention provides method for assessing an AI- based computer system, the method comprising: a) providing a set of preset questions to administrators of said Al-based computer system, said set of preset questions relating to a plurality of topics regarding said Al-based computer system; b) receiving answers from said administrators to said set of preset questions; c) determining a plurality of trust category scores, each of said plurality of trust category scores corresponding to one of a set of trust categories and each one trust category score being based on said answers and said questions as said answers and questions relate to a particular one of said set of trust categories;d) determining one or more remediation measures to implement based on at least one of: said answers and one or more of said plurality of trust category scores, said remediation measures being measures to be applied to said Al-based computer system to thereby prevent said Al-based computer system from being used in a particular way or to prevent said Al-based computer system from functioning in a particular way; e) causing an implementation of said one or more remediation measures for said Al-based computer system.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The embodiments of the present invention will now be described by reference to the following figures, in which identical reference numerals in different figures indicate identical elements and in which:FIGURE 1 is a block diagram of a system according to one aspect of the present invention.DETAILED DESCRIPTION

[0010] Referring to Fig. 1, a block diagram of a system according to one aspect of the present invention is illustrated. As can be seen, the system 10 includes a database 20 of preset questions. These preset questions are provided to administrators of an Al system 30 that is to be assessed, with the Al system 30 having a particular and specific use. The answers received from the administrators are first passed through suitable modules 35 that may apply NLP methods and other suitable language related Al processes to understand / process the answers. The preset questions are, along with the processed answers from the administrators of the Al system 30, sent to a trained sufficiency model 40 that assesses the administrators' answers to the questions for sufficiency (i.e., has the question been answered, are the answers sufficient as an answer to the question with suitable levels of detail, etc., etc.). Should any particular answer be deemed to be insufficient by the trained sufficiency model 40, this answer and the associated question are re-sent to the administrators to be redone / edited / added to.

[0011] Once the answers from the administrators are determined to be suitable, these answers and the associated questions are sent to multiple trained category models 50A-50E, each of which corresponds to a trust category. Each trust category is a topic / category of concern for Al systems and by which the Al system 30 is to be assessed. As examples, these categories may include:- compliance (compliance with applicable official rules and regulations)- data governance- privacy- explainability and interpretability (explainability and interpretability of results achieved using the Al system being assessed)- bias / faimess- ethics- technical performance

[0012] Each of the relevant trained category models ingest the questions and the answers from the administrators of the Al system and produce a score for the Al system for a specific category based on the answers received from the administrators. For each category, the score is already based on the relevance of each question to the category and of the suitability, applicability, and effect of the received answer on the trustworthiness of the Al system being assessed relative to the trust category. In one interpretation, an answer from the administrator is assessed as to whether the answer renders the Al system under assessment as more trustworthy or not relative to the specific category. As should be clear, in one implementation, the administrators' answers are assessed as they pertain to the specific use for that specific Al system 30.

[0013] The questions and the answers, and the various scores for each category, are also sent to a remediation trained model 60. The remediation trained model 60 operates with a module database 70 that contains guardrail modules that may be installed on the Al system 30 to remediate / address any concerns uncovered by the questions / answers and by the trained category models.

[0014] The remediation trained model 60 assesses each answer and its associated question and, based on the answer and the question, determines whether one or more guardrail modules are to be installed on the Al system 30 being assessed. Guardrail modulesare modules that, when installed in an Al based system, prevents or remediates a trust concern. Guardrail modules may, for example, prevent specific uses for the Al based system or may prevent the use of specific parameters. Such guardrail modules would be designed to constrain, as necessary, the application, use, or configuration of the Al based system to address one or more specific concerns as expressed or revealed by one or more specific questions. The installation of the guardrail modules may be triggered by, for example, an inadequate (but completely sufficient) response by the administrators to one or more questions or by, as another example, a suitably low score for one of the trust categories. In one example, if an Al based system (such as a facial recognition system) is used for a particular function but may also be used (as detailed by administrator answers) for human target recognition and acquisition by drones with an unfettered ability to discharge a weapon towards an acquired human target, one or more guardrail modules that blocks the unfettered ability to discharge the weapon without human intervention and that blocks target acquisition may be installed. Similarly, if an answer to a question indicates that changing a parameter past a certain value would violate a value or tenet for the Al based system, a guardrail module that prevents such a change may be installed.

[0015] As can be seen, the trained models and their behavior are configured based on their training. To generate the training data sets (and to therefore determine the behavior of the respective models), a database of the preset questions and previous answers to these preset questions are used. The preset questions and the previous answers to these preset questions in the database are each ranked and rated by curators / analysts. The ranked and rated questions and their respective answers may then be used as training datasets to train other relevant models.

[0016] It should be clear that, for the preset questions, each question is separately rated for relevance, weight, and importance to each of the trust categories. Thus a question that asks about whether the Al based system on facial recognition was created / tested for racial bias performance (e.g. was the Al based system tested to ensure accuracy when recognizing both white and non-white subjects?), then that question may be highly rated by analysts for relevance to the bias / faimess trust category (in terms of relevance / weight for that trust category) but may be lower rated in terms of relevance for the privacy trust category. The answer may also be differently rated for differenttrust categories. As an example, an administrator's answer to the testing of the Al based system to ensure accuracy when recognizing both white and non- white subjects may indicate that the system was optimized for facial recognition of male non-white subjects. Such an answer would be lower rated for the bias / faimess trust category (as in the answer tends to render the Al based system being assessed as less trustworthy) but might be higher rated for the tech performance trust category (as in, technically, there is optimization so the system tends to be more trustworthy from a technical performance viewpoint) and, again, be lower rated for the ethics trust category (i.e., the analyst may consider optimizing for one particular segment of the population to be unethical and thus renders the Al based system as being less trustworthy from an ethical viewpoint).

[0017] Accordingly, each question and each corresponding answer is separately rated for each trust category. Each question and answer may also be associated with an available remediation / guardrail module to be installed to address the question and / or the answer. Thus, each particular question and an associated answer would have its own rating for each trust category from an analyst as well as possibly an association with a remediation / guardrail module. The resulting database of preset questions and previous answers is thus populated with not just the preset questions and multiple answers but also with each question's rating for relevance to each trust category, each answer's tendency to render the Al system more or less trustworthy relative to each trust category, and, as necessary, remediation / guardrail module(s) that an analyst recommends to be installed to address a particular answer.

[0018] From the above, each question's relevance rating (from an analyst) for each trust category and each answer's trustworthiness rating (from an analyst) for each trust category are then used to train each trust category model. Thus, a trained trust category model can be provided with one of the preset questions and an administrator's answer to that question for a new Al system for assessment. The trained trust category model can then take that preset question's relevance to the trust category (based on the training dataset from previous questions that shows how relevant / irrelevant that specific question is to the trust category) and assess that new answer's trustworthiness relative to the trust category (based on previous answers to the training data set) to arrive at a number rating for the particular presetquestion / answer pair. Once all the question / answer pairs have been assessed by the trained trust category model, the number rating for all of the question / answer pairs are collated and can be processed using a suitable mathematical function (which may take into account each preset question's relevance rating to the specific trust category to accordingly weight each answer’s contribution to the overall trust category rating) to arrive at the Al system’s trust category rating for that trust category. This trust category rating can then be passed on to the trained remediation model.

[0019] To train the trained remediation model, each of the preset questions and each of the answers with any remediation modules associated with the answer are used in one or more training datasets. This therefore trains the remediation model such that, if an incoming answer is similar to how a preset question was answered in the past (with the answer in the past resulting in an analyst recommending the installation of a particular remediation module), this causes that same particular remediation module to be installed on the Al based system currently being assessed.

[0020] For the trained sufficiency model, this can be trained by, again, using previous answers from administrators of previous Al based computer systems and by the preset questions being answered. If a previous answer was deemed by analysts as being insufficient to a specific question, this previous answer and question pair can be used in a training dataset such that the sufficiency model “learns” what is insufficient. As an example, if a previous answer to a technical question did not have enough detail this can be flagged by an analyst as being insufficient. Similarly, if a question asks for multiple optional uses for the Al system and an administrator’s answer does not provide alternatives, an analyst can flag this as, again, insufficient. These insufficient answer / question pairs can thus be used for a training dataset for the sufficiency model. The trained sufficiency model can thus be trained to recognize answers to technical questions that do not have technical details or to recognize answers that do not have options when the question being answered need multiple options.

[0021] It should also be clear that, depending on the character of the Al based system being assessed, different sets of preset questions can be provided to the administrators. As an example, white box Al based systems would be provided with one set of questions while grey box Al based systems would be provided with another set of questions. Black box Al based systems would be provided with yet a third set of questions. Thewhite box and grey box sets of questions may include technical questions relating to the actual Al based systems. As can be imagined, the black box set of questions would not include any technical questions.

[0022] The assessment of an Al based system may be performed periodically and assessment results (i.e. trust category scores and overall scores) from different assessments may be compared. Each assessment may produce a report detailing the trust category scores for that assessment as well as which guardrail modules were installed due to the assessment.

[0023] In another aspect, the various aspects and objects of the present invention may be achieved by means of a software application designed and implemented with a view to codifying or encoding one or more desirable types of answers to the preset questions. For this implementation, one or more base Al systems will have their administrators queried with the set of preset questions. The answers from their administrators can then be used as a baseline or as a basis for a scoring system. As an example, if the preset questions are provided with a preset number of preset multiple options for answers, each of the preset multiple options for the answers can be associated with a preset score. Thus, for a specific question, selecting option A may give a score of 1, selecting option B may give a score of 2, etc., etc. The concept is that each option is provided with a preset score such that if the option is more likely to lead to a desirable result (e.g. the use of the Al based system accords with ethical / desirable ends / uses) then the option is given a higher score. If an option is less likely to lead to a desirable result, then the option is given a lower score. Conversely, the scoring system may be set up so that options that are less likely to lead to a desirable result is given a higher score while options that are more likely to lead to a desirable result is given a lower score.

[0024] Once all the answers have been received from the administrators of the base Al systems, the answers are scored and, for each of the trust categories, the resulting scores are aggregated to result in the trust category scores as explained above. The resulting trust category scores for these base Al systems can then be tweaked and / or adjusted as necessary based on consultations or feedback from the administrators or other interested parties. As part of this, the scoring for the specific answer options forspecific questions can also be tweaked / adjusted as necessary, again based on the feedback and / or consultations with interested parties and / or stakeholders.

[0025] After the trust category scores and the scoring for the specific questions / answer options have been adjusted accordingly, the preset questions and the now preset scoring rubric for the answer options can be codified / encoded in a software application. Such a software application would ingest answers to the preset questions from administrators of Al-based systems being assessed and, based on the hard coded scoring system for the specific answer options to the preset questions, would output trust category scores. The software application can also be programmed to recommend specific remediation measures based on one or more answers to specific questions and based on the resulting trust category scores. As an example, if answers to specific questions relating to the use of facial recognition systems are such that the answers are likely to lead to the use of the facial recognition system for law enforcement without any guardrails against racial / age / demographic bias, the software application can recommend specific measures to guard against such bias. Similarly, if answers to specific questions relating to military use of the Al -based system for target acquisition, assessment, and execution are such that it is likely that there will be no human in the decision making loop, the application can recommend specific measures to ensure the presence of a human in the decision making process.

[0026] After the trust category scores have been provided to the administrators of the AI- based system being assessed, the software application can, of course, provide the recommended remediation measures to the administrators. Such recommendations can include providing software modules to be installed in the Al-based systems, periodic re-assessments of the Al-based system, as well as potential education / awareness courses for the administrators to ensure awareness of the dangers of unmitigated, unguarded, and potentially unethical use of Al-based systems.

[0027] It should be clear that the various aspects of the present invention may be implemented as software modules in an overall software system. As such, the present invention may thus take the form of computer executable instructions that, when executed, implements various software modules with predefined functions.

[0028] Additionally, it should be clear that, unless otherwise specified, any references herein to 'image' or to 'images' refer to a digital image or to digital images, comprising pixels or picture cells. Likewise, any references to an 'audio file' or to 'audio files' refer to digital audio files, unless otherwise specified. 'Video', 'video files', 'data objects', 'data files' and all other such terms should be taken to mean digital files and / or data objects, unless otherwise specified.

[0029] The embodiments of the invention may be executed by a computer processor or similar device programmed in the manner of method steps, or may be executed by an electronic system which is provided with means for executing these steps. Similarly, an electronic memory means such as computer diskettes, CD-ROMs, Random Access Memory (RAM), Read Only Memory (ROM) or similar computer software storage media known in the art, may be programmed to execute such method steps. As well, electronic signals representing these method steps may also be transmitted via a communication network.

[0030] Embodiments of the invention may be implemented in any conventional computer programming language. For example, preferred embodiments may be implemented in a procedural programming language (e.g., "C" or "Go") or an object-oriented language (e.g., "C++", "java", "PHP", "PYTHON" or "C#"). Alternative embodiments of the invention may be implemented as pre-programmed hardware elements, other related components, or as a combination of hardware and software components.

[0031] Embodiments can be implemented as a computer program product for use with a computer system. Such implementations may include a series of computer instructions fixed either on a tangible medium, such as a computer readable medium (e.g., a diskette, CD-ROM, ROM, or fixed disk) or transmittable to a computer system, via a modem or other interface device, such as a communications adapter connected to a network over a medium. The medium may be either a tangible medium (e.g., optical or electrical communications lines) or a medium implemented with wireless techniques (e.g., microwave, infrared or other transmission techniques). The series of computer instructions embodies all or part of the functionality previously described herein. Those skilled in the art should appreciate that such computer instructions can be written in a number of programming languages for usewith many computer architectures or operating systems. Furthermore, such instructions may be stored in any memory device, such as semiconductor, magnetic, optical or other memory devices, and may be transmitted using any communications technology, such as optical, infrared, microwave, or other transmission technologies. It is expected that such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (e.g., shrink-wrapped software), preloaded with a computer system (e.g., on system ROM or fixed disk), or distributed from a server over a network (e.g., the Internet or World Wide Web). Of course, some embodiments of the invention may be implemented as a combination of both software (e.g., a computer program product) and hardware. Still other embodiments of the invention may be implemented as entirely hardware, or entirely software (e.g., a computer program product).

[0032] A person understanding this invention may now conceive of alternative structures and embodiments or variations of the above all of which are intended to fall within the scope of the invention as defined in the claims that follow.

Claims

We claim:

1. A system for assessing an Al -based computer system, the system comprising:- a database of preset questions, said preset questions being sent to administrators of said Al based computer system and received answers corresponding to said preset questions being received from said administrators;- plurality of trained category models, each corresponding to one of a set of trust categories, each trained category model being trained to:- receive said received answers and said preset questions,- for each received answer, to provide a score based on said received answer and based on a preset relevance of a corresponding question to a specific one of said trust categories, and- produce a trust category score based on scores from said received answers;- a trained remediation model to determine remediation measures to implement on said AI- based computer system based on said trust category scores and on received answers;- a database of guardrail modules, said guardrail modules being for installation on Al-based computer system as remediation measures, each guardrail module being installed on said AI- based computer system to prevent or remediate a trust concern; wherein said trust concern is based on at least one of: a received answer or a trust category score.

2. The system according to claim 1, wherein each trained category model is trained on training data sets that comprise previous answers and corresponding preset questions, said previous answers being rated by analysts and said corresponding preset questions have been rated by said analysts for relevance to each of said trust categories.

3. The system according to claim 2, wherein said remediation model is trained on training data sets that comprise said previous answers and corresponding preset questions and associated guardrail modules as selected by said analysts.

4. The system according to claim 1, wherein different preset questions are provided to administrators based on a character of said Al-based computer system.

5. The system according to claim 4, wherein said character is whether said Al -based computer system is a white box, a grey box, or a black box Al-based computer system.

6. A method for assessing an Al-based computer system, the method comprising: a) providing a set of preset questions to administrators of said Al-based computer system, said set of preset questions relating to a plurality of topics regarding said Al -based computer system; b) receiving answers from said administrators to said set of preset questions; c) determining a plurality of trust category scores, each of said plurality of trust category scores corresponding to one of a set of trust categories and each one trust category score being based on said answers and said questions as said answers and questions relate to a particular one of said set of trust categories; d) determining one or more remediation measures to implement based on at least one of: said answers and one or more of said plurality of trust category scores, said remediation measures being measures to be applied to said Al-based computer system to thereby prevent said AI- based computer system from being used in a particular way or to prevent said Al-based computer system from functioning in a particular way; e) causing an installation of one or more guardrail modules in said Al-based computer system, said one or more guardrail modules being implementations of at least one of said remediation measures determined in step d).

7. The method according to claim 6, wherein step c) is implemented by a plurality of trained category models, each corresponding to one of said set of trust categories, each trained category model being trained to:- receive received answers and said preset questions and,- for each received answer, to provide a score based on said received answer and based on a preset relevance of a corresponding question to a specific one of said trust categories, and- produce a trust category score based on scores from said received answers.

8. The method according to claim 6, wherein steps d) and e) are implemented by- a trained remediation model to determine said remediation measures to implement on said Al-based computer system based on said trust category scores and on received answers;- a database of said guardrail modules, said guardrail modules being for installation on AI- based computer system as remediation measures, each guardrail module being installed on said Al -based computer system to prevent or remediate a trust concern; wherein said trust concern is based on at least one of: a received answer or a trust category score.

9. The method according to claim 6, wherein different preset questions are provided to said administrators based on a character of said Al-based computer system.

10. The method according to claim 9, wherein said character is whether said Al -based computer system is a white box, a grey box, or a black box Al-based computer system.

11. The system according to claim 1, wherein said set of trust categories includes at least one of:- compliance;- data governance;- privacy;- explainability and interpretability;- bias / faimess;- ethics; and- technical performance.

12. The method according to claim 6, wherein said set of trust categories includes at least one of:- compliance;- data governance;- privacy;- explainability and interpretability;- bias / faimess;- ethics; and- technical performance.

13. A method for assessing an Al -based computer system, the method comprising: a) providing a set of preset questions to administrators of said Al-based computer system, said set of preset questions relating to a plurality of topics regarding said AI- based computer system; b) receiving answers from said administrators to said set of preset questions; c) determining a plurality of trust category scores, each of said plurality of trust category scores corresponding to one of a set of trust categories and each one trust category score being based on said answers and said questions as said answers and questions relate to a particular one of said set of trust categories; d) determining one or more remediation measures to implement based on at least one of: said answers and one or more of said plurality of trust category scores, said remediation measures being measures to be applied to said Al-based computer system to thereby prevent said Al-based computer system from being used in a particular way or to prevent said Al-based computer system from functioning in a particular way; e) causing an implementation of said one or more remediation measures for said AI- based computer system.

14. The method according to claim 13, wherein for step c), said plurality of trust category scores is determined based on previous answers to said questions.

15. The method according to claim 14, wherein said plurality of trust category scores is determined by assigning predetermined scores to specific answers to specific questions and aggregating scores for all of said answers received from said administrators.

16. The method according to claim 13, wherein step d) includes communicating said one or more remediation measures to said administrators along with said trust category scores.

17. An invention according to the attached figures and text.

Citation Information

Patent Citations

  • Methods and systems for the measurement of relative trustworthiness for technology enhanced with ai learning algorithms

    CA3048056A1

  • Artificial intelligence method and system for detecting anomalies in a computer network

    US10771489B1

  • Cyberrisk governance system and method to automate cybersecurity detection and resolution in a network

    WO2022205808A1

  • Framework for trustworthiness

    WO2022237963A1