A secured artificial intelligence system and a security assessment system for artificial intelligence model
Patent Information
- Application Number
- PCT/SG2024/050129
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-10-02
AI Technical Summary
Existing cybersecurity frameworks for artificial intelligence (AI) models, particularly Generative AI (GenAI) models, lack comprehensive security solutions to address the evolving and multifaceted vulnerabilities, including threats such as jailbreaking, prompt injection, and data poisoning, which can lead to unauthorized data extraction and generation of harmful content.
A holistic AI security framework comprising a model protection system, assessment system, and red teaming system, which includes input-based and output-based detection and prevention techniques, along with a wrapper to manage interactions, ensuring continuous security assessment and protection across the AI model's lifecycle.
The framework effectively assesses and enhances the security of AI models, identifying and mitigating vulnerabilities, providing continuous protection against evolving cybersecurity threats, and ensuring compliance with regulatory standards.
Smart Images

Figure SG2024050129_02102025_PF_FP_ABST
Abstract
Description
DESCRIPTIONTITLE OF INVENTION: [A SECURED ARTIFICIAL INTELLIGENCE SYSTEM AND A SECURITY ASSESSMENT SYSTEM FOR ARTIFICIAL INTELLIGENCE MODEL]TECHNICAL FIELD
[0001] The present disclosure relates generally to the field of artificial intelligence (Al) technologies, with a specific focus on the cyber security of Al models.BACKGROUND
[0002] The following discussion of the background to the invention is intended to facilitate an understanding of the present invention only. It should be appreciated that the discussion is not an acknowledgement or admission that any of the material referred to was published, known or part of the common general knowledge of the person skilled in the art in any jurisdiction as at the priority date of the invention.
[0003] The field of artificial intelligence (Al) is undergoing unprecedented growth, with the advent of new Al models, particularly Generative Al (GenAI) models, emerging at a brisk pace. These cutting-edge models are finding their niche across a broad spectrum of industries, revolutionizing the way tasks are approached and solutions are devised. The applications of Al models range widely, from enhancing customer experiences through sophisticated virtual assistants, automating and improving efficiency in software development via code generation, to sparking creativity in content creation among others. The potential uses for these models are not just vast but are in a state of constant evolution, with new applications being explored and realized on an ongoing basis, thereby broadening the horizon of possibilities that Al can offer.
[0004] Despite the substantial progress and expanding applications of Al, vulnerabilities in Al models, notably GenAI models, expose them to significant cybersecurity risks. Unprotected Al applications are prone to a variety of cyber threats, such as jailbreaking, prompt injection, and data poisoning. These threats can have dire consequences, including the unauthorized extraction of sensitive data, the creation of harmful or inappropriate content, and the generation of inaccurate or misleading information, among other negative outcomes. The cybersecurity landscape for Al models is complex, with threats not only varied but continually evolving, necessitatingdefenses that are both adaptive and comprehensive. However, the current landscape reveals a critical gap: a comprehensive security framework dedicated to the protection of Al models across their lifecycle is notably absent.
[0005] Existing literature, such as the work by Iqbal, U., Kohno, T., & Roesner, F. (2023) (LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI’s ChatGPT Plugins. arXiv preprint arXiv:2309. 10254.), which focuses on threat modeling within the context of user-plugin and plugin-LLM interactions, addresses only certain aspects of Al security.
[0006] US Patent Application US20230269272A1 “System and method for implementing an artificial intelligence security platform” discloses a system for implementing an artificial intelligence (Al) security platform, which comprises model security module configured to provide insights on the model performance from a security perspective. However, the so-called model security module is rather simplistic, including basic functions such as model awareness and abnormal model behaviour analytics, challenging to comprehensive address evolving cybersecurity issues faced by Al models.
[0007] Such conventional methodologies overlook the broader necessity for modelcentric security solutions that holistically address the multifaceted vulnerabilities Al models face.
[0008] Thus, there exists a need to develop a system that can efficiently assess and enhance the security of Al models, particularly GenAI models facing constantly evolving cybersecurity issues, and alleviate at least one of the aforementioned problems.SUMMARY
[0009] Specifically, to address the above-mentioned technical problems, the present invention specifically uses the following technical solutions:
[0010] Accordingly, in accordance with one aspect of the present invention, there is provided an artificial intelligence system, comprising a model, a first model protection system and a wrapper, wherein the wrapper is operable to receive a first input from a model user and to direct the first input to the first model protection system, wherein thefirst model protection system is operable to harden the first input to generate a second input, wherein the wrapper is operable to receive the second input and to further direct the second input to the model; wherein the model is operable to receive the second input and to further generate the first output that is further directed to the first model protection system; wherein the first model protection system is operable to receive and harden the first output to generate a second output, wherein the second output is directed to the wrapper; and wherein the wrapper is operable to receive the second output from the first model protection system and to further direct the second output to the model user.
[0011] In some embodiments, the first model protection system comprises at least one Input-based Detection and Prevention Technique (IDPT), wherein the at least one IDPT is operable to analyse the first input.
[0012] In some embodiments, the at least one IDPT is selected from a group consisting of poisoned data D&P, prompt injection D&P, and model theft or extraction D&P.
[0013] In some embodiments, the first model protection system comprises at least one Output-based Detection and Prevention Technique (ODPT), wherein the at least one ODPT is operable to analyse the first output.
[0014] In some embodiments, the at least one ODPT is selected from a group consisting of poisoned model D&P, model jailbreak D&P, data leakage D&P, and toxicity or harmful content D&P.
[0015] In some embodiments, the first model protection system comprises an application programming interface configured to be in communication with the wrapper.
[0016] In some embodiments, the first model protection system is extensible to incorporate one or more detection and prevention techniques.
[0017] In some embodiments, the model is a generative artificial intelligence model.
[0018] In accordance with another aspect of the present invention, there is provided a security assessment system for a target model, comprising an assessment system, a red teaming system; wherein the assessment system is configured to submit an assessment input to the target model for generating an assessment output, and tofurther receive the assessment output for analysis; wherein the red teaming system is configured to submit a red team input to the target model for generating a red team output, and to further receive the first red team output for analysis; wherein a first pair comprising the assessment input and the assessment output is configured to be provided to the red teaming system for improving attack coverage of the red teaming system; and, wherein a second pair comprising the red team input and the red team output is configured to be provided to the assessment system for improving assessment coverage of the assessment system.
[0019] In some embodiments, the security assessment system further comprising a second model protection system and a wrapper; wherein the red team input is first directed to the wrapper, which is configured to direct the red team input to the second model protection system; wherein the second model protection system is configured to receive and harden the red team input to generate a hardened red team input; wherein the wrapper is configured to receive the hardened red team input and to further direct the hardened red team input to the target model; wherein the target model is operable to receive the hardened red team input and to further generate the red team output, which is directed to the wrapper; wherein the wrapper is configured to receive the red team output and to further direct the red team output to the second model protection system; wherein the second model protection system is configured to receive the red team output and to further generate the hardened red team output, which is then directed to the wrapper; wherein the wrapper is configured to receive the hardened red team output and to further direct the hardened red team output to the red team system for analysis.
[0020] In some embodiments, the assessment system comprises at least one assessment technique.
[0021] In some embodiments, the at least one assessment technique is selected from a group consisting of prompt injection or jailbreak testing, training data reconstruction testing, and compliance assessment.
[0022] In some embodiments, the red teaming system comprises at least one red teaming technique.
[0023] In some embodiments, the at least one red teaming technique is promptfuzzing and penetration testing.
[0024] In some embodiments, the second model protection system comprises at least one Input-based Detection and Prevention Technique (IDPT), wherein the at least one IDPT is operable to analyse and harden the red team input.
[0025] In some embodiments, the at least one IDPT is selected from a group consisting of poisoned data D&P, prompt injection D&P, and model theft or extraction D&P.
[0026] In some embodiments, the second model protection system comprises at least one Output-based Detection and Prevention Technique (ODPT), wherein the at least one ODPT is operable to analyse and harden the red team output.
[0027] In some embodiments, the at least one ODPT is selected from a group consisting of poisoned model D&P, model jailbreak D&P, data leakage D&P, and toxicity or harmful content D&P.
[0028] In some embodiments, the assessment system and the red teaming system are extensible to incorporate additional techniques.
[0029] In some embodiments, the second model protection system is extensible to incorporate one or more additional detection and prevention techniques.
[0030] In some embodiments, the target model is a generative artificial intelligence model.
[0031] Other aspects and features of the present invention will become apparent to those of ordinary skill in the art upon review of the following description of specific embodiments of the invention in conjunction with the accompanying figures.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In the figures, which illustrate, by way of non-limiting examples only, embodiments of the present invention,
[0033] [Fig. 1]: illustrates, in a block diagram, an Al security framework according to various embodiments of the present invention.
[0034] [Fig. 2]: illustrates, in a block diagram, an Al security framework according to various embodiments of the present invention.
[0035] [Fig. 3A] : illustrates, in a block diagram, a first model protection system (e.g., GenAI protection system) according to various embodiments of the present invention.
[0036] [Fig. 3B]: illustrates, in a flowchart, the operation steps of the first model protection system (e.g., GenAI protection system) according to various embodiments of the present invention.
[0037] [Fig. 4A]: illustrates, in a block diagram, the assessment system (e.g., GenAI assessment system), the red teaming system (e.g., GenAI red teaming system) and the second model protection system (e.g., introspective GenAI protection system) according to various embodiments of the present invention.
[0038] [Fig. 4B]: illustrates, in a flowchart, the operation steps of the assessment system (e.g., GenAI assessment system) according to various embodiments of the present invention.
[0039] [Fig. 4C]: illustrates, in a flowchart, the operation steps of the red teaming system (e.g., GenAI red teaming system) according to various embodiments of the present invention.
[0040] [Fig. 5]: illustrates, in block diagram, the interactions among the first model protection system (e.g., GenAI protection system), the red teaming system (e.g., GenAI red teaming system) and the second model protection system (e.g., introspective GenAI protection system), according to various embodiments of the present invention.
[0041] [Fig. 6]: illustrates, in block diagram, an overview of the Al security framework, including its interaction with the Al model (e.g., GenAI model) in both production environment and testing environment, according to various embodiments of the present invention.DETAILED DESCRIPTION
[0042] Throughout this document, unless otherwise indicated to the contrary, the terms “comprising”, “consisting of”, “having” and the like, are to be construed as non- exhaustive, or in other words, as meaning “including, but not limited to”.
[0043] Furthermore, throughout the document, unless the context requires otherwise, the word “include” or variations such as “includes” or “including” will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.
[0044] Unless defined otherwise, all other technical and scientific terms used herein have the same meaning as is commonly understood by a skilled person to which the subject matter herein belongs.
[0045] According to various embodiments of the present invention, there is provided an artificial intelligence (Al) system, comprising an Al model, a first model protection system and a wrapper; wherein the wrapper is operable to receive a first input (also known as prompt) from a model user and to further direct the first input to the first model protection system, wherein the first model protection system is operable to harden the first input to generate a second input, wherein the wrapper is operable to receive the second input and to further direct the second input to the model; wherein the model is operable to receive the second input and to further generate the first output that is further directed to the first model protection system; wherein the first model protection system is operable to receive and harden the first output to generate a second output, wherein the second output is directed to the wrapper; and, wherein the wrapper is operable to receive the second output from the first model protection system and to further direct the second output to the model user.
[0046] According to various embodiments of the present invention, there is provided a security assessment system for a target model, comprising an assessment system, a red teaming system; wherein the assessment system is configured to submit an assessment input to the target model for generating an assessment output, and to further receive the assessment output for analysis; wherein the red teaming system is configured to submit a red team input to the target model for generating a red team output, and to further receive the first red team output for analysis; wherein a first pair comprising the assessment input and the assessment output is configured to be provided to the red teaming system for improving attack coverage of the red teamingsystem; and, wherein a second pair comprising the red team input and the red team output is configured to be provided to the assessment system for improving assessment coverage of the assessment system.
[0047] In some embodiments, the security assessment system further comprising a second model protection system and a wrapper; wherein the red team input is first directed to the wrapper, which is configured to direct the red team input to the second model protection system; wherein the second model protection system is configured to receive and harden the red team input to generate a hardened red team input; wherein the wrapper is configured to receive the hardened red team input and to further direct the hardened red team input to the target model; wherein the target model is operable to receive the hardened red team input and to further generate the red team output, which is directed to the wrapper; wherein the wrapper is configured to receive the red team output and to further direct the red team output to the second model protection system; wherein the second model protection system is configured to receive the red team output and to further generate the hardened red team output, which is then directed to the wrapper; wherein the wrapper is configured to receive the hardened red team output and to further direct the hardened red team output to the red team system for analysis.
[0048] The first model protection system, the assessment system, the red teaming system, and the second model protection system of the present invention collectively and synergistically function together to form an Al security framework 100. This Al security framework 100 is operable to not only assess and enhance the security of an Al model under development prior to its deployment but also to provide continuous protection to an Al model in the deployment phase. In other words, the present invention aims to provide a comprehensive and holistic security management framework that assesses and mitigates the risks associated with Al models, particularly large language models (LLM).
[0049] Referring to [Fig. 1] and [Fig. 2], in certain embodiments tailored for the safeguarding of Generative Al (GenAI) models, the Al security framework 100 is structured to include (i) a GenAI Protection System (that is, the first model protection system) 101 , dedicated to securing and safeguarding the GenAI model 105 when deployed in a customer's solution and during its active operation; (ii) a GenAIAssessment System (that is, the assessment system) 102, responsible for evaluating the security posture and compliance of the intended GenAI model (that is, the target model) before its deployment; (iii) a GenAI Red Teaming System (that is, the red teaming system) 103, designed for executing offensive attacks on the target GenAI model with the aim of identifying any previously undetected vulnerabilities; and (iv) an extensively introspective GenAI Protection System (that is, the second model protection system) 104, tasked with monitoring and understanding the protective capabilities in instances where the GenAI model 105 faces attacks orchestrated by the GenAI Red Teaming System 103.
[0050] More specifically, as depicted in Fig. 2, the subsystems within the Al Security Framework 100, which include the GenAI protection system, the GenAI assessment system, the GenAI red teaming system, and the introspective GenAI protection system, are engineered to safeguard the GenAI model 204 in the production environment 200. Additionally, these subsystems aim to evaluate and fortify the security posture of the GenAI model 204 during its development phase in the testing environment 201 .
[0051] Within the production environment 200, an End-User 202 can interact with the GenAI model by providing inputs (also referred to as prompts) to obtain the corresponding generated outputs. The inputs furnished by the End-User 202 are intended to be processed through a wrapper interface 205, which orchestrates the flow of inputs and outputs, thereby controlling the End-User’s interaction with the GenAI model.
[0052] Before the GenAI Model 204 is deployed to the production environment 200, a Developer 203 engages in comprehensive security assessments and testing of the newly developed or enhanced GenAI model within the testing environment 201. This phase is aimed at identifying unresolved vulnerabilities and potentially discovering new ones. Within this testing environment 201 , the GenAI Model 204 undergoes thorough scrutiny, through both the GenAI assessment system and the GenAI red teaming system. Furthermore, the GenAI red teaming system is configured to both directly probe the security of the GenAI or indirectly probe the security of the GenAI through a specifically configured wrapper interface 205.
[0053] Specifically, concerning the first model protection system, for instance, the GenAI Protection System depicted in Fig. 3A, designed for safeguarding the GenAImodel during the deployment phase, this system incorporates APIs (Application Programming Interfaces) 300. These APIs are crucial for facilitating the reception of model inputs and the delivery of model outputs. Beyond this, the GenAI Protection System is equipped with an array of GenAI Input-based Detection & Prevention Techniques (IDPT) 301 and GenAI Output-based Detection & Prevention Techniques (ODPT) 302. These techniques, which may alternatively be known as components or tools in certain embodiments, are integral for defending against both existing and emerging threats. They are crafted with the flexibility to incorporate new methods aimed at addressing threats in innovative ways or countering novel threats. Furthermore, in some embodiments, the IDPT and ODPT techniques are designed to operate interdependently, enabling the exchange of analysis results 303 as necessary. This interplay is critical for making informed decisions and directing the processes of sanitization or hardening effectively.
[0054] Examples of IDPT include poisoned data D&P, prompt injection D&P, and model theft or extraction D&P. Examples of ODPT include poisoned model D&P, model jailbreak D&P, Pll (Personally Identifiable Information) or data leakage D&P, and toxicity or harmful content D&P.
[0055] An overview of the data flow is as follows: an end-user initiates the process by submitting an input prompt 304, which is processed through a specially configured wrapper interface 308 and then forwarded 306 to the GenAI Protection System for a detailed analysis. Following this, the now sanitized or hardened input prompt is sent back 306 and introduced into the GenAI Model to generate the corresponding output. This output is subsequently directed 307 to the GenAI Protection System for another round of analysis. After being sanitized or hardened, the final output is transmitted 307 through the wrapper interface 308, ultimately reaching the end-user 305.
[0056] A detailed workflow of a secure Al system, as outlined in some embodiments of the present invention, is illustrated in Fig. 3B:
[0057] The production system (or application) collects the raw input prompt directly from the end-user, initially processed at the wrapper interface 310. Optionally, in certain embodiments, developers have the flexibility to enhance the wrapper to further process and / or log the raw input prompt for analytical purposes 31 1 .
[0058] The wrapper forwards the raw (or previously processed) input prompt to the GenAI Protection System via a specifically configured API 312.
[0059] Upon reception, the GenAI Protection System disseminates the raw input prompt among all input-based D&P components 313 for comprehensive scrutiny.
[0060] Each input-based D&P component conducts an in-depth analysis of the input prompt and communicates the findings back.
[0061] The GenAI Protection System compiles and processes these findings from the D&P components to enhance the security of the input prompt 314.
[0062] The hardened (sanitized) input prompt is then relayed back to the customer’s production system 315.
[0063] Upon receipt, the production system handles the hardened input prompt 316 and submits it to the GenAI model 317 for output generation.
[0064] The GenAI model processes the hardened input and generates the corresponding output.
[0065] The production system captures this generated output 318 and forwards it to the GenAI Protection System for analysis via a configured API 319.
[0066] The GenAI Protection System distributes the generated output to all outputbased D&P components 320 for examination, wherein each output-based D&P component evaluates the output and shares the analysis results.
[0067] The GenAI Protection System aggregates these results, utilizing them to further harden the output 321 .
[0068] This hardened output is then dispatched back to the customer’s production system 322.
[0069] The production system, upon receiving the hardened output 323, may optionally, in some embodiments, engage the wrapper to process and / or log this output for analytics 324.
[0070] Finally, the production system delivers the secured output back to the end-user 325, concluding the process.
[0071] Before the Al model is officially deployed in the production environment, a security assessment system can be implemented to evaluate and improve the security of the Al model while it is still in the development phase within a testing environment. This preemptive measure ensures that potential vulnerabilities are identified and mitigated early on.
[0072] Within certain embodiments of the present invention, this security assessment system includes a GenAI Assessment System (illustrated in Fig. 4A) (i.e., assessment system). This system features a suite of GenAI assessment techniques 400, which is designed to be scalable, allowing for the incorporation of new techniques as they are developed. These methodologies, which may also be referred to as components or tools, cover a range of tests and evaluations. Examples of such GenAI assessment techniques include prompt injection or jailbreak testing to evaluate model security boundaries, training data reconstruction testing to assess the model's resistance to data extraction attacks, and compliance assessment to ensure adherence to relevant standards and regulations.
[0073] In addition, some embodiments feature a GenAI Red Teaming System (i.e., the red teaming system). This GenAI red teaming system is comprised of various GenAI red teaming techniques 401 , with a framework that supports the addition of new and emerging techniques to stay ahead of potential threats. Among the red teaming techniques employed are prompt fuzzing, aimed at uncovering vulnerabilities through unexpected or malformed inputs, and penetration testing, which involves systematic attempts to breach the model’s defenses in order to identify weaknesses.
[0074] Regarding the data flow overview, techniques employed by the GenAI Assessment System involve submitting specifically crafted input prompts directly 403 to the GenAI model (bypassing any wrapper) to elicit output 404. The resulting inputoutput pairs are then analyzed to pinpoint potential weaknesses in the GenAI model.
[0075] The GenAI Red Teaming System's approach involves both (i) direct submission of crafted input prompts to the GenAI model to obtain raw generated output 405, and (ii) indirect submission through a configured wrapper interface 407. The latter directs the inputs / prompts to the GenAI model protected by the introspective GenAIProtection System 402, yielding a hardened set of outputs 406. Analysis of both sets of input-output pairs is helpful for identifying and addressing model flaws or vulnerabilities comprehensively.
[0076] A detailed work flow of GenAI Assessment System is as follows (Fig. 4B):
[0077] A comprehensive workflow of the GenAI Assessment System is detailed below (refer to Fig. 4B for visual representation):
[0078] The GenAI Assessment System is arranged to dispatch a notification to each GenAI assessment technique, indicating the security assessment target 410.
[0079] Subsequently, each GenAI assessment technique is arranged to prepare and submit specially crafted input prompts to the target GenAI model 411 , aiming to elicit outputs for analysis.
[0080] Upon receiving the generated outputs from the GenAI model 412, each GenAI assessment technique proceeds to the next phase of evaluation.
[0081] The GenAI assessment technique conducts a thorough evaluation of the input-output pairs to identify any potential weaknesses within the model 413. Following this evaluation, the findings are systematically reported back to the GenAI Assessment System 414.
[0082] The above procedure is repeated for a variety of crafted input prompts, continuing until the assessment is either completed or intentionally terminated 415.
[0083] Throughout this procedure, the GenAI Assessment System is operable to actively aggregate and process the security assessment findings as they are received from the various GenAI assessment techniques 416.
[0084] Once the security assessment phase is concluded, the GenAI Assessment System compiles a comprehensive security assessment report. This report encapsulates all the findings processed during the assessment 417.
[0085] Developers are then provided with the opportunity to thoroughly review the security assessment report. This crucial step enables them to gain insights into potential vulnerabilities and to strategize on further improvements for their GenAI model,ensuring a robust and secure Al solution.
[0086] A detailed workflow of the GenAI Red Teaming System is outlined as follows (refer to Fig. 4C):
[0087] The GenAI Red Teaming System initiates the process by informing each GenAI red teaming technique of the security testing target model 420.
[0088] For Direct Security Testing:
[0089] Each GenAI red teaming technique devises and dispatches specifically crafted input prompts directly to the target GenAI model 421 , bypassing any wrapper, with the aim of eliciting raw output for evaluation.
[0090] The raw generated output is then directly received from the GenAI model by the red teaming technique 422.
[0091] This technique evaluates the direct input-output pair to uncover potential vulnerabilities within the target model 423, subsequently reporting these findings to the GenAI Red Teaming System 424.
[0092] The above procedure is methodically repeated with various crafted input prompts until the direct testing phase is deemed complete or is otherwise terminated 425.
[0093] For Indirect Security Testing:
[0094] Each GenAI red teaming technique submits crafted input prompts indirectly to the target GenAI model 426 via a specially configured wrapper interface. This wrapper interface ensures that all raw input prompts and their corresponding generated outputs are processed through the introspective GenAI Protection System, resulting in hardened outputs.
[0095] The hardened outputs are then received from the configured wrapper interface by the red teaming technique 427.
[0096] Similar to the direct testing phase, the red teaming technique evaluates the hardened input-output pair for vulnerabilities 428 and reports these assessments to the GenAI Red Teaming System 429.
[0097] This indirect testing procedure is also repeated with different crafted input prompts until completion or termination 430.
[0098] Throughout both testing methods, the GenAI Red Teaming System diligently compiles and processes the security testing findings as they are reported by the red teaming techniques 431 .
[0099] At the conclusion of the security testing process, the GenAI Red Teaming System assembles a detailed security testing report. This report encompasses all findings that have been processed during the testing phases 432.
[0100] Developers are then presented with the opportunity to review the security testing report. This step is helpful for identifying and understanding the vulnerabilities within the GenAI model under development, thereby guiding the development of further enhancements to bolster the model's security posture.
[0101] The Al Security Framework is ingeniously designed to be inherently dynamic and adaptable, enabling it to effectively counteract the constantly changing and evolving threats that Al models face, as depicted in FIG. 5. This capability for ongoing evolution ensures that the framework remains updated of the defense mechanisms against the rapidly advancing landscape of cybersecurity threats to Al models.
[0102] In some embodiments of the present invention, within the production environment, when the GenAI Protection System (i.e., the first model protection system) encounters or identifies new instances of malicious or previously unseen input prompts, these discoveries become a source for the iterative enhancement of the GenAI Assessment System 500. This involves leveraging these real-world encounters to craft increasingly sophisticated and targeted prompts, thereby enhancing the system's ability to simulate and preempt potential threats more effectively.
[0103] Conversely, the advent of new compliance standards and assessment protocols leads to the generation of fresh input-output pairs. These innovative pairs are crucial for the continuous improvement of the Detection & Prevention (D&P) techniques within the GenAI Protection System, significantly bolstering its defenses against recognized threats 500.
[0104] Similarly, these newly minted input-output pairs could be used to facilitaterefining the red teaming techniques within the GenAI Red Teaming System 501. By expanding the range and depth of simulated attack scenarios, the red teaming system is arranged to cover the evolving types of attacks against Al models. In reverse, the emergence of new security insights from the input-output pairs evaluated during red teaming exercises provides a rich dataset for the GenAI Assessment System. These data points enrich the pool of scenarios used for stressing the GenAI models, thereby enhancing the ability to detect and address weaknesses 501 .
[0105] Furthermore, the activity logs captured during simulations of attacks by the GenAI Red Teaming System against a closely monitored GenAI Protection System are also helpful in facilitating the iteration of the security testing system. These activity logs shed light on any unexpected or irregular responses of D&P techniques under duress, laying the groundwork for targeted debugging and subsequent refinement of these strategies 502. Through this debugging and refinement process, the generation of improved quality input-output pairs ensues. These optimized pairs are then leveraged to further refine and strengthen the D&P techniques, ensuring an enhanced level of protection 502.
[0106] It is important to underscore that the Al security framework's feature of continuous improvement, particularly involving (1 ) inputs submitted by end-users and (2) outputs generated by the models, is designed with privacy considerations at its core. This ensures that the process of enhancing the system's capabilities is conducted within the bounds of the relevant privacy policies.
[0107] The holistic overview of the Al Security Framework, including its intricate interactions with the GenAI model across both the production and testing environments, is illustrated in Fig. 6. This detailed illustration encapsulates the Al security framework's comprehensive approach to security, highlighting its features of having a continuous cycle of evaluation, protection, enhancement, and adaptation to safeguard Al models against an ever-evolving array of cybersecurity threats.
[0108] It is important to note that a significant aspect of the present invention is its reliance on the input-output pairs generated by Al models for conducting thorough testing and providing robust protection. This innovative approach eliminates the need for direct access to the proprietary technical specifications (e.g., inner structure or working) of the Al models themselves. By focusing on the observable interactions,specifically, the inputs provided to and the outputs produced by these models, the systems of the present invention can effectively assess and enhance model security without encroaching on the confidentiality of the underlying proprietary technologies. This method broadens the potential applicability of the invention, making it a versatile and non-invasive solution for securing Al models across various platforms and applications, irrespective of their proprietary nature.
[0109] In summary, the Al security framework of the present invention spans the entire lifecycle of these models, from their initial development phase (e.g., the security assessment system) through to their deployment and beyond into their operational life (e.g., the secured Al system protected by the first model protection system), enhancing protection against the evolving landscape of cybersecurity threats.
[0110] It should be further appreciated by the person skilled in the art that variations and combinations of features described above, not being alternatives or substitutes, may be combined to form yet further embodiments falling within the intended scope of the invention.
[0111] As would be understood by a person skilled in the art, each embodiment, may be used in combination with other embodiment or several embodiments.
Claims
ClaimsClaim 1. An artificial intelligence system, comprising a model, a first model protection system and a wrapper, wherein the wrapper is operable to receive a first input from a model user and to direct the first input to the first model protection system, wherein the first model protection system is operable to harden the first input to generate a second input, wherein the wrapper is operable to receive the second input and to further direct the second input to the model; wherein the model is operable to receive the second input and to further generate the first output that is further directed to the first model protection system; wherein the first model protection system is operable to receive and harden the first output to generate a second output, wherein the second output is directed to the wrapper; and, wherein the wrapper is operable to receive the second output from the first model protection system and to further direct the second output to the model user.Claim 2. The artificial intelligence system according to Claim 1 , wherein the first model protection system comprises at least one Input-based Detection and Prevention Technique (IDPT), wherein the at least one IDPT is operable to analyse the first input.Claim 3. The artificial intelligence system according to Claim 2, wherein the at least one IDPT is selected from a group consisting of poisoned data D&P, prompt injection D&P, and model theft or extraction D&P.Claim 4. The artificial intelligence system according to anyone of Claim 1 - Claim 3, wherein the first model protection system comprises at least one Output-based Detection and Prevention Technique (ODPT), wherein the at least one ODPT is operable to analyse the first output.Claim 5. The artificial intelligence system according to Claim 4, wherein the at least one ODPT is selected from a group consisting of poisoned model D&P, model jailbreak D&P, data leakage D&P, and toxicity or harmful content D&P.Claim 6. The artificial intelligence system according to anyone of Claim 1 - Claim 5, wherein the first model protection system comprises an application programming interface configured to be in communication with the wrapper.Claim 7. The artificial intelligence system according to Claim 1 , wherein the first model protection system is extensible to incorporate one or more detection and prevention techniques.Claim 8. The artificial intelligence system according to Claim 1 , wherein the model is a generative Al model, classification Al model, regression Al model, supervised or unsupervised Al model.Claim 9. A security assessment system for a target model, comprising an assessment system, a red teaming system; wherein the assessment system is configured to submit an assessment input to the target model for generating an assessment output, and to further receive the assessment output for analysis; wherein the red teaming system is configured to submit a red team input to the target model for generating a red team output, and to further receive the first red team output for analysis; wherein a first pair comprising the assessment input and the assessment output is configured to be provided to the red teaming system for improving attack coverage of the red teaming system; and,wherein a second pair comprising the red team input and the red team output is configured to be provided to the assessment system for improving assessment coverage of the assessment system.Claim 10. The security assessment system according to Claim 9, wherein the security assessment system further comprising a second model protection system and a wrapper; wherein the red team input is first directed to the wrapper, which is configured to direct the red team input to the second model protection system; wherein the second model protection system is configured to receive and harden the red team input to generate a hardened red team input; wherein the wrapper is configured to receive the hardened red team input and to further direct the hardened red team input to the target model; wherein the target model is operable to receive the hardened red team input and to further generate the red team output, which is directed to the wrapper; wherein the wrapper is configured to receive the red team output and to further direct the red team output to the second model protection system; wherein the second model protection system is configured to receive the red team output and to further generate the hardened red team output, which is then directed to the wrapper; wherein the wrapper is configured to receive the hardened red team output and to further direct the hardened red team output to the red team system for analysis.Claim 11. The security assessment system according to Claim 9 or Claim 10, wherein the assessment system comprises at least one assessment technique.Claim 12. The security assessment system according to Claim 11 , wherein the at least one assessment technique is selected from a group consisting of prompt injection or jailbreak testing, training data reconstruction testing, and compliance assessment.Claim 13. The security assessment system according to anyone of Claim 9 - Claim 12, wherein the red teaming system comprises at least one red teaming technique.Claim 14. The security assessment system according to Claim 13, wherein the at least one red teaming technique is prompt fuzzing and penetration testing.Claim 15. The security assessment system according to Claim 10, wherein the second model protection system comprises at least one Input-based Detection and Prevention Technique (IDPT), wherein the at least one IDPT is operable to analyse and harden the red team input.Claim 16. The security assessment system according to Claim 15, wherein the at least one IDPT is selected from a group consisting of poisoned data D&P, prompt injection D&P, and model theft or extraction D&P.Claim 17. The security assessment system according to Claim 10, wherein the second model protection system comprises at least one Output-based Detection and Prevention Technique (ODPT), wherein the at least one ODPT is operable to analyse and harden the red team output.Claim 18. The security assessment system according to Claim 17, wherein the at least one ODPT is selected from a group consisting of poisoned model D&P, model jailbreak D&P, data leakage D&P, and toxicity or harmful content D&P.Claim 19. The security assessment system according to Claim 9, wherein the assessment system and the red teaming system are extensible to incorporate additional techniques.Claim 20. The security assessment system according to Claim 10, wherein the second model protection system is extensible to incorporate one or more additional detection and prevention techniques.Claim 21. The security assessment system according to anyone of Claim 9 - 20, wherein the target model is a generative Al model, classification Al model, regression Al model, supervised or unsupervised Al model.