Real-time threat determination method and device based on automatic artificial intelligence model vulnerability assessment
Patent Information
- Application Number
- KR1020240198169
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2026-09-09
- Estimated Expiration
- 2044-12-27
Smart Images

Figure 112024144817128-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an apparatus and method for improving the reliability of an artificial intelligence model. Background Technology
[0003] Jailbreak attacks, a representative example of integrity vulnerabilities, refer to injecting malicious input into an Artificial Intelligence model to generate harmful or prohibited content.
[0004] While technologies for evaluating the vulnerabilities of AI models have existed, existing adversarial attack-based techniques have all focused on a single type of AI model, making them difficult to apply to other threat types or data modalities where input generation methods differ.
[0005] In addition, existing technologies had limitations in responding to new vulnerabilities and presented the problem of having to train AI models with vulnerability data. Prior art literature
[0006] (Patent Document 0001) KR 2023-0126423 A1 The problem to be solved
[0007] The present invention is intended to provide a determination device and a determination method that automatically evaluate the vulnerability of an artificial intelligence model and determine threats targeting the artificial intelligence model in real time. means of solving the problem
[0009] The determination device of the present invention may include: an evaluation unit that generates vulnerable data that causes jailbreaking and determines a crash of an artificial intelligence model into which the generated vulnerable data is input; and a determination unit that determines the vulnerability of input data input to the artificial intelligence model.
[0010] Here, the evaluation unit can generate the vulnerability data by mutating the initial seed, add the vulnerability data that caused the crash to the seed queue, and determine the vulnerability in real time by comparing the vulnerability data added to the seed queue with the input data.
[0011] Here, the evaluation unit may include: a collection unit that collects an initial seed and selects a specific seed; a generation unit that generates vulnerable data that causes a jailbreak by mutating the specific seed selected by the collection unit and inputs it into an artificial intelligence model; and an evaluation unit that determines a crash involving a breach of confidentiality or a breach of integrity through the analysis of data output from the artificial intelligence model into which the vulnerable data is input.
[0012] Here, the system further includes a detection unit that detects newly generated coverage in an artificial intelligence model into which the above-mentioned vulnerable data is input, and the evaluation unit may apply coverage feedback when the new coverage is detected by the detection unit or when it is determined that the above-mentioned crash has occurred.
[0013] Here, the collection unit selects the specific seed using any one of a plurality of selection algorithms, the collection unit changes the selection algorithm applied to the selection of the specific seed according to a first condition, and the collection unit may continuously apply the currently applied selection algorithm when the coverage feedback is applied.
[0014] Here, if the coverage feedback is not applied, the collection unit may reapply the change in the selection algorithm according to the first condition after the first set time has elapsed from the point in time when the coverage feedback was not applied.
[0015] Here, the generator mutates the specific seed using any one of a plurality of mutation algorithms, the generator changes the mutation algorithm applied to the mutation of the specific seed according to a second condition, and the generator can continuously apply the currently applied mutation algorithm when the coverage feedback is applied.
[0016] Here, if the coverage feedback is not applied, the generation unit may reapply the change in the variation algorithm according to the second condition after the second set time has elapsed from the point in time when the coverage feedback was not applied.
[0017] Here, the judgment unit may further include: an input guard unit that determines whether the input data is vulnerable when input data targeting the artificial intelligence model is input; and an output guard unit that determines whether the output data output from the artificial intelligence model to which the input data is input is vulnerable.
[0018] Here, the input guard unit blocks the input data from being provided to the artificial intelligence model if the similarity between the pre-collected vulnerable input sample and the input data satisfies a first setting value, and if the similarity satisfies a second setting value that is lower than the first setting value, the input guard unit inserts tag information into the input data, and the input guard unit provides potential threat data corresponding to the input data with the inserted tag information to the artificial intelligence model, and if the similarity does not satisfy both the first setting value and the second setting value, the input guard unit may provide the input data to the artificial intelligence model as is.
[0019] Here, a tag unit positioned between the input guard unit and the artificial intelligence model is further included, and the tag unit can tune the artificial intelligence model to process the potential threat data within the scope of maintaining the original function of the artificial intelligence model.
[0020] The tag unit can add a new layer to the artificial intelligence model while maintaining the current layer structure and parameter values of the artificial intelligence model, and the tag unit can adjust the parameter values of the new layer using the vulnerable input sample.
[0021] Here, the system includes an embedding unit that collects an initial seed, which is the source of the vulnerable data that caused the crash, from the evaluation unit, and the embedding unit may provide the initial seed, which is the source of the vulnerable data, to the input guard unit as the vulnerable input sample.
[0022] Here, the output guard unit may include an embedding unit that blocks the output data from being provided to the user when the similarity between the pre-collected vulnerable output sample and the output data satisfies a third setting value, and provides the output result of the artificial intelligence model determined as a crash by the evaluation unit to the output guard unit as the vulnerable output sample.
[0023] The determination method of the present invention is a method performed by a computing device having one or more processors and a memory that stores one or more programs executed by said one or more processors, and may include: an evaluation step of generating vulnerable data that causes jailbreaking and determining a crash of an artificial intelligence model into which the generated vulnerable data is input; and a determination step of determining the vulnerability of input data input to said artificial intelligence model using at least one of an initial seed that is the source of the vulnerable data that caused the crash or an output result of said artificial intelligence model determined to be a crash. Effects of the invention
[0025] The determination device of the present invention can provide an environment in which a robust artificial intelligence model can be built from vulnerable data that induces the jailbreak of an artificial intelligence model.
[0026] For example, the judgment device can automatically generate data to be input into an artificial intelligence model and use it to evaluate the vulnerability of the artificial intelligence model to jailbreaking or detect initial seeds, such as text that threatens the artificial intelligence model.
[0027] The evaluation results or the detection results of the initial seed can be used as a means to determine the threat of input data subsequently entered by the user or to prevent jailbreaking caused by said input data.
[0028] The judgment device of the present invention can reinforce the vulnerabilities of an artificial intelligence model as needed, and in this case, a reinforcement method carried out within the scope of maintaining the original function of the artificial intelligence model can be proposed by the judgment device.
[0029] The decision device of the present invention can be formed with a structure that operates at the input and output terminals of an artificial intelligence model, instead of a method of training the artificial intelligence model itself. Through this, the decision device of the present invention is advantageous for maintaining the original functions of the artificial intelligence model and can be applied to various types of artificial intelligence models.
[0030] Consequently, the judgment device of the present invention can provide a coverage-based automatic evaluation system that is universally applicable to various artificial intelligence models.
[0031] In addition, according to the judgment device of the present invention, the scalability and accuracy of an input guard or output guard utilizing embedding information can be improved. Brief explanation of the drawing
[0033] FIG. 1 is a block diagram showing the judgment device of the present invention. Figure 2 is a schematic diagram showing an evaluation unit and a judgment unit included in a judgment device. Figure 3 is a flowchart illustrating the determination method of the present invention. Specific details for implementing the invention
[0034] Hereinafter, exemplary embodiments according to the present invention will be described in detail with reference to the contents described in the attached drawings. However, the present invention is not limited or restricted by exemplary embodiments. Unless otherwise defined, all terms used in this specification (including technical and scientific terms) shall be used in a meaning that is commonly understood by those skilled in the art to which this disclosure belongs, but this may vary depending on the intent of those skilled in the art, case law, the emergence of new technology, etc.
[0035] Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. In certain cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant explanatory sections. Accordingly, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.
[0036] Throughout this specification, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, the singular form used in this specification includes the plural form unless specifically stated otherwise. Additionally, the expression "at least one of a, b, and / or c" as used throughout this specification may encompass 'a alone', 'b alone', 'c alone', 'a and b', 'a and c', 'b and c', or 'a, b, and c all'.
[0037] Meanwhile, terms such as "first and / or second" used in this specification may be used to describe various components, but they are used solely for the purpose of distinguishing one component from another and are not intended to limit the scope to the components referred to by such terms. For example, without departing from the scope of the present invention, the first component may be named the second component, and the second component may also be named the first component.
[0038] Additionally, terms such as “…part,” “…module,” etc., as described in this specification refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software. Furthermore, embodiments of this disclosure may be represented in this specification by functional block configurations and various processing steps. These functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, embodiments of this disclosure may employ integrated circuit configurations such as memory, processing, logic, look-up tables, etc., which can execute various functions under the control of one or more microprocessors or other control devices.
[0039] In an embodiment according to the present disclosure, functions related to artificial intelligence may be implemented through a processor and memory. In this case, the processor may be any one of a general-purpose processor such as a CPU (Center Processing Unit), AP (Application Processor), DSP (Digital Signal Processor), a graphics-dedicated processor such as a GPU (Graphic Processing Unit) or VPU (Vision Processing Unit), and an artificial intelligence-dedicated processor such as an NPU (Neural Network Processing Unit). The processor may process input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the processor is an artificial intelligence-dedicated processor, the artificial intelligence-dedicated processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model. In some embodiments according to the present disclosure, functions related to artificial intelligence may be implemented through a plurality of processors.
[0040] In an embodiment according to the present disclosure, a predefined operation rule or artificial intelligence model may be configured to perform machine learning. Here, being configured to perform machine learning means that the predefined operation rule or artificial intelligence model is configured to perform a desired characteristic (or objective) by learning using a plurality of training data based on a learning algorithm. Such learning may be performed on the device itself in which the artificial intelligence according to the present disclosure is implemented, or it may be performed through a separate server and / or system.
[0041] Artificial intelligence models can be implemented as neural networks (or artificial neural networks) and can operate based on statistical learning algorithms that mimic biological neurons in machine learning and cognitive science. A neural network can refer to a model in which artificial neurons (nodes), which form a network through synaptic connections, change the strength of synaptic connections through learning to possess problem-solving capabilities. A neural network can be composed of multiple neural network layers; for example, a neural network may include an input layer, a hidden layer, and an output layer. Each of the multiple neural network layers may include at least one node and at least one weight, and neural network operations can be performed through operations between the results of previous (precious) layers and the weights. At least one weight possessed by the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, at least one weight may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Neural networks can infer a result to be predicted from an arbitrary input.
[0042] The learning methods of artificial intelligence models can be classified according to the learning approach into supervised learning, where input and output data are provided as training data and the correct answer (output data) corresponding to the problem (input data) is predetermined; unsupervised learning, where only input data is provided without output data and the correct answer (output data) corresponding to the problem (input data) is not predetermined; and reinforcement learning, where a reward is granted whenever an action is taken from the current state and learning proceeds in a direction that maximizes this reward. Alternatively, they can be classified according to the architecture, which is the structure of the learning model.
[0043] In the embodiments of the present disclosure, the artificial intelligence model is a Convolutional Neural Network (CNN) such as GoogleNet, AlexNet, VGG Network, Region with Convolutional Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based Deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restructured Boltzmann Machine (RBM), Fully Convolutional Network, Long Short-Term Memory Network (LSTM), Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA for Natural Language Processing, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics, Visual Understanding, Video Synthesis for Vision Processing, Anomaly Detection, Prediction, Time-Series Forecasting, Optimization, Recommendation for ResNet Data Intelligence, At least one of various artificial intelligence structures and algorithms, such as data creation, may be used. The examples described above are merely examples of artificial intelligence structures and algorithms used according to the embodiments of the present disclosure and do not limit the artificial intelligence structures and algorithms used according to the embodiments of the present disclosure.
[0044] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In describing the embodiments, technical details that are well known in the art to which the present invention pertains and are not directly related to the present invention will be omitted. This is to ensure that the essence of the present invention is conveyed more clearly without obscuring it by omitting unnecessary explanations. For the same reason, some components in the accompanying drawings may be exaggerated, omitted, or schematically depicted. Furthermore, the size of each component does not entirely reflect its actual size. Throughout this specification, the same reference numerals may refer to the same or corresponding components.
[0045] FIG. 1 is a block diagram showing a judgment device (1000) of the present invention. FIG. 2 is a schematic diagram showing an evaluation unit (100) and a judgment unit (200) included in the judgment device (1000).
[0046] Referring to FIG. 1, the judgment device (1000) may include an evaluation unit (100) and a judgment unit (200).
[0047] The evaluation unit (100) can generate vulnerable data that causes a jailbreak. Additionally, the evaluation unit (100) can determine a crash of the artificial intelligence model (90) into which the generated vulnerable data is input.
[0048] 'Jailbreak' can refer to injecting malicious input into an Artificial Intelligence model to generate harmful or prohibited content. Therefore, vulnerable data that triggers a jailbreak may include instructions to build a bomb in stages.
[0049] In this document, the term 'vulnerability' may refer to the degree to which an artificial intelligence model complies with jailbreak questions.
[0050] A crash of the artificial intelligence model (90) may mean a state in which the artificial intelligence model (90) executes the jailbreak as is. That is, a crash may mean a state in which the output result of the artificial intelligence model (90) violates confidentiality or integrity.
[0051] As used herein, the term "breach of confidentiality" may mean the disclosure of specific information that is to be kept confidential. As used herein, the term "breach of integrity" may mean the inclusion of data that should not be included.
[0052] The judgment unit (200) can determine the vulnerability of input data input into the artificial intelligence model (90) in real time.
[0053] At this time, the evaluation unit (100) can generate vulnerable data by mutating the initial seed.
[0054] In this document, the term 'seed' may refer to an initial value used to control randomness, and specifically, the 'initial seed' may be multimodal, which is various forms of data such as text, images, and voice, and multimodal may also include multistep, such as instructions to build a bomb step by step.
[0055] Vulnerable data may refer to data that caused a crash in the artificial intelligence model (90).
[0056] The evaluation unit (100) can add vulnerable data that caused the crash to the seed queue.
[0057] Specifically, the evaluation unit (100) may include a seed queue that starts fussing with an initial seed to receive data and stores an effective seed, and a seed selection module that selects an effective seed. The seed queue that stores the seed and the seed selection module may be integrated into a collection unit (110).
[0058] The evaluation unit (100) may use an algorithm for seed selection. Here, the algorithm for seed selection may be random selection, round robin, Monte Carlo Tree Search, etc.
[0059] The evaluation unit (100) may include both generative artificial intelligence model-based mutation modules and adversarial attack-based mutation modules. Accordingly, the evaluation unit (100) can effectively generate inputs that cause data reverse extraction or theft or adversarial disturbance.
[0060] For example, a generative AI-based mutation module can use a seed to generate vulnerable data as a jailbreak prompt.
[0061] For example, an adversarial attack-based mutation module can generate vulnerable data by passing a loss function as a gradient to the input layer to perturb the input data.
[0062] Additionally, the evaluation unit (100) can be configured to select an algorithm to be applied through coverage feedback.
[0063] The evaluation unit (100) can apply coverage measured in the hidden layer of the target artificial intelligence model (90) to analyze the internal state of the model and add vulnerable inputs to the seed queue.
[0064] The evaluation unit (100) can process the hidden layer output information according to the coverage measurement method and check whether new coverage has been found compared to the existing one.
[0065] Here, the coverage measurement method can use the neuron activation output value or gradient value of the hidden layer as coverage information, and for this purpose, the evaluation unit (100) can compare the distance, variance, etc. with the coverage value of the existing seed for the new coverage standard.
[0066] The evaluation unit (100) may apply an ensemble model of a confidentiality determination model and an integrity determination model to simultaneously and universally determine whether the confidentiality / integrity of the output of the same artificial intelligence model (90) has been violated.
[0067] When a confidentiality or integrity breach is determined by applying a previously trained confidentiality / integrity determination model, it is called a crash, and vulnerable data that caused the crash can be added to a seed queue.
[0068] When the evaluation unit (100) determines that the coverage is new, it analyzes the coverage information and can provide coverage feedback that is valid when compared with the seed queue only when the coverage is new or effective in determining confidentiality / integrity.
[0069] The judgment unit (200) can determine vulnerability in real time by comparing vulnerability data added to the seed queue with input data.
[0070] Specifically, the judgment unit (200) can efficiently block risk inputs in a rule-based manner in the embedding space and insert tag information into the input data identified as potential threats in a subsequent step to process the input data specially.
[0071] The judgment unit (200) can convert input data into an embedding vector and compare the similarity with the embedding vector of a vulnerable input sample collected in advance.
[0072] The judgment unit (200) can insert tag information by input data that is not highly similar but is potentially a threat.
[0073] The judgment unit (200) can closely analyze the output result of the artificial intelligence model (90) based on the input in the embedding space and combine it with the learned detection model to comprehensively utilize it for the final threat judgment.
[0074] The judgment unit (200) can convert the output of the target model into an embedding vector and compare the similarity with the embedding vector of a vulnerable output sample collected in advance. At this time, the judgment unit (200) can additionally utilize an artificial intelligence model (90) that has learned the embedding vectors of the vulnerable output samples. Consequently, the judgment unit (200) can determine whether there is a final threat by combining the embedding vector similarity information and the judgment result of the artificial intelligence model (90).
[0075] The judgment unit (200) can tune the model so that the model can process specific information (input data) tagged as a potential threat on its own without compromising the function of the target artificial intelligence model (90).
[0076] Specifically, the judgment unit (200) can maintain the current layer structure and parameter values of the target model. The judgment unit (200) can adjust the parameters by learning from the collected vulnerable samples after adding a new layer.
[0077] The judgment unit (200) can apply embedding space expansion and optimization to utilize embedding information of vulnerable inputs and outputs discovered through an automatic vulnerability assessment tool.
[0078] Whenever a new vulnerability is discovered, the judgment unit (200) can extract the input embedding for the sample from the embedding layer of the input variant model and the output embedding from the embedding layer of the judgment model and add them to the vector database.
[0079] Referring to FIG. 2, the evaluation unit (100) may include a collection unit (110), a generation unit (130), an evaluation unit (170), and a detection unit (190).
[0080] The collection unit (110) can collect initial seeds and select specific seeds. For example, the collection unit (110) can select all collected initial seeds or select some of the collected initial seeds.
[0081] The generating unit (130) can generate vulnerable data that causes a jailbreak by mutating a specific seed selected by the collecting unit (110) and input it into the artificial intelligence model (90).
[0082] The evaluation unit (170) can determine whether a crash involving a breach of confidentiality or a breach of integrity is present by analyzing the data output from the artificial intelligence model (90) into which vulnerable data is input.
[0083] The detection unit (190) can detect new coverage newly generated in the artificial intelligence model (90) into which vulnerable data is input. Coverage can represent an indicator that numerically indicates how much of the source code a test case of the software has executed.
[0084] The evaluation unit (170) can apply coverage feedback when a new coverage is detected or a crash is determined by the detection unit (190).
[0085] Various operations can be performed by the evaluation unit (100) using coverage feedback. For example, coverage feedback can affect the operation of the collection unit (110) or the operation of the generation unit (130).
[0086] The collection unit (110) can select a specific seed using any one of a plurality of selection algorithms.
[0087] The collection unit (110) can change the selection algorithm applied to the selection of a specific seed according to a first condition. For example, the first condition may include whether the set number of times is satisfied, whether the set time has elapsed, etc. Taking the case where the first condition is the set number of times as an example, the collection unit (110) can select a specific seed among the initial seeds by applying the first selection algorithm for the set time. When the set time has elapsed, the collection unit (110) can select a specific seed among the initial seeds using the second selection algorithm.
[0088] The collection unit (110) can continuously apply the currently applied selection algorithm when coverage feedback is applied. Through this, specific seeds with a similar tendency to the initial seed satisfying the coverage feedback can be continuously selected by the same selection algorithm. According to this, specific seeds effective for self-evaluation of jailbreak can be selected in large quantities.
[0089] If coverage feedback is not applied, the collection unit (110) can reapply the selection algorithm change according to the first condition after the first set time has elapsed from the point in time when coverage feedback is not applied. In other words, the collection unit (110) can operate in a manner similar to exploring a new area after all the minerals in a certain area have been mined.
[0090] The generation unit (130) can mutate the specific seed using any one of a plurality of mutation algorithms.
[0091] The generating unit (130) can change the mutation algorithm applied to the mutation of a specific seed according to a second condition. The second condition may include whether the set number of times is satisfied, whether the set time has elapsed, etc., similar to the first condition.
[0092] When coverage feedback is applied, the generation unit (130) can continuously apply the mutation algorithm currently being applied.
[0093] If coverage feedback is not applied, the generation unit (130) may reapply the mutation algorithm change according to the second condition after the second set time has elapsed from the point in time when coverage feedback was not applied. A situation in which coverage feedback is not applied may be similar, for example, to a state in which all minerals being mined in a certain area during a mineral mining operation have been depleted. In this case, it is advantageous to detect other areas, and accordingly, the mutation algorithm according to the second condition may be reapplied.
[0094] The judgment unit (200) illustrated in FIG. 2 may include an input guard unit (210), an output guard unit (230), a tag unit (250), and an embedding unit (270). FIG. 2 shows that one artificial intelligence model (90) is included in the evaluation unit (100) and one is included in the judgment unit (200), but the artificial intelligence model (90) may be formed as a single unit. In other words, the two artificial intelligence models (90) illustrated in FIG. 2 may correspond to a single model.
[0095] The input guard (210) can determine whether the input data is vulnerable when input data targeting the artificial intelligence model (90) is received.
[0096] The output guard (230) can determine whether the output data output from the artificial intelligence model (90) into which the input data is input is vulnerable.
[0097] The input guard (210) can block the input data from being provided to the artificial intelligence model (90) if the similarity between the input data and the previously collected vulnerable input samples (such as the vector similarity described above) satisfies the first setting value.
[0098] The input guard unit (210) can insert tag information into the input data if it satisfies a second setting value that is lower than the first setting value in terms of similarity.
[0099] The input guard unit (210) can provide potential threat data corresponding to the input data into which tag information is inserted to the artificial intelligence model (90).
[0100] If the similarity does not satisfy the first and second setting values, the input guard (210) can provide the input data as is to the artificial intelligence model (90) without the process of inserting separate tag information.
[0101] The tag section (250) can be placed between the input guard section (210) and the artificial intelligence model (90) in terms of data flow.
[0102] The tag unit (250) can tune the artificial intelligence model (90) to process potential threat data within the scope of maintaining the original function of the artificial intelligence model (90).
[0103] The tag section (250) can add a new layer to the artificial intelligence model (90) while maintaining the current layer structure and parameter values of the artificial intelligence model (90).
[0104] The tag section (250) can adjust the parameter values of the new layer using vulnerable input samples.
[0105] The evaluation unit (100) described above can generate vulnerable data that causes a jailbreak and determine a crash of the artificial intelligence model (90) into which the generated vulnerable data is input.
[0106] The embedding unit (270) can collect an initial seed that is the source of vulnerable data that caused the crash from the evaluation unit (100).
[0107] The embedding unit (270) can provide an initial seed that is a source of vulnerable data to the input guard unit (210) as a vulnerable input sample.
[0108] The output guard (230) can block the output data from being provided to the user if the similarity between the pre-collected vulnerable output sample and the output data satisfies a third setting value.
[0109] The embedding unit (270) can provide the output result of the artificial intelligence model (90) determined to be a crash from the evaluation unit (100) to the output guard unit (230) as a vulnerable output sample.
[0110] Figure 3 is a flowchart illustrating the determination method of the present invention.
[0111] The determination method of FIG. 3 may be performed by the determination device (1000) of FIG. 1. Alternatively, the determination method of FIG. 3 may be performed by a computing device having one or more processors and a memory that stores one or more programs executed by said one or more processors.
[0112] The judgment method may include an evaluation step (S510) and a judgment step (S520).
[0113] Step S510 may be a step of generating vulnerable data that causes a jailbreak and determining a crash of an artificial intelligence model (90) into which the generated vulnerable data is input. Step S510 may be performed by an evaluation unit (100).
[0114] Step S520 may be a step of determining the vulnerability of input data input to the artificial intelligence model (90) in real time using at least one of the initial seed that is the source of the vulnerable data that caused the crash or the output result of the artificial intelligence model (90) determined as a crash. Step S520 may be performed by a determination unit (200).
[0115] Although embodiments of the present invention have been described in detail above, the scope of the present invention is not limited thereto, and various modifications and improvements made by a person skilled in the art using the basic concept of the present invention as defined in the following claims also fall within the scope of the present invention.
Claims
Claim 1 An evaluation unit that generates vulnerable data that causes jailbreaking and determines a crash of an artificial intelligence model into which the generated vulnerable data is input; and a determination unit that determines the vulnerability of input data input into the artificial intelligence model; wherein the determination unit includes an input guard unit that determines whether the input data is vulnerable when input data targeting the artificial intelligence model is input; and an output guard unit that determines whether the output data output from the artificial intelligence model into which the input data is input is vulnerable; wherein the input guard unit blocks the input data from being provided to the artificial intelligence model if the similarity between the input data and a previously collected vulnerable input sample satisfies a first set value, and the input guard unit inserts tag information into the input data if the similarity satisfies a second set value that is lower than the first set value, and the input guard unit provides potential threat data corresponding to the input data into which the tag information is inserted to the artificial intelligence model, and the input guard unit provides the input data to the artificial intelligence model as is if the similarity does not satisfy both the first set value and the second set value. Claim 2 In claim 1, the evaluation unit generates vulnerability data by mutating an initial seed, adds the vulnerability data that caused the crash to a seed queue, and determines vulnerability through comparison between the vulnerability data added to the seed queue and the input data. Claim 3 In claim 1, the evaluation unit comprises: a collection unit that collects an initial seed and selects a specific seed; a generation unit that generates vulnerable data that causes a jailbreak by mutating the specific seed selected by the collection unit and inputs it into an artificial intelligence model; and an evaluation unit that determines a crash containing a breach of confidentiality or a breach of integrity through the analysis of data output from the artificial intelligence model into which the vulnerable data is input. Claim 4 In paragraph 3, the judgment device further includes a detection unit that detects newly generated coverage in an artificial intelligence model into which the above-mentioned vulnerable data is input, and the evaluation unit that applies coverage feedback when the above-mentioned new coverage is detected by the detection unit or when the above-mentioned crash is determined to have occurred. Claim 5 In paragraph 4, the collection unit selects the specific seed using any one of a plurality of selection algorithms, the collection unit changes the selection algorithm applied to the selection of the specific seed according to a first condition, and the collection unit continuously applies the currently applied selection algorithm when the coverage feedback is applied. Claim 6 In paragraph 5, the collection unit is a determination device that, if the coverage feedback is not applied, reapplies the selection algorithm change according to the first condition after a first set time has elapsed from the point in time when the coverage feedback was not applied. Claim 7 In paragraph 4, the generation unit mutates the specific seed using any one of a plurality of mutation algorithms, the generation unit changes the mutation algorithm applied to the mutation of the specific seed according to a second condition, and the generation unit continuously applies the currently applied mutation algorithm when the coverage feedback is applied. Claim 8 In claim 7, the generation unit is a determination device that, if the coverage feedback is not applied, reapplies the variation algorithm change according to the second condition after a second set time has elapsed from the point in time when the coverage feedback was not applied. Claim 9 delete Claim 10 delete Claim 11 In claim 1, a tag unit is further included between the input guard unit and the artificial intelligence model, and the tag unit is a determination device that tunes the artificial intelligence model to process potential threat data within the range of maintaining the original function of the artificial intelligence model. Claim 12 In claim 11, the tag unit adds a new layer to the artificial intelligence model while maintaining the current layer structure and parameter values of the artificial intelligence model as they are, and the tag unit adjusts the parameter values of the new layer using the vulnerable input sample. Claim 13 A determination device according to claim 1, comprising an embedding unit that collects an initial seed which is the source of vulnerable data that caused the crash from the evaluation unit, and wherein the embedding unit provides the initial seed which is the source of vulnerable data as the vulnerable input sample to the input guard unit. Claim 14 delete Claim 15 A method performed by a computing device having one or more processors and a memory storing one or more programs executed by said one or more processors, comprising: an evaluation step in which an evaluation unit generates vulnerable data that causes a jailbreak and determines a crash of an artificial intelligence model into which the generated vulnerable data is input; and a determination step in which a determination unit determines the vulnerability of input data input to said artificial intelligence model using at least one of an initial seed that is the source of the vulnerable data that caused the crash or an output result of said artificial intelligence model determined to be a crash; wherein the determination unit includes an input guard unit that determines whether the input data is vulnerable when input data targeting said artificial intelligence model is input; A determination method comprising: an output guard unit that determines whether the input data is vulnerable and output data is output from the artificial intelligence model into which the input data is input; wherein the input guard unit blocks the input data from being provided to the artificial intelligence model if the similarity between the input data and a previously collected vulnerable input sample satisfies a first set value, and the input guard unit inserts tag information into the input data if the similarity satisfies a second set value which is lower than the first set value, and the input guard unit provides potential threat data corresponding to the input data with the inserted tag information to the artificial intelligence model, and the input guard unit provides the input data to the artificial intelligence model as is if the similarity does not satisfy both the first set value and the second set value.