Large model attack evaluation method and device, electronic equipment and storage medium

CN122548731APending Publication Date: 2026-08-11BEIJING CHJ AUTOMOTIVE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本发明实施例所要解决的技术问题是提供一种大模型攻击评估方法、装置、电子设备、可读存储介质,以便解决评估耐药性现象和的技术问题

Benefits of technology

[0048]依据本发明实施例,通过根据攻击语料集和/或采集的外部信息,生成攻击检测集合,确定所述攻击检测集合对被测大模型进行攻击的攻击性评估结果,基于所述攻击性评估结果,根据所述攻击检测集合,对所述攻击语料集进行更新,使得通过攻击语料集的自发现和/或通过外部信息的补充,提供一个能够自我扩展和/或动态更新的攻击检测集合,从而反映未知的大模型安全风险,有效减少评估耐药性现象的出现,实现了实时监控与适应性调整机制,能够自动收集和整合新的攻击语料,及时对监管变化做出快速响应,并减少对人工的依赖,提高攻击检测能力的改进效率,继而实现自动的安全防护的持续增强和改进。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548731A_ABST
    Figure CN122548731A_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, electronic device, and storage medium for large-scale model attack assessment. The method includes: generating an attack detection set based on an attack corpus and / or collected external information; determining the attack assessment result of the attack detection set on the tested large-scale model; and updating the attack corpus based on the attack assessment result and the attack detection set. This allows for a self-expanding and / or dynamically updated attack detection set through self-discovery of the attack corpus and / or supplementation by external information, thereby reflecting unknown large-scale model security risks, effectively reducing the occurrence of assessment resistance, realizing a real-time monitoring and adaptive adjustment mechanism, automatically collecting and integrating new attack corpora, responding quickly to regulatory changes, reducing reliance on manual intervention, improving the efficiency of attack detection capability improvement, and ultimately achieving continuous enhancement and improvement of automatic security protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for evaluating large model attacks, a device for evaluating large model attacks, an electronic device, and a readable storage medium. Background Technology

[0002] Currently, artificial intelligence and large-scale model technology are developing rapidly. Large-scale models can be divided into two main categories: Large Language Models (LLMs) and Large Multimodal Models (LMMs). Security assessment of large-scale models is becoming increasingly important. A large-scale model is a deep learning model with a large number of parameters. While it demonstrates exceptional capabilities in processing massive amounts of data and performing complex tasks, it also carries potential security vulnerabilities and risks. The purpose of large-scale model security assessment is to ensure that the large-scale model is safe, reliable, and conforms to human value standards, expecting it to be a responsible form of artificial intelligence and avoiding security risks caused by attacks on large-scale models. Currently known large-scale model security assessment schemes are basically as follows: Figure 1 The architecture shown.

[0003] Through research, the inventors have discovered that in existing technologies, the safety detection set is static and can only be expanded manually or through procurement. This makes it difficult to reflect dynamic iteration issues, resulting in obvious phenomena in the assessment of drug resistance. Furthermore, it is impossible to discover and integrate the latest large-scale model safety information in a timely manner, especially for adaptive adjustments to the latest regulatory content. Only purely manual feedback can be provided, which consumes a lot of manpower and cannot respond quickly. Summary of the Invention

[0004] The technical problem to be solved by the embodiments of the present invention is to provide a method, apparatus, electronic device, and readable storage medium for evaluating large-scale attack phenomena, so as to solve the technical problem of evaluating drug resistance phenomena.

[0005] To address the above problems, this invention provides a method for evaluating large-scale model attacks, the method comprising:

[0006] Generate an attack detection set based on the attack corpus and / or collected external information;

[0007] Determine the attack assessment results of the attack detection set on the large model under test;

[0008] Based on the attack assessment results, the attack corpus is updated according to the attack detection set.

[0009] Optionally, generating an attack detection set based on the attack corpus and / or collected external information includes:

[0010] The attack corpus and / or collected external information are transformed using attack methods to obtain a set of attack prompt words.

[0011] An attack detection set is determined based on the attack prompt word set; wherein the attack detection set includes multiple attack detection inputs.

[0012] Optionally, determining the attack detection set based on the attack prompt word set includes:

[0013] The attack detection set is formed by combining the attack prompt word set and the attack scenario set, wherein the attack scenario set includes multiple attack scenario templates.

[0014] Optionally, determining the attack assessment result of the attack detection set on the tested large model includes:

[0015] The attack detection set is input into the tested large model and at least one reference large model to obtain the corresponding tested output and reference output.

[0016] The attack assessment result is determined based on the tested output and the reference output.

[0017] Optionally, determining the attack assessment result based on the tested output and the reference output includes:

[0018] The tested output and the reference output are input into the attack assessment model to obtain the attack assessment result. The attack assessment model is used to evaluate the similarity between the tested output and the reference output after semantic understanding, and to determine the attack assessment result based on the similarity.

[0019] Optionally, updating the attack corpus based on the attack assessment result and according to the attack detection set includes:

[0020] For any attack detection input in the attack detection set, if the attack evaluation result meets the target condition, all or part of the attack detection input is determined as new corpus.

[0021] The newly added corpus is added to the attack corpus set.

[0022] Optionally, the attack method includes at least one of the following: prefix and suffix injection, random root injection, and sentence expansion.

[0023] The present invention also provides a large-scale attack evaluation device, the device comprising:

[0024] The set generation module is used to generate an attack detection set based on the attack corpus and / or collected external information;

[0025] The result determination module is used to determine the attack assessment result of the attack detection set on the tested large model;

[0026] The corpus update module is used to update the attack corpus based on the attack assessment results and the attack detection set.

[0027] Optionally, the set generation module includes:

[0028] The transformation submodule is used to transform the attack corpus and / or collected external information using the attack method to obtain a set of attack prompt words;

[0029] The set determination submodule is used to determine the attack detection set based on the attack prompt word set; wherein, the attack detection set includes multiple attack detection inputs.

[0030] Optionally, the set-determining submodule includes:

[0031] The set determination unit is used to combine the attack prompt word set and the attack scenario set into an attack detection set, wherein the attack scenario set includes multiple attack scenario templates.

[0032] Optionally, the device further includes:

[0033] The scene recognition module is used to identify new attack scenarios;

[0034] The scenario addition module is used to add the new attack scenario to the attack scenario set.

[0035] Optionally, the result determination module includes:

[0036] The output acquisition submodule is used to input the attack detection set into the tested large model and at least one reference large model, and acquire the corresponding tested output and reference output.

[0037] The result determination submodule is used to determine the attack assessment result based on the tested output and the reference output.

[0038] Optionally, the result determination submodule includes:

[0039] The result acquisition unit is used to input the tested output and the reference output into the attack assessment model to obtain the attack assessment result, wherein the attack assessment model is used to evaluate the similarity between the tested output and the reference output after semantic understanding, and to determine the attack assessment result based on the similarity.

[0040] Optionally, the corpus update module includes:

[0041] The corpus filtering submodule is used to determine all or part of any attack detection input in the attack detection set as new corpus if the attack evaluation result meets the target condition.

[0042] The corpus addition submodule is used to add the new corpus to the attack corpus set.

[0043] Optionally, the attack method includes at least one of the following: prefix and suffix injection, random root injection, and sentence expansion.

[0044] This invention also discloses an electronic device, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0045] Memory, used to store computer programs;

[0046] The processor, when executing a program stored in memory, implements the method described above.

[0047] This invention also discloses a readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the method described above.

[0048] According to embodiments of the present invention, an attack detection set is generated based on an attack corpus and / or collected external information. The attack detection set is then used to determine the attack capability assessment result of the attack on the tested large model. Based on the attack assessment result, the attack corpus is updated according to the attack detection set. This allows for the self-discovery of the attack corpus and / or the supplementation of external information, providing a self-expanding and / or dynamically updated attack detection set. This reflects unknown security risks of large models, effectively reduces the occurrence of assessment resistance, and realizes a real-time monitoring and adaptive adjustment mechanism. It can automatically collect and integrate new attack corpora, respond quickly to regulatory changes, reduce reliance on manual intervention, and improve the efficiency of attack detection capability improvement, thereby achieving continuous enhancement and improvement of automatic security protection. Attached Figure Description

[0049] Figure 1 A schematic diagram of the architecture of currently known large-scale model security assessment schemes is shown;

[0050] Figure 2 A flowchart illustrating the steps of a large-model attack evaluation method provided by an embodiment of the present invention is shown.

[0051] Figure 3A flowchart illustrating the steps of a large model attack evaluation method provided by another embodiment of the present invention is shown.

[0052] Figure 4 A schematic diagram illustrating the working principle of the Blue Team security inspection is shown;

[0053] Figure 5 A schematic diagram of the blue team's security assessment process is shown;

[0054] Figure 6 This diagram illustrates a structural block diagram of an embodiment of a large-scale attack evaluation apparatus provided by an embodiment of the present invention;

[0055] Figure 7 A structural block diagram of an electronic device for evaluating large-scale attacks is shown according to an exemplary embodiment. Detailed Implementation

[0056] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0057] Reference Figure 2 The diagram illustrates a flowchart of a large-model attack evaluation method provided by an embodiment of the present invention, which may specifically include the following steps:

[0058] Step 101: Generate an attack detection set based on the attack corpus and / or collected external information.

[0059] In this embodiment of the invention, the attack corpus refers to a collection of attack corpora that cause a large model to output specific content (such as sensitive information). For example, an attack corpus used to store knowledge-based corpus information can provide the raw data for corpus processing for the blue team attack engine.

[0060] In this embodiment of the invention, the collected external information refers to information collected from outside the attack corpus, including but not limited to: information from third-party attack corpora, information scraped from the internet, etc. This embodiment of the invention does not impose any limitations on this. For example, in addition to its own attack corpus containing a large amount of data, it can also support third-party test corpora. The blue team attack engine can access offline and online third-party detection corpora, and the attack corpus has an extension interface for enhancement during question bank iteration. The blue team attack engine must also have web scraping capabilities to scrape dynamic information from the internet that can be used for compliant testing, such as current events, hot topics, and sensitive news.

[0061] In this embodiment of the invention, the attack detection set refers to the set of text prompts input when performing attack detection on a large model. That is, the attack detection set is a set of multiple attack detection inputs, where the attack detection inputs are the prompt words used to input the large model for the attack. When generating the attack detection set, the attack corpus and / or collected external information are used as the corpus. The attack detection set can be generated based on the attack corpus, or based on the collected external information, or based on both the attack corpus and the collected external information.

[0062] In this embodiment of the invention, the attack corpus and / or collected external information are processed to generate text prompts for attack detection, thereby obtaining an attack detection set. For example, prefix and suffix injection, random root injection, sentence expansion, etc., can be used; any applicable processing method can be employed, and this embodiment of the invention does not impose any limitations on this.

[0063] In this embodiment of the invention, during the process of generating attack detection input, after the data in the attack data set is processed, some new data that were not originally in the attack data set will be generated. For example, the original attack data in the attack data set is "ABCD", and after the attack method is transformed, the new data "AB%CD" is obtained.

[0064] In this embodiment of the invention, the collected external information may include new data that was not originally in the attack corpus, or it may generate new data that was not originally in the attack corpus after processing.

[0065] In one optional embodiment of the present invention, a specific implementation of generating an attack detection set based on an attack corpus and / or collected external information may include: transforming the attack corpus and / or collected external information using an attack method to obtain an attack prompt word set; and determining the attack detection set based on the attack prompt word set.

[0066] An attack method refers to the way in which attack prompts are generated. These prompts can be generated using an attack corpus and / or collected external information as seeds, and can be different from the seeds themselves. Specifically, any applicable attack method can be used, and this embodiment of the invention does not impose any limitations on this. Using the attack corpus and / or collected external information as seed information, and transforming it using the attack method, a set of attack prompts is generated, denoted as the attack prompt set. Then, based on the attack prompt set, an attack detection set is determined. The attack detection set includes multiple attack detection inputs. The specific implementation of determining the attack detection set based on the attack prompt set can include various methods, such as combining all or part of the attack prompts in the attack prompt set to form the attack detection set, etc. Specifically, any applicable implementation method can be used, and this embodiment of the invention does not impose any limitations on this.

[0067] For example, the Blue Team attack engine can use attack corpus and / or collected external information as input to execute attack methods, thereby generating an attack detection set.

[0068] By varying the attack methods, a wide variety of attack cue words are generated, thus creating a dynamic and effective attack detection set. This allows for more challenging inputs to the large model being tested, continuously maintaining the quality of the blue team's evaluation.

[0069] In an optional embodiment of the present invention, the attack method includes at least one of the following: prefix and suffix injection, random root injection, and sentence expansion.

[0070] The attack method refers to the way in which the attack prompt is generated, such as prefix / suffix injection, random root injection, sentence expansion, or any other applicable attack method. This embodiment of the invention does not limit this. The execution of the attack method can be completed within the Blue Team Attack Engine. The Blue Team Attack Engine can support loading external attack methods through plugins, such as: artificial intelligence methods for large or small models, regular expressions, combinations with web scraping information, combinations with random information, fuzzing, and other attack methods. The attack methods are diverse and can be continuously expanded to adapt to new attack methods, avoiding the problem of insufficient scalability of attack methods.

[0071] Step 102: Determine the attack assessment results of the attack detection set on the tested large model.

[0072] In this embodiment of the invention, the large model under test refers to a large model whose security performance needs to be tested. An attack is performed on the large model under test based on an attack detection set. Specifically, the attack detection input from the attack detection set can be input into the large model under test, and the output of the large model under test can be obtained, denoted as the test output.

[0073] In this embodiment of the invention, the effectiveness of the attack is evaluated to obtain an attack evaluation result of the attack detection set. This attack evaluation result reflects the strength of the attack on the tested large model by the attack detection input in the attack detection set. For example, the similarity between the tested output and the corresponding fixed answer is used as the attack evaluation result. Another example is to input the attack detection input from the attack detection set into the tested large model and at least one reference large model to obtain the corresponding tested output and reference output. The similarity between the tested output and the reference output is used as the attack evaluation result. Higher similarity indicates weaker attack, i.e., more ineffective attack; lower similarity indicates stronger attack, i.e., more effective attack. Any applicable implementation method can be used, and this embodiment of the invention does not limit this.

[0074] For example, the Blue Team attack engine uses the generated attack detection set to attack the large model under test, and sends the output of the large model under test to the Blue Team evaluation engine, which evaluates the attack and obtains the attack evaluation result.

[0075] Step 103: Based on the attack assessment results, update the attack corpus according to the attack detection set.

[0076] In this embodiment of the invention, the attack evaluation result reflects the attack power of the attack detection input in the attack detection set on the tested large model.

[0077] In this embodiment of the invention, when the attack assessment result indicates that the attack detection input in the attack detection set is highly aggressive against the tested large model and meets the attack requirements, the attack detection input in the attack detection set is effective in attacking the tested large model. Therefore, this attack detection input in the attack detection set is helpful for the attack corpus. The attack corpus can be updated based on all or part of this attack detection input, thereby enhancing the attack corpus based on the feedback from the attack assessment. For example, corpora not present in the attack corpus can be selected from the attack detection set and added to the attack corpus.

[0078] In this embodiment of the invention, when the attack evaluation results indicate that the attack detection input in the attack detection set is weakly offensive to the tested large model and does not meet the attack requirements, the attack detection input in the attack detection set is ineffective against the tested large model. Therefore, the attack detection input in the attack detection set is not helpful to the attack corpus and there is no need to update the attack corpus.

[0079] For example, the blue team assessment engine and the blue team attack engine achieve collaborative feedback and optimization, establishing an effective feedback mechanism through assessment and alerts. Compared to existing solutions that lack a feedback mechanism between assessment and attack, where assessment and attack are isolated and fail to achieve autonomous, continuous capability iteration and enhancement, this invention tightly integrates the blue team assessment engine and the blue team attack engine, enabling self-learning and continuous optimization.

[0080] According to embodiments of the present invention, an attack detection set is generated based on an attack corpus and / or collected external information. The attack detection set is then used to determine the attack capability assessment result of the attack on the tested large model. Based on the attack assessment result, the attack corpus is updated according to the attack detection set. This allows for the self-discovery of the attack corpus and / or the supplementation of external information, providing a self-expanding and / or dynamically updated attack detection set. This reflects unknown security risks of large models, effectively reduces the occurrence of assessment resistance, and realizes a real-time monitoring and adaptive adjustment mechanism. It can automatically collect and integrate new attack corpora, respond quickly to regulatory changes, reduce reliance on manual intervention, and improve the efficiency of attack detection capability improvement, thereby achieving continuous enhancement and improvement of automatic security protection.

[0081] Reference Figure 3 The diagram illustrates a flowchart of a large-model attack evaluation method according to another embodiment of the present invention, which may specifically include the following steps:

[0082] Step 201: The attack corpus and / or collected external information are transformed using the attack method to obtain the attack prompt word set.

[0083] In this embodiment of the invention, the specific implementation of this step can be found in the description of the foregoing embodiments, and will not be repeated here.

[0084] In an optional embodiment of the present invention, a specific implementation of determining the attack detection set based on the attack prompt word set includes: combining the attack prompt word set and the attack scenario set to form an attack detection set, wherein the attack scenario set includes multiple attack scenario templates.

[0085] Step 202: Combine the attack prompt word set and the attack scenario set into an attack detection set, wherein the attack scenario set includes multiple attack scenario templates.

[0086] In this embodiment of the invention, the attack scenario refers to the different input methods or attack processes used in the attack on the large model. In this scenario, the attack corpus can be combined to produce threatening security effects, such as target hijacking, code attack, role-playing, etc.

[0087] In this embodiment of the invention, attack scenarios exist as templates within the attack scenario set. Security personnel can extract new scenario templates from ongoing attack research to expand the set. Simultaneously, external knowledge bases can be accessed to discover and organize new attack scenarios, and then corresponding attack scenario templates can be created and added to the attack scenario set.

[0088] In this embodiment of the invention, attack prompts from the attack prompt word set and attack scenario templates from the attack scenario set are combined, and the combined attack prompts are filled into the attack scenario templates to obtain attack detection input, thereby generating a dynamic and effective attack detection set.

[0089] In this embodiment of the invention, multiple combinations can generate attack detection sets, so that attack corpus (including attack corpus set and / or collected external information), attack method, and attack scenario can be combined to form a variety of attack detection sets. By adding more difficult inputs (attack corpus, attack method, attack scenario), the quality of the blue team evaluation can be maintained continuously.

[0090] In an optional embodiment of the present invention, it may further include: identifying new attack scenarios; and adding the new attack scenarios to the attack scenario set.

[0091] During attacks on the tested large model, attack scenarios not found in the attack scenario set are selected, corresponding attack scenario templates are generated, and added to the attack scenario set. By continuously identifying and adding new attack scenarios and adapting to emerging attack scenarios, the Blue Team Attack Engine can expand to cover new and growing attack scenarios, thus increasing the diversity of attack scenarios.

[0092] In an optional embodiment of the present invention, a specific implementation of determining the attack assessment result of the attack detection set attacking the tested large model may include: inputting the attack detection set into the tested large model and at least one reference large model to obtain the corresponding tested output and reference output; and determining the attack assessment result based on the tested output and the reference output.

[0093] Step 203: Input the attack detection set into the tested large model and at least one reference large model to obtain the corresponding tested output and reference output.

[0094] In this embodiment of the invention, the reference large model refers to a large model similar to the large model under test, used as a reference. Specifically, the attack detection set can be input into the reference large model to obtain its output, which is denoted as the reference output. The attack detection set can then be input into the large model under test to obtain the corresponding test output. Multiple reference large models may be included, each with a corresponding reference output.

[0095] For example, attack detection assemblies, acting as a benchmark, provide unified input to multiple large model services. This unified approach shields the internal details of various large model services, thereby supporting the integration of diverse model services.

[0096] This invention supports various large models and service forms, such as basic models, model services, and vertical domain models. Regardless of whether content security filtering and policies are added to the model service, it can be implemented and its effectiveness evaluated.

[0097] Step 204: Determine the attack assessment result based on the tested output and the reference output.

[0098] In this embodiment of the invention, the security of the tested output is evaluated by comprehensively analyzing the tested output and the reference output, thereby obtaining the attack assessment result of the attack detection set. The evaluation strategy or weights can be customized, such as a strategy of repeatedly selecting the worst value, or assigning high weights to values.

[0099] For example, an attack score is determined based on the similarity between the tested output and multiple reference outputs. A high attack score indicates weak attack, while a low attack score indicates high attack. Specifically, higher similarity results in a higher attack score, meaning a more ineffective attack; lower similarity results in a lower attack score, meaning a more effective attack. The Blue Team evaluation engine has an alert mechanism that automatically issues alerts for tested outputs with abnormally low attack scores. Attack information from the corresponding attack detection set is then sent to the attack corpus. Through continuous and timely feedback, the engine continuously improves its security assessment performance.

[0100] In this embodiment of the invention, the specific implementation of determining the attack assessment result based on the tested output and the reference output may include a variety of methods, and this embodiment of the invention does not limit this.

[0101] For example, the Blue Team evaluation engine can be based on a large model. This large model used for evaluation is a pre-trained, small-scale, large language model that primarily performs semantic understanding on several inputs and then scores their similarity to obtain an attack rating. The test output and multiple reference outputs are input into this large model used for evaluation, and the test output and multiple reference outputs are comprehensively analyzed to assess the security of the test output.

[0102] By analyzing and scoring various input sources of the large model used for evaluation, the evaluation engine can stably and reliably quantify the security level of the tested output, ensuring the consistency and accuracy of the results.

[0103] For example, such as Figure 4The diagram illustrates the working principle of Blue Team Security Detection. Blue Team Security Detection primarily comprises four core modules: an attack corpus, an attack scenario library, a Blue Team attack engine, and a Blue Team evaluation engine. The attack scenario library can access external knowledge bases to expand the attack scenarios. The attack corpus provides the Blue Team attack engine with attack corpus sets, and the attack scenario library provides the Blue Team attack engine with attack scenario sets. Based on the attack corpus and attack scenario sets, the Blue Team attack engine generates an attack set, i.e., an attack detection set. The attack detection set is input into the model under test, along with reference large model 1 (e.g., GPT4) and reference large model 2 (e.g., Tongyi 1000 Questions), to obtain the tested answer, reference answer 1, and reference answer 2. The tested answer, reference answer 1, and reference answer 2 are then input into the Blue Team evaluation engine. The Blue Team evaluation engine receives alarm outputs and reference answers, and provides feedback enhancement to the attack corpus.

[0104] In an optional embodiment of the present invention, a specific implementation of determining the attack assessment result based on the tested output and the reference output may include: inputting the tested output and the reference output into an attack assessment model to obtain the attack assessment result, wherein the attack assessment model is used to evaluate the similarity between the tested output and the reference output after semantic understanding, and to determine the attack assessment result based on the similarity.

[0105] An aggression assessment model is a trained method used to perform semantic understanding on the tested output and the reference output, then evaluate their similarity, and determine the aggression assessment result based on the similarity. The aggression assessment result can be an aggression score, an aggression level, or any other applicable result; this embodiment of the invention does not impose any limitations on this.

[0106] The system inputs the output to be tested and at least one reference output into the attack assessment model, which outputs the attack assessment result. This allows the system to continuously acquire knowledge and progress from the reference model, such as when the reference model adjusts its response to regulations. The system can quickly capture and respond to these changes, thus reducing reliance on human intervention.

[0107] In an optional embodiment of the present invention, a specific implementation of updating the attack corpus based on the attack assessment result and the attack detection set may include: for any attack detection input in the attack detection set, if the attack assessment result meets the target condition, determining all or part of the any attack detection input as new corpus; and adding the new corpus to the attack corpus set.

[0108] Step 205: For any attack detection input in the attack detection set, if the attack evaluation result meets the target condition, all or part of the attack detection input is determined as new corpus.

[0109] In this embodiment of the invention, the target condition is set according to actual needs. For example, when the aggression assessment result is an aggression score, the target condition is a preset threshold. This embodiment of the invention does not impose any restrictions on this.

[0110] In this embodiment of the invention, for any attack detection input in the attack detection set, if the attack evaluation result meets the target condition, it indicates that the attack detection input in the attack detection set is effective against the tested large model. In this case, all or part of the attack detection input is determined as new corpus. The specific implementation methods for determining new corpus can include various approaches, and this embodiment of the invention does not limit them.

[0111] For example, based on the attack corpus, non-existent corpora are filtered from the attack detection input and added as new corpora.

[0112] Step 206: Add the newly added corpus to the attack corpus set.

[0113] In this embodiment of the invention, the newly selected corpus is added to the attack corpus set, thereby expanding and updating the attack corpus set.

[0114] By adding new corpora to the attack corpus, new attack corpora can be automatically collected and integrated, enabling rapid responses to regulatory changes, reducing reliance on manual intervention, improving the efficiency of attack detection capabilities, and ultimately achieving continuous enhancement and improvement of automated security protection.

[0115] In an optional embodiment of the present invention, it may further include: iteratively executing the large model attack evaluation method until m iterations are reached, where m is a positive integer greater than 1.

[0116] The above-mentioned large-model attack evaluation method is iteratively executed. After m rounds of iteration, the large-model attack evaluation is completed, realizing self-learning and continuous optimization.

[0117] For example, such as Figure 5The diagram illustrates the blue team security assessment process. The LLM blue team security assessment begins by obtaining attack data from an attack corpus and attack scenarios from an attack scenario library. It then integrates and synchronizes the latest information from an external knowledge base, including collected external information and the latest attack scenarios. The attack data, attack scenarios, and latest information are input into the blue team attack engine. Based on the attack methods and the input information, the blue team attack engine generates a dynamic attack detection set. The tested large model and n reference large models then execute the attack detection set, obtaining the tested response and n reference responses. The tested response and n reference responses are integrated and input into the attack assessment model for analysis and judgment, resulting in an attack effectiveness assessment result. Based on the attack effectiveness assessment result, it is determined whether the attack detection set achieves the desired attack effect. If the attack effectiveness is not achieved, the assessment result is output; if the attack effectiveness is achieved, the attack corpus is augmented based on the corresponding attack detection set. After m rounds of iterative execution, the attack corpus becomes increasingly more aggressive towards the tested large model.

[0118] According to embodiments of the present invention, by transforming an attack corpus and / or collected external information through an attack method, an attack prompt word set is obtained. This attack prompt word set and an attack scenario set are combined to form an attack detection set. The attack detection set is input into the tested large model and at least one reference large model to obtain corresponding tested outputs and reference outputs. Based on the tested outputs and reference outputs, the attack assessment result is determined. For any attack detection input in the attack detection set, if the attack assessment result meets the target conditions, all or part of the attack detection input is identified as new corpus, and the new corpus is added to the attack corpus set. This allows for a self-expanding and / or dynamically updated attack detection set through the self-discovery of the attack corpus set and / or the supplementation of external information. This reflects unknown security risks of the large model, effectively reduces the occurrence of evaluation resistance, realizes a real-time monitoring and adaptive adjustment mechanism, automatically collects and integrates new attack corpus, responds quickly to regulatory changes, reduces reliance on manual intervention, improves the efficiency of attack detection capability improvement, and ultimately achieves continuous enhancement and improvement of automatic security protection.

[0119] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0120] Reference Figure 6The diagram illustrates a structural block diagram of a large-model attack evaluation device according to another embodiment of the present invention, which may specifically include the following modules:

[0121] The set generation module 301 is used to generate an attack detection set based on the attack corpus and / or collected external information;

[0122] The result determination module 302 is used to determine the attack assessment result of the attack detection set on the tested large model;

[0123] The corpus update module 303 is used to update the attack corpus based on the attack assessment results and the attack detection set.

[0124] Optionally, the set generation module includes:

[0125] The transformation submodule is used to transform the attack corpus and / or collected external information using the attack method to obtain a set of attack prompt words;

[0126] The set determination submodule is used to determine the attack detection set based on the attack prompt word set; wherein, the attack detection set includes multiple attack detection inputs.

[0127] Optionally, the set-determining submodule includes:

[0128] The set determination unit is used to combine the attack prompt word set and the attack scenario set into an attack detection set, wherein the attack scenario set includes multiple attack scenario templates.

[0129] Optionally, the result determination module includes:

[0130] The output acquisition submodule is used to input the attack detection set into the tested large model and at least one reference large model, and acquire the corresponding tested output and reference output.

[0131] The result determination submodule is used to determine the attack assessment result based on the tested output and the reference output.

[0132] Optionally, the result determination submodule includes:

[0133] The result acquisition unit is used to input the tested output and the reference output into the attack assessment model to obtain the attack assessment result, wherein the attack assessment model is used to evaluate the similarity between the tested output and the reference output after semantic understanding, and to determine the attack assessment result based on the similarity.

[0134] Optionally, the corpus update module includes:

[0135] The corpus filtering submodule is used to determine all or part of any attack detection input in the attack detection set as new corpus if the attack evaluation result meets the target condition.

[0136] The corpus addition submodule is used to add the new corpus to the attack corpus set.

[0137] Optionally, the attack method includes at least one of the following: prefix and suffix injection, random root injection, and sentence expansion.

[0138] According to embodiments of the present invention, an attack detection set is generated based on an attack corpus and / or collected external information. The attack detection set is then used to determine the attack capability assessment result of the attack on the tested large model. Based on the attack assessment result, the attack corpus is updated according to the attack detection set. This allows for the self-discovery of the attack corpus and / or the supplementation of external information, providing a self-expanding and / or dynamically updated attack detection set. This reflects unknown security risks of large models, effectively reduces the occurrence of assessment resistance, and realizes a real-time monitoring and adaptive adjustment mechanism. It can automatically collect and integrate new attack corpora, respond quickly to regulatory changes, reduce reliance on manual intervention, and improve the efficiency of attack detection capability improvement, thereby achieving continuous enhancement and improvement of automatic security protection.

[0139] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0140] Figure 7 This is a structural block diagram of an electronic device 700 for large-scale attack evaluation, according to an exemplary embodiment. For example, the electronic device 700 may be a domain controller, an in-vehicle computer, an in-vehicle controller, a battery management system (BMS), a fully self-driving system (FSD), etc.

[0141] Reference Figure 7 The electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.

[0142] Processing component 702 typically controls the overall operation of electronic device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.

[0143] Memory 704 is configured to store various types of data to support the operation of device 700. Examples of this data include instructions for any application or method operating on electronic device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0144] Power supply component 706 provides power to various components of electronic device 700. Power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 700.

[0145] Multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0146] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when electronic device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.

[0147] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0148] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of electronic device 700. For example, sensor assembly 714 may detect the on / off state of device 700, the relative positioning of components such as the display and keypad of electronic device 700, changes in position of electronic device 700 or a component of electronic device 700, the presence or absence of user contact with electronic device 700, orientation or acceleration / deceleration of electronic device 700, and temperature changes of electronic device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0149] Communication component 716 is configured to facilitate wired or wireless communication between electronic device 700 and other devices. Electronic device 700 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0150] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0151] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of an electronic device 700 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0152] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a terminal's processor, enables the terminal to perform a large-scale attack evaluation method, the method comprising:

[0153] Generate an attack detection set based on the attack corpus and / or collected external information;

[0154] Determine the attack assessment results of the attack detection set on the large model under test;

[0155] Based on the attack assessment results, the attack corpus is updated according to the attack detection set.

[0156] Optionally, generating an attack detection set based on the attack corpus and / or collected external information includes:

[0157] The attack corpus and / or collected external information are transformed using attack methods to obtain a set of attack prompt words.

[0158] An attack detection set is determined based on the attack prompt word set; wherein the attack detection set includes multiple attack detection inputs.

[0159] Optionally, determining the attack detection set based on the attack prompt word set includes:

[0160] The attack detection set is formed by combining the attack prompt word set and the attack scenario set, wherein the attack scenario set includes multiple attack scenario templates.

[0161] Optionally, determining the attack assessment result of the attack detection set on the tested large model includes:

[0162] The attack detection set is input into the tested large model and at least one reference large model to obtain the corresponding tested output and reference output.

[0163] The attack assessment result is determined based on the tested output and the reference output.

[0164] Optionally, determining the attack assessment result based on the tested output and the reference output includes:

[0165] The tested output and the reference output are input into the attack assessment model to obtain the attack assessment result. The attack assessment model is used to evaluate the similarity between the tested output and the reference output after semantic understanding, and to determine the attack assessment result based on the similarity.

[0166] Optionally, updating the attack corpus based on the attack assessment result and according to the attack detection set includes:

[0167] For any attack detection input in the attack detection set, if the attack evaluation result meets the target condition, all or part of the attack detection input is determined as new corpus.

[0168] The newly added corpus is added to the attack corpus set.

[0169] Optionally, the attack method includes at least one of the following: prefix and suffix injection, random root injection, and sentence expansion.

[0170] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0171] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0172] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0173] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0175] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0176] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0177] The foregoing has provided a detailed description of a large-scale attack assessment method, a large-scale attack assessment device, an electronic device, and a readable storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for evaluating large-scale model attacks, characterized in that, The method includes: Generate an attack detection set based on the attack corpus and / or collected external information; Determine the attack assessment results of the attack detection set on the large model under test; Based on the attack assessment results, the attack corpus is updated according to the attack detection set.

2. The method according to claim 1, characterized in that, The step of generating an attack detection set based on the attack corpus and / or collected external information includes: The attack corpus and / or collected external information are transformed using attack methods to obtain a set of attack prompt words. An attack detection set is determined based on the attack prompt word set; wherein the attack detection set includes multiple attack detection inputs.

3. The method according to claim 2, characterized in that, The step of determining the attack detection set based on the attack prompt word set includes: The attack detection set is formed by combining the attack prompt word set and the attack scenario set, wherein the attack scenario set includes multiple attack scenario templates.

4. The method according to any one of claims 1-3, characterized in that, The determination of the attack assessment results of the attack detection set on the tested large model includes: The attack detection set is input into the tested large model and at least one reference large model to obtain the corresponding tested output and reference output. The attack assessment result is determined based on the tested output and the reference output.

5. The method according to claim 4, characterized in that, Determining the attack assessment result based on the tested output and the reference output includes: The tested output and the reference output are input into the attack assessment model to obtain the attack assessment result. The attack assessment model is used to evaluate the similarity between the tested output and the reference output after semantic understanding, and to determine the attack assessment result based on the similarity.

6. The method according to any one of claims 1-5, characterized in that, The step of updating the attack corpus based on the attack assessment results and the attack detection set includes: For any attack detection input in the attack detection set, if the attack evaluation result meets the target condition, all or part of the attack detection input is determined as new corpus. The newly added corpus is added to the attack corpus set.

7. The method according to claim 2, characterized in that, The attack methods include at least one of the following: prefix and suffix injection, random root injection, and sentence expansion.

8. A large-scale attack evaluation device, characterized in that, The device includes: The set generation module is used to generate an attack detection set based on the attack corpus and / or collected external information; The result determination module is used to determine the attack assessment result of the attack detection set on the tested large model; The corpus update module is used to update the attack corpus based on the attack assessment results and the attack detection set.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method as described in any one of claims 1-7.

10. A readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-7.