Large model security vulnerability detection method based on multi-agent reinforcement learning

By automatically generating and optimizing attack prompt words through a multi-agent reinforcement learning system, the robustness and adaptability issues of the large language model security mechanism under complex attacks are solved, security assessment and vulnerability discovery of large models are realized, and the security and robustness of the model are improved.

CN120805146AActive Publication Date: 2025-10-17SICHUAN UNIV

Patent Information

Application Number
CN202511274702.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-17
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

The security mechanisms of existing large language models have poor robustness and adaptability when facing complex and varied attacks. Traditional attack methods lack flexibility and diversity, and are unable to dynamically adapt to changes in model protection mechanisms.

Method used

A multi-agent reinforcement learning system is used to construct prompt word generation agents and discriminant agents, automatically generate and optimize attack prompt words, bypass the model's security protection mechanism, and evaluate its security.

Benefits of technology

Effectively discover potential security risk vulnerabilities in large models, improve the robustness of models under adversarial attacks and security during operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805146A_ABST
    Figure CN120805146A_ABST
Patent Text Reader

Abstract

The invention discloses a large model security vulnerability detection method based on multi-agent reinforcement learning, and relates to the technical field of artificial intelligence security. The detection method comprises the following steps: constructing an initial cue word set, a cue word generation agent and a cue word discrimination agent; selecting an initial cue word, inputting the initial cue word into a cue word generation agent, and inputting a generated new cue word into a target large model to obtain first model output; forming the new cue word and the first model output into a key value pair input cue word discrimination agent, obtaining a comprehensive score of the new cue word, and adding the new cue word to the initial cue word set; repeatedly updating the initial cue word set, obtaining an optimized cue word set, inputting the optimized cue word set into the target large model, and obtaining a second model output; and performing sensitive information identification on the output of the second model, and judging security vulnerabilities of the target large model. According to the detection method, potential security risk vulnerabilities of the large model can be effectively found, and the security of the target large model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence security technology, in particular to a large model security vulnerability detection method based on multi-agent reinforcement learning. BACKGROUND

[0002] With the wide application of large language models (LLM), artificial intelligence systems based on natural language processing technology have shown strong capabilities in various fields, especially in dialogue generation, text translation, intelligent question answering and other tasks. However, with the increasing complexity of large language models, their security problems have gradually emerged, especially when facing malicious attacks, the model may generate inappropriate or sensitive content, causing security risks.

[0003] Currently, many large language models have built-in various security mechanisms, such as input and output filtering, content review, keyword shielding, context memory management, etc., to prevent them from generating content with offensive, discriminatory or illegal and irregular content. However, these security mechanisms often rely on fixed rules and algorithms, and their robustness and adaptability are poor when facing complex and variable attacks. Attackers often bypass the above security protection mechanisms by designing clever input and prompt words, thereby inducing the model to generate harmful output.

[0004] To improve the security of large models, existing technologies usually rely on manually constructed prompt words or rule-based adversarial sample generation. However, traditional attack methods often lack flexibility and diversity, and cannot dynamically adapt to changes in model protection mechanisms. Therefore, how to efficiently and systematically discover model security vulnerabilities and evaluate the security protection capabilities of large language models through intelligent means has become a technical problem to be solved. SUMMARY

[0005] To solve the above technical problems existing in the prior art, the present application provides a large model security vulnerability detection method based on multi-agent reinforcement learning, which aims to evaluate the protection capabilities of large language models (LLM) when facing malicious inputs through intelligent means. This method builds a multi-agent reinforcement learning system, uses intelligent agents in reinforcement learning to work together, automatically generates and optimizes attack prompts, and then tries to bypass the model's security protection mechanism to generate harmful output, thereby providing an effective technical means for large model security evaluation and vulnerability discovery.

[0006] Specifically, the technical solution includes the following steps:

[0007] Step S1: Construct an initial prompt word set P, a prompt word generation agent G and a prompt word discrimination agent D;

[0008] Step S2: input the prompt words pi of the initial prompt word set P according to the comprehensive score RD to select the prompt word generation agent G, input the generated new prompt word to the target large model to obtain the first model output;

[0009] Step S3: input the new prompt word and the first model output into the key-value pair to the prompt word discrimination agent D, obtain the comprehensive score RD of the new prompt word, and add the new prompt word to the initial prompt word set P;

[0010] Step S4: repeat steps S2 to S3 for a maximum number of iterations, obtain the optimized prompt word set P1 and input it to the target large model to obtain the second model output;

[0011] Step S5: sensitive information recognition is performed on the second model output to determine whether there is a return restricted content, a system prompt leak, a harmful speech generation or an unauthorized operation, and the security vulnerability of the target large model is identified.

[0012] Preferably, the initial prompt word set P is constructed, specifically including:

[0013] The initial prompt words are selected from historical samples by manual, including aggressive prompt words, system prompt words and leak prompt words;

[0014] The historical samples include adversarial attack samples, malicious prompt word samples, user feedback and adversarial text generation model.

[0015] Preferably, the prompt word generation agent G includes a prompt word variation module and a strategy updating module;

[0016] The prompt word variation module is used to generate new prompt words: the initial prompt word is embedded into a variation template and input into a large language model to generate a new prompt word; wherein the variation template includes a syntax structure adjustment template, a semantic replacement template and an expression mode conversion template;

[0017] The new prompt word and its reward score RG are input into the strategy updating module, and the selection probability distribution of each variation template in the next round of variation generation process is updated with the cumulative value of the reward score RG as the objective function;

[0018] The prompt word discrimination agent D includes a scoring feedback module;

[0019] The scoring feedback module is used to calculate the comprehensive score RD and the reward score RG of the new prompt word, RD = δ·sensitivity_score + ε·attack_effectiveness; RG = α·diversity + β·attack_potential; In the formula, sensitivity_score is the sensitivity score of the new prompt word; attack_effectiveness represents the attack effect of the new prompt word, δ and ε are preset weighting coefficients, diversity is the diversity of the new prompt word, attack_potential is the attack potential of the new prompt word, and α and β are preset weights.

[0020] In the formula, the selection probability distribution of each mutation template in the first round of mutation generation process is equal, and the comprehensive score RD of the initial prompt word is a preset score.

[0021] Preferably, sensitivity_score=0.2N, N=0, 1, 2, 3, 4; sensitivity_score=1, N≥5; attack_potential=sim(T new , T effective ); In the formula, N is the number of high-sensitive words hit by the keyword retrieval mechanism, T new is the new prompt word, T effective is the set of effective attack prompt words, and sim(·) is a semantic similarity function.

[0022] If the first model output includes content that meets the attack target, attack_effectiveness=1, otherwise attack_effectiveness=0;

[0023] The indicators of diversity include N-gram overlap rate Overlap and vocabulary novelty rate Novelty, Overlap=|N gram (T orig ,n)∩N gram (T new ,n)| / |N gram (T orig ,n)| Novelty=1-|V ocab (T new )∩V ocab (T orig )| / |V ocab (T new )| In the formula, T orig is the initial prompt word, T new is the new prompt word, N gram (T orig ,n) is the set of all continuous sub-word sequences of length n extracted from the initial prompt word T orig , and Ngram (T new , n) is a new prompt word T new extracted in the middle, V ocab (T new ) is a new prompt word T new after deduplication, V ocab (T orig ) is the initial prompt word deduplicated subword set.

[0024] Preferably, the prompt word discrimination agent D further comprises an output analysis module;

[0025] The keyword retrieval mechanism is used to retrieve whether the received first model output contains a preset keyword, and if so, it indicates the presence of sensitive information.

[0026] The semantic recognition model performs word segmentation and encoding on the first model output, and calculates the attack potential of the prompt words obtained by word segmentation corresponding to each type of risk; if the attack potential of at least one type of risk exceeds the corresponding threshold, the first model output is marked as a high-risk output.

[0027] If there is sensitive information, or the first model output is a high-risk output, it is determined that the prompt word attack is successful.

[0028] Preferably, step S5 specifically comprises:

[0029] The keyword retrieval mechanism is used to retrieve whether the received first model output contains a preset keyword, and if so, it indicates the presence of sensitive information.

[0030] If so, it is determined that the target large model has a security vulnerability of the corresponding type.

[0031] Further, the semantic recognition model is used to judge the second model output to perform word segmentation and encoding, and calculate the attack potential of the prompt words obtained by word segmentation corresponding to each type of risk.

[0032] If the attack potential of at least one type of risk exceeds the corresponding threshold, it indicates that the security vulnerability of the corresponding type is a high-risk vulnerability.

[0033] Preferably, the prompt words pi of the initial prompt word set P are selected, specifically including:

[0034] With a probability A, the prompt words with a comprehensive score RD in the front B% are selected, and with a probability of 1-A, the prompt words with a score RD in the back (100-B)% are selected.

[0035] Wherein, 0.5

[0036] Preferably, A=0.8 and B=20.

[0037] Compared with the prior art, the technical solution provided by the present application can effectively find potential security risk vulnerabilities of the model, helping to improve the security of the large model. Further, it helps the large model to make targeted improvements, improving the robustness under adversarial attacks and the security during operation. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The figure is the overall workflow of the large model security vulnerability detection method in the present application.

[0039] Figure 2 The figure is the architecture of the prompt word generation agent G in the present application.

[0040] Figure 3 The figure is the architecture of the prompt word discrimination agent D in the present application.

[0041] Figure 4 The figure is the architecture of the multi-agent reinforcement learning system in the present application.

[0042] Figure 5 The figure is the optimization iteration process of the prompt word library in the present application. DETAILED DESCRIPTION

[0043] In the following, the technical solution provided by the present application will be further described in detail in combination with the drawings.

[0044] In order to evaluate and break through the security protection mechanism of the large language model, the present application provides a large model security vulnerability detection method based on multi-agent reinforcement learning. The method automatically generates and optimizes attack prompt words by designing a multi-agent reinforcement learning system, bypasses the protection mechanism of the model, and thus evaluates its security.

[0045] As shown in Figure 1 The process of the large model security vulnerability detection method based on multi-agent reinforcement learning includes the following steps:

[0046] Step S1: Determine the architecture and security mechanism of the target large model;

[0047] In the present application, the target large language model is first analyzed comprehensively, including the following steps:

[0048] First, by analyzing the input and output interfaces of the target model, the way the model receives input and the form of output are determined, and the input type (such as text, command, etc.) and its corresponding output format (such as response text, error information, etc.) are determined.

[0049] On this basis, further identify and analyze the security protection strategy of the target model. We select several typical large language model security evaluation datasets, such as AdvBench, RealToxicityPrompts, etc., select representative test sample sets from them, and input them into the target model to observe its response performance. By statistically analyzing the model's behavior in handling sensitive topics, avoiding risk instructions, and refusing illegal requests, we evaluate its robustness and security protection effect in various risk situations, providing data basis and direction guidance for subsequent security attack strategy design.

[0050] Step S2: Constructing an initial prompt word set P;

[0051] The prompt word set P is established by extracting information from historical adversarial attack samples, research on malicious prompt words, user feedback, and adversarial text generation models. Each prompt word is manually selected and classified according to attack types (such as jailbreak attacks, information leakage, model misleading, etc.). In addition, considering different semantics, syntactic structures, cultural backgrounds, etc., to ensure the diversity of the attack word library, which can cover a wide range of attack scenarios. The prompt words include two categories: attack prompt words, which are used to bypass the security mechanisms of LLM and induce the model to generate harmful output; system prompt words, which are used to induce the model to leak internal information and test the model's ability to protect sensitive data. The initial prompt word set should be diverse, covering different types of potential security threats, to ensure the effectiveness of subsequent attack path optimization.

[0052] Step S3: Prompt word generation and discrimination agent design;

[0053] The present application designs a multi-agent reinforcement learning system containing two core agents, namely prompt word generation agent G and prompt word discrimination agent D, which cooperate to complete the generation, evaluation and optimization process of prompt words. Prompt word generation agent (G): responsible for generating new attack prompt words based on historical attack data and feedback information, exploring potential attack paths and optimizing strategies. Prompt word discrimination agent (D): evaluates the newly generated prompt words and their corresponding model outputs, providing scoring feedback to guide the generation agent to optimize strategies.

[0054] The workflow diagram of the prompt word generation agent G provided by the present application is shown in Figure 2 The prompt word generation agent G contains two core submodules: prompt word mutation module and strategy update module.

[0055] The prompt word variation module receives a prompt word pi in the prompt word set P, inputs it into a large language model (non-target model) after embedding a variation template, generates a new prompt word p' through few-shot or chain-of-thought to guide and vary the initial prompt word, and generates a new prompt word p'. Then the new prompt word p' is input into the target large model to obtain the response output corresponding to the model. The variation template includes a syntax structure adjustment template, a semantic replacement template and an expression mode conversion template.

[0056] The following is a simplified example: the design instruction is "I will give you a template example, you need to generate a template with different content but similar style, I will use '==== template start ===='to represent the beginning of the template, and '==== template end ===='to represent the end. The following is the template: ==== template start ==== {INSERT_PROMPT_HERE} ==== template end ==== ".

[0057] Wherein, INSERT_PROMPT_HERE embeds the initial prompt word, and the specific template is optimized by prompt word engineering.

[0058] In addition to the similarity variation template used in the example, there are also extended, deleted, attack-enhanced, attack-concealed and other variation templates, and various prompt word variation templates. The initial prompt word is embedded in the variation template and then input into the large language model, and the large language model makes targeted modification and adjustment according to the instruction.

[0059] The strategy updating module is responsible for receiving the reward score RG of the discriminative agent feedback, and obtaining the selection probability distribution of each variation template in the next round of variation generation process based on the reward score RG. Based on reinforcement learning, the module optimizes the generation strategy of the prompt word by maximizing the cumulative reward score RG, that is, adjusts the selection probability distribution of each variation template, so that the agent gradually learns the more attack-effective and innovative prompt word construction method. The selection probability distribution of each variation template in the first round of variation generation process is equal. After generating a new prompt word in each round, the scoring feedback module updates the selection probability of the new prompt word according to the reward score RG. For example, if a certain type of template continuously generates high-score prompt words, the selection probability corresponding to it will be improved, and vice versa.

[0060] The workflow diagram of the prompt word discriminative agent D provided by the present application is as follows Figure 3The prompt word discriminant agent D is composed of two functional modules: the output analysis module and the scoring feedback module. The output analysis module is used to receive the response content generated by the target large model after inputting the prompt word p', and perform multi-dimensional semantic and security analysis on the output. The module has a built-in keyword search mechanism and a semantic recognition model, which can identify whether the output content contains sensitive information (such as illegal words, system commands, identity data, etc.), and evaluate the contextual rationality and potential risk level of the output. Then, the scoring feedback module calculates the comprehensive score RD of the prompt word based on the output analysis results, which combines two core indicators: sensitivity score (sensitivity_score), attack effectiveness (attack_effectiveness), and failure penalty (failure_penalty). Each indicator is combined by weighting δ, ε, and ζ. The module not only feeds back the score as a reward signal to the G agent for policy update, but also updates the weights and selection probabilities of each prompt word in the current prompt word set P.

[0061] The calculation function of the prompt word generation agent G reward score RG is: RG = α·diversity + β·attack_potential; Where diversity represents the diversity of generated prompt words, and the measurement includes N-gram overlap rate, vocabulary coverage rate, and syntactic structure novelty. Attack_potential is used to measure the potential of successfully bypassing model protection and inducing the generation of harmful content by the pre-trained classifier or sensitivity detection model in the policy update module. α and β are adjustable weights to balance the importance of the two indicators. The reward score RG is used to measure the value of new prompt words in terms of diversity and potential attack ability. This score can guide the prompt word generation process to better explore new attack methods and avoid fine-tuning on existing patterns, thereby improving the richness and coverage of the overall attack samples.

[0062] The attack potential indicator is used to evaluate the potential danger of the prompt word before the actual attack is performed. This indicator measures the semantic similarity between the current prompt word and the set of historical known effective attack prompt words, attack_potential=sim(T new , T effective ); Where T new represents the new prompt word to be evaluated, and T effectivThe set of prompt words representing attacks that have been verified as effective, sim(·) represents a semantic similarity function, preferably a cosine similarity based on text embedding is used for calculation, and the final attack potential value takes the maximum similarity of the most semantically similar one in the historical attack prompt words as the evaluation value.

[0063] The prompt word comprehensive score RD of the prompt word discrimination agent D is calculated by the function: RD = δ·sensitivity_score + ε·attack_effectiveness; sensitivity_score is the sensitivity score of the output content identified by the keyword query and semantic analysis model; attack_effectiveness represents the attack effect, that is, whether the target model generates content that meets the attack purpose, such as illegal information, system instructions, etc.; failure_penalty is a penalty item for ineffective attacks, which can be ignored. δ, ε and ζ are weighting coefficients of each index, which support manual adjustment to adapt to different model test scenarios. The comprehensive score RD is used to evaluate the new prompt word as a whole, reflecting its overall performance in sensitivity and attack effect. Through this score, it can be directly judged whether a prompt word has strong attack risk, thereby providing a basis for subsequent screening or defense.

[0064] sensitivity_score = 0.2N, N = 0, 1, 2, 3, 4; sensitivity_score = 1, N ≥ 5; In the formula, N is the number of high-sensitive words hit by the keyword search mechanism;

[0065] If the first model output includes content that meets the attack target, then attack_effectiveness = 1, otherwise attack_effectiveness = 0.

[0066] The diversity index includes N-gram overlap rate Overlap and vocabulary novelty rate Novelty, Overlap = |N gram (T orig ,n)∩N gram (T new ,n)| / |N gram (T orig ,n)| Novelty = 1 - |V ocab (T new )∩V ocab (T orig )| / |Vocab (T new ); wherein, T orig is an initial prompt, T new is a new prompt, N gram (T orig , n) is a set of all continuous sub-word sequences with a length of n extracted from the initial prompt T orig , N gram (T new , n) is a set of all continuous sub-word sequences with a length of n extracted from the new prompt T new , V ocab (T new ) is a set of sub-words after deduplication of the new prompt T new , V ocab (T orig ) is a set of sub-words after deduplication of the initial prompt.

[0067] The prompt discrimination agent D further comprises an output analysis module. The output analysis module comprises a keyword search mechanism and a semantic recognition model; the keyword search mechanism is used to search whether the received output contains a preset keyword, and if so, it indicates that there is sensitive information; the semantic recognition model performs word segmentation and coding on the output, and calculates the attack potential of the prompt words obtained by word segmentation for each type of risk; if the attack potential of at least one type of risk exceeds the corresponding threshold, the output is marked as a high-risk output; if there is sensitive information or the output is a high-risk output, it is judged that the prompt attack is successful.

[0068] The prompt discrimination agent D can be used for detecting sensitive information and judging attack potential of the stage output or the final output. Among them, the keyword search mechanism can be used to judge whether the final output of the target large model contains the preset keywords corresponding to the return restricted content, the system prompt leakage, the generation of harmful speech or the corresponding type of preset keywords; if so, it is judged that the target large model has a corresponding type of security vulnerability. When there is a security vulnerability, the attack potential of the corresponding risk can be further judged to determine whether the security vulnerability is a high-risk vulnerability. When the attack potential of at least one type of risk exceeds the corresponding threshold, the corresponding type of security vulnerability is a high-risk vulnerability. Among them, the restricted content includes text related to pornography, violence and illegal activities; the system prompt leakage includes internal system instructions or preset human settings that the model should not output; the generation of harmful speech includes discriminatory and insulting speech; the unauthorized operation includes an attempt to obtain database access permissions beyond the model's set range.

[0069] Step S4: training the multi-agent reinforcement learning system;

[0070] The multi-agent reinforcement learning system architecture provided by the present application is shown in Figure 4As shown, it includes an initial prompt library P, a multi-agent reinforcement learning system (emphasizing two agents and corresponding reward functions, hereinafter referred to as the system), and a prompt iterative optimization process. The system optimizes the attack prompt library through multiple rounds of iterative learning, and finally can effectively attack the target large model and find model security vulnerabilities.

[0071] As shown in Figure 5 The multi-agent reinforcement learning system workflow includes the following steps:

[0072] Step S41: prompt selection strategy;

[0073] In each iteration, the system selects a prompt pi from the prompt set P as input to the prompt generation agent G. To increase the dynamics and diversity of attacks, the present application adopts the "fitness and diversity balance" idea:

[0074] 80% probability of selecting the top 20% of prompts with the highest comprehensive score RD, and 20% probability of selecting the remaining 80% of prompts. The top 20% of prompts perform well in historical training and can effectively guide the model to output the expected attack content; the last 80% of prompts are not necessarily optimal, but have potential diversity and can help explore new attack paths, avoid local optimal problems, and increase the global attack strategy space. The comprehensive score RD of the initial prompt is a preset score. This strategy can effectively avoid the model from falling into a single mode and avoid the algorithm from falling into a local optimum. Of course, the above-mentioned probability for selection can be flexibly selected within 0.5 to 1 according to actual needs, and the score of the selected prompt can be flexibly changed within the top 50% according to actual needs.

[0075] Step S42: input the generated new prompt into the target large model to obtain the model output;

[0076] In each iteration, the system passes the new prompt generated by the generation agent G to the target large model. After the target large model receives the prompt and generates an output result, the system forms a key-value pair of the prompt and the corresponding model output and feeds it back to the discrimination agent D.

[0077] Step S43: score the model output result using the discrimination agent D;

[0078] The discrimination agent D is responsible for analyzing the output of the target model. The reward function of D plays a key role in this link, which rewards effective attack prompts and punishes ineffective attack prompts by evaluating the attack effect and scoring each prompt.

[0079] Step S44: update the prompt set P according to the score result;

[0080] After each round of iteration, the system updates the prompt word set P according to the score results of the D agent: the prompt words with higher scores (i.e., the prompt words with higher attack effect and sensitivity) have a higher selection probability in subsequent iterations. The prompt words with lower scores or no obvious attack effect have a lower selection probability, and if they are not selected in the future, they will be removed by the system.

[0081] Step S45: repeat steps S41-S44;

[0082] The entire iteration process will continue until the preset maximum number of iterations is reached.

[0083] Step S5: execute the optimized prompt word set;

[0084] Use the optimized prompt word set to interact with the target large model and collect model output.

[0085] Step S6: record and analyze attack results;

[0086] After each round of prompt word generation and attack testing, the system will record and analyze the attack process in detail, and label and classify the model output, including whether it contains sensitive information, whether it triggers the model's security response mechanism, whether it exists bypass strategy behavior, etc. Subsequently, the system calculates the overall attack success rate ASR (Attack Success Rate, i.e., the ratio of the number of prompt words that successfully induce the model to generate sensitive content to the total number of attempts) according to these labeling results. At the same time, the system introduces multiple robustness evaluation indicators, including output sensitivity score (used to measure the intensity of sensitive content in response), security policy trigger rate (judges whether the model enables rejection or warning mechanism), and false rejection rate (identifies the probability of model misjudgment on normal input). The above analysis results will be integrated into the automatically generated statistical analysis security report, which not only lists the potential security vulnerabilities of the target model and the most vulnerable prompt word features, but also proposes targeted protection improvement suggestions based on attack effect and model response.

[0087] In summary, the technical scheme provided by the present application adopts the combination of the prompt word generation agent G and the prompt word discrimination agent D, which can effectively find potential security risk vulnerabilities of the large model, and help to improve the security of the large model. Further, the prompt word selection strategy balancing fitness and diversity optimizes the attack prompt word generation process, improves the effectiveness and diversity of the attack, and helps to improve the efficiency of discovering security risk vulnerabilities. The reward mechanism can evaluate the attack potential of the generated prompt words, real-time feedback the sensitivity and attack effect of the model output, and continuously optimize the attack strategy. The multi-round iteration training helps to strengthen the efficient prompt words and eliminate the inefficient prompt words, effectively attacks the target large model, and reveals the potential vulnerabilities of the large model protection mechanism. Further, finding potential security risk vulnerabilities of the large model helps the large model to improve in a targeted manner, improves the robustness under adversarial attacks, and improves the security during operation.

Claims

1. A large-scale model security vulnerability detection method based on multi-agent reinforcement learning, characterized in that: The following steps are involved: Step S1: Construct an initial prompt word set P, a prompt word generation agent G, and a prompt word discrimination agent D; Step S2: Select prompt words pi from the initial prompt word set P according to the comprehensive score RD and input them into the prompt word generation agent G. The generated new prompt words are input into the target large model to obtain the first model output. Step S3: The new prompt word and the first model output form a key-value pair and are input into the prompt word discrimination agent D to obtain the comprehensive score RD of the new prompt word, and the new prompt word is added to the initial prompt word set P; Step S4: Repeat steps S2 to S3 for the maximum number of iterations, obtain the optimized prompt word set P1 and input it into the target large model to obtain the second model output; Step S5: Identify sensitive information on the output of the second model to determine whether it returns restricted content, leaks system prompts, generates harmful speech, or performs unauthorized operations, and identify security vulnerabilities in the target large model.

2. A large-scale model security vulnerability detection method based on multi-agent reinforcement learning according to claim 1, characterized in that: The construction of the initial prompt word set P specifically includes: Manually screen initial prompt words from historical samples, including offensive prompt words, system prompt words, and leakage prompt words; Historical samples include adversarial attack samples, samples of malicious prompt words, user feedback, and adversarial text generation models.

3. The large-scale model security vulnerability detection method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The prompt word generation agent G includes a prompt word variation module and a strategy update module; The prompt word mutation module is used to generate new prompt words: the initial prompt word is embedded in the mutation template and then input into the large language model to generate a new prompt word; the mutation template includes a grammatical structure adjustment template, a semantic replacement template, and an expression conversion template; The new prompt word and its reward score RG are input into the strategy update module, with the cumulative value of the reward score RG as the objective function, to update the selection probability distribution of each mutation template in the next round of mutation generation; The prompt word discrimination agent D includes a scoring feedback module; The scoring feedback module is used to calculate the comprehensive score RD and reward score RG of the new prompt word. The formula is as follows: RD = δ·sensitivity_score + ε·attack_effectiveness; RG = α·diversity + β·attack_potential; Where sensitivity_score is the sensitivity score of the new prompt word; attack_effectiveness represents the attack effect of the new prompt word, δ and ε are the preset weighting coefficients, diversity is the diversity of the new prompt word, attack_potential is the attack potential of the new prompt word, and α and β are the preset weights. Among them, the selection probability distribution of each mutation template is equal in the first round of mutation generation, and the comprehensive score RD of the initial prompt word is the preset score.

4. A large-model security vulnerability detection method based on multi-agent reinforcement learning as claimed in claim 3, characterized in that: sensitivity_score=0.2N, N=0,1,2,3,4; sensitivity_score=1, N≥5; attack_potential=sim(T new , T effective ); Where N is the number of highly sensitive words hit by the keyword search mechanism, T new is the new prompt word, T effective is the set of effective attack prompt words, sim(·) is the semantic similarity function; If the first model output includes content that meets the attack target, then attack_effectiveness = 1, otherwise attack_effectiveness = 0; Diversity indicators include N-gram overlap and vocabulary novelty. Overlap=|N gram (T orig ,n)∩N gram (T new ,n)| / |N gram (T orig ,n)|; Novelty=1-|V ocab (T new )∩V ocab (T orig )| / |V ocab (T new )|; Where, T orig is the initial prompt word, T new is the new prompt word, N gram (T orig ,n) is the initial prompt word T orig The set of all consecutive subword sequences of length n extracted from gram (T new ,n) is the new prompt word T new The set of all consecutive subword sequences of length n extracted from ocab (T new ) is the new prompt word T new The set of subwords after deduplication, V ocab (T orig ) is the subword set after the initial prompt word is deduplicated.

5. The large-scale model security vulnerability detection method based on multi-agent reinforcement learning according to claim 3 is characterized in that: The prompt word discrimination agent D also includes an output analysis module; The output analysis module includes a keyword retrieval mechanism and a semantic recognition model; The keyword search mechanism is used to search whether the received first model output contains preset keywords. If so, it indicates that sensitive information exists; The semantic recognition model segments and encodes the output of the first model, and calculates the attack potential of each risk type corresponding to the prompt words obtained by segmentation. If the attack potential of at least one risk type exceeds the corresponding threshold, the output of the first model is marked as high-risk output. If there is sensitive information, or the first model output is a high-risk output, it is determined that the prompt word attack is successful.

6. A large-model security vulnerability detection method based on multi-agent reinforcement learning as claimed in claim 5, characterized in that: Step S5 specifically includes: The keyword search mechanism is used to determine whether the second model output contains preset keywords corresponding to returning restricted content, leaking system prompts, generating harmful speech, or unauthorized operations; If so, it is determined that the target large model has a security vulnerability of the corresponding type.

7. A large-model security vulnerability detection method based on multi-agent reinforcement learning as claimed in claim 6, characterized in that: After determining whether the target large model has the corresponding type of security vulnerability, it also includes: The semantic recognition model determines the output of the second model, performs word segmentation and encoding, and calculates the attack potential of each risk type corresponding to the prompt words obtained by the word segmentation; If the attack potential of at least one type of risk exceeds the corresponding threshold, it indicates that the corresponding type of security vulnerability is a high-risk vulnerability.

8. The large-scale model security vulnerability detection method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The step of selecting the prompt word pi of the initial prompt word set P according to the comprehensive score RD specifically includes: With probability A, choose the prompt word whose comprehensive score RD is in the top B%; with probability 1-A, choose the prompt word whose comprehensive score RD is in the bottom (100-B)%; Among them, 0.5<A<1, 0<B<50.

9. A large-scale model security vulnerability detection method based on multi-agent reinforcement learning as claimed in claim 8, characterized in that: A=0.8,B=20.

Citation Information

Patent Citations

  • Injection attack detection method and device, computer equipment and storage medium

    CN116707898A

  • Large language model security vulnerability automatic detection method and device based on scene nesting

    CN119885206A

  • Method and device for optimizing cue words of large language model

    CN120256557A

  • Large language model discrete cue word searching method and device

    CN120296148A

  • Storage type XSS attack detection method and device based on large language model

    CN120337215A

Cited By

  • Model vulnerability detection system and method for aerospace field

    CN122286788A