Large language model vulnerability testing method based on working memory tree

By employing a working memory tree-based vulnerability testing method, and utilizing a tree structure and multi-prompt combination strategy, the generation of adversarial prompts is optimized. This addresses the issue of limited vulnerability testing coverage in multi-prompt combination attacks using large language models, achieving more efficient and comprehensive vulnerability detection and security assessment.

CN120910871APending Publication Date: 2025-11-07ZHEJIANG UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511156094.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing large language models suffer from limited vulnerability testing coverage, single attack paths, and high computational costs when facing multi-prompt combined attacks, especially when dealing with jailbreak attacks, making it difficult to fully uncover potential security vulnerabilities.

Method used

This paper adopts a vulnerability testing method based on working memory trees. By constructing a tree structure model, each node represents an adversarial prompt word. Combining data category analysis and multi-prompt combination strategies, multiple adversarial prompt words are generated for parallel input. The prompt generation process is optimized by utilizing working memory theory and cognitive psychology principles, so as to achieve multi-angle and multi-path attacks on large language models.

Benefits of technology

It significantly improves the coverage and attack success rate of vulnerability testing for large language models, enhances the concealment and diversity of adversarial prompts, can more realistically simulate potential failure risks in actual deployment scenarios, and improves the security and credibility assessment capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910871A_ABST
    Figure CN120910871A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model vulnerability testing method based on a working memory tree. The method comprises the following steps of 1, collecting malicious semantic texts; 2, constructing an automatic confrontation prompt; 3, testing vulnerabilities of the large language model; 4, evaluating a test result; and 5, carrying out iterative optimization on confrontation prompt. According to the method, antagonism prompts with concealment and attack diversity can be accurately constructed, and the limitation that a traditional large language model vulnerability testing method is limited in coverage range and insufficient in multi-prompt exploratory performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a large language model vulnerability testing method based on a working memory tree, and belongs to the technical fields of Internet and artificial intelligence. BACKGROUND

[0002] With the continuous evolution of artificial intelligence technology, large language models have shown strong capabilities in the field of natural language processing. Thanks to the introduction of the Transformer structure, the scale of the model has rapidly expanded from the early GPT-1 to the GPT-4 with a trillion-level parameter. Such models have made breakthroughs in text generation, language understanding, question and answer dialogue, and have gradually expanded to complex scenarios such as multi-modal understanding and cross-language processing, promoting the rapid landing of intelligent applications in education, healthcare, finance and other industries. At the same time, large language models represented by PaLM and LLaMA are widely deployed globally, and domestic Chinese language models such as ChatGLM, Qwen and DeepSeek have also been introduced. With the evolution of technology, these models have shown human-level intelligence in knowledge reasoning, logical judgment and language generation.

[0003] However, as its influence expands, large language models have exposed many security risks in practical applications. The model may learn biases and stereotypes from the source data during the pre-training phase, resulting in harmful content or discriminatory expressions in gender aspects in the generated content. In order to improve the safety and credibility of the model output, researchers have introduced alignment techniques including supervised fine-tuning and reinforcement learning based on human feedback. These methods aim to make the model better understand user intent, follow human values and reduce the risk of inappropriate content output.

[0004] Even if the model is aligned, there is still a possibility of being "bypassed" or "misled", especially vulnerable to jailbreak attacks. Such attacks, through carefully designed input prompts, can induce the model to generate illegal, harmful or sensitive content without touching the internal structure or weight parameters of the model. Early jailbreak attacks mainly rely on prompts shared by users on social platforms, or specific attack instructions designed by security experts. However, with the deepening of research, jailbreak technology is gradually evolving towards systematization and automation. For example, through role-playing, prompt nesting, coding avoidance and other strategies, attackers can induce the model to break the security boundary and generate offensive or discriminatory content. However, despite the progress made in jailbreak attack technology, existing methods still have many limitations. On the one hand, token-based attacks, although achieving automatic generation, often lack semantic interpretability, making them easily identified and intercepted by confusion or language fluency detection mechanisms, and the computational cost is high in searching for the optimal suffix. On the other hand, automated prompt attacks have broken through the bottleneck of manual design to some extent, but current research mainly focuses on single prompt attacks, and has not yet explored the complex risks that may arise from the coordinated triggering of multiple prompts.

[0005] In view of the deficiencies of existing research in large language model vulnerability testing, especially the limitations in dealing with multi-prompt combination attacks, the present invention proposes a large language model vulnerability testing method based on working memory tree. First, the present invention constructs a tree structure model, where each node represents an adversarial prompt word, and iteratively optimizes it through a "thought chain" reasoning mechanism. Second, a data category analysis module is designed to identify potential security vulnerabilities in different categories. Under the guidance of working memory theory, the present invention constructs context-related examples to optimize the generation process of child node prompt words. By effectively integrating these context examples with the memory information in the parent node, the generation quality of child node prompt words can be improved, resulting in more context-related and attack-effective adversarial prompt words. To further improve vulnerability mining capabilities, the present invention also introduces a multi-prompt word combination strategy, which inputs multiple adversarial prompt words in combination to guide the large language model to generate potential harmful outputs, thereby uncovering more potential security risks. SUMMARY

[0006] To address the problems and shortcomings of existing technologies, this invention proposes a large language model vulnerability testing method based on working memory trees. By introducing working memory theory and data category analysis mechanisms, it achieves efficient generation and optimization of adversarial prompts, significantly improving the success rate and diversity of large language model attacks. Specifically, this invention constructs a tree-structured generation framework based on a "thinking chain" reasoning mechanism, where each node represents an adversarial prompt generated by the large language model through optimization. Nodes are linked through logical reasoning and contextual information to form semantic association chains, thereby achieving layer-by-layer optimization and semantic refinement of the prompts. Simultaneously, combined with a data category analysis module, it identifies potential model vulnerabilities in different categories of data, providing directional guidance for the prompt generation process. Furthermore, to further enhance the diversity and versatility of attack strategies, this invention also designs a multi-prompt combination strategy, integrating multiple semantically related or complementary adversarial prompts and combining them with diverse technologies to generate combined adversarial prompts, which induce the large language model to generate potentially harmful or jailbreak-related outputs. This strategy effectively expands the attack boundary of the large language model and improves the coverage of large language model vulnerability testing.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows: a method for testing vulnerabilities in a large language model based on a working memory tree, the method comprising the following steps:

[0008] Step 1: Malicious semantic text collection;

[0009] Step 2: Build an automated system to counter suggestion prompts;

[0010] Step 3: Large language model vulnerability testing;

[0011] Step 4: Evaluation of test results;

[0012] Step 5: Iterative optimization against prompts.

[0013] As an improvement to this invention, step 1: Malicious semantic text collection. First, information text containing malicious semantics is automatically collected from the Internet using web crawling technology to construct an initial dataset of malicious questions X = {X1, X2, ..., X...}. N}, where X i This dataset represents an initial malicious instruction or question used in the subsequent construction of adversarial prompts and vulnerability testing. It covers a variety of sensitive content, demonstrating diversity and representativeness.

[0014] As an improvement of this invention, step 2: Automatic adversarial prompt construction. Using the malicious dataset collected in step 1, the final adversarial prompt is automatically constructed through data category analysis, prompt scenario construction, and multi-prompt combination strategies for large language model vulnerability testing. The implementation of this step can be divided into the following sub-steps:

[0015] Sub-step 2-1, data category analysis. Given the wide coverage of the adopted dataset, covering multiple fields such as government privacy, dangerous activities, fraud, malware, and other harmful information, it is difficult to effectively respond to the differentiated characteristics and potential vulnerabilities of each category if a unified adversarial prompt design strategy is adopted. Therefore, this step first classifies the original dataset systematically and divides it into six categories: “Illegal Activity”, “Hate Speech”, “Malware”, “Physical Harm”, “Fraud” and “Privacy Violence”. Subsequently, with the help of a large language model, the semantic and attack characteristics between harmful instructions in the initial dataset and their corresponding categories are analyzed, and the key differences between each category compared to other categories are identified. Through such in-depth analysis, the unique model vulnerability characteristics of each category can be further identified and utilized in the subsequent generation of adversarial prompts. This strategy can effectively improve the attack efficiency of the prompt and ensure that it is more tailored to the inherent weaknesses of each category, thereby improving the coverage and relevance of the overall vulnerability test.

[0016] Sub-step 2-2, prompt scenario construction. Based on the category knowledge analysis of step 2-1, the invention uses a working memory tree structure to construct the scenario prompt. At each node, the model will integrate the adversarial prompts, optimization suggestions, and example demonstrations of its parent nodes to jointly construct a complete working memory, thereby optimizing the subsequent prompt generation process. This design fully combines the primacy effect and working memory theory in cognitive psychology. The primacy effect states that humans have stronger memory and recall abilities for the initial information received, while the working memory theory indicates that there is a capacity limit in information processing. In analogy to the processing mechanism of large language models, the former is reflected in the model's attention mechanism preference for initial input, and the latter is represented by the constraint of the context window size. At the same time, the model prioritizes the response to early input when allocating limited attention resources. To depict the primacy bias in this attention allocation process, the invention uses an exponential decay function to model the attention weight of information at different time steps, which is expressed as follows:

[0017] α t =α0·e -λt

[0018] where α t represents the attention intensity of the model at time step t, α0 is the attention baseline value of the initial input, and λ is the hyperparameter controlling the attention decay rate.

[0019] In the prompt construction process, in order to further improve the expression quality and the confrontation concealment, the application introduces three types of enhancement elements: example demonstration, optimization suggestion and role-playing situation, which synergistically drive the prompt generation process of each node. Among them, the example demonstration is used to provide format specification and output reference; the optimization suggestion proposes explicit guidance information to optimize the model generation path at the current node, aiming at the potential deviation in the model generation process; the role-playing situation aims to provide a logical and semantically coherent context shell for the adversarial prompt, and through the design of specific roles, dialogue settings or task backgrounds, it can hide the direct exposure of malicious goals and enhance the concealment and rationality of the prompt. In order to depict the comprehensive performance of various enhancement elements in prompt generation, the application constructs the following weighted combination:

[0020] S total =w1·E+w2·G+w3·R

[0021] Among them, E, G and R represent the performance scores of example demonstration, optimization suggestion and role-playing, and w1, w2 and w3 are the corresponding weight parameters.

[0022] Substep 2-3, multi-prompt combination strategy. In order to further improve the attack strength and concealment of adversarial prompt generation, the application introduces a multi-prompt combination strategy in the prompt construction process to break through the limitations of traditional single prompt method in attack path diversity and model vulnerability coverage. Traditional jailbreak attack methods usually rely on a single prompt template, which limits their ability to fully explore potential vulnerabilities of large language models. Inspired by multi-sample attack and multi-round jailbreak technology, the application innovatively designs a combination mechanism based on multiple adversarial prompts in parallel to construct a prompt set with greater attack depth and coverage. Specifically, based on model behavior analysis and historical attack records, the application constructs a set of complementary final adversarial prompts P = {P1, P2,..., P n}, where P i represents the i-th adversarial prompt generated by substep 2-2. Then, all prompts are uniformly scheduled and input into the target large language model in parallel to form a multi-angle, multi-path parallel attack scene, which enhances the triggering ability of multiple vulnerable points of the language model.

[0023] As an improvement of the application, step 3: large language model vulnerability testing, uses the final adversarial prompts obtained in step 2 to test the vulnerability of the target large language model, to identify the jailbreak vulnerabilities or output boundary vulnerability defects exposed by the large language model when facing malicious input. Specifically, given a target large language model T to be attacked and the malicious semantic data set X = {X1, X2,..., X N} in step 1, if X iWhen the input is sent to the target large language model, the model will usually reject to answer according to its alignment mechanism, and output responses such as "I'm sorry, but I can't assist with that." To break through this defense line, the present application uses the final set of adversarial prompts P generated in step 2 to maximize the probability of inducing the target large language model T to output the target harmful response R while concealing the original problem attack intention, thereby achieving comprehensive testing and evaluation of potential security vulnerabilities of the large language model. The generation behavior of the large language model in this step can be expressed in the form of the following conditional probability distribution formula: i The hidden package maximizes the probability of inducing the target large language model T to output the target harmful response R while concealing the original problem attack intention, thereby achieving comprehensive testing and evaluation of potential security vulnerabilities of the large language model. The generation behavior of the large language model in this step can be expressed in the form of the following conditional probability distribution formula:

[0024]

[0025] where n represents the number of tokens in the adversarial prompt set P, R represents the total number of tokens in the model response R, x n+1:n+L represents the token sequence of response R, x 1:n represents the token sequence of the adversarial prompt set P, and P(A|B) represents the probability of the model generating the next token sequence A given the token sequence B.

[0026] As an improvement of the present application, step 4: test result evaluation, the response results obtained in step 3 are systematically evaluated to determine whether the current adversarial prompt successfully induces the target large language model to generate the expected harmful content, thereby achieving the goal of vulnerability testing. The present application designs two evaluation methods, namely the keyword-based evaluation method and the GPT-4 model-based evaluation method, to realize multi-dimensional judgment of the response output. This step can be divided into the following sub-steps:

[0027] Sub-step 4-1, keyword-based evaluation. Keyword-based evaluation refers to string matching analysis of the response R generated by the language model in step 3, specifically by identifying whether there are predefined rejection phrases (such as "I'm sorry, but I can't assist with that.") and other content to determine whether this round of testing is intercepted by the model alignment mechanism. If the model response contains the above rejection statements or safety interception symbolic expressions, it is determined that the adversarial prompt node has failed to successfully break through the model defense line, and is considered as a failure of this round of vulnerability testing, which can enter the subsequent optimization.

[0028] Sub-step 4-2, GPT-4-based evaluation. To compensate for the limitations of the keyword-based evaluation method in the semantic understanding level, the present application further introduces a GPT-4-based evaluation mechanism. This method utilizes the capabilities of large language model GPT-4 in natural language understanding and value judgment as an automated evaluation tool to qualitatively and quantitatively score the response R generated in step 3 in the semantic level. In this evaluation process, GPT-4 acts as an objective scorer, evaluating the potential harmfulness of the response content according to the pre-set scoring standard and outputting the corresponding score. The higher the score, the more likely the response contains sensitive, rule-breaking or harmful content. To standardize the evaluation decision, the present application pre-sets a threshold value. If the score result reaches the set threshold value, it is considered that the current round of adversarial prompt has successfully broken through the target model's defense line, and it is determined as an effective jailbreak vulnerability test. If the score is lower than the threshold value, it is considered that the current round of test fails, and the subsequent iterative optimization process is triggered.

[0029] As an improvement of the present application, step 5: adversarial prompt iterative optimization. According to the design of the previous four steps, the present application proposes an adversarial prompt iterative optimization strategy. After the end of step 4 test evaluation, if the current adversarial prompt node fails to successfully induce the target large language to generate harmful responses, the system will automatically start the iterative generation process based on the chain-of-thought reasoning to gradually build its sub-nodes. This process includes analyzing the reasons for the failure of the previous round of prompt node, deducing the possible improvement direction of the prompt, and generating new adversarial prompt nodes in combination with example demonstrations. In the process of generating improvement suggestions, the present application provides a pre-set prompt optimization template for large language models {“If the jailbreak fails, you will need to provide aimprovement. The improvement value contains a few sentences interpreting the language model’s response and how the prompt should be modified to achieve the goal.”} This template guides the large language model to generate targeted improvement suggestions aimed at solving the failure reasons in the current stage. After receiving this template prompt, the large language model will propose feasible optimization strategies or adjustment schemes to address the specific problems identified. This iterative process will continue until the prompt node successfully achieves vulnerability testing or reaches the pre-set maximum tree depth, marking the termination of the prompt path exploration.

[0030] An electronic device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the large language model vulnerability testing method based on the working memory tree when executing the program.

[0031] A storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the large language model vulnerability testing method based on the working memory tree.

[0032] Compared with the prior art, the advantages of the present application are as follows:

[0033] (1) The present application introduces a large language model vulnerability testing method based on working memory tree, which adopts a design framework combining tree structure and iterative optimization mechanism. By introducing working memory theory and data category analysis mechanism, the present application can realize targeted prompt construction and dynamic path optimization, effectively improve the attack adaptability and generation quality of prompts at different nodes, and significantly improve the identification ability and test coverage of potential vulnerabilities of large language models.

[0034] (2) The present application further proposes a multi-prompt combination strategy, which can detect the security boundary of the model from multiple angles through the construction of parallel adversarial prompt set and prompt diversification mechanism, significantly enhancing the diversity and concealment of adversarial prompt attack test.

[0035] (3) The present application introduces a context prompt construction mechanism based on working memory tree structure, combined with category differentiation analysis and role playing and other prompt enhancement strategies, so that the system can accurately simulate the attack path in real use scenarios in the multi-round reasoning process, thereby more realistically reproducing the potential failure risk of large language models in actual deployment. At the same time, the multi-prompt combination strategy proposed by the present application breaks through the limitation of traditional single prompt attack path, and can parallelly schedule multiple adversarial prompts to stimulate the potential vulnerabilities of language models in different context from multiple angles and paths, greatly improving the comprehensiveness and attack efficiency of vulnerability testing. The above design not only significantly enhances the context coherence and attack concealment of adversarial prompts, but also provides a high adaptability, high coverage, and high efficiency vulnerability evaluation and red team testing tool for large language model deployment in key fields such as government systems, financial security, content review, and intelligent customer service, effectively improving the security guarantee capability and credibility evaluation depth of the model before going online, and further promoting the reliable application of large models in complex real scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 The overall model diagram of the embodiment of the present application is shown in the figure;

[0037] Figure 2This is a flowchart illustrating the method processing of an embodiment of the present invention. Detailed Implementation

[0038] To enhance understanding of the present invention, the invention will be further explained below with reference to specific embodiments.

[0039] Example 1: A method for testing vulnerabilities in large language models based on working memory trees. By introducing working memory theory and data category analysis mechanisms, this method achieves efficient generation and optimization of adversarial prompts, significantly improving the success rate and diversity of attacks on large language models. Specifically, this invention constructs a tree-structured generation framework based on a "thinking chain" reasoning mechanism. Each node represents an adversarial prompt generated by optimization of the large language model, and the nodes form semantic association chains through logical reasoning and contextual information, thereby achieving layer-by-layer optimization and semantic refinement of the prompts. Simultaneously, combined with a data category analysis module, potential model vulnerabilities in different categories of data are identified, providing directional guidance for the prompt generation process. Furthermore, to further enhance the diversity and versatility of attack strategies, this invention also designs a multi-prompt combination strategy, integrating multiple semantically related or complementary adversarial prompts and combining them with diverse technologies to generate combined adversarial prompts, which are used to induce the large language model to generate potentially harmful or jailbreak-related outputs. This strategy can effectively expand the attack boundary of large language models and improve the coverage of large language model vulnerability testing.

[0040] For the specific model, please refer to [link / details]. Figure 1 The detailed implementation steps are as follows:

[0041] Step 1: Malicious Semantic Text Collection. First, use web crawling technology to automatically collect information text containing malicious semantics from the Internet, and construct an initial dataset of malicious questions X = {X1, X2, ..., X...}. N}, where X i This dataset represents an initial malicious instruction or question used in the subsequent construction of adversarial prompts and vulnerability testing. It covers a variety of sensitive content, demonstrating diversity and representativeness.

[0042] Step 2: Automated Adversarial Hint Construction. Utilizing the malicious dataset collected in Step 1, the final adversarial hints are automatically constructed through data category analysis, hint scenario construction, and multi-hint combination strategies for large language model vulnerability testing. This step can be divided into the following sub-steps:

[0043] Sub-step 2-1, data category analysis. Given the wide coverage of the adopted dataset, covering multiple fields such as government privacy, dangerous activities, fraud, malware, and other harmful information, it is difficult to effectively respond to the differentiated characteristics and potential vulnerabilities of each category if a unified adversarial prompt design strategy is adopted. Therefore, this step first classifies the original dataset systematically and divides it into six categories: “Illegal Activity”, “Hate Speech”, “Malware”, “Physical Harm”, “Fraud” and “Privacy Violence”. Subsequently, with the help of a large language model, the semantic and attack characteristics between harmful instructions in the initial dataset and their corresponding categories are analyzed, and the key differences between each category compared to other categories are identified. Through such in-depth analysis, the unique model vulnerability characteristics of each category can be further identified and utilized in the subsequent generation of adversarial prompts. This strategy can effectively improve the attack efficiency of the prompt and ensure that it is more tailored to the inherent weaknesses of each category, thereby improving the coverage and relevance of the overall vulnerability test.

[0044] Sub-step 2-2, prompt scenario construction. Based on the category knowledge analysis of step 2-1, the invention uses a working memory tree structure to construct the scenario prompt. At each node, the model will integrate the adversarial prompts, optimization suggestions, and example demonstrations of its parent nodes to build a complete working memory, thereby optimizing the subsequent prompt generation process. This design fully combines the primacy effect and working memory theory in cognitive psychology. The primacy effect indicates that humans have stronger memory and recall ability for the initial information received, while the working memory theory indicates that there is a capacity limit in information processing. In analogy to the processing mechanism of large language models, the former is reflected in the preference of the model's attention mechanism for the initial input, and the latter is reflected in the constraint of the context window size. At the same time, the model prioritizes the response to early inputs when allocating limited attention resources. To depict the primacy bias in this attention allocation process, the invention uses an exponential decay function to model the attention weight of information at different time steps, which is expressed as follows:

[0045] α t = α0·e -λt

[0046] where α t represents the attention intensity of the model at time step t, α0 is the attention baseline value of the initial input, and λ is the hyperparameter that controls the attention decay rate.

[0047] In the prompt construction process, in order to further improve the expression quality and the confrontation concealment, the application introduces three types of enhancement elements: example demonstration, optimization suggestion and role-playing situation, which synergistically drive the prompt generation process of each node. Among them, the example demonstration is used to provide format specification and output reference; the optimization suggestion proposes explicit guidance information to optimize the model generation path at the current node, aiming at the potential deviation in the model generation process; the role-playing situation aims to provide a logical and semantically coherent context shell for the adversarial prompt, which hides the direct exposure of malicious goals and enhances the concealment and execution rationality of the prompt. In order to depict the comprehensive performance of various enhancement elements in prompt generation, the application constructs the following weighted combination:

[0048] S total =w1·E+w2·G+w3·R

[0049] Wherein, E, G, R represent the performance scores of example demonstration, optimization suggestion and role-playing respectively, and w1, w2, w3 are the corresponding weight parameters.

[0050] Substep 2-3, multi-prompt combination strategy. In order to further improve the attack strength and concealment of the adversarial prompt generation, the application introduces a multi-prompt combination strategy in the prompt construction process to break through the limitations of traditional single prompt method in attack path diversity and model vulnerability coverage. Traditional jailbreak attack method usually relies on a single prompt template, which limits its comprehensive mining ability of potential vulnerabilities of large language model. Inspired by multi-sample attack and multi-round jailbreak technology, the application innovatively designs a combination mechanism based on parallel cooperation of multiple adversarial prompts to construct a prompt set with more attack depth and coverage. Specifically, based on model behavior analysis and historical attack records, the application constructs a set of complementary final adversarial prompts P = {P1, P2,..., P n}, wherein P i represents the i-th adversarial prompt generated by substep 2-2. Then, all prompts are uniformly scheduled and input into the target large language model in parallel to form a multi-angle, multi-path parallel attack scene, which strengthens the triggering ability of multiple vulnerable points of the language model.

[0051] Step 3: Large language model vulnerability testing, using the final adversarial prompts obtained in step 2 to test the target large language model for vulnerabilities, to identify the jailbreak vulnerabilities or output boundary vulnerability defects exposed by the large language model when facing malicious input. Specifically, given a target large language model T to be attacked and the malicious semantic data set X = {X1, X2,..., X N} in step 1, if X iWhen the input is sent to the target large language model, the model will usually reject to answer according to its alignment mechanism, and output responses such as "I'm sorry, but I can't assist with that." To break through this defense line, the present application uses the final set of adversarial prompts P generated in step 2 to maximize the probability of inducing the target large language model T to output the target harmful response R while concealing the original problem attack intention, thereby achieving comprehensive testing and evaluation of potential security vulnerabilities of the large language model. The generation behavior of the large language model in this step can be expressed in the form of the following conditional probability distribution formula: i The hidden package maximizes the probability of inducing the target large language model T to output the target harmful response R while concealing the original problem attack intention, thereby achieving comprehensive testing and evaluation of potential security vulnerabilities of the large language model. The generation behavior of the large language model in this step can be expressed in the form of the following conditional probability distribution formula:

[0052]

[0053] where n represents the number of tokens in the adversarial prompt set P, R represents the total number of tokens in the model response R, x n+1:n+L represents the token sequence of response R, x 1:n represents the token sequence of the adversarial prompt set P, and P(A|B) represents the probability of the model generating the next token sequence A given the token sequence B.

[0054] Step 4: Test result evaluation, based on the response results obtained in step 3, systematic evaluation is carried out to determine whether the current adversarial prompt successfully induces the target large language model to generate the expected harmful content, thereby achieving the goal of vulnerability testing. The present application designs two evaluation methods, namely the keyword-based evaluation method and the GPT-4 model-based evaluation method, to realize multi-dimensional judgment of the response output. This step can be divided into the following sub-steps:

[0055] Sub-step 4-1, keyword-based evaluation. Keyword-based evaluation refers to string matching analysis of the response R generated by the language model in step 3, specifically by identifying whether there are predefined rejection phrases (such as "I'm sorry, but I can't assist with that.") and other content to determine whether this round of testing is intercepted by the model alignment mechanism. If the model response contains the above rejection statements or safety interception symbolic expressions, it is determined that the adversarial prompt node has failed to successfully break through the model defense line, and is considered as a failure of this round of vulnerability testing, which can enter the subsequent optimization.

[0056] Sub-step 4-2, GPT-4-based evaluation. To compensate for the limitations of the keyword-based evaluation method in the semantic understanding level, the present invention further introduces a GPT-4-based evaluation mechanism. This method utilizes the capabilities of large language model GPT-4 in natural language understanding and value judgment as an automated evaluation tool to qualitatively and quantitatively score the response R generated in step 3 in the semantic level. In this evaluation process, GPT-4 acts as an objective scorer, evaluating the potential harmfulness of the response content according to the pre-set scoring standard and outputting the corresponding score. The higher the score, the more likely the response contains sensitive, rule-breaking or harmful content. To standardize the evaluation decision, the present invention pre-sets a threshold value. If the score result reaches the set threshold value, it is considered that the current round of adversarial prompt has successfully broken through the target model's defense line, and it is determined as an effective jailbreak vulnerability test. If the score is lower than the threshold value, it is considered that the current round of test fails, and the subsequent iterative optimization process is triggered.

[0057] Step 5: Adversarial prompt iterative optimization. Based on the design of the previous four steps, the present invention proposes an adversarial prompt iterative optimization strategy. After the test evaluation in step 4 is completed, if the current adversarial prompt node fails to successfully induce the target large language to generate harmful responses, the system will automatically start the iterative generation process based on the reasoning of thought chain to gradually build its sub-nodes. This process includes analyzing the reasons for the failure of the previous round of prompt nodes, deducing possible prompt improvement directions, and generating new adversarial prompt nodes in combination with example demonstrations. In the process of generating improvement suggestions, the present invention provides a pre-set prompt optimization template for large language models {“If the jailbreak fails, you will need to provide an improvement. The improvement value contains a few sentences interpreting the language model’s response and how the prompt should be modified to achieve the goal.”} This template guides the large language model to generate targeted improvement suggestions aimed at solving the failure reasons in the current stage. The large language model will propose feasible optimization strategies or adjustment schemes to address the specific problems identified after receiving the template prompt. This iterative process will continue until the prompt node successfully achieves vulnerability testing or reaches the pre-set maximum tree depth, marking the termination of the prompt path exploration.

[0058] Based on the same inventive concept, the application discloses a large language model vulnerability testing method and device based on a working memory tree, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the computer program is loaded into the processor, the above-mentioned large language model vulnerability testing method based on the working memory tree is realized.

[0059] Those skilled in the art will appreciate that the embodiments described herein are presented for the purpose of aiding the reader in understanding the principles of the application and that the embodiments are intended to be illustrative only and are not intended to limit the scope of the application, which is defined in the claims.

Claims

1. A method for vulnerability testing of large language models based on working memory trees, characterized by, The method comprises the following steps: Step 1: Malicious semantic text collection; Step 2: Automatic adversarial prompt construction; Step 3: Large language model vulnerability testing; Step 4: Test result evaluation; Step 5: Adversarial prompt iterative optimization.

2. The method of claim 1, wherein, Step 1: Malicious semantic text collection, first automatically collect information text containing malicious semantics from the Internet through crawler technology, construct the initial malicious question data set X = {X1, X2,..., XN}, where X N represents an original malicious instruction or question, used for subsequent construction of adversarial prompts and vulnerability testing process. i ​ 3. The method of claim 1, wherein, Step 2: Automatic adversarial prompt construction, comprising the following sub-steps: Sub-step 2-1, data category analysis, first, the original data set is systematically classified, and it is inducted and divided into six categories of "Illegal Activity", "Hate Speech", "Malware", "Physical Harm", "Fraud" and "Privacy Violence", then, with the help of a large language model, the semantics and attack characteristics between the harmful instructions in the initial data set and their corresponding categories are analyzed, and the key differences between the categories compared with other categories are identified; Sub-step 2-2, prompt scene construction, based on the category knowledge analysis of step 2-1, the attention weight of different time step information is modeled using an exponential decay function, which is expressed as follows: α t =α0·e -λt where α t The attention strength of the expression model at time step t, α0is the attention baseline value of the initial input, and λ is a hyperparameter that controls the decay rate of attention. The following weighted combination is constructed: S total = w1 · E + w2 · G + w3 · R Where E, G, and R represent the effectiveness scores of example demonstration, optimization suggestion, and role-playing, respectively, and w1, w2, and w3 are the corresponding weight parameters, Sub-step 2-3, multi-prompt combination strategy, as follows, based on model behavior analysis and historical attack records, a set of complementary final adversarial prompt set P = {P1, P2,..., P n} is constructed, where P i represents the i-th adversarial prompt generated through sub-step 2-2, then all prompts are uniformly scheduled and input into the target large language model in parallel, forming a multi-angle, multi-path parallel attack scene, and strengthening the triggering ability of multiple vulnerable points of the language model.

4. The method of claim 1, wherein, Step 3: Large Language Model Vulnerability Testing. Using the final adversarial hints obtained in Step 2, vulnerability testing is performed on the target large language model to identify jailbreak vulnerabilities or output boundary vulnerabilities exposed when the large language model faces malicious input. Specifically, given a target large language model T to be attacked and the malicious semantic dataset X = {X1, X2, ..., X...} from Step 1... N If X is directly... i When the input is fed into the target large language model, the model typically rejects the response based on its alignment mechanism and outputs a response. Using the final adversarial cue set P generated in step two, the original malicious input X is then... i The covert packaging maximizes the probability of inducing the target large language model T to output a harmful response R while concealing the original attack intent. This enables comprehensive testing and evaluation of potential security vulnerabilities in the large language model. The generative behavior of the large language model in this step is formally expressed by the following conditional probability distribution formula: where n denotes the number of tokens in the adversarial prompt set P, R denotes the total number of tokens in the model response R, x n+1:n+L denotes the token sequence of the response R, x 1:n denotes the token sequence of the adversarial prompt set P, P(A|B) denotes the probability that the model generates the next token sequence as A given the condition of the token sequence B.

5. The method of claim 1, wherein, Step 4: Test result evaluation, comprising the following sub-steps: Sub-step 4-1, keyword-based evaluation, keyword-based evaluation refers to string matching analysis of the response R generated by the language model in step 3, specifically by identifying whether there are predefined rejection phrases, to determine whether the test is intercepted by the model alignment mechanism, if the model response contains the above rejection statements or safety interception symbolic expressions, it is determined that the adversarial prompt node has failed to break through the model defense line, and is considered as a failed vulnerability test, which can enter the subsequent optimization, Sub-step 4-2, GPT-4-based evaluation, using the large language model GPT-4's ability in natural language understanding and value judgment as an automated evaluation tool, the response R generated in step 3 is qualitatively and quantitatively scored at the semantic level, in this evaluation process, GPT-4 as an objective scorer, according to the pre-set scoring standard, evaluates the potential harmfulness of the response content and outputs the corresponding score, the higher the score, the more likely the response contains sensitive, illegal or harmful content, a threshold is preset, if the score reaches the threshold, it is considered that the adversarial prompt has successfully broken through the defense line of the target model, and is determined as an effective jailbreak vulnerability test; if the score is lower than the threshold, it is considered that the test fails, and the subsequent iterative optimization process is triggered.

6. The method of claim 1, wherein, Step 5: Adversarial prompt iterative optimization, according to the design of the previous four steps, after the test evaluation in step 4 is completed, if the current adversarial prompt node fails to successfully induce the target large language to generate harmful responses, the system will automatically start the iterative generation process based on the reasoning of the thought chain, to gradually build its sub-nodes, this process includes analyzing the reasons for the failure of the previous prompt node, deducing the possible improvement direction of the prompt, and generating new adversarial prompt nodes combined with example demonstration.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: The processor implements the working memory tree-based large language model vulnerability testing method in any one of claims 1-6 when executing the program.

8. A storage medium having stored thereon a computer program, characterized in that: The computer program implements the working memory tree-based large language model vulnerability testing method in any one of claims 1-6 when executed by the processor.

Citation Information

Cited By

  • Steganographic confrontation test method and system for large language model

    CN121278736A