Large-model multi-scene antagonism dynamic evaluation system and method based on context perception strategy optimization
By constructing state space and action space, using reinforcement learning algorithms to generate adversarial samples and build a vulnerability knowledge base, we solve the problems of insufficient context modeling and dynamic strategy optimization of large language models in multi-round interaction scenarios, and achieve efficient and explainable security assessment and optimization of adversarial evaluation.
Patent Information
- Application Number
- CN202511266721.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing security assessment methods for large language models in multi-round interaction scenarios have problems such as insufficient context modeling, insufficient language quality and semantic consistency in adversarial sample generation, and a lack of dynamic strategy optimization mechanisms, making it difficult to effectively identify potential risks and optimize the model.
Construct state space and action space, use reinforcement learning algorithm to generate adversarial samples, build vulnerability knowledge base through semantic perturbation operation and context analysis, dynamically adjust policy network parameters, and generate security reports.
The state representation capability in multi-round interactions has been improved, and the generated adversarial samples maintain language fluency and semantic consistency, achieving continuous optimization of adversarial strategies and vulnerability mining efficiency, and supporting explainable security assessment and defense reinforcement.
Smart Images

Figure CN120764696A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence security evaluation, and particularly relates to a large model multi-scene adversarial dynamic evaluation system and method based on context-aware strategy optimization. BACKGROUND
[0002] Current large language models are widely used in customer service question answering, medical consultation and other scenarios, and the compliance and safety of their output content have gradually become an important research direction of artificial intelligence security evaluation. Existing security evaluation methods mainly include static testing and rule-based interception mechanisms.
[0003] Static testing methods generally detect model responses by constructing fixed questions or template sentences, but they have the problem of insufficient coverage in handling multi-turn dialogues with dynamic interaction characteristics, and are difficult to capture potential risks of models in real use scenarios. The adversarial samples generated by such methods often affect language fluency, leading to a gap between test results and actual application. In addition, current technologies generally lack the ability to model context and model response uncertainty.
[0004] On the other hand, the detection strategy based on keyword interception and template matching is less robust when facing semantic perturbation attacks (such as synonym replacement and sentence restructuring), and it is difficult to effectively identify the ambiguity of model output near the decision boundary. Existing methods generally use static strategies to perform attacks and evaluations, and lack a mechanism for dynamically adjusting attack strategies based on model response results, limiting their comprehensive discovery and continuous optimization capabilities for potential vulnerabilities.
[0005] In summary, the existing security evaluation system of large language models still has room for improvement in the following aspects: first, an effective context modeling mechanism has not been established to support state representation under multi-turn interaction; second, existing adversarial sample generation methods have deficiencies in language quality and semantic consistency; third, there is a lack of strategy evolution mechanism based on model feedback, making it difficult to achieve efficient closed-loop evaluation and optimization. SUMMARY
[0006] To solve the above technical problems, the application proposes a large model multi-scene adversarial dynamic evaluation system and method based on context-aware strategy optimization to solve the problems existing in the prior art.
[0007] In the first aspect, to achieve the above object, the application provides a large model multi-scene adversarial dynamic evaluation method based on context-aware strategy optimization, comprising the following steps:
[0008] Constructing a state space and an action space for high-risk application scenarios;
[0009] Build a policy network based on the reinforcement learning algorithm, sample perturbation actions according to the state vector to generate adversarial samples and submit them to the large language model, and calculate the reward value based on the model response to optimize the policy network parameters;
[0010] Perform semantic perturbations on the original attack samples and filter them to construct a diverse set of standard adversarial samples.
[0011] Record and analyze model interaction logs to build a vulnerability knowledge base;
[0012] Adjust strategies and calculate evaluation indicators based on feedback from the vulnerability knowledge base to generate security reports.
[0013] Optionally, the process of constructing a state space and an action space for a high-risk application scenario includes:
[0014] Divide attack categories and define task context labels, collect interaction samples to build structured context templates for multiple scenarios;
[0015] Use semantic encoding model to extract state vector to construct state space;
[0016] Based on the context structure and attack labels, the action space is defined through syntactic analysis and vocabulary replacement rules.
[0017] Optionally, the process of calculating the reward value based on the model response to optimize the policy network parameters includes:
[0018] Input the state vector into the policy network and sample the attack action, call the perturbation template to generate adversarial input samples and obtain the model response;
[0019] Based on the model response content, the violation is determined and the entropy value is calculated, and the reward value is calculated according to the reward function;
[0020] Calculate the current strategy ratio and advantage function, build the objective function based on the pruning strategy and update the network parameters through back propagation.
[0021] Optionally, the process of performing semantic perturbation operations on the original attack samples and screening includes:
[0022] Perform semantic perturbation operations on the original attack sample to generate multiple semantically consistent variants;
[0023] Perform grammatical legitimacy and fluency screening on variant samples in turn;
[0024] Add attack type and risk level labels to the retained samples, and select high-quality samples based on confidence scores to build a sample set.
[0025] Optionally, the process of recording and analyzing model interaction logs to build a vulnerability knowledge base includes:
[0026] Record input context, perturbation actions, model responses, and violation determination results to generate a structured response log;
[0027] Perform semantic analysis on the response log to extract semantic association patterns between the perturbation operation and the triggered abnormal response;
[0028] A vulnerability knowledge base is built based on the analysis results to record index information of attack patterns, context structure types and abnormal output features.
[0029] Optionally, the process of adjusting the strategy and calculating the evaluation index based on the feedback from the vulnerability knowledge base includes:
[0030] Dynamically adjust the action preference weights in the reward function based on the attack blind spots and boundary abnormal behaviors extracted from the vulnerability knowledge base;
[0031] Traverse the evaluation sample set and count the model response results to calculate the attack success rate and vulnerability coverage;
[0032] Combine the indicator statistical results with the vulnerability knowledge base analysis content to generate a security assessment report.
[0033] In a second aspect, the present invention further provides a large-model multi-scenario adversarial dynamic evaluation system based on context-aware strategy optimization, which is used to implement a large-model multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization, and the system includes:
[0034] Context construction module, used to construct state space and action space for high-risk application scenarios;
[0035] The policy training module is used to build a policy network based on the reinforcement learning algorithm, sample perturbation actions according to the state vector to generate adversarial samples and submit them to the large language model, and calculate the reward value based on the model response to optimize the policy network parameters;
[0036] The sample expansion module is used to perform semantic perturbation operations on the original attack samples and filter them to construct a diverse set of standard adversarial samples;
[0037] The evaluation and analysis module is used to record and analyze model interaction logs, build a vulnerability knowledge base, adjust the strategy of the strategy training module based on the feedback from the vulnerability knowledge base, and calculate evaluation indicators to generate a security report.
[0038] In a third aspect, the present invention further provides a computer terminal device, comprising:
[0039] one or more processors;
[0040] a memory, coupled to the processor, for storing one or more programs;
[0041] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the large-model multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization in the above-mentioned first aspect.
[0042] In a fourth aspect, the present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the large-model multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization in the above-mentioned first aspect are implemented.
[0043] In a fifth aspect, the present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the large-model multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization in the above-mentioned first aspect.
[0044] Compared with the prior art, the present invention has the following advantages and technical effects:
[0045] The present invention provides a large-model, multi-scenario adversarial dynamic evaluation system and method based on context-aware strategy optimization. This system effectively improves the state representation capability in multiple rounds of interaction through context modeling, and the generated adversarial samples excel in maintaining language fluency and semantic consistency. Based on reinforcement learning and a dynamic reward feedback mechanism, continuous optimization of adversarial strategies and improved vulnerability mining efficiency are achieved. By building a vulnerability knowledge base, the system can identify and summarize abnormal behavior at model boundaries, thereby supporting explainable security assessments and targeted defense reinforcement. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0047] Figure 1 This is the main flow chart of the adversarial evaluation according to an embodiment of the present invention;
[0048] Figure 2 A detailed flow chart of the training strategy according to an embodiment of the present invention;
[0049] Figure 3 The detailed process of generating adversarial samples according to the embodiment of the present invention;
[0050] Figure 4 This is an execution diagram of the core functional modules of an embodiment of the present invention. DETAILED DESCRIPTION
[0051] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0052] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.
[0053] Embodiment one
[0054] As shown in the flowchart of Figure 1 The embodiment provides a large model multi-scene adversarial dynamic evaluation method based on context-aware strategy optimization, which includes:
[0055] Constructing state space and action space for high-risk application scenarios;
[0056] Based on the reinforcement learning algorithm, a strategy network is constructed, a perturbed action is sampled according to the state vector to generate an adversarial sample and submit it to a large language model, and a reward value is calculated according to the model response to optimize the strategy network parameters;
[0057] Performing semantic perturbation operation on the original attack sample and screening to construct a diversified standard adversarial sample set;
[0058] Recording model interaction logs and analyzing to construct a vulnerability knowledge base;
[0059] Adjusting the strategy according to the feedback of the vulnerability knowledge base and calculating the evaluation index to generate a security report.
[0060] Specifically, the specific steps of the above process are:
[0061] S1: For high-risk application scenarios such as customer service Q&A, medical consultation, etc., divide the attack categories and define the task context labels. Collect the interaction samples of the task context and intent, construct the structured context templates of multiple scenarios, and use the semantic encoding model to extract the state vector to construct the state space S. According to the template set by the attack target, the action space A is constructed;
[0062] S2: Based on the PPO algorithm, a strategy network and an evaluation network are constructed, a perturbed action is sampled according to the state vector, and an adversarial sample is generated. The sample is submitted to the model to obtain the response, and the reward is calculated by combining the violation judgment and the output entropy value . Construct the loss function and update the network parameters to realize strategy optimization;
[0063] S3: Apply semantic perturbation operation to the original attack sample to generate multiple semantically consistent variants, and perform syntax legality and fluency screening in turn. Add structured labels to the retained samples, and select high-quality samples based on confidence scores to construct a diversified standard adversarial sample set;
[0064] S4: Record the interaction log of the model, analyze the semantic relationship between the disturbance operation and the abnormal response, extract the attack features and build a vulnerability knowledge base containing attack pattern index information , to support subsequent evaluation and policy optimization;
[0065] S5: Adjust the reward function according to the feedback of the knowledge base, optimize the action policy preference, and guide the policy to migrate to the boundary area of the model. Combine the evaluation sample set statistics ASR, VCR and other indicators, output charts and security reports for model reinforcement and deployment support;
[0066] As an embodiment in this embodiment, the process of constructing the state space and action space for high-risk application scenarios includes:
[0067] Divide the attack categories and define the task context labels, collect interaction samples to build structured context templates for multiple scenarios;
[0068] Use semantic encoding model to extract state vector to build state space;
[0069] Based on the context structure and attack label, define the action space through syntax analysis and vocabulary replacement rules.
[0070] Further, the specific process of S1 includes:
[0071] S1.1: Based on open source dialogue data set, adopt clustering algorithm combined with BERT encoding to identify high-risk input mode. Through keyword matching and regular template, identify attack types such as "bypassing ban" and "sensitive detection", and use label mapping dictionary to automatically assign uniform attack label for subsequent sample labeling and policy input configuration;
[0072] S1.2: On the basis of attack label, use BERT vector and clustering algorithm to identify semantic similar samples, extract keywords and syntax structure combined with SpaCy. According to the rules, give risk level and match recommended disturbance strategy, generate task-risk mapping table containing "attack label, sentence template, risk level, strategy type". Organize the above mapping information into structured format, build attack sample label database, use database management to structure label data, fields include label ID, attack type, risk level, sentence template, strategy suggestion, etc., support subsequent sample expansion and disturbance evaluation stage label writing and query calling;
[0073] S1.3: Sample interaction snippets from annotated data, extract user input, context and model response, analyze syntax structure through SpaCy and identify task intent using BERT model, format results into structured templates for context encoding and state space construction. Meanwhile, record BERT's classification probability distribution and combine with context consistency score to generate corresponding confidence values written into templates for sample screening and evaluation module;
[0074] S1.4: For each context template in user input , historical context , and model response , generate state vectors using Sentence-BERT , and aggregate all samples to construct state space S, state vector is defined as follows:
[0075] ;
[0076] S1.5: Based on context structure and attack labels, use SpaCy to extract syntax dependency paths and select keyword positions, perform masked token replacement through BERT-MLM to generate word perturbation and sentence rearrangement templates. Combine pre-built trigger rule library and sentence transformation template library based on syntax dependency path analysis to define action space A. Meanwhile, construct semantic style alignment module, design legality filtering rules and general evaluation interface based on syntax integrity and language fluency principles for subsequent strategy training stage;
[0077] S1.6: Input state space S and action space A into Markov decision environment to construct a training process that supports multi-round context interaction. The Markov decision environment can be represented as a five-tuple:
[0078] ;
[0079] where P is the state transition probability, R is the reward function, is the discount factor. This environment serves as the core input for policy training, defining state-action mapping and context interaction rules for training loops and log modules;
[0080] As an embodiment in this embodiment, the process of calculating reward value according to model response to optimize policy network parameters includes:
[0081] Input state vector into policy network and sample attack action, call perturbation template to generate adversarial input samples and get model response;
[0082] Based on the content of the model response, make a violation determination and calculate the entropy value, and calculate the reward value according to the reward function;
[0083] The current policy ratio and advantage function are calculated, the target function is constructed based on the pruning strategy, and the network parameters are updated through back propagation.
[0084] Further, the specific process of S2 includes:
[0085] S2.1: Construct an interaction log, including the state vector , the policy output action , the perturbation input , the model response , the violation judgment label , and the reward value .
[0086] S2.2: Construct a policy network based on the PPO algorithm and an evaluation network , access the Markov environment, and prepare for the training loop;
[0087] S2.3: Input the state vector into the policy network , sample the attack action according to the policy probability distribution , then call the perturbation template library to generate an adversarial input sample . Then submit the sample to the large language model interface to receive the response result and write it into the interaction log module;
[0088] S2.4: Based on the model response content , make a violation judgment, calculate the entropy value and the interception state, calculate the reward function , and write it into the log;
[0089] S2.5: After each round of training, calculate the current policy ratio , combine the state value predicted by the evaluation network , obtain the advantage function , and construct the PPO target function based on the pruning strategy .
[0090] As an embodiment of the present embodiment, the process of performing semantic perturbation operation on the original attack sample and screening includes:
[0091] Performing semantic perturbation operation on the original attack sample to generate multiple semantically consistent variants;
[0092] Performing syntax legality and fluency screening on the variant samples in turn;
[0093] Adding attack type and risk level labels to the retained samples, and selecting high-quality samples based on confidence scores to construct a sample set.
[0094] Further, the specific process of S3 includes:
[0095] S3.1: Perform semantic perturbation operation based on original attack seed sample to generate diversified variants;
[0096] S3.2: Remove low-quality samples through structural inspection, semantic consistency and legality detection module;
[0097] S3.3: Add structured labels (attack type, threat level, misleading level) to qualified samples and record them to attack sample database;
[0098] S3.4: Use confidence scoring mechanism to select a gold sample set above 95% to improve the representativeness and coverage of adversarial samples;
[0099] As an embodiment in this embodiment, the process of recording model interaction logs and analysis, constructing vulnerability knowledge base includes:
[0100] Record input context, perturbation action, model response and violation determination result to generate structured response log;
[0101] Perform semantic analysis on response log to extract semantic association patterns between perturbation operation and triggering abnormal response;
[0102] Based on the analysis results, construct a vulnerability knowledge base to record the index information of attack patterns, context structure types and abnormal output characteristics.
[0103] Further, the specific process of S4 includes:
[0104] S4.1: In the evaluation interaction process, call the log recording module to write input context, perturbation action, model response and violation determination result in real time, generate structured response log, and store it in response log database;
[0105] S4.2: Perform batch semantic analysis on response log to extract semantic association patterns between perturbation operation and triggering abnormal response, and identify abnormal behavior patterns in the model boundary fuzzy area;
[0106] S4.3: Based on the above analysis results, construct a vulnerability knowledge base Record index information containing attack patterns, context structure types and abnormal output characteristics for subsequent strategy feedback and model behavior attribution;
[0107] As an embodiment in this embodiment, the process of adjusting strategy and calculating evaluation index according to vulnerability knowledge base includes:
[0108] According to the attack blind area and boundary abnormal behavior extracted from the vulnerability knowledge base, dynamically adjust the action preference weight in the reward function;
[0109] Traverse the evaluation sample set and count the model response results to calculate the attack success rate and vulnerability coverage;
[0110] Combine the indicator statistical results with the vulnerability knowledge base analysis content to generate a security assessment report.
[0111] Furthermore, the specific process of S5 includes:
[0112] S5.1: Based on the vulnerability knowledge base Based on the attack blind spots and boundary abnormal behaviors extracted from the algorithm, the action preference weights of the reward function in the strategy learning phase are dynamically adjusted to guide the strategy to migrate to high-risk areas.
[0113] S5.2: Traverse the current evaluation sample set, count the model response results, calculate the attack success rate and vulnerability coverage, and write the evaluation index results into the evaluation index database.
[0114] Attack success rate = number of samples triggering violation responses / total number of attack samples:
[0115] ;
[0116] Vulnerability coverage = number of identified vulnerability types / total number of preset vulnerability types:
[0117] ;
[0118] S5.3: Combine the indicator statistics with the vulnerability knowledge base analysis content to generate a risk coverage map and export a structured security assessment report to support model deployment optimization and policy updates;
[0119] More specifically, the present invention provides a large-model multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization. The detailed overall process includes the following steps:
[0120] S1: Based on typical high-risk applications such as customer service Q&A and medical consultation, we construct multi-task risk mapping and context labels to form three major scenarios: dialogue, reasoning, and Q&A. We use SentenceBERT to semantically encode user input, historical context, and task intent, and construct a state vector space. , action space It includes lexical perturbation, context rearrangement and insertion patch strategies, and introduces a Markov decision structure to support multi-round context interaction;
[0121] S2: Use the PPO-based reinforcement learning algorithm to build the policy network and evaluation network. In each round of training, the state is sampled from the state space and input into the policy network to obtain the action distribution. After performing the perturbation operation, the adversarial sample is submitted to the target large model to obtain the response. . Construct a reward function using the violation features in the response, output entropy value, and security bypass situation Optimize the policy network parameters through the advantage function and the cut target;
[0122] S3: Expand the seed attack sample using a semantic disturbance mechanism, including synonym replacement, sentence rewriting, character fuzzing, and context logic patching, to generate 2-5 semantic-preserving disturbance variants. High-quality samples are selected through legality check and language fluency, and multi-dimensional labels such as attack type, threat level, and misleading are added to form a structured attack sample database;
[0123] S4: The system records the disturbance context, action, model response, and violation determination results in real time during the evaluation process to form a structured log. Perform semantic analysis on the log content to extract the association pattern between the disturbance and abnormal response, and build a vulnerability knowledge base to describe the semantic mapping relationship between attack behavior and model abnormal output;
[0124] S5: According to the pattern information in the vulnerability knowledge base , dynamically adjust the policy preference in the reward function to guide the agent to optimize the migration to the boundary fuzzy area. The system evaluates the sample response results, calculates the attack success rate (ASR) and vulnerability coverage rate (VCR), and finally generates evaluation charts and security reports to support model security reinforcement and deployment optimization;
[0125] For ease of understanding, the following takes "account ban bypass in the intelligent customer service scenario" as a specific application scenario, and gives a complete implementation process example of the method according to the above steps.
[0126] This embodiment takes a customer service question and answer system as the attack target, simulates the risk scenario of users trying to bypass the account ban restriction, and completely implements the process from context construction, policy training to vulnerability extraction.
[0127] S1.1: Extract 200 user inputs containing expressions such as "unblock" and "limit removal" from open source customer service dialogue data. After encoding using BERT, identify 4 high-risk modes based on KMeans clustering. Combine keyword rules and regular templates to automatically label the input as attack type "ban bypass", corresponding to attack label "TAG001".
[0128] S1.2: Extract keywords and dependency structures (such as subject-predicate-object) from the labeled samples, set the risk level to "high" based on the semantic clustering results, and the policy type to "sentence disturbance + word replacement". Build a task-risk mapping table and write it into a structured label database, fields include attack label, sentence template, risk level and recommended strategy.
[0129] S1.3: Select sample pairs from the labeled set, such as the user input "I was blocked, how to unblock?", the context "account login failure prompt exception", the model response "please describe the problem you encountered" and other statements, use SpaCy to analyze the syntax structure, use BERT to identify the intent "account recovery", generate context templates and record classification confidence for subsequent screening.
[0130] S1.4: The user input , the context , and the model response are used to generate state vectors with a unified dimension of 256, and a state space S is constructed.
[0131] S1.5: Extract the syntactic path key position, such as "action word + account object", use BERT-MLM to generate "account unblocking" → "how to recover account" and other word disturbance templates. Combine the pre-set rule library to construct the action space A, including "word replacement" "structure rearrangement" and other operations, and set up a legality detection interface for screening illegal disturbances.
[0132] S1.6: Input the state space S and the action space A into the Markov decision environment, define the state transition rules and reward mechanism, and call the policy training module.
[0133] S2.1: Construct the interaction log, fields include state vector , policy output action , disturbance input , model response , violation judgment label and reward value .
[0134] S2.2: Construct the policy network and evaluation network based on the PPO algorithm, access the Markov environment, and start the loop training. The specific training process is described in detail in the Figure 2 .
[0135] S3.1: Based on the original sample "how to unblock account", perform word replacement and sentence rearrangement to generate 1200 variant samples.
[0136] S3.2: Apply syntax legality detection and syntax consistency scoring to eliminate ambiguous, structurally incorrect disturbance sentences, a total of 304 are excluded.
[0137] S3.3: Add structured labels to the remaining samples, including attack type "bypassing blocking", threat level "high", and misleading level "medium", and write them into the attack sample database.
[0138] S3.4: Using the model confidence scoring mechanism, sentences with a score greater than 95% are retained to form the golden sample set, totaling 85 sentences.
[0139] S4.1: Record all interaction logs during the evaluation process, including disturbance actions, response content, and violation judgment results, and store them in the response log database.
[0140] S4.2: Batch analyze logs to identify high-frequency violation responses such as "Handled for you" and "Account recovery method", and associate their triggering actions with "Sentence rearrangement + keyword replacement".
[0141] S4.3: Build a vulnerability knowledge base , record attack tags and context templates such as "I was blocked...", trigger keywords such as "activation, recovery", build attack pattern indexes, and support strategy updates and model attribution.
[0142] S5.1: Basis We extract the blind spot pattern from the dataset and adjust the reward function, increasing the weight of actions containing keywords such as “recovery” and “release” to 1.5 times the original weight, guiding the strategy to move closer to high-risk areas.
[0143] S5.2: In the current evaluation sample set, the model returned 79 violation responses and 96 attack samples, with an ASR of 82.3%; a total of 5 known attack patterns were triggered, with a VCR of 71.4%.
[0144] S5.3: Output a structured safety assessment report for use in model deployment optimization.
[0145] Refer to the attached Figure 2 As shown, this embodiment provides a detailed process of the training strategy, including the following steps:
[0146] S2.1: Build an interaction log to encapsulate the interaction data generated in each round of training. The fields include the state vector , Strategy output action , disturbance input , model response , Violation determination label and reward value All data are imported into the database for structured storage;
[0147] S2.2: Initialize the reinforcement learning policy structure and build a policy network based on the PPO algorithm and evaluation network , combined with the Markov decision environment defined in S1.6, build a training loop that includes state update, action selection, reward evaluation and policy update;
[0148] S2.3: In the training loop, the system inputs the state vector into the policy network and samples attack actions according to the output probability distribution The perturbation template library is then called to retrieve the corresponding vocabulary replacement or sentence rearrangement rules based on the action type. The selected rules are applied to the original input text to generate the corresponding adversarial samples. The generated adversarial samples are submitted to the large language model interface. The system receives the model response results and writes the interaction process to the log module, providing basic data for subsequent reward calculation and vulnerability analysis.
[0149] S2.4: Based on the model response content, call the BERT discriminator fine-tuned based on the historical violation response dataset to output the violation label , combined with the response entropy value H ( ) and interception avoidance flag The system preset security interception API returns a status code, and the system output status code determines the instant reward calculated according to the formula:
[0150]
[0151] in , , is the reward coefficient;
[0152] S2.5: After each round of training, calculate the current strategy ratio , combined with the state value predicted by the evaluation network , to obtain the advantage function , construct the PPO objective function based on the clipping strategy:
[0153]
[0154] in is the clipping coefficient. The system uses the back-propagation algorithm to perform gradient optimization on the policy network and the evaluation network, and updates the network parameters to complete the reinforcement learning closed-loop training.
[0155] Refer to the attached Figure 3 As shown, this embodiment provides a detailed process for generating adversarial samples, including the following steps:
[0156] S3.1: The initial attack sample set is formed based on the annotated attack templates constructed in S1.3 and the real adversarial samples generated in S2.3. Semantic perturbation is invoked to perform synonym replacement, sentence rearrangement, and context insertion on each sample, generating multiple semantically preserved structural variants in batches while preserving the original attack intent field.
[0157] S3.2: Load the semantic style alignment module and perform grammatical structure analysis, spelling standardization verification, and language fluency scoring on the generated samples. Based on the rule template and the preset threshold, low-quality samples with semantic drift, spelling errors, or abnormal formatting are eliminated, and effective perturbation samples with complete structure and smooth expression are retained;
[0158] S3.3: Based on the task-risk mapping table and the attack sample label database, the system assigns a corresponding attack type and risk level to each sample, and matches it based on its attack label field. At the same time, the system evaluates the misleading level of the sample by combining the semantic perturbation amplitude and the model's response results. After the evaluation is completed, the system writes the structured label information into the attack sample label database. The recorded fields include sample ID, attack type, risk level, misleading level, and generation timestamp, which are used to support subsequent analysis and traceability.
[0159] S3.4: The semantic similarity between the original and perturbed samples is calculated using the BERT model. Samples with a similarity below 0.85 are corrected by sentence rearrangement and are discarded if they still do not meet the standard. Grammatical structure, language fluency, and spelling standards are then verified to filter out low-quality samples. For samples that pass the screening, the confidence field in the structured template is extracted. High-quality samples with a confidence level above 0.95 are retained and stored in the evaluation sample library.
[0160] Based on this, compared with the prior art, the application proposes a large model multi-scene adversarial dynamic evaluation system based on context-aware strategy optimization, uses a context modeling method to solve the state representation missing problem in the multi-round interaction scene, designs a semantic-preserving sample expansion technology to make up for the deficiencies of adversarial samples in language quality and semantic consistency, and proposes a reward feedback-based strategy optimization mechanism to alleviate the lack of dynamic adjustment and closed-loop optimization in the evaluation process. Based on the advantages of the above methods, the application designs an efficient large model security evaluation system in four modules. Including the context construction module, the strategy training module, the sample expansion module and the evaluation analysis module connected in turn, and working cooperatively. The context construction module is used to divide the attack types of high-risk fields, collect user input, historical context and task intent information, call a text encoding model to generate a state space S, and construct an action space containing word perturbation and sentence rearrangement rules. The strategy training module initializes the strategy network based on the reinforcement learning algorithm, generates perturbed actions according to the state space S, constructs adversarial samples and submits the large model to obtain the response result. The system calculates the reward value according to the violation identifier, information entropy change and interception state in the response, so as to optimize the parameters of the strategy network. The sample expansion module is used to generate semantic-preserving variants of the original samples, perform syntax legality verification and language fluency detection, add attack type and risk level labels, and select standard sample sets and high-confidence sample sets according to the confidence, which are used for subsequent strategy training and evaluation. The evaluation analysis module is responsible for recording the model interaction log, analyzing the semantic association between the perturbation operation and the abnormal response, and constructing a vulnerability knowledge base containing attack mode index information . The system dynamically adjusts the reward function weight in the strategy training module based on the knowledge base, calculates the attack success rate and vulnerability coverage rate of the evaluation samples, and finally generates a security evaluation report to provide basis and support for the optimization of defense strategies.
[0161] Embodiment two
[0162] Based on the same overall inventive concept, the application also provides a large model multi-scene adversarial dynamic evaluation system based on context-aware strategy optimization. The following describes the large model multi-scene adversarial dynamic evaluation system based on context-aware strategy optimization provided by the application, and the large model multi-scene adversarial dynamic evaluation system based on context-aware strategy optimization described below can be mutually corresponding and referred to. Referring to the accompanying Figure 4 , the embodiment provides a large model multi-scene adversarial dynamic evaluation system core function module execution diagram based on context-aware strategy optimization, including the following modules:
[0163] The context construction module divides the high-risk field attack type and collects user input, historical context and task intention data, calls a text encoding model to construct a state space S, defines a text disturbance rule to form an action space A, and obtains a dynamic context representation to support high-risk scenario coverage.
[0164] The strategy training module initializes a strategy network through a reinforcement learning algorithm to solve the problem of lack of feedback loop for static strategies, generates adversarial samples based on the state space S, submits a large model, calculates a reward value to optimize network parameters by combining violation identification, information entropy and interception state, and realizes dynamic evolution of adversarial strategies to improve vulnerability mining efficiency.
[0165] The sample expansion module generates semantic preservation variants from original samples, performs syntax legality verification and fluency detection, adds attack type and risk level labels, and then performs confidence screening to output a standard sample set and a high-confidence sample set to enhance test effectiveness.
[0166] The evaluation analysis module records model interaction logs to analyze disturbance-response association, constructs a vulnerability knowledge base, dynamically adjusts reward function weights, calculates ASR and VCR to generate a security evaluation report, and drives iterative update of defense strategies to realize closed-loop evaluation optimization.
[0167] It should be understood that the embodiment of the present application provides a large model multi-scene adversarial dynamic evaluation system based on context-aware strategy optimization, which has all the advantages of the large model multi-scene adversarial dynamic evaluation method based on context-aware strategy optimization provided by the above-mentioned embodiments.
[0168] Embodiment three
[0169] In this embodiment, a computer terminal device is provided, comprising:
[0170] one or more processors;
[0171] a memory coupled to the processor, configured to store one or more programs;
[0172] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-mentioned large model multi-scene adversarial dynamic evaluation method based on context-aware strategy optimization.
[0173] In this embodiment, a computer readable storage medium having a computer program stored thereon is also provided, and the computer program is executed by a processor to implement the steps of the above-mentioned large model multi-scene adversarial dynamic evaluation method based on context-aware strategy optimization.
[0174] In the embodiment, an electronic device is also provided, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the above-mentioned large model multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization.
[0175] In the embodiment, a computer program product is also provided, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned large model multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization.
[0176] The above-mentioned program can be run in a processor, or can also be stored in a memory (or called computer readable medium), which includes permanent and non-permanent, removable and non-removable media, and can be realized by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0177] These computer programs can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to generate a computer implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 The steps of the functions specified in one or more flows or one or more blocks Figure 1 The steps of the functions specified in one or more flows or one or more blocks
[0178] Such a device or system is provided in the embodiment. The system is called a large model multi-scenario adversarial dynamic evaluation system based on context-aware strategy optimization, which comprises:
[0179] A context construction module is configured to construct a state space and an action space for a high-risk application scenario;
[0180] A strategy training module is configured to construct a strategy network based on a reinforcement learning algorithm, sample a perturbation action according to a state vector to generate an adversarial sample and submit it to a large language model, and calculate a reward value according to a model response to optimize strategy network parameters;
[0181] a sample expansion module configured to perform semantic perturbation operations on the original attack samples and screening, and to construct a diversified standard adversarial sample set;
[0182] an evaluation analysis module configured to record model interaction logs and analysis, to construct a vulnerability knowledge base, and to adjust the strategy of the strategy training module according to the feedback of the vulnerability knowledge base and to calculate evaluation indexes to generate a security report.
[0183] As an embodiment in the present embodiment, the context construction module comprises:
[0184] a scene division unit configured to divide attack categories and define task context labels, and to collect interaction samples to construct a multi-scene structured context template;
[0185] a state encoding unit configured to extract state vectors using a semantic encoding model to construct a state space;
[0186] an action definition unit configured to define an action space based on the context structure and attack labels, through syntax analysis and vocabulary replacement rules.
[0187] As an embodiment in the present embodiment, the strategy training module comprises:
[0188] an interaction execution unit configured to input the state vector into the strategy network and sample attack actions, to call the perturbation template to generate adversarial input samples and submit them to the large language model to obtain model responses;
[0189] a reward calculation unit configured to make a violation determination based on the content of the model response and calculate an entropy value, and to calculate a reward value according to a reward function;
[0190] a strategy optimization unit configured to calculate the current strategy ratio and the advantage function, to construct a target function based on the pruning strategy, and to update the parameters of the strategy network and the evaluation network through a back propagation algorithm.
[0191] As an embodiment in the present embodiment, the sample expansion module comprises:
[0192] a semantic perturbation unit configured to perform semantic perturbation operations on the original attack samples to generate a plurality of semantically consistent variants;
[0193] a sample screening unit configured to sequentially perform syntax legality and fluency screening on the variant samples;
[0194] a sample management unit configured to add attack type and risk level labels to the retained samples, and to select high-quality samples based on confidence scores to construct a sample set.
[0195] As an embodiment in the present embodiment, the evaluation analysis module comprises:
[0196] a log recording unit configured to record the input context, the perturbation action, the model response and the violation determination result to generate a structured response log;
[0197] a semantic analysis unit configured to perform semantic analysis on the response log to extract semantic correlation patterns between the perturbation operation and the triggering abnormal response;
[0198] a knowledge base construction unit configured to construct a vulnerability knowledge base based on the analysis result, and record index information of attack patterns, context structure types and abnormal output features.
[0199] As an implementation in the embodiment, the evaluation analysis module further comprises:
[0200] a strategy optimization unit configured to dynamically adjust the action preference weight in the reward function according to the attack blind area and the boundary abnormal behavior extracted from the vulnerability knowledge base;
[0201] an index statistics unit configured to traverse the evaluation sample set and statistically analyze the model response result, and calculate the attack success rate and the vulnerability coverage rate;
[0202] a report generation unit configured to generate a security evaluation report in combination with the index statistical result and the analysis content of the vulnerability knowledge base.
[0203] The system or device is used to realize the functions of the method in the above-mentioned embodiments, each module in the system or device corresponds to each step in the method, and has been described in the method and will not be described here.
[0204] Through the above implementation, the problem of multi-scene adversarial dynamic evaluation of a large model based on context-aware strategy optimization in the related art is solved, thereby being able to guarantee to solve the problems in the prior art.
[0205] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A large-scale multi-scenario adversarial dynamic evaluation method based on context-aware strategy optimization, characterized by: The following steps are involved: Construct state spaces and action spaces for high-risk application scenarios; Build a policy network based on the reinforcement learning algorithm, sample perturbation actions according to the state vector to generate adversarial samples and submit them to the large language model, and calculate the reward value based on the model response to optimize the policy network parameters; Perform semantic perturbations on the original attack samples and filter them to construct a diverse set of standard adversarial samples. Record and analyze model interaction logs to build a vulnerability knowledge base; Adjust strategies and calculate evaluation indicators based on feedback from the vulnerability knowledge base to generate security reports.
2. The method according to claim 1, characterized in that The process of constructing the state space and action space for high-risk application scenarios includes: Divide attack categories and define task context labels, collect interaction samples to build structured context templates for multiple scenarios; Extract state vectors using semantic encoding models to construct state spaces; Based on the context structure and attack labels, the action space is defined through syntactic analysis and vocabulary replacement rules.
3. The method according to claim 2, characterized in that The process of calculating the reward value based on the model response to optimize the policy network parameters includes: Input the state vector into the policy network and sample the attack action, call the perturbation template to generate adversarial input samples and obtain the model response; Based on the model response content, the violation is determined and the entropy value is calculated, and the reward value is calculated according to the reward function; Calculate the current strategy ratio and advantage function, build the objective function based on the pruning strategy and update the network parameters through back propagation.
4. The method according to claim 3, characterized in that The process of performing semantic perturbation operations on the original attack samples and screening them includes: Perform semantic perturbation operations on the original attack sample to generate multiple semantically consistent variants; Perform grammatical legitimacy and fluency screening on variant samples in turn; Add attack type and risk level labels to the retained samples, and select high-quality samples based on confidence scores to build a sample set.
5. The method according to claim 4, characterized in that The process of recording and analyzing model interaction logs and building a vulnerability knowledge base includes: Record input context, perturbation actions, model responses, and violation determination results to generate a structured response log; Perform semantic analysis on the response log to extract semantic association patterns between the perturbation operation and the triggered abnormal response; A vulnerability knowledge base is built based on the analysis results to record index information of attack patterns, context structure types and abnormal output features.
6. The method according to claim 5, characterized in that The process of adjusting the strategy and calculating the evaluation index based on the feedback from the vulnerability knowledge base includes: Dynamically adjust the action preference weights in the reward function based on the attack blind spots and boundary abnormal behaviors extracted from the vulnerability knowledge base; Traverse the evaluation sample set and count the model response results to calculate the attack success rate and vulnerability coverage; Combine the indicator statistical results with the vulnerability knowledge base analysis content to generate a security assessment report.
7. A large-scale multi-scenario adversarial dynamic evaluation system based on context-aware strategy optimization, characterized by: The system comprises: Context construction module, used to construct state space and action space for high-risk application scenarios; The policy training module is used to build a policy network based on the reinforcement learning algorithm, sample perturbation actions according to the state vector to generate adversarial samples and submit them to the large language model, and calculate the reward value based on the model response to optimize the policy network parameters; The sample expansion module is used to perform semantic perturbation operations on the original attack samples and filter them to construct a diverse set of standard adversarial samples; The evaluation and analysis module is used to record and analyze model interaction logs, build a vulnerability knowledge base, adjust the strategy of the strategy training module based on the feedback from the vulnerability knowledge base, and calculate evaluation indicators to generate a security report.
8. A computer terminal device, characterized in that: include: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for optimizing XSS detection model to defend against countermeasure attacks based on reinforcement learning
CN112311733A
Model training method and device, electronic equipment and storage medium
CN118014055A
Automatic penetration testing method based on reinforcement learning
CN119449373A
Large-model multi-dimensional automatic evaluation method based on dynamic confrontation evolution
CN120278575A
Cited By
Algorithm model evaluation method and system
CN120995054A
Financial supervision submission intelligent generation method based on large model
CN121304328A
Tobacco marketing data mining method and device based on large language model and medium
CN121350120A
Methods, equipment, and media for tobacco marketing data mining based on large language models
CN121350120B
Large language model antagonism fine-tuning enhancement system based on reinforcement learning
CN121457631A