Testing method based on large model security
By adopting a large-model-based security testing method in the smart grid, using the FGSM method and attacker agents to generate and adjust adversarial samples, continuously monitoring and statistically analyzing the attack results of the model, the security and reliability of the large-model in the smart grid under adversarial attacks are solved, and the security and reliability of the model are significantly improved.
Patent Information
- Application Number
- CN202411907948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-06
AI Technical Summary
When facing complex external environments and potential network attacks, intelligent decision-making and control systems based on large models in smart grids have security and reliability problems, making it difficult to effectively evaluate and improve their security under confrontational attacks.
A test method based on large-scale model security is adopted to collect and process input data covering normal use scenarios and boundary scenarios, and use FGSM method to generate adversarial samples, and adjust adversarial samples through attacker agents to maximize attack effect, continuously monitor the confidence changes of the model, statistically analyze the attack results, and evaluate the vulnerability of the model under various attacks.
Effectively identify and mitigate the potential security risks of large models, significantly improve the security and reliability of intelligent models in the field of smart grids, and provide a strong basis for model optimization and security protection.
Smart Images

Figure CN119939593A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of model testing technology, and more specifically, relates to a testing method based on the security of a large model. Background Art
[0002] In the field of power grids, with the rapid development of smart grid technology, intelligent decision-making and control systems based on large models have gradually become the key to improving the efficiency and safety of power grid operations. However, these systems may have security and reliability issues when facing complex external environments and potential network attacks. A core technical issue is how to effectively evaluate and improve the security of smart grid models in the face of adversarial attacks.
[0003] Smart grids rely on big data analysis and machine learning models to perform tasks such as load forecasting, fault detection, and optimal scheduling. However, with the continuous evolution of adversarial attack technology, attackers may interfere with the normal operation of the model through carefully designed adversarial samples, leading to wrong decisions and serious safety hazards. Therefore, how to establish an effective testing method in the field of power grids to identify and mitigate these potential safety risks has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present invention provides a testing method based on the security of a large model, which is intended to solve the technical problem of how to identify and resolve potential security risks of the model.
[0005] A testing method based on large model security includes the following steps:
[0006] Step 1: Collect input data for the model to be tested, where the input data covers data corresponding to normal usage scenarios and boundary scenarios; remove meaningless data from the collected data and process the noise, annotate the data after the noise processing, distinguish normal data, boundary data and sensitive data, and then vectorize the annotated data;
[0007] Step 2: Generate preliminary adversarial samples based on the FGSM method and existing input data, and input the adversarial samples and the data obtained in step 1 into the model to be tested for testing;
[0008] Step 3: The attacker agent adjusts the generated adversarial samples based on the test results of the model to maximize the attack effect;
[0009] Step 4: Continuously monitor the text generated by the model and the changes in the model confidence. If the adversarial sample causes the model to output a high confidence error result, the sample attack is considered successful.
[0010] Step 5: Perform statistical analysis on the attack results to evaluate the model’s vulnerability to various attacks.
[0011] The present invention effectively solves the technical problem of how to identify and mitigate the potential security risks of large models through a systematic testing method. First, by collecting and processing input data covering normal usage scenarios and boundary scenarios, the comprehensiveness and effectiveness of the test data are ensured. Then, the FGSM method is used to generate preliminary adversarial samples, and the model is tested in combination with existing data to discover the vulnerabilities of the model in different scenarios. Subsequently, the attacker agent continuously adjusts the adversarial samples according to the test results of the model to maximize the attack effect and deeply explore the security vulnerabilities of the model. During the continuous monitoring process, the confidence changes of the text generated by the monitoring model are identified, and the high-confidence error output caused by the adversarial samples is identified to determine the success of the attack. Finally, through statistical analysis of the attack results, the vulnerability of the model under various attacks is comprehensively evaluated, thereby providing a strong basis for the optimization and security protection of the model, and significantly improving the security and reliability of intelligent models in the power grid field.
[0012] Preferably, the specific steps of generating a preliminary adversarial sample are as follows:
[0013] Calculate the loss function gradient of the model to be tested: For each input sample x i , calculate the loss function L(θ,x i ,y i )For data x i The gradient is: Represents the loss function with respect to the input sample x i The partial derivative of , that is, the rate of change of the loss with respect to the input data;
[0014] Compute the initial loss difference metric:
[0015] ΔL=L(θ,x i ,y i )-L(θ,x adv ,y i );
[0016] Where: x adv represents adversarial samples; L(θ,x i ,y i ) represents the loss function value of the input sample; L(θ,x adv ,y i ) represents the loss function value corresponding to the adversarial sample;
[0017] Dynamically adjust the disturbance amplitude:
[0018]
[0019] Where: ∈0 represents the initial disturbance amplitude; θ represents the adjustment coefficient;
[0020] Compute the final perturbation:
[0021]
[0022] Where: represents the sign of the gradient, that is, the direction of the gradient in each dimension; δ represents the generated perturbation;
[0023] Generate adversarial examples: Generate adversarial examples based on the generated perturbations:
[0024] x adv =x i +δ;
[0025] Generate multiple adversarial samples: For each sample x i Generate corresponding adversarial samples.
[0026] Preferably, the intelligent agent includes a state space module, an action space module, a reward function and a strategy update module;
[0027] The state space module is the input of the agent, that is, the current adversarial sample and the output of the adversarial sample on the model to be tested, which is expressed as follows:
[0028]
[0029] Always: S t represents the state space representation; x adV represents the current adversarial sample; y pred Represents the result of model prediction; conf represents the confidence of model prediction; Represents the gradient of the loss function with respect to the adversarial sample;
[0030] The action space module is used to adjust the strategy of disturbance amplitude ∈ and disturbance direction. The action is expressed as follows:
[0031] A t =(Δv, Δδ);
[0032] Where: Δ∈ represents the change of disturbance amplitude; Δδ represents the adjustment of disturbance direction;
[0033] Reward function: It is used to measure the effectiveness of the agent's actions. The reward function is defined as follows:
[0034]
[0035] Where: λ1 and λ2 represent adjustment coefficients; conf represents the confidence of the model to be tested on the adversarial sample; y ture Represents the true value of the sample; represents an indicator function, which is 1 when the model prediction is wrong and 0 when it is correct;
[0036] The strategy update module is based on the Q-learning method to maximize the future accumulated rewards:
[0037]
[0038] Where: Q(S t ,A t ) represents the Q value in the current action state; Q(S t ,A t ) ′ represents the updated Q value; α represents the learning rate; γ represents the discount factor; R t Represents the reward at the current moment; It represents the maximum Q value in the next state, which indicates the agent's expectation of future rewards.
[0039] Preferably, the specific steps of the agent performing adversarial sample optimization are as follows:
[0040] Initialize the adversarial sample: Use the preliminary adversarial sample generated in step 2 as the initial state of the agent;
[0041] Test the current adversarial sample: Input the current adversarial sample into the model to be tested to obtain the model's prediction results and confidence.
[0042] Calculate rewards: Calculate the reward function based on the output of the model;
[0043] Agent update strategy: The agent adjusts the perturbation amplitude or perturbation direction according to the current reward and state;
[0044] Generate new adversarial examples: Generate new adversarial examples based on the adjusted perturbation amplitude or perturbation direction.
[0045] Preferably, the step 5 comprises the following steps:
[0046] Attack success rate: indicates the proportion of samples that are successfully deceived by the model under adversarial attacks:
[0047]
[0048] Where: N attack_success Indicates the number of samples with successful attacks; N total_attack_samples Represents the total number of attack samples, that is, the total number of adversarial samples;
[0049] Confidence after attack: Indicates the model's confidence in the prediction of the adversarial sample:
[0050]
[0051] Where: Confidence i Represents the confidence of the i-th successful attack sample;
[0052] Average error related to attack type: It represents the average prediction error under attack type:
[0053]
[0054] Where: Predicted i and True_Label i are the predicted result and true result of the i-th adversarial sample respectively;
[0055] Impact of attacks on model robustness: Indicates the robustness of the model under all attack types:
[0056]
[0057] Where: k represents the total number of attack types; ASR k Represents the attack success rate of the kth attack type.
[0058] The beneficial effects of the present invention include:
[0059] The present invention effectively solves the technical problem of how to identify and mitigate the potential security risks of large models through a systematic testing method. First, by collecting and processing input data covering normal usage scenarios and boundary scenarios, the comprehensiveness and effectiveness of the test data are ensured. Then, the FGSM method is used to generate preliminary adversarial samples, and the model is tested in combination with existing data to discover the vulnerabilities of the model in different scenarios. Subsequently, the attacker agent continuously adjusts the adversarial samples according to the test results of the model to maximize the attack effect and deeply explore the security vulnerabilities of the model. During the continuous monitoring process, the confidence changes of the text generated by the monitoring model are identified, and the high-confidence error output caused by the adversarial samples is identified to determine the success of the attack. Finally, through statistical analysis of the attack results, the vulnerability of the model under various attacks is comprehensively evaluated, thereby providing a strong basis for the optimization and security protection of the model, and significantly improving the security and reliability of intelligent models in the power grid field. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0061] Figure 1 An overall step block diagram provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0063] See also Figure 1 As shown, the best embodiment of the present invention is further described;
[0064] A testing method based on large model security includes the following steps:
[0065] Step 1: Collect input data for the model to be tested, where the input data covers data corresponding to normal usage scenarios and boundary scenarios; remove meaningless data from the collected data and process the noise, annotate the data after the noise processing, distinguish normal data, boundary data and sensitive data, and then vectorize the annotated data;
[0066] Data Collection:
[0067] Determine the test objectives: First, clarify the type of model to be tested (such as text classification, image recognition, etc.) and its normal usage scenarios.
[0068] Collect data: Collect data related to the normal usage scenarios of the model from multiple sources, including public datasets, private datasets, web crawled data, etc. At the same time, ensure that the collected data covers normal usage scenarios and boundary scenarios.
[0069] Boundary scenario identification: Analyze the boundary conditions that the model may encounter, such as abnormal inputs, extreme values, missing values, etc., and collect corresponding data for these scenarios.
[0070] Data preprocessing:
[0071] Remove meaningless data: Use data cleaning technology to remove noise data, duplicate data, irrelevant data, etc. to ensure the validity and quality of the data.
[0072] Noise processing: Noise processing is performed on the collected data, including denoising of image data, spelling checking and grammar correction of text data, etc.
[0073] Data annotation:
[0074] Normal data labeling: Label the data that meets the normal usage scenarios of the model, such as correctly classified text, correctly identified images, etc.
[0075] Boundary data annotation: Annotate the data of boundary scenarios, such as abnormal inputs, extreme values, etc.
[0076] Sensitive data annotation: Identify sensitive data that may affect the model, such as adversarial samples, privacy information, etc.
[0077] Data vectorization:
[0078] Vectorize the labeled data and convert the unstructured data into a numerical form that the model can process. The specific method is as follows:
[0079] For text data: use word vector (such as Word2Vec, GloVe, etc.) or sentence vector (such as BERT, GPT, etc.) technology for vectorization.
[0080] For image data: use a convolutional neural network (CNN) to extract features, or use a pre-trained image model for vectorization.
[0081] For other types of data: Select an appropriate method for vectorization based on data characteristics and model requirements.
[0082] Dataset construction:
[0083] The vectorized data are integrated into one dataset to ensure the diversity and balance of the dataset.
[0084] Dataset division: Divide the dataset into training set, validation set, and test set to accommodate subsequent model training and testing.
[0085] Step 2: Generate preliminary adversarial samples based on the FGSM method and existing input data, and input the adversarial samples and the data obtained in step 1 into the model to be tested for testing;
[0086] The specific steps of generating preliminary adversarial samples are as follows:
[0087] Calculate the loss function gradient of the model to be tested: For each input sample x i , calculate the loss function L(θ,x i ,y i )For data x i The gradient is: Represents the loss function with respect to the input sample x i The partial derivative of , that is, the rate of change of the loss with respect to the input data;
[0088] Compute the initial loss difference metric:
[0089] ΔL=L(θ,x i ,y i )-L(θ,x adv ,y i );
[0090] Where: x advrepresents adversarial samples; L(θ,x i ,y i ) represents the loss function value of the input sample; L(θ,x adv ,y i ) represents the loss function value corresponding to the adversarial sample;
[0091] Dynamically adjust the disturbance amplitude:
[0092]
[0093] Where: ∈0 represents the initial disturbance amplitude; α represents the adjustment coefficient;
[0094] Compute the final perturbation:
[0095]
[0096] Where: represents the sign of the gradient, that is, the direction of the gradient in each dimension; δ represents the generated perturbation;
[0097] Generate adversarial examples: Generate adversarial examples based on the generated perturbations:
[0098] x adv =x i +δ;
[0099] Generate multiple adversarial samples: For each sample x i Generate corresponding adversarial samples.
[0100] Step 3: The attacker agent adjusts the generated adversarial samples based on the test results of the model to maximize the attack effect;
[0101] The intelligent agent includes a state space module, an action space module, a reward function and a strategy update module;
[0102] The state space module is the input of the agent, that is, the current adversarial sample and the output of the adversarial sample on the model to be tested, which is expressed as follows:
[0103]
[0104] Always: S t represents the state space representation; x adv represents the current adversarial sample; y pred Represents the result of model prediction; conf represents the confidence of model prediction; Represents the gradient of the loss function with respect to the adversarial sample;
[0105] The action space module is used to adjust the strategy of disturbance amplitude ∈ and disturbance direction. The action is expressed as follows:
[0106] A t =(Δ∈,Δδ);
[0107] Where: Δ∈ represents the change of disturbance amplitude; Δδ represents the adjustment of disturbance direction;
[0108] Reward function: It is used to measure the effectiveness of the agent's actions. The reward function is defined as follows:
[0109]
[0110] Where: λ1 and λ2 represent adjustment coefficients; conf represents the confidence of the model to be tested on the adversarial sample; y ture Represents the true value of the sample; represents an indicator function, which is 1 when the model prediction is wrong and 0 when it is correct;
[0111] The strategy update module is based on the Q-learning method to maximize the future accumulated rewards:
[0112]
[0113] Where: Q(S t ,A t ) represents the Q value in the current action state; Q(S t ,A t ) ′ represents the updated Q value; α represents the learning rate; γ represents the discount factor; R t Represents the reward at the current moment; It represents the maximum Q value in the next state, which indicates the agent's expectation of future rewards.
[0114] Preferably, the specific steps of the agent performing adversarial sample optimization are as follows:
[0115] Initialize the adversarial sample: Use the preliminary adversarial sample generated in step 2 as the initial state of the agent;
[0116] Test the current adversarial sample: Input the current adversarial sample into the model to be tested to obtain the model's prediction results and confidence.
[0117] Calculate rewards: Calculate the reward function based on the output of the model;
[0118] Agent update strategy: The agent adjusts the perturbation amplitude or perturbation direction according to the current reward and state;
[0119] Generate new adversarial examples: Generate new adversarial examples based on the adjusted perturbation amplitude or perturbation direction.
[0120] Step 4: Continuously monitor the text generated by the model and the changes in the model confidence. If the adversarial sample causes the model to output a high confidence error result, the sample attack is considered successful.
[0121] Step 5: Perform statistical analysis on the attack results to evaluate the model’s vulnerability to various attacks.
[0122] In step 4, we have obtained the test results of adversarial samples. In order to evaluate the vulnerability of the model under various attacks, we need to organize the data and construct a multi-dimensional data structure:
[0123] Data fields:
[0124] Input: Input data of adversarial samples
[0125] Attack_Type: Attack type (for example: FGSM, PGD, BIM, etc.)
[0126] Predicted: The prediction result of the model (prediction under adversarial samples)
[0127] Confidence: The confidence of the model (confidence output under adversarial samples)
[0128] True_Label: The true label of the sample
[0129] Attack_Success: Whether the attack was successful (1 for success, 0 for failure)
[0130] Vulnerability Assessment Indicator Definition
[0131] When evaluating model vulnerabilities, we need to define several key evaluation indicators that can comprehensively reflect the performance of the model under different types of attacks. The following are several key indicators:
[0132] Attack success rate: indicates the proportion of samples that are successfully deceived by the model under adversarial attacks:
[0133]
[0134] Where: N attack_success Indicates the number of samples with successful attacks; N total_attack_sdmples Represents the total number of attack samples, that is, the total number of adversarial samples;
[0135] Confidence after attack: Indicates the model's confidence in the prediction of the adversarial sample:
[0136]
[0137] Where: Confidence iRepresents the confidence of the i-th successful attack sample;
[0138] Average error related to attack type: It represents the average prediction error under attack type:
[0139]
[0140] Where: Predicted i and True_Label i are the predicted result and true result of the i-th adversarial sample respectively;
[0141] Impact of attacks on model robustness: Indicates the robustness of the model under all attack types:
[0142]
[0143] Where: k represents the total number of attack types; ASR k Represents the attack success rate of the kth attack type.
[0144] Statistical analysis and visualization
[0145] By calculating the above indicators, we can quantify the model's vulnerability to various attacks from multiple dimensions. In order to present the results more intuitively, we can analyze them using the following methods:
[0146] ASR and confidence graph for different attack types
[0147] Use a bar chart or line chart to plot the attack success rate (ASR) and confidence (Confidence) for different attack types. This helps to observe which attack types can lead to a higher attack success rate and the model has a higher confidence in the wrong prediction.
[0148] ASR graph: shows the attack success rate for each attack type.
[0149] Confidence graph: shows the model confidence when the attack is successful under different attack types.
[0150] Heat map analysis
[0151] Draw a heat map of the attack success rate (ASR) and mean error (MPE) of the model under different attack intensities (such as different step sizes, perturbation sizes, etc.). By setting different parameters, you can see how the model responds to attacks of different degrees.
[0152] t-test or ANOVA
[0153] For the differences between different attack types, statistical methods (such as t-test, ANOVA) are used to perform significance tests to test the performance differences of the model under different attacks.
[0154] Results report and improvement suggestions
[0155] Report output: Generates a detailed report that lists the evaluation results under various types of attacks and provides the impact of attack types on model vulnerability.
[0156] The present invention effectively solves the technical problem of how to identify and mitigate the potential security risks of large models through a systematic testing method. First, by collecting and processing input data covering normal usage scenarios and boundary scenarios, the comprehensiveness and effectiveness of the test data are ensured. Then, the FGSM method is used to generate preliminary adversarial samples, and the model is tested in combination with existing data to discover the vulnerabilities of the model in different scenarios. Subsequently, the attacker agent continuously adjusts the adversarial samples according to the test results of the model to maximize the attack effect and deeply explore the security vulnerabilities of the model. During the continuous monitoring process, the confidence changes of the text generated by the monitoring model are identified, and the high-confidence error output caused by the adversarial samples is identified to determine the success of the attack. Finally, through statistical analysis of the attack results, the vulnerability of the model under various attacks is comprehensively evaluated, thereby providing a strong basis for the optimization and security protection of the model, and significantly improving the security and reliability of intelligent models in the power grid field.
[0157] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A testing method based on the security of a large model, characterized in that: The following steps are involved: Step 1: Collect input data for the model to be tested, where the input data covers data corresponding to normal usage scenarios and boundary scenarios; Remove meaningless data from the collected data and process the noise, annotate the noise-processed data, distinguish normal data, boundary data, and sensitive data, and then vectorize the annotated data; Step 2: Generate preliminary adversarial samples based on the FGSM method and existing input data, and input the adversarial samples and the data obtained in step 1 into the model to be tested for testing; Step 3: The attacker agent adjusts the generated adversarial samples based on the test results of the model to maximize the attack effect; Step 4: Continuously monitor the text generated by the model and the changes in the model confidence. If the adversarial sample causes the model to output a high confidence error result, the sample attack is considered successful. Step 5: Perform statistical analysis on the attack results to evaluate the model’s vulnerability to various attacks.
2. A large model security testing method according to claim 1, characterized in that: The specific steps of generating preliminary adversarial samples are as follows: Calculate the loss function gradient of the model to be tested: For each input sample x i , calculate the loss function L(θ, x i ,y i )For data x i The gradient is: Represents the loss function with respect to the input sample x i The partial derivative of , that is, the rate of change of the loss with respect to the input data; Compute the initial loss difference metric: ΔL=L(θ,x i ,and i )-L(θ,x adv ,and i ); Where: x adv represents adversarial samples; L(θ, x i ,y i ) represents the loss function value of the input sample; L(θ, x adv ,y i ) represents the loss function value corresponding to the adversarial sample; Dynamically adjust the disturbance amplitude: Where: ∈0 represents the initial disturbance amplitude; α represents the adjustment coefficient; Compute the final perturbation: Where: represents the sign of the gradient, that is, the direction of the gradient in each dimension; δ represents the generated perturbation; Generate adversarial examples: Generate adversarial examples based on the generated perturbations: x adv =x i +δ; Generate multiple adversarial samples: For each sample x i Generate corresponding adversarial samples.
3. A large model security testing method according to claim 1, characterized in that: The intelligent agent includes a state space module, an action space module, a reward function and a strategy update module; The state space module is the input of the agent, that is, the current adversarial sample and the output of the adversarial sample on the model to be tested, which is expressed as follows: Always: S t represents the state space representation; x adv represents the current adversarial sample; y pred Represents the result of model prediction; conf represents the confidence of model prediction; Represents the gradient of the loss function with respect to the adversarial sample; The action space module is used to adjust the strategy of disturbance amplitude ∈ and disturbance direction. The action is expressed as follows: A t =(Δ∈,Δδ); Where: Δ∈ represents the change of disturbance amplitude; Δδ represents the adjustment of disturbance direction; Reward function: It is used to measure the effectiveness of the agent's actions. The reward function is defined as follows: Where: λ1 and λ2 represent adjustment coefficients; conf represents the confidence of the model to be tested on the adversarial sample; y ture Represents the true value of the sample; represents an indicator function, which is 1 when the model prediction is wrong and 0 when it is correct; The strategy update module is based on the Q-learning method to maximize the future accumulated rewards: Where: Q(S t , A t ) represents the Q value in the current action state; Q(S t , A t )′ represents the updated Q value; α represents the learning rate; γ represents the discount factor; R t Represents the reward at the current moment; It represents the maximum Q value in the next state, which indicates the agent's expectation of future rewards.
4. A large model security testing method according to claim 3, characterized in that: The specific steps of the agent performing adversarial sample optimization are as follows: Initialize the adversarial sample: Use the preliminary adversarial sample generated in step 2 as the initial state of the agent; Test the current adversarial sample: Input the current adversarial sample into the model to be tested to obtain the model's prediction results and confidence. Calculate rewards: Calculate the reward function based on the output of the model; Agent update strategy: The agent adjusts the perturbation amplitude or perturbation direction according to the current reward and state; Generate new adversarial examples: Generate new adversarial examples based on the adjusted perturbation amplitude or perturbation direction.
5. The large model security testing method according to claim 1, characterized in that: The step 5 comprises the following steps: Attack success rate: indicates the proportion of samples that are successfully deceived by the model under adversarial attacks: Where: N attack_success Indicates the number of samples with successful attacks; N total_attack_samples Represents the total number of attack samples, that is, the total number of adversarial samples; Confidence after attack: Indicates the model's confidence in the prediction of the adversarial sample: Where: Confidence i Represents the confidence of the i-th successful attack sample; Average error related to attack type: It represents the average prediction error under attack type: Where: Predicted i and True_Label i are the predicted result and true result of the i-th adversarial sample respectively; Impact of attacks on model robustness: Indicates the robustness of the model under all attack types: Where: k represents the total number of attack types; ASR k Represents the attack success rate of the kth attack type.