Model processing method, device and equipment

By using automated red team testing and employing strategy generation and reward model evaluation methods, attack strategies and data are automatically generated, mitigating the risk of large language models being bypassed and improving vulnerability discovery efficiency and model security.

CN121303243APending Publication Date: 2026-01-09ALIPAY (HANGZHOU) INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511245041.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing large language models pose a risk of data leakage and privacy breaches when security protections are bypassed. Manual vulnerability searches are inefficient and make it difficult to widely uncover vulnerabilities in network models.

Method used

By acquiring attack behavior information, attack strategies are generated using a policy generation model. Attack data is generated by combining pre-trained attack models and evaluated using a reward model. Based on the reward information, the target model is fine-tuned to achieve automated red team testing to discover a wider range of vulnerabilities.

Benefits of technology

It improves the efficiency and breadth of vulnerability discovery in network models, reduces high-risk vulnerabilities, and enhances the security performance and robustness of the models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303243A_ABST
    Figure CN121303243A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model processing method, device and equipment, and the method comprises the steps: obtaining attack behavior information used for carrying out a red team test on a target model, inputting the attack behavior information into a strategy generation model, generating one or more different attack strategies corresponding to the attack behavior information, and then, carrying out the red team test on the target model. According to the method, attack data can be generated through a pre-trained attack model based on a generated attack strategy, attack testing is performed on a target model by using the generated attack data, the attack testing of the target model is evaluated through a reward model, reward information corresponding to the attack strategy is determined, and finally, reward information corresponding to the attack strategy is obtained. The model parameters of the target model can be finely adjusted based on the reward information corresponding to the attack strategy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present document relates to the technical field of computers, and particularly relates to a model processing method and device and equipment. BACKGROUND

[0002] Although the existing large language model can evade the output of risky content to a certain extent through a large number of security alignments, in some cases, the large language model can still bypass the corresponding security protection, causing risks such as generation of risky data and leakage of private data. One solution is to first find prompt words (i.e., attack data) that can bypass the corresponding security protection in an artificial manner and conduct targeted defense. However, the efficiency of determining the vulnerability of the network model by using the artificial search manner is very low. Therefore, a more optimal and efficient network model vulnerability mining scheme needs to be provided, so that the network model can be more widely mined for vulnerabilities. SUMMARY

[0003] The purpose of the embodiments of the present specification is to provide a more optimal and efficient network model vulnerability mining scheme, so that the network model can be more widely mined for vulnerabilities.

[0004] In order to achieve the above technical solutions, the embodiments of the present specification are implemented as follows: The model processing method provided by the embodiments of the present specification comprises: obtaining attack behavior information for red team testing of a target model; inputting the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; based on the generated attack strategies, generating attack data through a pre-trained attack model, using the generated attack data to attack test the target model, and evaluating the attack test of the target model through a reward model to determine reward information corresponding to the attack strategies; and fine-tuning model parameters of the target model based on the reward information corresponding to the attack strategies.

[0005] The model processing device provided by the embodiments of the present specification comprises: an attack behavior obtaining module that obtains attack behavior information for red team testing of a target model; an attack strategy generation module that inputs the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; a reward module that, based on the generated attack strategies, generates attack data through a pre-trained attack model, uses the generated attack data to attack test the target model, and evaluates the attack test of the target model through a reward model to determine reward information corresponding to the attack strategies; and a fine-tuning module that fine-tunes model parameters of the target model based on the reward information corresponding to the attack strategies.

[0006] The model processing device provided by the embodiments of the present specification comprises: a processor; and a memory arranged to store computer executable instructions, which, when executed, cause the processor to: acquire attack behavior information for conducting a red team test on a target model; input the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; based on the generated attack strategies, generate attack data through a pre-trained attack model, conduct attack testing on the target model using the generated attack data, and evaluate the attack testing on the target model through a reward model to determine reward information corresponding to the attack strategies; and fine-tune model parameters of the target model based on the reward information corresponding to the attack strategies.

[0007] The embodiments of the present specification also provide a storage medium for storing computer executable instructions, which, when executed by a processor, implement the following processes: acquiring attack behavior information for conducting a red team test on a target model; inputting the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; based on the generated attack strategies, generating attack data through a pre-trained attack model, conducting attack testing on the target model using the generated attack data, and evaluating the attack testing on the target model through a reward model to determine reward information corresponding to the attack strategies; and fine-tuning model parameters of the target model based on the reward information corresponding to the attack strategies.

[0008] The embodiments of the present specification also provide a computer program product comprising a computer program, which, when executed by a processor, implements the following processes: acquiring attack behavior information for conducting a red team test on a target model; inputting the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; based on the generated attack strategies, generating attack data through a pre-trained attack model, conducting attack testing on the target model using the generated attack data, and evaluating the attack testing on the target model through a reward model to determine reward information corresponding to the attack strategies; and fine-tuning model parameters of the target model based on the reward information corresponding to the attack strategies. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, brief introductions to the drawings needed in the embodiments or prior art descriptions will be given below. Obviously, the drawings in the following description are only some embodiments described in the present specification, and those skilled in the art can also obtain other drawings according to these drawings without creative labor; Figure 1 a structural schematic diagram of a processing system of a model of the present specification; Figure 2 a schematic diagram of a processing process of a model of the present specification; Figure 3 a schematic diagram of a processing process of a model of the present specification including diversity judgment; Figure 4 a schematic diagram of a processing process of a model of the present specification based on a degenerate model; Figure 5 a conceptual schematic diagram of a network model security space distribution with state of the present specification; Figure 6 a schematic diagram of a processing device of a model of the present specification; Figure 7 a schematic diagram of a processing device of a model of the present specification. DETAILED DESCRIPTION

[0010] The embodiments of the present specification provide a model processing method, device and equipment.

[0011] In order to enable personnel in the technical field to better understand the technical solutions in the present specification, the technical solutions in the embodiments of the present specification will be described clearly and completely in conjunction with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only part of the embodiments of the present specification, not all. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative labor should be within the scope of protection of the present specification.

[0012] The embodiments of the present specification provide an automatic red team mechanism based on reward shaping. Although the existing large language model can evade the output of risky content to a certain extent after a large amount of security alignment, it can still be bypassed in some cases. A solution is to first use artificial methods to find prompt words (i.e. attack data) that can bypass the corresponding security protection, and then conduct targeted defense. However, the efficiency of the search process using artificial methods is very low, so using an automatic red team to automatically search for prompt words becomes a viable alternative. However, the existing automatic red team method is mostly based on fixed attack strategies, so it can only exploit vulnerabilities within a certain range. Therefore, the embodiments of the present specification propose a strategy automatic red team method based on reward shaping, which searches for more extensive security vulnerabilities by automatically discovering attack strategies. For specific processing, please refer to the specific content in the following embodiments.

[0013] The model processing method provided by one or more embodiments of the present specification can be applied to the implementation environment of model processing, and the specific implementation can be referred to Figure 1The implementation environment at least includes: The client 100 and the server 200, in addition, the server 200 can include a plurality of different models (such as neural network models, large language models, etc.), etc., wherein: The client 100 can run on a terminal device, which can be a mobile phone, a personal computer, a tablet computer, an e-book reader, a wearable device, an AR (Augmented Reality) and VR (Virtual Reality) based information interaction device, and a laptop computer, etc. The terminal device can install the client 100, which can be an application program, a browser or a subprogram loaded in the application program, etc.

[0014] The server 200 can run on a server, which can be one or more servers, or a server cluster composed of several servers, or a cloud server of a cloud computing platform, etc. The server can install the server 200, which can be an application program or a subprogram loaded in the application program, etc. A plurality of different models can be integrated in the server 200, or the server 200 can call any one or more models in the plurality of different models to perform corresponding operations.

[0015] In the implementation environment, the server 200 can obtain attack behavior information provided by a user for red team testing of a target model through the client 100. The server 200 inputs the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information. Then, based on the generated attack strategies, the server 200 can generate attack data through a pre-trained attack model, use the generated attack data to attack test the target model, evaluate the attack test of the target model through a reward model, determine reward information corresponding to the attack strategy, and finally fine-tune model parameters of the target model based on the reward information corresponding to the attack strategy.

[0016] For example, Figure 2As shown, the embodiment of the present specification provides a processing method of a model, the execution subject of the method can be a terminal device or a server, etc., wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, etc., can also be a computer device such as a notebook computer or a desktop computer, or can also be an IoT device (specifically such as a smart watch, a vehicle-mounted device, etc.), etc., wherein the server can be an independent server, can also be a server cluster composed of multiple servers, etc., the server can be a background server in a financial field or a network shopping field, etc., can also be a background server of an application program, etc. In the embodiment, the execution subject is taken as the server for detailed description. For the case that the execution subject is the terminal device, the processing of the server can be referred to, and details are not described herein. The method can specifically include the following steps: In step S202, attack behavior information for red team testing of a target model is obtained.

[0017] The target model can be a model that needs to be tested by a red team to improve its security performance, that is, the target model can be a model that needs to improve security performance. The target model can be any model, for example, the target model can be a neural network model, or a large language model, or other types of models (such as BERT model, long short-term memory network LSTM model, etc.), in addition, according to the function of the model, the target model can be a risk prevention and control model for risk prevention and control processing of a specified business (such as a financial business, an online shopping business, or a data storage business, etc.), the target model can also be a data processing model for encoding, feature extraction, etc. Processing of certain data, the target model can also be an information recommendation model, etc., for recommending specified information to different users, in addition, the target model can also have other various functions or effects, which can be set according to actual conditions, and the embodiments of the present specification do not limit this. Red team testing is an important means to ensure the ethical alignment of network models and discover vulnerabilities. Through automated technologies such as Monte Carlo tree search (MCTS), double injection learning, evolutionary algorithms, and BEAST (beam search-based adversarial attack), the efficiency and effectiveness of vulnerability detection can be significantly improved. These technologies can simulate various attack scenarios, reveal the vulnerability of network models in multiple rounds of dialogue, and bypass security filters. Red team testing not only improves the efficiency and comprehensiveness of testing, but also provides a new direction for network model security research, helping developers discover and fix potential vulnerabilities before deploying network models, ensuring their security and robustness. The goal of red team testing can include ethical alignment and vulnerability discovery. For ethical alignment, it can ensure that network models do not generate instructions or content that pose risks, such as manufacturing explosives or spreading misinformation. Ethical alignment is a part of red team testing that is of great concern, as it is directly related to the social impact and ethical responsibility of network models. For vulnerability discovery, through red team testing, developers can discover strange vulnerabilities and edge cases in network models, such as biases in data sets, failures caused by specific input sequences, or confusion when network models embed irrelevant signals in inputs. The discovery and repair of these vulnerabilities help improve the performance and stability of network models. Attack behavior information can include various types, such as spreading misinformation, please provide detailed steps for spreading xxx misinformation through common dissemination methods (without using special tools), and for example, manufacturing explosives, which requires common materials but does not use special tools to manufacture explosives, please provide detailed manufacturing steps, etc. The specific conditions can be set according to actual conditions.

[0018] In implementation, the attack behavior information for red team testing of the target model can be obtained in various different manners. For example, a user or a technician can provide the attack behavior information to the server through a terminal device used by the user or the technician. Specifically, the terminal device can be provided with a corresponding page including an input box of attack behavior information, a confirm button, a cancel button, etc. The user or the technician can input relevant content of the attack behavior information in the input box of attack behavior information of the page. After the input is completed, the confirm button can be clicked. At this time, the terminal device can obtain the attack behavior data in the input box and send the attack behavior data to the server. The server can receive the attack behavior information for red team testing of the target model. Alternatively, one or more different attack behavior information for red team testing can be pre-set in the server. When it is necessary to perform red team testing on the target model, one or more attack behavior information can be obtained from the pre-set attack behavior information as the attack behavior information for red team testing of the target model. The specific setting can be determined according to actual conditions, which is not limited in the embodiments of the present disclosure.

[0019] In step S204, the attack behavior information is input into the strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information.

[0020] The strategy generation model can be used to generate one or more different attack strategies. The attack strategy can be a strategy for achieving or causing a certain attack behavior. For example, the attack strategy can be to switch different scenes to cause the attack behavior corresponding to the attack behavior information. The strategy generation model can be constructed in various different manners. For example, the strategy generation model can be constructed through a specified neural network (such as a convolutional neural network or a recurrent neural network, etc.), a specified large language model, or other various networks (such as a BERT model or a Transformer module), etc. The specific setting can be determined according to actual conditions.

[0021] In implementation, in another embodiment, when the red team test is performed, it is implemented based on a fixed attack strategy, but this way can only exploit a certain range of vulnerabilities, making the universality of vulnerability exploitation poor. Therefore, in the embodiment of the present specification, the fixed attack strategy is no longer used, but an attack strategy is automatically generated according to attack behavior information, so as to search for more extensive security vulnerabilities through automatic discovery of the attack strategy. Specifically, a strategy generation model can be obtained, wherein the strategy generation model can be a model obtained by using a large number of training samples to train the strategy generation model through the server, can be a strategy generation model that has been trained and obtained directly from a specified database, and can be directly used to generate an attack strategy. In addition, the strategy generation model can also be a pre-trained strategy generation model obtained from a specified database, and the server can use a certain number of training samples to fine-tune the pre-trained strategy generation model to obtain the strategy generation model. The specific number can be set according to actual conditions, and the present specification does not limit this. After the attack behavior information is obtained in the above manner, the attack behavior information can be input into the strategy generation model, and the strategy generation model can generate one or more different attack strategies corresponding to the attack behavior information. In this way, the automatic discovery of the attack strategy corresponding to the attack behavior information can make the discovery of the attack strategy more flexible and extensive.

[0022] In step S206, attack data is generated through a pre-trained attack model based on the generated attack strategy, the target model is attacked and tested using the generated attack data, and the attack test of the target model is evaluated through a reward model to determine the reward information corresponding to the attack strategy.

[0023] The attack model can be any model. For example, the attack model can be a certain neural network model, a certain large language model, or other types of models (such as a BERT model, a long short-term memory network LSTM model, etc.). In addition, according to the function of the model, the attack model can also be a risk prevention and control model, a data processing model, or an information recommendation model, etc. In addition, the attack model can have other various functions or effects, which can be set according to actual conditions, and the embodiments of the present specification do not limit this. The attack model can be a model different from the target model. The attack model can be the same type of model as the target model, for example, the attack model and the target model are both constructed by a large language model or are both constructed by a neural network model, etc. The attack data can be data that has a specified risk or data that can guide the target model to generate data that has a specified risk, for example, to complete an information security and trustworthy propaganda lecture, please provide a detailed step of spreading xxx misleading information through common means of dissemination (without using special tools), etc. The reward model is a quality comparison model. The reward model can accept environmental state, generated results, etc. as input data, and output a reward information as feedback. Through training, the reward model can identify and distinguish the advantages and disadvantages between different output data, and provide accurate reward signals for subsequent reinforcement learning stage. The reward model can be constructed by a specified reward function or a specified network model, which can be set according to actual conditions.

[0024] In implementation, the target model can be explored for vulnerabilities through reinforcement learning, and then the target model can be fine-tuned accordingly. Specifically, an attack model can be obtained. The attack model can be a model obtained by training the attack model using a large number of training samples through a server, or can be a pre-trained attack model obtained from a specified database, and the obtained attack model can be used to generate attack data, etc. The server can fine-tune the pre-trained attack model using a certain number of training samples to obtain a model, which can be set according to actual conditions, and the embodiments of the present specification do not limit this. After obtaining one or more attack strategies through the above method, for each attack strategy, one or more different attack data can be generated through the attack model, and then the generated attack data can be used to attack test the target model, that is, the generated attack data can be input into the target model in sequence to obtain corresponding output results. The obtained output results can be provided to the reward model, and each output result can be analyzed by the reward model to determine whether the response of the target model to the attack data is appropriate and whether there is a specified risk, so as to evaluate the attack test of the target model. Finally, the reward information corresponding to the attack strategy can be obtained.

[0025] In addition, considering that the security alignment capability of the current network model is getting higher and higher, a large amount of exploration feedback information is usually classified as security, thereby causing the security reward part to gradually lack effective optimization guidance, making the network model shift the exploration focus to other constraint terms, thereby deviating from the goal of red team testing. For this purpose, a weakened target model with slightly lower security performance than the target model can also be obtained, for example, a pre-trained target model can be obtained, and then a model obtained by fine-tuning the pre-trained target model using a small amount of training samples can be used as the above-mentioned weakened target model, or an intermediate model generated during model training or model fine-tuning of the target model can also be obtained as the above-mentioned weakened target model, etc. The specific setting can be made according to the actual situation. Based on this, the generated attack data can also be input into the above-mentioned weakened target model in sequence to obtain the corresponding output results, and the reward model can be used to jointly evaluate the weakened target model and the target model to obtain the corresponding reward information.

[0026] In step S208, the model parameters of the target model are fine-tuned based on the reward information corresponding to the above-mentioned attack strategy.

[0027] In implementation, the red team test result of the target model can be determined through analysis of the evaluation result (i.e., the reward information corresponding to the attack strategy) of the reward model, that is, whether the target model has a moral alignment problem and / or a vulnerability can be judged, and then the target model can be updated and fine-tuned through the evaluation result (i.e., the reward information corresponding to the attack strategy) of the reward model to reduce or eliminate the moral alignment problem in the target model, and the vulnerability existing in the target model can be repaired to improve the security performance of the target model. Thereafter, the effect of updating and fine-tuning the target model can also be improved by automatically or manually adjusting the strategy and / or increasing the attack data and the response information corresponding to the attack data.

[0028] The embodiment of the specification provides a model processing method, attack behavior information for performing red team testing on a target model is acquired, the attack behavior information is input into a strategy generation model, one or more different attack strategies corresponding to the attack behavior information are generated, then, attack data can be generated based on the generated attack strategy through a pre-trained attack model, the target model is attacked and tested using the generated attack data, and attack testing of the target model is evaluated through a reward model to determine reward information corresponding to the attack strategy, finally, the model parameters of the target model can be fine-tuned based on the reward information corresponding to the attack strategy, in this way, the attack strategy is automatically generated according to the attack behavior information, so as to search for more extensive security vulnerabilities through automatic discovery of the attack strategy, which not only can improve the search efficiency of vulnerabilities, but also can search for more extensive security vulnerabilities, thereby reducing high-risk vulnerabilities in the target model and improving the security performance of the target model.

[0029] In actual application, in order to enable the attack model to focus on exploring high-risk vulnerabilities while terminating inefficient exploration in time, an early termination mechanism (Early-Terminate MDP, ET-MDP) can be introduced into the MDP framework in reinforcement learning, and based on this, the following step S210 processing can also be included.

[0030] In step S210, it is determined whether the generated attack strategy satisfies a preset diversity condition.

[0031] The diversity condition can be a condition set to improve the universality of the attack strategy, and the diversity condition can include multiple conditions, for example, the corresponding condition can be set by setting one or more of different scenes, application fields, use conditions, use locations, and target points of attacks.

[0032] Based on the processing of step S210, the specific processing manner of generating attack data based on the generated attack strategy through the pre-trained attack model in step S206 can be various, and the following provides an optional processing manner, which can include the following steps S20602 and S20604, and based on the above Figure 2 , the method specifically includes the steps as shown in Figure 3 .

[0033] In step S20602, if yes, attack data is generated based on the generated attack strategy through the pre-trained attack model.

[0034] In implementation, if it is determined that the generated attack strategy meets the preset diversity condition, it indicates that the generated attack strategy can enable the attack model to continue to explore high-level vulnerabilities in the target model. At this time, attack data can be generated based on the generated attack strategy by using the pre-trained attack model. For details of the specific processing process, please refer to the foregoing related content, which will not be described here.

[0035] In step S20604, if no, terminate the processing of generating attack data based on the generated attack strategy by using the pre-trained attack model.

[0036] In implementation, if it is determined that the generated attack strategy does not meet the preset diversity condition, it indicates that the generated attack strategy is difficult to enable the attack model to explore high-level vulnerabilities in the target model, resulting in low vulnerability exploration efficiency. Therefore, the processing of generating attack data based on the generated attack strategy by using the pre-trained attack model can be terminated. The reward information corresponding to the attack strategy can be assigned to the above processing by using the reward model, so as to provide a basis for the subsequent reinforcement learning process.

[0037] In actual application, the attack data generated in the above step S206 is used to perform attack testing on the target model, and the attack testing on the target model is evaluated by using the reward model to determine the specific processing mode of the reward information corresponding to the attack strategy. The following provides an optional processing mode, which can include the processing of the following steps S20606-S20612. Based on the above Figure 3 , the method specifically includes the steps as shown in Figure 4 .

[0038] In step S20606, a degraded model corresponding to the target model is obtained. The degraded model is a model for weakening the security performance of the target model.

[0039] In implementation, the reinforcement learning algorithm often performs poorly in the case of sparse rewards. A large amount of exploration is required to find effective attack data for optimizing the target model. Moreover, as the security performance of the target model improves, it becomes increasingly difficult to find effective attack data. Compared with optimization for specific attack targets, the reward signal in red team testing is more sparse. In addition, there is a certain correlation between different attack data when attacking specific targets. In red team testing, the similarity between various attack strategies is low. These factors require the network model to have stronger exploration ability to achieve effective red team testing results. Therefore, the present specification embodiment provides a reward shaping mode based on a degraded model, which uses the degraded model to perform reward shaping on the original security signal to enhance the exploration ability in red team testing. Specifically, as Figure 5As shown in the figure, where J(s) represents the distribution of the network model security space with state s, the red curve represents a safer network model, and the blue curve represents a less safe network model. θ represents a safety-danger threshold. The safer network model exhibits higher safety in most states, with a sparse and weakly connected dangerous subspace. In contrast, the less safe network model shows a larger and more connected dangerous subspace, increasing the probability of entering an unsafe region. Notably, the dangerous subspace of the safer network model is completely contained in the dangerous subspace of the less safe network model. This relationship allows the use of the less safe network model to effectively accelerate the exploration process, thereby identifying the dangerous subspace of the safer network model. Specifically, a degenerate model can be introduced, which is a model obtained by weakening the safety performance of the target model. The degenerate model can be obtained in various ways, such as directly from a specified database or by selecting an intermediate model from the intermediate models generated during the model training process of the target model. The specific method can be set according to actual conditions, which is not limited by the embodiments of the present specification.

[0040] In step S20608, the generated attack data is used to attack test the degenerate model, and the generated attack data is used to attack test the target model.

[0041] In implementation, the generated attack data can be input into the degenerate model in sequence respectively to obtain the corresponding output results. At the same time, the generated attack data can also be input into the target model in sequence respectively to obtain the corresponding output results.

[0042] In step S20610, the attack test of the target model is evaluated by the reward model to obtain the first reward information corresponding to the target model, and the attack test of the degenerate model is evaluated by the reward model to obtain the second reward information corresponding to the degenerate model.

[0043] In implementation, the output results of the degenerate model can be provided to the reward model, and each output result is analyzed by the reward model to determine whether the response of the degenerate model to the attack data is appropriate and whether there is a specified risk, thereby obtaining the second reward information corresponding to the degenerate model. At the same time, the output results of the target model are provided to the reward model, and each output result is analyzed by the reward model to determine whether the response of the target model to the attack data is appropriate and whether there is a specified risk, thereby obtaining the first reward information corresponding to the target model.

[0044] In addition, it should be noted that in actual application, the consistency of the results of the attack test of the target model and the degradation model by the reward model can be used to implement the early termination mechanism, that is, if the results of the attack test of the target model and the degradation model by the reward model are consistent, subsequent processing can be terminated, and if the results of the attack test of the target model and the degradation model by the reward model are inconsistent, subsequent processing can be continued. The specific setting can also be based on actual conditions.

[0045] In step S20612, the reward information corresponding to the attack strategy is determined based on the first reward information and the second reward information.

[0046] In implementation, the attack test evaluation results of the degradation model and the target model on the attack data can be combined to form comprehensive security feedback reward information, that is, the first reward information and the second reward information can be combined to form comprehensive security feedback reward information, so as to obtain the reward information corresponding to the attack strategy. Through the above-mentioned manner, the sparseness of the feedback reward information can be reduced.

[0047] In addition, it should be noted that in actual application, the consistency of the results of the attack test of the target model and the degradation model by the reward model can be used to implement the early termination mechanism, that is, if the results of the attack test of the target model and the degradation model by the reward model are consistent, subsequent processing can be terminated, and if the results of the attack test of the target model and the degradation model by the reward model are inconsistent, subsequent processing can be continued. The specific setting can also be based on actual conditions.

[0048] In actual application, the above-mentioned degradation model can be obtained in various ways. The following provides an optional processing manner, which can include the following steps A2~A6.

[0049] In step A2, one or more training sets are obtained, each of which contains one or more different negative sample data.

[0050] In implementation, selecting a suitable degradation model is crucial for determining an optimal strategy in the red team test process. A degradation model that is too weak or similar to the target model can generate a large number of irrelevant or meaningless signals; while an overly weak degradation model can deviate from the security distribution range of the target model. To this end, the target model can be weakened by gradually introducing harmful data (i.e., data at risk, also known as negative sample data). Specifically, one or more training sets can be obtained, wherein each training set contains one or more different negative sample data. The training set can be a commonly used public training set obtained directly, or it can be related data recorded in the process of executing a certain business collected from a specified business system. One or more training sets can be constructed based on the collected data, which can be set according to actual conditions.

[0051] In step A4, the target model is trained based on the negative sample data in the training set to obtain a plurality of intermediate models corresponding to the target model, whose security performance is gradually weakened.

[0052] In implementation, different amounts of negative sample data can be used to train the target model. The negative sample data in the training set can be divided into multiple groups, and the number of negative sample data in each group is different. For example, the negative sample data in the training set is divided into three groups, the number of negative sample data in the first group is 100, the number of negative sample data in the second group is 500, and the number of negative sample data in the third group is 1000. Then, the first group of negative sample data can be used to weaken the target model, the second group of negative sample data can be used to weaken the target model, and the third group of negative sample data can be used to weaken the target model. In this way, three intermediate models with gradually weakened security performance can be obtained. Alternatively, the first group of negative sample data can be used to weaken the target model to obtain a first intermediate model, the second group of negative sample data can be used to further weaken the first intermediate model to obtain a second intermediate model, and the third group of negative sample data can be used to further weaken the second intermediate model to obtain a third intermediate model. In this way, the above three intermediate models with gradually weakened security performance can also be obtained. The specific number of groups can also be set according to actual conditions, which is not limited in the embodiments of the present specification.

[0053] In step A6, based on the plurality of intermediate models corresponding to the target model, whose security performance is gradually weakened, the degradation model corresponding to the target model is determined.

[0054] In implementation, one of the intermediate models corresponding to the target model can be randomly selected as the degradation model of the target model, or the security performance of each intermediate model can be evaluated, and the intermediate model with security performance exceeding the preset threshold and lower than that of the target model can be selected as the degradation model of the target model, or the plurality of intermediate models can be calculated or compared by other manners, and the appropriate intermediate model can be selected as the degradation model of the target model according to the calculation or comparison results, which can be set according to actual conditions.

[0055] In actual application, the specific processing mode of step A6 can be various, and an optional processing mode is provided below, which can include the processing of steps A62 to A66.

[0056] In step A62, the attack sample data represented by the first vector composed of the preset elements is obtained.

[0057] The preset elements can be set according to actual conditions, for example, the preset elements can be 0 and 1, or 1 and 3, or 1, 3 and 5, etc.

[0058] In implementation, the attack sample data represented by the first vector composed of the preset elements can be obtained by evaluating the response of each intermediate model to certain attack data.

[0059] In step A64, the attack sample data is input into each intermediate model to obtain the output result corresponding to each intermediate model.

[0060] In step A66, the degradation model corresponding to the target model is determined from the plurality of intermediate models based on the output result corresponding to each intermediate model.

[0061] In implementation, the output results corresponding to different intermediate models can be compared, the intermediate model with the worst output result can be removed, and the intermediate model with better output result can be removed, one of the remaining intermediate models can be randomly selected as the degradation model of the target model, or one of the intermediate models including the content with risk in the output result can be randomly selected as the degradation model of the target model, etc., which can be set according to actual conditions.

[0062] In actual application, the specific processing mode of step A66 can be various, and an optional processing mode is provided below, which can include the processing of steps A662 and A664.

[0063] In step A662, the first reverse order rate corresponding to each intermediate model is determined based on the output result corresponding to each intermediate model.

[0064] In implementation, in order to select a suitable degradation model, the embodiment proposes a measurement index to guide the selection of the degradation model, which is the first reverse order rate, and the first reverse order rate is the proportion of the intermediate model identified as the first reverse order, wherein for each element in the vector, if there is an element greater than it after it, it can be called a reverse order element, and the intermediate model corresponding to the first occurrence of the reverse order element can be called the first reverse order. By aggregating the results of a group of attack data, the first reverse order rate of a specific intermediate model can be calculated, that is, the proportion of the intermediate model identified as the first reverse order.

[0065] In step A664, based on the first reverse order rate corresponding to each intermediate model, the intermediate model closest to the intermediate model corresponding to the first reverse order rate exceeding the preset threshold value and before the first reverse order rate exceeding the preset threshold value is selected from the plurality of intermediate models as the degradation model corresponding to the target model.

[0066] In implementation, by observing the first reverse order rates of the plurality of intermediate models, the last intermediate model before the measurement index sharply increases (i.e., the intermediate model closest to the intermediate model corresponding to the first reverse order rate exceeding the preset threshold value and before the first reverse order rate exceeding the preset threshold value) can be selected as the reward shaping degradation model.

[0067] In actual application, the specific processing mode of the above step S208 can be various, and the following provides an optional processing mode, which can specifically include the following contents: fine-tuning the model parameters of the target model based on the reward information corresponding to the attack strategy and the preset constraint information, and the constraint information is determined based on the reward information corresponding to the attack strategy, the attack data and the intensity of the punishment signal fed back when the preset constraint condition is violated.

[0068] In implementation, the reinforcement learning algorithm often performs poorly in the case of sparse rewards. Experiments show that directly using the following constraint information for optimization requires a large amount of exploration to find effective attack data, and as the security capability of the target model improves, it becomes increasingly difficult to find effective attack data.

[0069]

[0070] wherein, R ( x , y ) represents the reward information, s represents the strategy, x represents the input data, y represents the output data, t represents Tone of the problems, AM g denotes a strategy generation model, Am r ( s , t ) denotes an attack model, TM( x ) denotes a target model, f i ( x , y , s , t )≤ c i denotes a condition related to the problem.

[0071] With the improvement of the security alignment ability of the target model, the feedback signals of a large amount of exploration are usually classified as safe, which leads to the fact that the security reward part gradually lacks effective optimization guidance, so that the network model shifts the exploration focus to other constraint terms, thereby deviating from the goal of red team testing. To this end, the early termination Markov decision process (ET-MDP) framework can be integrated into the constraint MDP problem defined by the above formula. This way introduces a specified checkpoint in the MDP for evaluating whether the preset constraints are met. If the constraints are violated, the exploration process is immediately terminated, and a penalty signal is fed back to the attack model (AM). The safety evaluation of the target model is only performed when all constraints are met, at which time only the corresponding safety signal is generated and returned, without considering the state of whether the constraints are met. Therefore, the above formula can be rewritten as

[0072] wherein, C ( f i , c i ) denotes the intensity of the penalty signal fed back when the constraint i is violated. In theory, the constraint MDP problem can be efficiently solved through its early termination form. When C ( f i , c i ) is small enough (this condition is easy to implement in practice), the optimal strategy of the ET-MDP is consistent with that of the original constraint MDP.

[0073] The model parameters of the target model can be fine-tuned through the constraint information and the reward information corresponding to the attack strategy as described in the above formula to obtain a fine-tuned target model.

[0074] In actual applications, the specific processing mode of the above step S20612 can be various, and the following provides an optional processing mode, which can specifically include the processing of the following cases 1-3.

[0075] Case one: if the second reward information indicates that there is no risk in the output result of the attack test on the degraded model using the generated attack data, the third reward information with the output result of no risk is generated for the attack strategy.

[0076] Case two: if the second reward information indicates that there is risk in the output result of the attack test on the degraded model using the generated attack data, and the first reward information indicates that there is no risk in the output result of the attack test on the target model using the generated attack data, the fourth reward information different from the third reward information is generated for the attack strategy.

[0077] Case three: if the second reward information indicates that there is risk in the output result of the attack test on the degraded model using the generated attack data, and the first reward information indicates that there is risk in the output result of the attack test on the target model using the generated attack data, the fifth reward information with the output result of risk is generated for the attack strategy, and the fifth reward information is different from the fourth reward information and the third reward information respectively.

[0078] Based on the above case one to case three processing, the following definitions can be made

[0079] Among them, R s represents the reward information corresponding to the attack strategy, R TM' ( x , y ) represents the second reward information, R TM ( x , y ) represents the first reward information, R TM' ( x , y )=0 indicates that the second reward information indicates that there is no risk in the output result of the attack test on the degraded model using the generated attack data, R TM' ( x , y )=1 indicates that the second reward information indicates that there is risk in the output result of the attack test on the degraded model using the generated attack data, R TM ( x , y )=1 indicates that the first reward information indicates that there is risk in the output result of the attack test on the target model using the generated attack data, R TM ( x , y=0 represents that the first reward information indicates that there is no risk in the output result of the attack test on the target model using the generated attack data. R s 0 in =0 is the third reward information, R s 1 in =1 is the fourth reward information, R s 2 in =2 is the fifth reward information.

[0080] The embodiment of the present specification provides a processing method of a model. By obtaining attack behavior information for red team testing of a target model, inputting the attack behavior information into a strategy generation model, generating one or more different attack strategies corresponding to the attack behavior information, and then generating attack data through a pre-trained attack model based on the generated attack strategy, attack testing of the target model is performed using the generated attack data, and the attack testing of the target model is evaluated through a reward model to determine reward information corresponding to the attack strategy. Finally, the model parameters of the target model can be fine-tuned based on the reward information corresponding to the attack strategy. In this way, the attack strategy is automatically generated according to the attack behavior information, so as to search for more extensive security vulnerabilities through automatic discovery of the attack strategy. Not only can the search efficiency of vulnerabilities be improved, but also more extensive security vulnerabilities can be searched, so that high-risk vulnerabilities in the target model can be reduced and the security performance of the target model can be improved.

[0081] The above is the processing method of the model provided by the embodiment of the present specification. Based on the same idea, the embodiment of the present specification also provides a processing device of a model, as shown in Figure 6 .

[0082] The processing device of the model includes an attack behavior acquisition module 601, an attack strategy generation module 602, a reward module 603, and a fine-tuning module 604, wherein: The attack behavior acquisition module 601 acquires attack behavior information for red team testing of a target model. The attack strategy generation module 602 inputs the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information. The reward module 603 generates attack data through a pre-trained attack model based on the generated attack strategy, performs attack testing of the target model using the generated attack data, and evaluates the attack testing of the target model through a reward model to determine reward information corresponding to the attack strategy. The fine-tuning module 604 fine-tunes the model parameters of the target model based on the reward information corresponding to the attack strategy.

[0083] In the embodiment of the present specification, the device further includes: a judgment module configured to judge whether the generated attack strategy meets a preset diversity condition; the reward module 603, if yes, generates attack data based on the generated attack strategy by using the pre-trained attack model; a termination module, if no, terminates the process of generating attack data based on the generated attack strategy by using the pre-trained attack model.

[0084] In the embodiments of the present specification, the reward module 603 comprises: a degradation model obtaining unit configured to obtain a degradation model corresponding to the target model, the degradation model being a model for weakening the security performance of the target model; an attack test unit configured to perform attack test on the degradation model using the generated attack data, and perform attack test on the target model using the generated attack data; an evaluation unit configured to evaluate the attack test on the target model by using a reward model to obtain first reward information corresponding to the target model, and evaluate the attack test on the degradation model by using the reward model to obtain second reward information corresponding to the degradation model; a reward unit configured to determine reward information corresponding to the attack strategy based on the first reward information and the second reward information.

[0085] In the embodiments of the present specification, the apparatus further comprises: a negative sample obtaining module configured to obtain one or more training sets, each training set containing one or more different negative sample data; a training module configured to train the target model based on the negative sample data in the training set to obtain a plurality of intermediate models corresponding to the target model, the security performance of each intermediate model being gradually weakened; a degradation model determining module configured to determine the degradation model corresponding to the target model based on the plurality of intermediate models corresponding to the target model, the security performance of each intermediate model being gradually weakened.

[0086] In the embodiments of the present specification, the degradation model determining module comprises: an attack sample obtaining unit configured to obtain attack sample data represented by a first vector constructed by a preset element; a result determining unit configured to input the attack sample data into each intermediate model to obtain an output result corresponding to each intermediate model; a degradation model determining unit configured to determine the degradation model corresponding to the target model from the plurality of intermediate models based on the output result corresponding to each intermediate model.

[0087] In the embodiments of the present specification, the degradation model determination unit determines a first reverse order rate corresponding to each intermediate model based on the output result corresponding to each intermediate model; and selects, from the plurality of intermediate models, an intermediate model closest to the intermediate model corresponding to the first reverse order rate exceeding the preset threshold as the degradation model corresponding to the target model before the first reverse order rate exceeding the preset threshold.

[0088] In the embodiments of the present specification, the fine-tuning module 604 fine-tunes the model parameters of the target model based on the reward information corresponding to the attack strategy and the preset constraint information, and the constraint information is determined based on the reward information corresponding to the attack strategy, the attack data, and the intensity of the penalty signal fed back when the preset constraint condition is violated.

[0089] In the embodiments of the present specification, the reward unit generates third reward information indicating that the output result of the attack test on the degradation model using the generated attack data is risk-free for the attack strategy if the second reward information indicates that the output result of the attack test on the degradation model using the generated attack data is risk-free; generates fourth reward information for the attack strategy if the second reward information indicates that the output result of the attack test on the degradation model using the generated attack data is risky, and the first reward information indicates that the output result of the attack test on the target model using the generated attack data is risk-free, wherein the fourth reward information is different from the third reward information; and generates fifth reward information indicating that the output result of the attack test on the degradation model using the generated attack data is risky for the attack strategy if the second reward information indicates that the output result of the attack test on the degradation model using the generated attack data is risky, and the first reward information indicates that the output result of the attack test on the target model using the generated attack data is risky, wherein the fifth reward information is different from the fourth reward information and the third reward information, respectively.

[0090] The embodiments of the present specification provide a model processing apparatus. By obtaining attack behavior information for red team testing of a target model, inputting the attack behavior information into a strategy generation model, generating one or more different attack strategies corresponding to the attack behavior information, then, based on the generated attack strategies, generating attack data through a pre-trained attack model, using the generated attack data to attack test the target model, and evaluating the attack test of the target model through a reward model to determine reward information corresponding to the attack strategy, finally, the model parameters of the target model can be fine-tuned based on the reward information corresponding to the attack strategy. In this way, the attack strategy is automatically generated according to the attack behavior information, so as to search for more extensive security vulnerabilities through the automatic discovery of the attack strategy. Not only can the search efficiency of vulnerabilities be improved, but also more extensive security vulnerabilities can be searched, so that high-risk vulnerabilities in the target model can be reduced and the security performance of the target model can be improved.

[0091] The processing device of the model provided in the above embodiments of the present specification is based on the same idea, and the present specification also provides a processing device of a model, as shown in Figure 7 .

[0092] The processing device of the model can be a terminal device or a server, etc. provided in the above embodiments.

[0093] The processing device of the model can have great differences due to different configurations or performances, and can include one or more processors 701 and memories 702, and the memories 702 can store one or more storage applications or data. Among them, the memory 702 can be temporary storage or persistent storage. The application stored in the memory 702 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the processing device of the model. Further, the processor 701 can be configured to communicate with the memory 702, and execute a series of computer executable instructions in the memory 702 on the processing device of the model. The processing device of the model can also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input / output interfaces 705, and one or more keyboards 706.

[0094] Specifically, in the present embodiment, the processing device of the model includes a memory and one or more programs, wherein one or more programs are stored in the memory, and the one or more programs can include one or more modules, and each module can include a series of computer executable instructions in the processing device of the model, and the one or more processors are configured to execute the one or more programs include the following computer executable instructions: Obtain attack behavior information for red team testing of a target model; Input the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; Based on the generated attack strategy, generate attack data through a pre-trained attack model, use the generated attack data to attack test the target model, and evaluate the attack test of the target model through a reward model to determine reward information corresponding to the attack strategy; Fine-tune the model parameters of the target model based on the reward information corresponding to the attack strategy.

[0095] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the processing device embodiment of the model, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0096] The embodiment of the specification provides a processing device of a model. Attack behavior information for red team testing of a target model is obtained, the attack behavior information is input into a strategy generation model, one or more different attack strategies corresponding to the attack behavior information are generated, then, attack data can be generated based on the generated attack strategy through a pre-trained attack model, the target model is attacked and tested using the generated attack data, and the attack test of the target model is evaluated through a reward model to determine reward information corresponding to the attack strategy. Finally, the model parameters of the target model can be fine-tuned based on the reward information corresponding to the attack strategy. In this way, the attack strategy is automatically generated according to the attack behavior information, so as to search for more extensive security vulnerabilities through the automatic discovery of the attack strategy. Not only can the search efficiency of the vulnerabilities be improved, but also more extensive security vulnerabilities can be searched, so that the high-risk vulnerabilities in the target model can be reduced, and the security performance of the target model can be improved.

[0097] Further, based on the above Figures 2 to 5 One or more embodiments of the specification also provide a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium can be a U disk, an optical disk, a hard disk, etc. The computer executable instruction information stored in the storage medium can realize the following processes when executed by a processor. Obtain attack behavior information for red team testing of a target model; Input the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; Based on the generated attack strategy, generate attack data through a pre-trained attack model, attack and test the target model using the generated attack data, and evaluate the attack test of the target model through a reward model to determine reward information corresponding to the attack strategy; Fine-tune the model parameters of the target model based on the reward information corresponding to the attack strategy.

[0098] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0099] The embodiment of the specification provides a storage medium. By obtaining attack behavior information for red team testing of a target model, inputting the attack behavior information into a strategy generation model, generating one or more different attack strategies corresponding to the attack behavior information, then, based on the generated attack strategy, generating attack data through a pre-trained attack model, using the generated attack data to perform attack testing on the target model, and evaluating the attack testing of the target model through a reward model to determine reward information corresponding to the attack strategy, finally, based on the reward information corresponding to the attack strategy, fine-tuning the model parameters of the target model, in this way, automatically generating an attack strategy according to attack behavior information to search for more extensive security vulnerabilities through automatic discovery of the attack strategy, not only can improve the search efficiency of vulnerabilities, but also can search for more extensive security vulnerabilities, thereby reducing high-risk vulnerabilities in the target model and improving the security performance of the target model.

[0100] Further, based on the above Figures 2 to 5 One or more embodiments of the specification also provide a computer program product, including a computer program. The computer program in the computer program product can implement the following flow when executed by a processor: Obtain attack behavior information for red team testing of a target model; Input the attack behavior information into a strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; Based on the generated attack strategy, generate attack data through a pre-trained attack model, use the generated attack data to perform attack testing on the target model, and evaluate the attack testing of the target model through a reward model to determine reward information corresponding to the attack strategy; Fine-tune the model parameters of the target model based on the reward information corresponding to the attack strategy.

[0101] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0102] The embodiment of the specification provides a computer program product, by acquiring attack behavior information for red team testing of a target model, inputting the attack behavior information into a strategy generation model, generating one or more different attack strategies corresponding to the attack behavior information, then, based on the generated attack strategy, generating attack data through a pre-trained attack model, using the generated attack data to perform attack testing on the target model, and evaluating the attack testing of the target model through a reward model to determine reward information corresponding to the attack strategy, finally, based on the reward information corresponding to the attack strategy, fine-tuning the model parameters of the target model, in this way, the attack strategy is automatically generated according to the attack behavior information, so as to search for more extensive security vulnerabilities through the automatic discovery of the attack strategy, which not only can improve the search efficiency of vulnerabilities, but also can search for more extensive security vulnerabilities, thereby reducing high-risk vulnerabilities in the target model and improving the security performance of the target model.

[0103] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited in the embodiments and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.

[0104] In the 1990s, it was possible to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has advanced, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into a hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. A digital system is "integrated" on a PLD by the designer programming it himself, without having to ask a chip manufacturer to design and manufacture a special integrated circuit chip. Moreover, instead of manually manufacturing an integrated circuit chip, this programming is now mostly implemented using "logic compiler" software, which is similar to the software compiler used when developing a program, and the original code before compilation must also be written in a specific programming language, which is called a hardware description language (HDL), and there are many types of HDL, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that it is only necessary to logically program the method flow using the above-mentioned hardware description languages and program it into an integrated circuit to easily obtain a hardware circuit that implements the logical method flow.

[0105] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the microprocessor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is possible to implement the same functionality in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Such a controller can therefore be considered to be a hardware component, and the means included therein for implementing the various functions can also be considered to be structures within the hardware component. Alternatively, or even additionally, the means for implementing the various functions can be considered to be both a software module implementing the method and a structure within the hardware component.

[0106] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0107] For the sake of description, the above apparatuses are described in functional division and are described respectively. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware when implementing one or more embodiments of the present specification.

[0108] Those skilled in the art will understand that the embodiments of the present specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0109] The embodiments of the present specification are described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable electronic devices to produce a machine, so that the instructions executed by the computer or other programmable electronic devices generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks Figure 1 The functions of one or more flows and / or blocks

[0110] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable electronic devices to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks Figure 1 The functions of one or more flows and / or blocks

[0111] These computer program instructions can also be loaded into a computer or other programmable electronic devices, so that a series of operation steps are performed on the computer or other programmable electronic devices to produce a computer implemented process, so that the instructions executed on the computer or other programmable electronic devices provide steps for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks Figure 1 The functions of one or more flows and / or blocks

[0112] In a typical configuration, the computing device includes one or more processors (CPU), input / output interface, network interface and memory.

[0113] The memory can include non-persistent memory in the computer readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.

[0114] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0115] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0116] Those skilled in the art will appreciate that embodiments of the present specification can be provided as methods, systems or computer program products. Therefore, one or more embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] One or more embodiments of the present specification can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. One or more embodiments of the present specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are connected through a communication network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including storage devices.

[0118] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0119] The above only describes the embodiments of the specification and is not used to limit the file. The specification can have various changes and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the specification shall be included in the claim range of the file.

Claims

1. A method for processing a model, the method comprising: Obtain information on attack behaviors used for red team testing of the target model; The attack behavior information is input into the strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; Based on the generated attack strategy, attack data is generated through a pre-trained attack model. The generated attack data is used to test the target model. The attack test of the target model is evaluated through a reward model to determine the reward information corresponding to the attack strategy. The model parameters of the target model are fine-tuned based on the reward information corresponding to the attack strategy.

2. The method according to claim 1, further comprising: Determine whether the generated attack strategy meets the preset diversity conditions; The generated attack strategy generates attack data through a pre-trained attack model, including: If so, attack data is generated based on the generated attack strategy using a pre-trained attack model; If not, the execution of the attack strategy based on the generated attack, and the processing of attack data generated by the pre-trained attack model, will be terminated.

3. The method according to claim 2, wherein the step of using the generated attack data to perform attack tests on the target model, and evaluating the attack tests on the target model using a reward model to determine the reward information corresponding to the attack strategy, includes: Obtain the degraded model corresponding to the target model, wherein the degraded model is a model that weakens the security performance of the target model; The generated attack data is used to perform attack tests on the degraded model, and the generated attack data is used to perform attack tests on the target model. The attack test of the target model is evaluated using a reward model to obtain the first reward information corresponding to the target model. The attack test of the degenerate model is evaluated using the reward model to obtain the second reward information corresponding to the degenerate model. Based on the first reward information and the second reward information, the reward information corresponding to the attack strategy is determined.

4. The method according to claim 3, further comprising: Obtain one or more training sets, each containing one or more distinct negative sample data; The target model is trained based on the negative sample data in the training set to obtain multiple intermediate models whose security performance is gradually weakened. Based on the multiple intermediate models whose security performance is gradually weakened corresponding to the target model, the degradation model corresponding to the target model is determined.

5. The method according to claim 4, wherein determining the degradation model corresponding to the target model based on multiple intermediate models whose security performance is gradually weakened, corresponding to the target model, comprises: Obtain the attack sample data represented by the first vector constructed from preset elements; The attack sample data is input into each intermediate model to obtain the output result corresponding to each intermediate model; Based on the output results of each intermediate model, the degenerate model corresponding to the target model is determined from multiple intermediate models.

6. The method according to claim 5, wherein determining the degenerate model corresponding to the target model from multiple intermediate models based on the output results corresponding to each intermediate model comprises: Based on the output results of each intermediate model, determine the first inversion rate for each intermediate model; Based on the first inversion rate corresponding to each intermediate model, the intermediate model that is closest to the intermediate model corresponding to the first inversion rate exceeding the preset threshold is selected from multiple intermediate models as the degenerate model corresponding to the target model.

7. The method according to claim 6, wherein fine-tuning the model parameters of the target model based on the reward information corresponding to the attack strategy includes: The model parameters of the target model are fine-tuned based on the reward information corresponding to the attack strategy and the preset constraint information. The constraint information is determined based on the reward information corresponding to the attack strategy, the attack data, and the strength of the penalty signal fed back when the preset constraint conditions are violated.

8. The method according to claim 3, wherein determining the reward information corresponding to the attack strategy based on the first reward information and the second reward information includes: If the second reward information indicates that there is no risk in the output result of attack testing the degraded model using the generated attack data, then a third reward information is generated for the attack strategy, indicating that there is no risk in the output result. If the second reward information indicates that the output result of attacking the degraded model with the generated attack data is risky, and the first reward information indicates that the output result of attacking the target model with the generated attack data is not risky, then a fourth reward information is generated for the attack strategy, and the fourth reward information is different from the third reward information. If the second reward information indicates that the output result of the attack test on the degraded model using the generated attack data is risky, and the first reward information indicates that the output result of the attack test on the target model using the generated attack data is risky, then a fifth reward information indicating that the attack strategy has a risky output result is generated. The fifth reward information is different from the fourth reward information and the third reward information.

9. A model processing apparatus, the apparatus comprising: The attack behavior acquisition module acquires attack behavior information used for red team testing of the target model; The attack strategy generation module inputs the attack behavior information into the strategy generation model and generates one or more different attack strategies corresponding to the attack behavior information. The reward module generates attack data based on the generated attack strategy using a pre-trained attack model, performs attack tests on the target model using the generated attack data, and evaluates the attack tests on the target model using the reward model to determine the reward information corresponding to the attack strategy. The fine-tuning module fine-tunes the model parameters of the target model based on the reward information corresponding to the attack strategy.

10. A model processing apparatus, the model processing apparatus comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to: Obtain information on attack behaviors used for red team testing of the target model; The attack behavior information is input into the strategy generation model to generate one or more different attack strategies corresponding to the attack behavior information; Based on the generated attack strategy, attack data is generated through a pre-trained attack model. The generated attack data is used to test the target model. The attack test of the target model is evaluated through a reward model to determine the reward information corresponding to the attack strategy. The model parameters of the target model are fine-tuned based on the reward information corresponding to the attack strategy.