Large model parameter protection method, device and equipment based on deep learning

By generating and deploying gradient signature perturbation vectors and mapping tables, and dynamically inserting perturbation protection for large model parameters, the problem of complex inference side channel attacks is solved, achieving effective protection and security improvement for large model parameters.

CN120654254APending Publication Date: 2025-09-16EVERSEC BEIJING TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510928109.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively prevent complex inference side-channel attacks, resulting in an increased risk of deep learning model parameter theft. Especially in the environment of large-scale cloud deployment and open API calls, model parameter theft, reverse analysis and attribution are difficult to determine.

Method used

A large model parameter protection method based on deep learning is adopted. By calling a secure random source to generate a gradient signature perturbation vector, the Markov state transfer mechanism is used to generate a gradient signature perturbation mapping table. The gradient signature perturbation gradient is superimposed during the training phase, and additional perturbations are generated when deployed to the inference environment to protect the model parameters.

Benefits of technology

It achieves effective protection of large model parameters. By integrating inference query behavior monitoring, risk scoring and additional disturbance injection, it accurately identifies high-risk queries and dynamically inserts targeted disturbances to disrupt malicious inference results and improve the security of model parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654254A_ABST
    Figure CN120654254A_ABST
Patent Text Reader

Abstract

The invention discloses a message communication method, device and equipment based on a virtualization platform, and the method comprises the steps: superposing gradient signature perturbation vectors in a gradient signature perturbation vector library to corresponding gradient channels batch by batch according to a gradient signature perturbation mapping table in a training stage of a large model, and obtaining a gradient signature perturbation gradient; perturbing the precision change of the gradient calculation model based on the gradient signature; the weight snapshots and the safety-related auxiliary data serve as large model parameters to be deployed to a reasoning environment; when it is determined that the external query request obtained based on the reasoning environment is a high-risk request, additional disturbance is added to reasoning output of the large model so as to protect parameters of the large model. Reasoning query behavior monitoring, risk scoring and additional disturbance injection are integrated, external API calling behaviors are analyzed in real time through multi-dimensional features, high-risk query of parameter stealing tendency is accurately recognized, directional disturbance is dynamically inserted, malicious reasoning results are effectively disturbed, and large model parameters are effectively protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data theft protection technology, and in particular to a large model parameter protection method, device and equipment based on deep learning. Background Art

[0002] With the widespread application of deep learning models in natural language processing and computer vision, Transformer-type models based on large-scale parameters are increasingly deployed in commercial and scientific research scenarios. At the same time, the intellectual property protection and parameter security issues of the models are becoming increasingly prominent. In the environment of large-scale cloud deployment and open API calls, the risks of model parameters being stolen, reverse analyzed, and difficult to determine attribution have increased significantly.

[0003] Existing technologies for protecting against deep learning model parameter theft primarily focus on access control, API throttling, and encrypted deployment security measures. Common protection methods include strict limits on the number of external interface calls, encrypted storage of model weights, and encrypted calculations during inference. While traditional methods can, to a certain extent, hinder malicious users from directly downloading or brute-forcing model parameters, they struggle to effectively prevent complex inference side-channel attacks, which utilize high-frequency API calls, cleverly designed input samples, and monitoring model output to gradually infer and restore core parameters or reconstruct the original model. Therefore, there is an urgent need for more targeted and intelligent protection methods against the risk of large model parameter theft to enhance model security. Summary of the Invention

[0004] The present invention provides a large model parameter protection method based on deep learning to achieve the protection of large model parameters

[0005] According to a first aspect of the present invention, a large model parameter protection method based on deep learning is provided, comprising: calling a secure random source and generating a gradient signature perturbation vector based on a private key seed, and saving the generated gradient signature perturbation vector to a gradient signature perturbation vector library;

[0006] A Markov state transfer mechanism is used to generate a gradient signature perturbation mapping table, wherein the gradient signature perturbation mapping table includes a mapping relationship between the gradient channel of the large model in each training batch and the gradient signature perturbation vector;

[0007] In the training phase of the large model, the gradient signature perturbation vectors in the gradient signature perturbation vector library are superimposed on the corresponding gradient channels in batches according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradient;

[0008] Calculating the accuracy change of the model based on the gradient signature perturbation gradient, and iteratively optimizing the large model according to the accuracy change to obtain a large model weight snapshot;

[0009] Using the gradient signature perturbation vector library and the gradient signature perturbation mapping table as security-related auxiliary data, and deploying the weight snapshot and the security-related auxiliary data as large model parameters to an inference environment;

[0010] When it is determined that the external query request obtained based on the reasoning environment is a high-risk request, additional disturbance is generated according to the security-related auxiliary data, and the additional disturbance is added to the reasoning output of the large model to protect the large model parameters.

[0011] According to another aspect of the present invention, a large model parameter protection device based on deep learning is provided, comprising: a gradient signature perturbation vector library construction module, configured to call a secure random source and generate a gradient signature perturbation vector based on a private key seed, and save the generated gradient signature perturbation vector to the gradient signature perturbation vector library;

[0012] A gradient signature perturbation mapping table generation module is used to generate a gradient signature perturbation mapping table using a Markov state transfer mechanism, wherein the gradient signature perturbation mapping table includes a mapping relationship between the gradient channel of the large model in each training batch and the gradient signature perturbation vector;

[0013] A gradient signature perturbation gradient acquisition module is used to, during the training phase of the large model, superimpose the gradient signature perturbation vectors in the gradient signature perturbation vector library onto the corresponding gradient channels in batches according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradient;

[0014] A large model weight snapshot acquisition module is used to calculate the accuracy change of the large model based on the gradient signature perturbation gradient, and iteratively optimize the large model according to the accuracy change to obtain a large model weight snapshot;

[0015] An inference environment deployment module, configured to use the gradient signature perturbation vector library and the gradient signature perturbation mapping table as security-related auxiliary data, and deploy the weight snapshot and the security-related auxiliary data as large model parameters to an inference environment;

[0016] An additional disturbance adding module is used to generate additional disturbances based on the security-related auxiliary data when it is determined that the external query request obtained based on the reasoning environment is a high-risk request, and add the additional disturbances to the reasoning output of the large model to protect the parameters of the large model.

[0017] According to another aspect of the present invention, an electronic device is provided, comprising:

[0018] at least one processor; and

[0019] a memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method described in any embodiment of the present invention when executed.

[0022] The technical solution of the embodiment of the present invention integrates reasoning query behavior monitoring, risk scoring and additional disturbance injection. It can analyze external API call behavior in real time through multi-dimensional features, accurately identify high-risk queries with parameter theft tendencies, and dynamically insert targeted disturbances to effectively disrupt malicious reasoning results, thereby effectively protecting large model parameters.

[0023] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0025] Figure 1 This is a flowchart of a large model parameter protection method based on deep learning provided according to the first embodiment of the present invention;

[0026] Figure 2 This is a flowchart of a large model parameter protection method based on deep learning provided according to the second embodiment of the present invention;

[0027] Figure 3 1 is a structural diagram of a message communication device based on a virtualization platform provided according to a third embodiment of the present invention;

[0028] Figure 4 It is a structural diagram of an electronic device provided by the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] Example 1

[0032] Figure 1 A flowchart of a large model parameter protection method based on deep learning is provided for the first embodiment of the present invention. This embodiment is applicable to the case where large model parameters are protected from theft. The method can be executed by a large model parameter protection device based on deep learning, which can be implemented in the form of hardware and / or software. Figure 1 As shown, the method includes:

[0033] Step S101: call a secure random source and generate a gradient signature perturbation vector based on a private key seed, and save the generated gradient signature perturbation vector to a gradient signature perturbation vector library.

[0034] Optionally, call a secure random source and generate a gradient signature perturbation vector based on the private key seed, including:

[0035] Call the hardware-level secure random source to generate an initial random seed, input the initial random seed and the private key seed into the hash function to obtain the signature base seed; initialize the pseudo-random number generator according to the signature base seed to obtain the basis vector, and perform Hadamard orthogonal transformation on the basis vector to generate the initial gradient signature perturbation vector; perform amplitude normalization on the initial gradient signature perturbation vector to obtain the gradient signature perturbation vector.

[0036] Specifically, in this embodiment, an initial random seed is generated by calling a hardware-level secure random source. The initial random seed is used to enhance the uniqueness and unpredictability of the gradient signature perturbation vector. The initial random seed and the private key seed are input into the hash function together to output the signature base seed. The signature base seed is used to initialize the pseudo-random number generator, and a basis vector of dimension d is output. The length of the basis vector is consistent with the gradient channel dimension of a large model, such as a Transformer-type model in a single batch training. The basis vector is subjected to a Hadamard orthogonal transformation to generate an initial gradient signature perturbation vector. The orthogonal transformation makes the perturbation vector orthogonal in the gradient space, which is beneficial for maintaining consistency and robustness of the signature distribution in multiple gradient directions. The initial gradient signature perturbation vector is amplitude normalized to control the second norm to be equal to the perturbation amplitude upper limit parameter, where the perturbation amplitude upper limit parameter is the maximum perturbation tolerance set according to the task loss constraint threshold. The normalized gradient signature perturbation vector g is then converted to sug Store in the gradient signature perturbation vector library GS lib , and assign a unique vector identifier and generation timestamp to the gradient signature perturbation vector.

[0037] It should be noted that the number of vectors in the gradient signature perturbation vector library in this embodiment is related to the number of batches and gradient channels of model training. Usually, each training batch and each gradient channel selected for perturbation corresponds to an independent gradient signature perturbation vector. Therefore, there are as many corresponding vectors as there are perturbation operations on the channel during training. In this embodiment, the number of vectors in the gradient signature perturbation vector library is not limited.

[0038] Step S102: Generate a gradient signature perturbation mapping table using a Markov state transfer mechanism.

[0039] Optionally, a Markov state transition mechanism is used to generate a gradient signature perturbation mapping table, including: establishing a state space set, wherein each state element in the state space set is used to indicate whether each gradient channel is mapped by a gradient signature perturbation vector during a single batch training of a large model; constructing a Markov state transition probability matrix, wherein each probability element in the Markov state transition probability matrix is ​​used to indicate the state transition probability of the gradient signature perturbation vector in the current training batch of the large model; generating a gradient signature perturbation mapping sub-table corresponding to each training batch of the large model according to the state space set and the Markov state transition probability matrix; and merging each gradient signature perturbation mapping sub-table in the order of the training batches to generate a gradient signature perturbation mapping table.

[0040] Specifically, in this embodiment, when generating the gradient signature perturbation mapping table, a state space set S is first established, and each state s in the state space set S iIndicates whether the i-th gradient channel in the single batch training process of the Transformer model is perturbed by the gradient signature vector g sig Mapping, state s i The value is 0 or 1, 0 means that the gradient channel does not map the gradient signature perturbation vector g sig , 1 means that the gradient channel has mapped the gradient signature perturbation vector g sig , the state space set S is used to represent the perturbation mapping state of all gradient channels. In addition, in this embodiment, a Markov state transition probability matrix P is constructed. The element p in the Markov state transition probability matrix P is ij Indicates that in the current training batch of the Transformer class model, the gradient signature perturbation vector g sig From the state i Transfer to state s j The sum of the transition probabilities in each row is 1, and the transition probability between any states is a real number between 0 and 1. The Markov state transition probability matrix P is used to control the dynamic transfer process of the gradient channel perturbation state between different batches.

[0041] It should be noted that the mapping in this embodiment is actually to establish a corresponding relationship between a batch, a channel and a perturbation vector. The specific perturbation vector in the gradient signature perturbation vector library by which the i-th gradient is perturbed is determined by the gradient signature perturbation mapping table. Each training batch will specify whether the i-th gradient channel is superimposed with perturbations based on the perturbation state vector. If necessary, the perturbation vector corresponding to the channel and batch in the gradient signature perturbation vector library will be selected for injection to achieve the dispersion and uniqueness of the perturbation. And at this stage, it is not known whether each gradient channel will be mapped to the gradient signature perturbation vector. State s i The value of 0 or 1 is only used to initialize and describe the mapping relationship. Whether each batch is actually mapped is dynamically determined by the Markov state transition mechanism and the perturbation state vector. The specific mapping situation is gradually generated and updated during the training process.

[0042] In addition, the Markov state transition probability matrix P includes the state transition probability of the gradient signature perturbation vector. The reason why the gradient signature perturbation vector needs to undergo state transition is to dynamically adjust the distribution of perturbations between gradient channels and training batches through the Markov mechanism, making the perturbation injection process random and unpredictable. This can effectively prevent attackers from restoring perturbation characteristics through regularity analysis, thereby improving the robustness and security of model parameter protection.

[0043] Among them, the state space set S in this embodiment is used to represent the state of whether each gradient channel in each training batch of the Transformer model is mapped and disturbed. Subsequently, the initial perturbation state vector can be iterated through the Markov state transition probability matrix P to obtain the perturbation state vector of each batch, and the mapping relationship between each channel and the perturbation vector can be judged accordingly. For example, set the initial perturbation state vector s (0) , the initial perturbation state vector s (0) Indicates whether all gradient channels in the initial training batch of the Transformer class model are mapped to the gradient signature perturbation vector g sig The initial perturbation state vector s is processed using the Markov state transition probability matrix P. (0) Perform iterative updates to obtain the perturbation state vector s during the t-th batch training process (t) , the perturbation state vector s (t) Indicates whether each gradient channel in the current batch is mapped to the gradient signature perturbation vector g sig , the perturbation state vector s (t) It is determined by the previous batch of perturbation state vectors and the Markov state transition probability matrix. Then, according to the perturbation state vector s (t) Determine whether each gradient channel is mapped to the gradient signature perturbation vector g during the tth batch training of the Transformer model sig , get the mapping relationship between the gradient channel and the perturbation vector under the current batch, organize the mapping relationship into a gradient signature perturbation mapping subtable, and the gradient signature perturbation mapping subtable records each gradient channel in the current batch and the state of whether the perturbation is mapped. And the gradient channel mapping perturbation means that during the model training process, according to the pre-generated perturbation state vector, it is determined whether each gradient channel should be superimposed with a specific gradient signature perturbation vector. If mapped, the gradient of the channel will be perturbated, thereby introducing specific identifiable features when the parameters are updated, and realizing the traceability and protection of the model. After obtaining the gradient signature perturbation mapping subtables corresponding to all training batches, the gradient signature perturbation mapping subtables generated by all training batches will be merged in batch order to form a complete gradient signature perturbation mapping table. Therefore, the complete gradient signature perturbation mapping table defines the gradient channel and gradient signature perturbation vector g in all training batches of the Transformer model during the entire training cycle. sig The mapping relationship between them.

[0044] Step S103: During the training phase of the large model, the gradient signature perturbation vectors in the gradient signature perturbation vector library are superimposed batch by batch on the corresponding gradient channels according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradients.

[0045] Optionally, during the training phase of the large model, the gradient signature perturbation vectors in the gradient signature perturbation vector library are superimposed batch by batch on the corresponding gradient channel according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradient, including: constructing a set of original gradient vectors in the training batch of the year during the training phase of the large model, wherein each element in the original gradient vector set represents the original gradient value of the gradient channel under the current training batch; extracting the gradient signature perturbation mapping sub-table corresponding to the current training batch from the gradient signature perturbation mapping table, and extracting the perturbation state vector of the current training batch from the gradient signature perturbation mapping sub-table; and linearly superimposing the original gradient value with the gradient signature perturbation vector under the corresponding gradient channel based on the perturbation state vector of the current training batch.

[0046] Specifically, in this implementation, the original gradient vector set g of the Transformer model in the current t-th training batch is constructed. (t) , the original gradient vector set g (t) Each element Represents the original gradient value of the i-th gradient channel in the current batch. The total number of gradient channels is d. Each gradient channel corresponds to a gradient tensor with a tensor dimension of n. i Extract the gradient signature perturbation mapping subtable corresponding to the t-th training batch from the gradient signature perturbation mapping table, and extract the perturbation state vector s of the current batch (t) , the perturbation state vector s (t) Each element in Indicates whether the i-th gradient channel is mapped to the gradient signature perturbation vector Then according to the perturbation state vector s (t) The gradient signature perturbation vector corresponding to each selected gradient channel Construct the gradient signature perturbation gradient vector set of the current batch Each element Represents the gradient value after superimposing the disturbance in the i-th channel, specifically the original gradient value Perturbation vector with gradient signature The linear superposition of the gradient signature perturbation gradient is obtained, and only when the perturbation state When performing superposition operation, the gradient signature perturbation vector The second norm of must satisfy the upper limit constraint of the disturbance amplitude, that is, the modulus is equal to the maximum disturbance tolerance set according to the task loss constraint threshold.

[0047] It should be noted that in this embodiment, the constructed gradient signature perturbs the gradient vector set Input optimizer, used to update the weight parameter θ corresponding to the current training batch of the Transformer model (t), the optimizer performs weight parameter update with the learning rate parameter η as the step size, and obtains the updated weight parameter θ (t+1) , it traverses all training batches in the entire training cycle, completes the gradient signature perturbation injection and weight parameter update of each batch according to the batch sequence, and obtains the Transformer-type large model weight snapshot embedded with the complete voxel signature perturbation. However, the large model weight snapshot at this time is not the final model parameter result, but the intermediate result that is continuously updated during the training process. It is mainly used to record the parameter state of each batch with superimposed perturbations. Therefore, during the training phase, the system will select the corresponding gradient signature perturbation vector according to the gradient channel of the current training batch based on the perturbation state vector and the perturbation mapping table, and linearly superimpose it with the original gradient value to form a gradient vector containing perturbations. These perturbed gradient vectors are then input into the optimizer to perform parameter update operations, thereby obtaining a new round of model parameters. The goal of this stage is to inject signatures with perturbation features into the model parameters, provide a traceable structure for subsequent identification and attribution verification, and control the impact on the original performance of the model.

[0048] Step S104: calculate the accuracy change of the gradient calculation model based on the gradient signature perturbation, and iteratively optimize the large model according to the accuracy change to obtain a large model weight snapshot.

[0049] Optionally, the accuracy change of the large model is calculated based on the gradient signature perturbation gradient, and the large model is iteratively optimized according to the accuracy change to obtain a large model weight snapshot, including: constructing a composite task loss function and setting a task loss constraint threshold; calculating the accuracy change of the large model according to the composite task loss function for the gradient signature perturbation gradient; when it is determined that the accuracy change exceeds the task loss constraint threshold, the large model is iteratively optimized to obtain a large model weight snapshot embedded with the gradient signature perturbation.

[0050] Specifically, in this implementation, a composite task loss function L is constructed. task (θ), where the composite task loss function is composed of the main task cross entropy loss function L main (θ), gradient signature preserving loss function L sig (θ) and the model stealability suppression loss function L ext (θ) is weighted, as shown in the following formula (1):

[0051] L task (θ)=L main (θ)+λ sig L sig (θ)+λ ext L ext (θ) (1)

[0052] Among them, λ sig is the weight parameter of the gradient signature holding term, λext is the weight parameter of the stealability suppression item. main (θ) is mainly used to calculate the cross entropy loss value of the main task, which is used to measure the prediction performance of the Transformer model for the main classification task under the current weight parameter θ. The main task cross entropy loss value is obtained by summing the logarithmic product between the true label and the predicted probability of all categories and taking the negative value. L sig (θ) is mainly used to calculate the gradient signature preservation loss value, which is used to measure the stability of the current gradient signature perturbation vector in the frequency domain features of the model output. The gradient signature preservation loss value is obtained by comparing the output spectrum of each channel in the set of gradient channel indexes that have been mapped and perturbed in the current batch. The model output vector before the perturbation and the model output vector after the perturbation are Fourier transformed respectively. After extracting the frequency domain expression, the spectrum difference of the corresponding channel is squared and accumulated. The smaller the spectrum difference, the stronger the perturbation consistency. The smaller the gradient signature preservation loss value, the more stable the signature feature embedding. L ext (θ) is used to calculate the model stealability suppression loss value, which is used to measure the Transformer-type model's ability to respond to local perturbations of the query input under the current weight parameter θ. The model stealability suppression loss value is obtained by traversing the input sample set in the current training batch, adding a small perturbation vector to each input sample, and calculating the residual between the model output and its first-order linear approximation. The residual is obtained by square the difference between the model's true output and the linear response result of the Jacobian matrix.

[0053] Specifically, in this implementation, a task loss constraint threshold is also defined. Dynamically decay with the training batch, where α>0 is the initial tolerance, β>0 is the decay coefficient, and L is the current batch number. After each training batch, the composite task loss function that introduces gradient signature perturbation is calculated. and the loss function of the unperturbed composite task Get the batch accuracy change value, which is specifically: If the batch accuracy change value Δ (t) Greater than the task loss constraint threshold The disturbance amplitude adaptive adjustment mechanism is triggered, and the current batch disturbance vector The upper limit parameter ε of the amplitude (t,i) Perform scaling ε (t,i) ←γ·ε (t,i) , where 0<γ<1 is the convergence coefficient, and the above-mentioned amplitude upper limit parameter ε (t,i)It is used to limit the maximum value of the intensity of each gradient signature perturbation vector to ensure that the perturbation does not affect the model accuracy too much. Then the weight parameters of the gradient signature preservation term and the stealability suppression term in the composite task loss function will be updated. The update method is to add a fixed forward step size δ to the current training batch. sig and δ ext , where the step size δ sig Used to enhance the influence of gradient signature preservation constraints in the loss function, step size δ ext It is used to enhance the influence of the stealability suppression constraint in the loss function. The updated weight parameters participate in the gradient calculation in the next round of perturbation amplitude adjustment and compound task loss function minimization. Therefore, in this implementation, after each training batch, the weight parameters of the gradient signature preservation term and the stealability suppression term are dynamically adjusted according to the actual performance of the model. By increasing the respective forward step sizes, the constraints on the gradient signature stability and the anti-parameter stealing ability are further strengthened, thereby optimizing the model protection effect in subsequent training. Then, the updated amplitude upper limit parameter ε is used. (t,i) And the weight parameters of the gradient signature preservation term and the stealability suppression term renormalize the gradient signature perturbation vector And based on the updated composite task loss function L task (θ) The linear superposition of the gradient value and the gradient signature perturbation vector and the parameter update are performed again. Through continuous iterative optimization process, until Or the number of iterations exceeds the preset upper limit, confirming that the current batch has met the dual constraints of accuracy and protection. And by traversing all subsequent training batches, executing the same dynamic threshold test and perturbation amplitude, adaptive weight adjustment process until the training process converges, and finally obtaining the Transformer-like large model weight snapshot θ embedded in the gradient signature perturbation under the composite task loss constraint (*) .

[0054] It should be noted that in this embodiment, the composite task loss function is mainly used to inject the gradient of the gradient perturbation into the large model, and then verify the actual impact of the perturbation on the overall performance of the model, and mainly the actual impact of the task accuracy and security indicators, and adjust the perturbation strategy based on the following three aspects: First, task accuracy monitoring: by calculating the composite task loss function before and after the introduction of the perturbation, it is determined whether the current batch perturbation is too strong and whether the main task performance of the model is degraded beyond an acceptable range; second, the perturbation amplitude adjustment: if the batch accuracy change value is greater than the batch constraint threshold, that is, the model performance deteriorates, then the adaptive contraction mechanism of the perturbation amplitude is triggered to reduce the injection intensity of the perturbation to ensure the stability of the model in the next round of updates. The third is the loss function weight update: the system will adjust the weight ratio of the "gradient signature preservation term" and the "model stealability suppression term" in the loss function according to the current batch training performance to achieve dual optimization of perturbation embedding stability and protection effect. Therefore, the role of the composite task loss function in this embodiment is to achieve a dynamic balance and optimal adjustment between model accuracy and security while ensuring model accuracy through the weighted combination of the main task loss, gradient signature preservation loss and model theft suppression loss, so as to stably embed the gradient signature features and improve the protection capability against parameter theft attacks.

[0055] In step S105 , the gradient signature perturbation vector library and the gradient signature perturbation mapping table are used as security-related auxiliary data, and the weight snapshot and the security-related auxiliary data are deployed as large model parameters to the inference environment.

[0056] Among them, through the above perturbation injection and training process, the large model has embedded a unique gradient signature perturbation and obtained the final weight snapshot. At this time, it is necessary to synchronously deploy these weights and security-related auxiliary data consisting of the gradient signature perturbation vector library and the gradient signature perturbation mapping table to the inference environment, and initialize the inference query monitoring unit to ensure that all requests and responses during the inference process can be tracked. When deploying the inference environment, specifically upload the final weight snapshot θ of the Transformer class large model with the complete gradient signature perturbation embedded (*) To the deployment server, the final weight snapshot θ of the Transformer class large model (*) Represents the final parameter set after the training process is completed, which includes the gradient signature perturbation vectors injected by all batches and all channels. The weight snapshot will run in the inference server as the core model carrier of the inference service. Construct an inference environment parameter synchronization unit, which is used to synchronize the security-related auxiliary data structures generated during the training phase to the deployment server. Initialize the inference query monitoring unit deployed on the inference server. The inference query monitoring unit is used to capture the query request sequence submitted by external users in real time after the large model is put into actual use, and based on the deployed gradient signature perturbation vector library GS lib, gradient signature perturbation map Map and batch constraint threshold Extract the sensitive behavioral features of the model in the inference response. Perturb the gradient signature vector library GS lib , Gradient signature perturbation mapping table Map and batch loss constraint threshold The inference query monitoring module is persistently saved as a set of security parameters to the inference server, and an encrypted index and digital signature are established to support the call of key links in request risk identification, dynamic intervention response and model attribution verification in the subsequent inference stage, ensuring that the deployed model has fully traceable parameter theft protection capabilities and behavior auditing capabilities.

[0057] It should be noted that the security-related auxiliary data in this embodiment mainly includes the gradient signature perturbation vector library GS lib , Gradient signature perturbation mapping table Map and batch loss constraint threshold Among them, the gradient signature perturbation vector library GS lib Contains the gradient signature perturbation vectors used in all training batches, each vector Corresponding to a specific training batch number t and gradient channel index i, the vector's two-norm is equal to the perturbation amplitude upper limit parameter ε (t,i) The perturbation amplitude upper limit parameter is used to limit the perturbation intensity to ensure that the model accuracy does not decrease. The gradient signature perturbation mapping table Map represents the matching relationship between each gradient channel and the injected perturbation vector in all training batches. The gradient signature perturbation mapping table consists of multiple batch mapping sub-tables. Each batch mapping sub-table contains the perturbed gradient channel number in the batch and its corresponding perturbation vector index information. Batch loss constraint threshold Each value in represents the maximum accuracy drop tolerance allowed when training the t-th batch. The dynamic task loss constraint threshold is controlled by the initial tolerance parameter α and the decay coefficient β, and changes exponentially with the training batch number t.

[0058] Step S106: When it is determined that the external query request obtained based on the reasoning environment is a high-risk request, additional disturbance is generated according to the security-related auxiliary data, and the additional disturbance is added to the reasoning output of the large model to protect the large model parameters.

[0059] Optionally, when it is determined that the external query request obtained based on the inference environment is a high-risk request, additional disturbances are generated based on security-related auxiliary data, including: obtaining the external query request input by the user in the inference phase based on the inference environment, and extracting the behavioral characteristics of the external query request; mapping the behavioral characteristics into risk level values ​​according to the risk scoring function, and when it is determined that the risk level value is greater than the risk response trigger threshold, the external query request is determined to be a high-risk request; selecting a target gradient signature disturbance vector with a spatial direction close to the high-risk request from the gradient signature disturbance vector library; determining the target gradient channel corresponding to the target gradient signature disturbance vector from the gradient signature disturbance mapping table; and generating additional disturbances in the inference phase based on the target gradient signature disturbance vector and the target gradient channel.

[0060] In this embodiment, during the actual reasoning, the system monitors the external query behavior in real time, conducts risk assessment by analyzing the behavior characteristics, and automatically injects targeted additional disturbances into high-risk requests to effectively interfere with malicious parameter reasoning. Specifically, the query request sequence submitted by the external user in the reasoning phase is captured in real time. The query request sequence consists of multiple query requests, and each query request q j Represents an input sample submitted by the user in the inference interface. The input sample can be in the form of feature vector or natural language. Then extract each query request q j The query request sequence behavior feature vector Query request sequence behavior feature vector It consists of five parts: the time density value of the current request, which is used to indicate whether the request is in a high-frequency continuous access interval; the input disturbance sensitivity response change value, which is used to measure the impact of small input disturbances on the output prediction results; the probability prediction entropy value of the model output It is used to measure the uncertainty of the current model output distribution. The value is obtained by performing a logarithmic transformation on each component in the prediction probability vector of the model output and performing a weighted summation. The larger the entropy value, the more uncertain the output. The distribution density of the output results in the historical prediction space is used to measure whether the prediction results fall into the high-density area of ​​the conventional output distribution. The continuous request residual change trend indicator is used to determine whether the current output shows an abnormal oscillation pattern in the time dimension. Of course, this implementation is only an example and does not limit the specific type of behavioral characteristics. In addition, these features are used to comprehensively evaluate the risk level of each query and promptly identify potential parameter theft tendencies, thereby providing a decision-making basis for subsequent dynamic intervention, additional disturbance injection, and behavior tracing, significantly improving the security protection capabilities of the reasoning stage.

[0061] Specifically, in this embodiment, a risk scoring function R(q j ), the risk scoring function is used to transform the query request sequence behavior feature vector fq j Mapped to a risk level value R(qj ), risk level value R(q j ) is a real number ranging from 0 to 1. The larger the value, the closer the query request is to parameter stealing attack behavior. The value is obtained by jointly inputting behavioral features of multiple dimensions into the scoring function. When using the risk scoring function to map behavioral features to risk level values, specifically mapping the behavioral feature vector of the query request to a risk level value is achieved by constructing a risk scoring function. The function receives as input a feature vector consisting of five indicators: time density, disturbance sensitivity, output entropy, historical distribution density, and residual trend. It uses a weighted linear combination, a multi-layer perceptron, or other lightweight neural network for joint modeling and outputs a real value between 0 and 1 as the risk level. The higher the value, the closer the request is to parameter stealing attack characteristics in terms of behavior, thereby achieving a quantitative judgment of high-risk behavior.

[0062] In addition, in this embodiment, a risk response trigger threshold R is also set. th , risk response trigger threshold R th It is a fixed risk score limit used to divide normal requests from high-risk requests. j The risk level value R(q j ) is greater than or equal to the risk response trigger threshold R th When , it means that the query request triggers the response intervention condition, and the system will intervene in the output adjustment of the request. When the external query request is determined to be a high-risk request, the gradient signature perturbation vector library GS lib Select a group of high-risk query requests q j Gradient signature perturbation vectors with similar spatial directions. The similar spatial directions mentioned here mean that when selecting additional perturbation vectors for high-risk query requests, it is necessary to screen out perturbation vectors with directions close to the current query request input features or model output results in the vector space from the gradient signature perturbation vector library. The direction here can be understood as the normalized direction of the high-dimensional feature vector, that is, the unit vector. Similar spatial directions mean that the angle between the two vectors is small, or their cosine similarity is high. Selecting perturbation vectors in this way can make the injected perturbation act more effectively on the model output corresponding to the current query input, improve the interference efficiency of parameter theft attacks, and reduce the impact on the normal function of the model, so as to achieve a more accurate protection strategy. After selecting the gradient signature perturbation vector from the gradient signature perturbation vector library according to the spatial direction, it will be further combined with the perturbation channel information corresponding to the query request in the mapping table Map, and the additional perturbation vector δ in the inference stage will be generated through the perturbation compression remapping operation. infer , and perform amplitude constraint control on it, adding a disturbance vector δ infer The second norm of does not exceed the disturbance intensity threshold εadaptive , disturbance intensity threshold ε adaptive Represents the maximum output perturbation intensity allowed in the inference phase. Among them, the above-mentioned perturbation channel information refers to which model gradient channels are injected with gradient signature perturbations in each batch during the training process, as well as the corresponding injected perturbation vector identifiers. There are detailed records in the gradient signature perturbation mapping table: the mapping table records the correspondence between each training batch, each gradient channel and its corresponding perturbation vector. During the inference phase, when additional perturbations need to be generated for high-risk requests, the system will retrieve the perturbation channels and perturbation vectors related to the current request from the mapping table. The purpose of this information is to ensure that the perturbations injected in the inference phase can accurately correspond to the positions and features of the signatures embedded during training, realize the effective inheritance and dynamic adjustment of the perturbations, and thus enhance the protection and traceability of the model. The above-mentioned additional perturbation vector δ infer It is generated by selecting a perturbation vector that is close to the high-risk query request in the spatial direction from the gradient signature perturbation vector library, combining the perturbation channel information corresponding to the request in the mapping table, using perturbation compression and remapping operations, and constraining its amplitude to ensure that the final additional perturbation vector does not exceed the preset perturbation intensity threshold. Finally, in this implementation, the perturbation vector δ infer Inject the model output vector of the high-risk query request by performing an offset operation on the model prediction result in the output space along the direction of the perturbation vector to generate the perturbed output result vector The Euclidean distance between the perturbed output result vector and the original output result vector shall not exceed the perturbation intensity threshold ε adaptive , in order to control the prediction deviation within an acceptable range.

[0063] Example 2

[0064] Figure 2 The second embodiment of the present invention provides a flowchart of a large model parameter protection method based on deep learning. This embodiment is based on the above embodiment. After adding additional disturbances to the reasoning output of the large model, it also includes: obtaining the log data recorded by high-risk requests during the reasoning stage, and saving the log data to the reasoning security log database; performing attribution verification based on security-related auxiliary data and the security log database to generate a legal evidence chain and large model attribution. Figure 3 As shown, the method includes:

[0065] Step S201: Call a secure random source and generate a gradient signature perturbation vector based on a private key seed, and save the generated gradient signature perturbation vector to a gradient signature perturbation vector library.

[0066] Optionally, call a secure random source and generate a gradient signature perturbation vector based on the private key seed, including:

[0067] Call the hardware-level secure random source to generate an initial random seed, input the initial random seed and the private key seed into the hash function to obtain the signature base seed; initialize the pseudo-random number generator according to the signature base seed to obtain the basis vector, and perform Hadamard orthogonal transformation on the basis vector to generate the initial gradient signature perturbation vector; perform amplitude normalization on the initial gradient signature perturbation vector to obtain the gradient signature perturbation vector.

[0068] Step S202: Generate a gradient signature perturbation mapping table using a Markov state transfer mechanism.

[0069] Optionally, a Markov state transition mechanism is used to generate a gradient signature perturbation mapping table, including: establishing a state space set, wherein each state element in the state space set is used to indicate whether each gradient channel is mapped by a gradient signature perturbation vector during a single batch training of a large model; constructing a Markov state transition probability matrix, wherein each probability element in the Markov state transition probability matrix is ​​used to indicate the state transition probability of the gradient signature perturbation vector in the current training batch of the large model; generating a gradient signature perturbation mapping sub-table corresponding to each training batch of the large model according to the state space set and the Markov state transition probability matrix; and merging each gradient signature perturbation mapping sub-table in the order of the training batches to generate a gradient signature perturbation mapping table.

[0070] Step S203: During the training phase of the large model, the gradient signature perturbation vectors in the gradient signature perturbation vector library are superimposed on the corresponding gradient channels in batches according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradients.

[0071] Optionally, during the training phase of the large model, the gradient signature perturbation vectors in the gradient signature perturbation vector library are superimposed batch by batch on the corresponding gradient channel according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradient, including: constructing a set of original gradient vectors in the training batch of the year during the training phase of the large model, wherein each element in the original gradient vector set represents the original gradient value of the gradient channel under the current training batch; extracting the gradient signature perturbation mapping sub-table corresponding to the current training batch from the gradient signature perturbation mapping table, and extracting the perturbation state vector of the current training batch from the gradient signature perturbation mapping sub-table; and linearly superimposing the original gradient value with the gradient signature perturbation vector under the corresponding gradient channel based on the perturbation state vector of the current training batch.

[0072] Step S204: The accuracy change of the gradient calculation model is calculated based on the gradient signature perturbation, and the large model is iteratively optimized according to the accuracy change to obtain a large model weight snapshot.

[0073] Optionally, the accuracy change of the large model is calculated based on the gradient signature perturbation gradient, and the large model is iteratively optimized according to the accuracy change to obtain a large model weight snapshot, including: constructing a composite task loss function and setting a task loss constraint threshold; calculating the accuracy change of the large model according to the composite task loss function for the gradient signature perturbation gradient; when it is determined that the accuracy change exceeds the task loss constraint threshold, the large model is iteratively optimized to obtain a large model weight snapshot embedded with the gradient signature perturbation.

[0074] In step S205 , the gradient signature perturbation vector library and the gradient signature perturbation mapping table are used as security-related auxiliary data, and the weight snapshot and the security-related auxiliary data are deployed as large model parameters to the inference environment.

[0075] Step S206: When it is determined that the external query request obtained based on the reasoning environment is a high-risk request, additional disturbance is generated according to the security-related auxiliary data, and the additional disturbance is added to the reasoning output of the large model to protect the parameters of the large model.

[0076] Optionally, when it is determined that the external query request obtained based on the inference environment is a high-risk request, additional disturbances are generated based on security-related auxiliary data, including: obtaining the external query request input by the user in the inference phase based on the inference environment, and extracting the behavioral characteristics of the external query request; mapping the behavioral characteristics into risk level values ​​according to the risk scoring function, and when it is determined that the risk level value is greater than the risk response trigger threshold, the external query request is determined to be a high-risk request; selecting a target gradient signature disturbance vector with a spatial direction close to the high-risk request from the gradient signature disturbance vector library; determining the target gradient channel corresponding to the target gradient signature disturbance vector from the gradient signature disturbance mapping table; and generating additional disturbances in the inference phase based on the target gradient signature disturbance vector and the target gradient channel.

[0077] Step S207: Obtain log data recorded during the reasoning phase for high-risk requests, and save the log data to a reasoning security log database.

[0078] Specifically, in the inference phase, this embodiment will number the query request q j , its corresponding risk level value R(q j ), the additional perturbation vector δ injected into the inference phase infer , output result vector after disturbance The data is recorded in the inference security log database to continuously record and analyze high-risk operations and regularly generate attack risk indexes. After the inference query monitoring unit completes the risk level determination and additional disturbance injection of the high-risk query request, it will number the query request q. j , corresponding risk level value R(q j ), additional perturbation vector δ in the inference phase infer, and the perturbation model output result vector Information is composed of structured record entries, which are uniformly stored in the security log database. The security log database organizes all high-risk query logs in timestamp order. When a dispute occurs, the embedded gradient signature perturbation traceability is used to verify the model ownership, forming an evidence chain that can be used for legal evidence, and realizing the protection of model parameters and ownership traceability throughout the entire process.

[0079] Step S208: perform attribution verification based on security-related auxiliary data and the security log database to generate a legal evidence chain and a large model attribution.

[0080] Specifically, in this embodiment, all high-risk query requests in the security log database are periodically statistically analyzed to calculate the attack risk index sequence. i It represents the comprehensive index of the proportion of high-risk requests recorded in the i-th time window, the output response spectrum offset value and the disturbance trigger rate, as shown in the following formula (2):

[0081]

[0082] in, is the number of high-risk requests in the current time window, is the total number of queries in the time window, represents the average predicted entropy change, Γ (i) represents the perturbation insertion ratio, ρ1, ρ2, ρ3 are the risk weighting coefficients set by experience. i Exceeds the preset alarm threshold A th When the system marks the inference service as being in a state that may be subject to parameter theft attack and generates an event label. When the large model encounters a suspicious external call or a third-party detection tool triggers the attribution verification process, the attribution verification module automatically calls the deployed gradient signature perturbation vector library GS libTogether with the gradient signature perturbation mapping table Map, a specific set of samples is selected from the system's built-in challenge sample set to initiate an inference request to the suspicious target model. At this point, the response output of the suspicious target model to the challenge sample set is collected to form a target model output vector set. Each output vector in the set is matched with the corresponding spectral feature template in the gradient signature perturbation vector library for similarity. The matching index uses the frequency domain correlation coefficient, which reflects whether the output response contains the frequency domain signature of the embedded perturbation vector. If the frequency domain correlation coefficient exceeds the set judgment threshold in multiple output responses, it is confirmed that the suspicious target model is embedded with the source model's unique gradient signature perturbation spectrum structure, and an attribution judgment label is established. The attribution judgment label and the event label jointly constitute a complete attribution authentication data packet, which is then hashed and sealed for storage, bound to a unique number and timestamp to form a verifiable legal evidence chain.

[0083] Among them, this embodiment introduces a Markov state transition mechanism to dynamically control the injection method of gradient signature perturbations in each training batch and each gradient channel, achieving an adaptive distribution of gradient signature perturbations during the training process. Unlike traditional random perturbation or fixed channel injection methods, the Markov mechanism can accurately set the perturbation transition probability based on the gradient channel state and historical perturbation mapping, making the gradient signature distribution random and robust, greatly increasing the difficulty of malicious inference or parameter reconstruction attacks, and effectively delaying and interfering with the attacker's parameter theft process. In addition, the present invention constructs a composite task loss function that integrates the main task loss, gradient signature preservation loss, and model stealability suppression loss, and designs a dynamic constraint threshold and adaptive perturbation amplitude adjustment mechanism to dynamically monitor model accuracy changes throughout the training cycle and perform real-time adaptive adjustment of the perturbation amplitude and loss weight. Compared with traditional static protection or one-time injection perturbation methods, the present invention can timely reduce the perturbation when the accuracy loss is too large, ensuring the stability of core business indicators, while maximizing the identifiability of gradient signature embedding and improving the model stealing suppression capability. The present invention integrates an integrated module of reasoning query behavior monitoring, risk scoring and additional disturbance injection at the model inference deployment end. It can analyze external API call behavior in real time through multi-dimensional features, accurately identify high-risk queries with parameter theft tendencies, and dynamically insert targeted disturbances to effectively disrupt malicious reasoning results. Through the joint traceability mechanism of query behavior logs and gradient signature responses, once parameter leakage or ownership disputes occur, a legal evidence chain with frequency domain signatures and timestamps can be automatically generated. The reasoning security protection solution can achieve an improvement in attack detection rate.

[0084] In one specific implementation, an operations engineer at an AI cloud center was monitoring the real-time security logs of a large-model inference API. A red alert popped up on the log monitoring platform, indicating that the risk score of API requests from IP address "13.251.102.118" had suddenly increased to 0.87, far exceeding the threshold of 0.72. System tracing revealed that the IP address was located in A, submitting input samples with an interval of less than 0.25 seconds, and all inputs were perturbations of a specific vector with extremely similar content. A query command was issued to retrieve the feature distribution of requests related to "13.251.102.118" within the last 10 minutes. The data showed that the input perturbation sensitivity index corresponding to the requests was higher than 0.85, and the average entropy of the inference output probability distribution was 3.13, significantly higher than the platform's average of 2.07. Furthermore, the logs captured 1,987 intensive API accesses by the IP address within a short period of time. The parameters of each request varied minimally, demonstrating continuous gradient perturbations. The inference security unit automatically activated the dynamic intervention process, calling the perturbation vector library and injecting a perturbation of 0.0019 in real time at the output of high-risk requests. Backend data showed that after the perturbation injection, the Euclidean distance between the model output and the historical unperturbed model output for the IP request increased from a mean of 0.017 to 0.093. After nine consecutive high-risk requests, the system detected a significant anomaly in the IP's response pattern. The automated attack script automatically stopped batching requests after receiving inconsistent data feedback. During post-incident tracing, the security team compared response data from a control group (unprotected model) with similar attack behaviors. When the unprotected model encountered a simulated attack from a similar IP address "103.85.33.120," the attacker successfully recovered the model's layer 9 parameters after 12,000 consecutive requests, achieving a recovery rate of 37.5%. In contrast, under the protected model, the correlation between all derived gradient sequence features and the signature perturbation was less than 0.28, resulting in a parameter recovery rate of less than 11.2%. During the daytime peak hours, the API detected that the IP address "222.186.50.171" from a domestic IDC began submitting similar templated input requests at a high frequency, making 3,078 consecutive calls, with a risk score fluctuating between 0.79 and 0.93. The security system automatically assigned a perturbation channel, dynamically adjusted the perturbation component in the response content, and wrote the following to the inference log:

[0085] 2024-04-26 10:03:52, IP "222.186.50.171", request number #104223, disturbance amplitude 0.0024, response vector frequency domain offset 0.039, risk level HIGH.

[0086] At 10:04:01 AM on April 26, 2024, the API responded to an abnormal fluctuation detection. The system injected a gradient disturbance component [GS_0418_031], requesting an increase in the output spectrum change from 1.8 to 5.2. The system pushed a security alert to the operation and maintenance email and the administrator's WeChat account.

[0087] At 10:04:17 AM on April 26, 2024, the trend of the residuals of consecutive API requests turned negative. The attack script became invalid after the 3121st request and was automatically terminated.

[0088] When compiling daily security incident reports, the platform compared monitoring results from a traditional API model over the same time period. The traditional model suffered attacks from eight high-risk IP segments within a 24-hour period, all of which were passively logged, with no dynamic perturbation injection or active response mechanisms. Parameter leakage risk alerts were triggered zero times, and 89 exception responses were recorded. However, under the protection system of our invention, seven IP segments were identified as high-risk, all of which underwent real-time intervention and perturbation insertion. None of the parameter leakage attempts were successful, and 2,074 exception responses were recorded, all of which were caused by controllable system interventions. During the one-week security operations cycle, the inference security logs generated the following real-world data: total API requests: 975,000; high-risk behavior identification: 2,943 (0.3%); active perturbation insertion successes: 2,789; suspicious IP attribution verification challenges: 5; correlation coefficients exceeding 0.84; all legal evidence packages were archived; model key business function accuracy: 89.33% (protection group) and 89.58% (traditional group); parameter theft attack recovery rate: highest 11.2% in the protection group and 37.5% in the traditional group; attacker automatic termination rate: 98% in the protection group and less than 40% in the traditional group. Security analysts manually reviewed the system logs and confirmed that all high-risk attacks were dynamically identified and blocked in real time, with a clear and traceable chain of evidence. No substantial parameter leaks occurred that week. Company A's team archived the logs and chain of evidence as annual compliance materials for platform operations and model ownership protection.

[0089] In the implementation of the present application, by integrating reasoning query behavior monitoring, risk scoring and additional disturbance injection, it is possible to analyze external API call behavior in real time through multi-dimensional features, accurately identify high-risk queries with parameter theft tendencies, and dynamically insert targeted disturbances to effectively disrupt malicious reasoning results, thereby effectively protecting large model parameters.

[0090] Example 3

[0091] Figure 3 This is a schematic diagram of the structure of a large model parameter protection device based on deep learning provided in Example 3 of the present invention. Figure 3 As shown, the device includes: a gradient signature perturbation vector library construction module 310, a gradient signature perturbation mapping table generation module 320, a gradient signature perturbation gradient acquisition module 330, a large model weight snapshot acquisition module 340, an inference environment deployment module 350 and an additional perturbation adding module 360.

[0092] The gradient signature perturbation vector library construction module 310 is used to call a secure random source and generate a gradient signature perturbation vector based on a private key seed, and save the generated gradient signature perturbation vector to the gradient signature perturbation vector library;

[0093] A gradient signature perturbation mapping table generation module 320 is configured to generate a gradient signature perturbation mapping table using a Markov state transition mechanism, wherein the gradient signature perturbation mapping table includes a mapping relationship between the gradient channels of the large model in each training batch and the gradient signature perturbation vector;

[0094] The gradient signature perturbation gradient acquisition module 330 is used to superimpose the gradient signature perturbation vectors in the gradient signature perturbation vector library to the corresponding gradient channels in batches according to the gradient signature perturbation mapping table during the training phase of the large model to obtain the gradient signature perturbation gradient;

[0095] A large model weight snapshot acquisition module 340 is used to calculate the accuracy change of the large model based on the gradient signature perturbation gradient, and iteratively optimize the large model according to the accuracy change to obtain a large model weight snapshot;

[0096] An inference environment deployment module 350 is used to use the gradient signature perturbation vector library and the gradient signature perturbation mapping table as safety-related auxiliary data, and to deploy the weight snapshot and safety-related auxiliary data as large model parameters to the inference environment;

[0097] The additional disturbance adding module 360 ​​is used to generate additional disturbances based on security-related auxiliary data when it is determined that the external query request obtained based on the reasoning environment is a high-risk request, and add the additional disturbances to the reasoning output of the large model to protect the parameters of the large model.

[0098] Optionally, a gradient signature perturbation vector library building module is used to call a hardware-level secure random source to generate an initial random seed, and input the initial random seed and the private key seed into a hash function to obtain a signature base seed;

[0099] Initialize the pseudo-random number generator according to the signature base seed to obtain the basis vector, and perform Hadamard orthogonal transformation on the basis vector to generate the initial gradient signature perturbation vector;

[0100] The initial gradient signature perturbation vector is amplitude normalized to obtain the gradient signature perturbation vector.

[0101] Optionally, a gradient signature perturbation mapping table generation module is used to establish a state space set, where each state element in the state space set is used to indicate whether each gradient channel is mapped by a gradient signature perturbation vector during a single batch training process of a large model;

[0102] Construct a Markov state transition probability matrix, where each probability element in the Markov state transition probability matrix is ​​used to represent the state transition probability of the gradient signature perturbation vector in the current training batch of the large model;

[0103] Generate the gradient signature perturbation mapping sub-table corresponding to each training batch of the large model according to the state space set and the Markov state transition probability matrix;

[0104] Each gradient signature perturbation mapping sub-table is merged in the order of training batches to generate a gradient signature perturbation mapping table.

[0105] Optional, gradient signature perturbation gradient acquisition module, used to construct the original gradient vector set in the training batch of the current year during the training phase of the large model, where each element in the original gradient vector set represents the original gradient value of the gradient channel in the current training batch;

[0106] Extracting the gradient signature perturbation mapping sub-table corresponding to the current training batch from the gradient signature perturbation mapping table, and extracting the perturbation state vector of the current training batch from the gradient signature perturbation mapping sub-table;

[0107] Based on the perturbation state vector of the current training batch, the original gradient value is linearly superimposed with the gradient signature perturbation vector under the corresponding gradient channel to obtain the gradient signature perturbation gradient.

[0108] Optional, large model weight snapshot acquisition module, used to construct composite task loss functions and set task loss constraint thresholds;

[0109] Calculate the accuracy change of the large model based on the composite task loss function for the gradient signature perturbation gradient;

[0110] When it is determined that the accuracy change exceeds the task loss constraint threshold, the large model is iteratively optimized to obtain a snapshot of the large model weights embedded with the gradient signature perturbation.

[0111] Optionally, an additional perturbation adding module is used to obtain the external query request input by the user in the reasoning phase based on the reasoning environment and extract the behavioral characteristics of the external query request;

[0112] The behavior characteristics are mapped to a risk level value according to the risk scoring function. When the risk level value is determined to be greater than the risk response trigger threshold, the external query request is determined to be a high-risk request;

[0113] Select a target gradient signature perturbation vector from the gradient signature perturbation vector library, which is close to the spatial direction of the high-risk request;

[0114] Determine a target gradient channel corresponding to a target gradient signature perturbation vector from a gradient signature perturbation mapping table;

[0115] The additional perturbations in the inference stage are generated according to the target gradient signature perturbation vector and the target gradient channel.

[0116] Optionally, the apparatus further includes an attribution verification module for obtaining log data recorded by the high-risk request during the reasoning phase and saving the log data to a reasoning security log database;

[0117] Attribution verification is performed based on security-related auxiliary data and the security log database to generate a legal evidence chain and large-scale model attribution.

[0118] The large model parameter protection device based on deep learning provided by an embodiment of the present invention can execute the large model parameter protection method based on deep learning provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0119] Example 4

[0120] Figure 4 The present invention is a block diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0121] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0122] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0123] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the large model parameter protection method based on deep learning.

[0124] In some embodiments, the large model parameter protection method based on deep learning can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the large model parameter protection method based on deep learning described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the large model parameter protection method based on deep learning in any other appropriate manner (for example, by means of firmware).

[0125] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0126] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0127] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0129] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0130] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0131] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0132] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A large model parameter protection method based on deep learning, characterized in that: include: Calling a secure random source and generating a gradient signature perturbation vector based on a private key seed, and saving the generated gradient signature perturbation vector to a gradient signature perturbation vector library; A Markov state transfer mechanism is used to generate a gradient signature perturbation mapping table, wherein the gradient signature perturbation mapping table includes a mapping relationship between the gradient channel of the large model in each training batch and the gradient signature perturbation vector; In the training phase of the large model, the gradient signature perturbation vectors in the gradient signature perturbation vector library are superimposed on the corresponding gradient channels in batches according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradient; Calculating the accuracy change of the model based on the gradient signature perturbation gradient, and iteratively optimizing the large model according to the accuracy change to obtain a large model weight snapshot; Using the gradient signature perturbation vector library and the gradient signature perturbation mapping table as security-related auxiliary data, and deploying the weight snapshot and the security-related auxiliary data as large model parameters to an inference environment; When it is determined that the external query request obtained based on the reasoning environment is a high-risk request, additional disturbance is generated according to the security-related auxiliary data, and the additional disturbance is added to the reasoning output of the large model to protect the large model parameters.

2. The method according to claim 1, characterized in that The calling of a secure random source and generating a gradient signature perturbation vector based on a private key seed includes: Calling the hardware-level secure random source to generate an initial random seed, and inputting the initial random seed and the private key seed into a hash function to obtain a signature base seed; Initializing a pseudo-random number generator according to the signature base seed to obtain a basis vector, and performing a Hadamard orthogonal transform on the basis vector to generate an initial gradient signature perturbation vector; Perform amplitude normalization processing on the initial gradient signature disturbance vector to obtain the gradient signature disturbance vector.

3. The method according to claim 1, characterized in that The Markov state transfer mechanism is used to generate a gradient signature perturbation mapping table, including: Establishing a state space set, wherein each state element in the state space set is used to indicate whether each gradient channel is mapped by the gradient signature perturbation vector during a single batch training process of a large model; Constructing a Markov state transition probability matrix, wherein each probability element in the Markov state transition probability matrix is ​​used to represent the state transition probability of the gradient signature perturbation vector in the current training batch of the large model; Generate a gradient signature perturbation mapping subtable corresponding to each training batch of the large model according to the state space set and the Markov state transition probability matrix; The gradient signature perturbation mapping sub-tables are merged in the order of training batches to generate the gradient signature perturbation mapping table.

4. The method according to claim 3, characterized in that In the training phase of the large model, the gradient signature perturbation vectors in the gradient signature perturbation vector library are superimposed batch by batch on the corresponding gradient channel according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradient, including: During the training phase of the large model, a set of original gradient vectors in the current training batch is constructed, wherein each element in the set of original gradient vectors represents the original gradient value of the gradient channel in the current training batch; Extracting a gradient signature perturbation mapping sub-table corresponding to the current training batch from the gradient signature perturbation mapping table, and extracting a perturbation state vector of the current training batch from the gradient signature perturbation mapping sub-table; Based on the perturbation state vector of the current training batch, the original gradient value is linearly superimposed with the gradient signature perturbation vector under the corresponding gradient channel to obtain a gradient signature perturbation gradient.

5. The method according to claim 1, characterized in that The calculating the accuracy change of the large model based on the gradient signature perturbation gradient, and iteratively optimizing the large model according to the accuracy change to obtain a large model weight snapshot, includes: Construct a composite task loss function and set the task loss constraint threshold; Calculating the accuracy change of the large model according to the composite task loss function for the gradient signature perturbation gradient; When it is determined that the accuracy change exceeds the task loss constraint threshold, the large model is iteratively optimized to obtain a large model weight snapshot embedded with the gradient signature perturbation.

6. The method according to claim 1, characterized in that When determining that the external query request obtained based on the reasoning environment is a high-risk request, generating an additional disturbance according to the security-related auxiliary data includes: Acquiring the external query request input by the user in the reasoning phase based on the reasoning environment, and extracting behavioral features of the external query request; Mapping the behavior feature into a risk level value according to the risk scoring function, and determining that the external query request is the high-risk request when it is determined that the risk level value is greater than a risk response trigger threshold; Selecting a target gradient signature perturbation vector from the gradient signature perturbation vector library, which is close to the spatial direction of the high-risk request; Determining a target gradient channel corresponding to the target gradient signature perturbation vector from the gradient signature perturbation mapping table; The additional perturbation in the inference stage is generated according to the target gradient signature perturbation vector and the target gradient channel.

7. The method according to any one of claims 1 to 6, characterized in that After adding the additional disturbance to the inference output of the large model, the method further includes: Obtaining log data recorded during the reasoning phase of the high-risk request, and saving the log data to a reasoning security log database; Attribution verification is performed based on the security-related auxiliary data and the security log database to generate a legal evidence chain and a large model attribution.

8. A large model parameter protection device based on deep learning, characterized in that: include: A gradient signature perturbation vector library construction module is used to call a secure random source and generate a gradient signature perturbation vector based on a private key seed, and save the generated gradient signature perturbation vector to a gradient signature perturbation vector library; A gradient signature perturbation mapping table generation module is used to generate a gradient signature perturbation mapping table using a Markov state transfer mechanism, wherein the gradient signature perturbation mapping table includes a mapping relationship between the gradient channel of the large model in each training batch and the gradient signature perturbation vector; A gradient signature perturbation gradient acquisition module is used to, during the training phase of the large model, superimpose the gradient signature perturbation vectors in the gradient signature perturbation vector library onto the corresponding gradient channels in batches according to the gradient signature perturbation mapping table to obtain the gradient signature perturbation gradient; A large model weight snapshot acquisition module is used to calculate the accuracy change of the large model based on the gradient signature perturbation gradient, and iteratively optimize the large model according to the accuracy change to obtain a large model weight snapshot; An inference environment deployment module, configured to use the gradient signature perturbation vector library and the gradient signature perturbation mapping table as security-related auxiliary data, and deploy the weight snapshot and the security-related auxiliary data as large model parameters to an inference environment; An additional disturbance adding module is used to generate additional disturbances based on the security-related auxiliary data when it is determined that the external query request obtained based on the reasoning environment is a high-risk request, and add the additional disturbances to the reasoning output of the large model to protect the parameters of the large model.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program to be executed by the at least one processor, where the computer program is executed by the at least one processor so as to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 7 when executed.

Citation Information

Cited By

  • Active defense method and system based on large model

    CN121396685A

  • A large model-based active defense method and system

    CN121396685B