Large model security monitoring method and device, electronic equipment and storage medium

By building a multivariate Gaussian distribution model in the feature space of the large model and calculating the Mahayana distance, the backdoor attack detection problem of large-scale language models in the industrial environment is solved, real-time security monitoring and defense of the large model is realized, and security and reliability are improved.

CN120296742APending Publication Date: 2025-07-11北京华控智加科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510350996.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and defend backdoor attacks from large-scale language models in industrial environments, especially in industrial control systems with high real-time and diversity requirements, and traditional methods are difficult to work.

Method used

By constructing a multivariate Gaussian distribution model in each layer of feature space of the large model, the feature distribution and feature representation of the sample are calculated, and the anomalies of the sample are judged by using the Mahayana distance, real-time security monitoring of the large model and defense against potential attacks.

Benefits of technology

Real-time security monitoring of large models is realized, and it can effectively detect and defend against potential attacks, improving the security and reliability of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296742A_ABST
    Figure CN120296742A_ABST
Patent Text Reader

Abstract

The invention provides a large model security monitoring method and device, electronic equipment and a storage medium. The method comprises the steps of inputting an unpolluted first sample into a trained large model, and determining feature distribution of the first sample in each layer of feature space of the large model; obtaining a second sample input into the large model in real time, determining the feature representation of the second sample in each layer of feature space, and determining the mahalanobis distance between the feature representation and the feature distribution; and based on the mahalanobis distance, determining a global abnormal score of the large model, and based on the global abnormal score, identifying whether a polluted sample exists in the second sample so as to monitor the security of the large model. Therefore, according to the scheme, potential attack behaviors can be detected and defended in real time, and a powerful guarantee is provided for the safety and reliability of a large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of artificial intelligence and model monitoring technology, and particularly to a method, device, electronic device and storage medium for monitoring the security of large models. Background Art

[0002] In recent years, large language models (LLMs) have gradually become the core technology for industrial fault detection and predictive maintenance due to their powerful data learning ability and generalization performance. However, with the wide application of LLMs in industrial environments, their security issues have become increasingly prominent. As an emerging security threat, backdoor attacks pose a serious challenge to the application of LLMs. Backdoor attacks do not require real-time interference during the model deployment phase, but rather implant malicious data during the training phase, making the attack behavior more difficult to detect.

[0003] Currently, the security research of industrial large models mainly focuses on adversarial attacks and privacy protection fields. Most of the existing defense methods are designed for general scenarios and are difficult to meet the special requirements of industrial environments. The diversity and dynamism of industrial data make traditional backdoor detection methods ineffective; the real-time requirements of industrial control systems also limit the application of complex defense technologies. Summary of the Invention

[0004] The purpose of this application is to solve at least one of the technical problems in the related technologies to some extent.

[0005] To this end, the first purpose of this application is to propose a method for monitoring the security of large models to achieve real-time detection and defense of potential attack behaviors, providing a strong guarantee for the security and reliability of large models.

[0006] The second purpose of this application is to propose a device for monitoring the security of large models.

[0007] The third purpose of this application is to propose an electronic device.

[0008] The fourth purpose of this application is to propose a computer-readable storage medium.

[0009] The fifth purpose of this application is to propose a computer program product.

[0010] To achieve the above object, an embodiment of the first aspect of the present application provides a method for monitoring the security of a large model, including: inputting an uncontaminated first sample into a trained large model, and determining the feature distribution of the first sample in the feature space of each layer of the large model; obtaining a second sample that is input into the large model in real time, determining the feature representation of the second sample in the feature space of each layer, and determining the Mahalanobis distance between the feature representation and the feature distribution; based on the Mahalanobis distance, determining the global anomaly score of the large model, and based on the global anomaly score, identifying whether there is a contaminated sample in the second sample, so as to monitor the security of the large model.

[0011] To achieve the above object, an embodiment of the second aspect of the present application provides a device for monitoring the security of a large model, including: a first determination module, configured to input an uncontaminated first sample into a trained large model, and determine the feature distribution of the first sample in the feature space of each layer of the large model; a second determination module, configured to obtain a second sample that is input into the large model in real time, determine the feature representation of the second sample in the feature space of each layer, and determine the Mahalanobis distance between the feature representation and the feature distribution; an identification module, configured to determine the global anomaly score of the large model based on the Mahalanobis distance, and based on the global anomaly score, identify whether there is a contaminated sample in the second sample, so as to monitor the security of the large model.

[0012] To achieve the above object, an embodiment of the third aspect of the present application provides an electronic device, including: a processor; and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor can execute the method for monitoring the security of the large model according to the embodiment of the first aspect above.

[0013] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer instructions are used to make the computer execute the method for monitoring the security of the large model according to the embodiment of the above aspect.

[0014] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for monitoring the security of the large model according to the embodiment of the above aspect.

[0015] The monitoring method, device, electronic device and storage medium for large model security provided by this application can calculate the Mahalanobis distance according to the feature representation and feature distribution by determining the feature distribution of the first sample in the feature space of each layer of the large model and determining the feature representation of the second sample in the feature space of each layer of the large model. Further, it can judge whether the second sample is a contaminated sample according to the Mahalanobis distance, so as to realize the monitoring of the security of the large model. Thus, potential attack behaviors can be detected and defended in real time, providing a strong guarantee for the security and reliability of the large model.

[0016] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0018] Figure 1 is a schematic flowchart of a monitoring method for large model security provided by an embodiment of this application;

[0019] Figure 2 is a schematic flowchart of another monitoring method for large model security provided by an embodiment of this application;

[0020] Figure 3 is a schematic flowchart of another monitoring method for large model security provided by an embodiment of this application;

[0021] Figure 4 is a schematic flowchart of the training process of the large model in a monitoring method for large model security provided by an embodiment of this application;

[0022] Figure 5 is a schematic structural diagram of a monitoring device for large model security provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The embodiments of this application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are intended to explain this application, and should not be construed as a limitation to this application.

[0024] The monitoring method and device for large model security according to the embodiments of this application will be described below with reference to the drawings.

[0025] Figure 1 is a flowchart of a monitoring method for large model security provided by an embodiment of this application, as Figure 1As shown in the figure, the method for monitoring the security of the large model according to the embodiment of the present application includes, but is not limited to, the following steps:

[0026] S101, input the uncontaminated first sample into the trained large model, and determine the feature distribution of the first sample in the feature space of each layer of the large model.

[0027] It should be noted that the execution subject of the method for monitoring the security of the large model provided by the embodiment of the present application is an electronic device, and the electronic device may be a terminal device. Optionally, the terminal device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a personal computer (PC), a television, etc. The embodiment of the present application does not make specific limitations.

[0028] It should be noted that the uncontaminated sample in the embodiment of the present application refers to a sample that has not been attacked by a backdoor, and the contaminated sample refers to a sample that has been attacked by a backdoor.

[0029] In some embodiments, by obtaining the uncontaminated first sample and inputting the first sample into the trained large model, and then constructing a multivariate Gaussian distribution model according to the first sample in the feature space of each layer of the large model, the feature distribution of the first sample in the feature space of each layer of the large model can be calculated according to the multivariate Gaussian distribution model.

[0030] In some embodiments, the multivariate Gaussian distribution model can calculate the mean vector and covariance matrix of the first sample as the feature distribution of the first sample. That is to say, if there are 3 layers of feature space in the large model, each layer of the 3 layers of feature space corresponds to a feature distribution of the first sample.

[0031] In some embodiments, the uncontaminated data can be obtained from different data sources, and operations such as data cleaning and data augmentation are performed on the data to obtain the first sample.

[0032] S102, obtain the second sample input into the large model in real time, determine the feature representation of the second sample in each layer of the feature space, and determine the Mahalanobis distance between the feature representation and the feature distribution.

[0033] In some embodiments, the samples input into the large model in real time can be used as the second samples, where the second samples may contain contaminated samples. By extracting the feature representations of the second samples in the feature space of each layer and calculating the Mahalanobis distance between the feature representations and the feature distributions.

[0034] Optionally, by extracting the feature representations of the second samples in the feature space of any layer and calculating the Mahalanobis distance between the feature representations and the feature distributions in the feature space of that layer.

[0035] In some embodiments, the number of Mahalanobis distances is the same as the number of layers of the large model's feature space. For example, if the large model has 3 layers of feature space, the first samples correspond to 3 feature distributions, the second samples correspond to 3 feature representations, and then calculate the Mahalanobis distances between the corresponding feature representations and feature distributions in each layer of the feature space, so that 3 Mahalanobis distances can be obtained.

[0036] S103, based on the Mahalanobis distance, determine the global anomaly score of the large model, and based on the global anomaly score, identify whether there are contaminated samples in the second samples, so as to monitor the security of the large model.

[0037] In some embodiments, by normalizing the Mahalanobis distances corresponding to each layer of the feature space and selecting the largest Mahalanobis distance from the normalized Mahalanobis distances as the global anomaly score of the large model.

[0038] In some embodiments, by judging whether the global anomaly score is greater than the anomaly score threshold to identify whether there are contaminated samples in the second samples. If the global anomaly score is greater than the anomaly score threshold, it can be determined that there are contaminated samples in the second samples, realizing the monitoring of the security of the large model. Furthermore, the contaminated samples can be isolated, so that the large model can use the uncontaminated samples for subsequent processing work, providing a guarantee for the security of the large model.

[0039] In some embodiments, the output behavior of the large model can also be monitored to determine whether the large model is attacked, ensuring the accuracy and timeliness of the large model's defense.

[0040] In the method for monitoring the security of the large model provided by the embodiments of the present application, by determining the feature distributions of the first samples in the feature space of each layer of the large model and determining the feature representations of the second samples in the feature space of each layer of the large model, the Mahalanobis distance can be calculated according to the feature representations and the feature distributions. Further, according to the Mahalanobis distance, it is judged whether there are contaminated samples in the second samples, realizing the monitoring of the security of the large model. Thus, potential attack behaviors can be detected and defended in real time, providing a strong guarantee for the security and reliability of the large model.

[0041] Figure 2 is a flowchart of a method for monitoring the security of a large model provided by an embodiment of the present application, asFigure 2 As shown in Figure 2 , the monitoring method for large model security according to the embodiments of the present application includes, but is not limited to, the following steps:

[0042] S201, input the uncontaminated first sample into the trained large model.

[0043] In the embodiments of the present application, the implementation manner of step S201 can be implemented by any one of the embodiments of the present application, and no limitation is made here and will not be elaborated further.

[0044] S202, construct a multivariate Gaussian distribution model in the feature space of each layer of the large model.

[0045] S203, based on the multivariate Gaussian distribution model, calculate the mean vector and covariance matrix of the first sample, and use the mean vector and covariance matrix as the feature distribution.

[0046] In some embodiments, a multivariate Gaussian distribution model can be constructed based on the first sample in the feature space of each layer of the large model, and the mean vector and covariance matrix of the first sample can be calculated using the multivariate Gaussian distribution model, so that the mean vector and covariance matrix can be used as the feature distribution.

[0047] In some embodiments, the mean vector and covariance matrix of each class of samples in each layer of the feature space can be calculated.

[0048] In some embodiments, the formula for calculating the mean vector is as follows:

[0049]

[0050] Where, represents the mean vector of the j-th class in the i-th layer of the large model, N j represents the number of the first samples of the j-th class, D clean represents the first sample, f i (x) represents the mean vector of a sample.

[0051]

[0052] Where, ∑ i represents the covariance matrix, N represents the number of the first samples, and C represents the type of the first samples.

[0053] S204, obtain the second sample input to the large model in real time, determine the feature representation of the second sample in the feature space of each layer, and determine the Mahalanobis distance between the feature representation and the feature distribution.

[0054] In the embodiments of the present application, step S204 can be implemented in any one of the embodiments of the present application, and no limitation is made here and no further description is given.

[0055] S205. Determine the global anomaly score of the large model based on the Mahalanobis distance, and identify whether there are contaminated samples in the second sample based on the global anomaly score to monitor the security of the large model.

[0056] In the embodiments of the present application, step S205 can be implemented in any one of the embodiments of the present application, and no limitation is made here and no further description is given.

[0057] In the method for monitoring the security of a large model provided by the embodiments of the present application, a multivariate Gaussian distribution model is constructed in the feature space of each layer of the large model, and the mean vector and covariance matrix of the first sample are determined according to the multivariate Gaussian distribution model as the feature distribution. Thus, the feature relationship of the first sample can be comprehensively captured, the accuracy of subsequent calculation of the Mahalanobis distance can be improved, and the accuracy of the large model in detecting whether the input sample is contaminated can be improved, providing a strong guarantee for the security and reliability of the large model.

[0058] Figure 3 is a flowchart of a method for monitoring the security of a large model provided by the embodiments of the present application. As Figure 3 shown, the method for monitoring the security of the large model in the embodiments of the present application includes but is not limited to the following steps:

[0059] S301. Input the uncontaminated first sample into the trained large model, and determine the feature distribution of the first sample in the feature space of each layer of the large model.

[0060] In the embodiments of the present application, step S301 can be implemented in any one of the embodiments of the present application, and no limitation is made here and no further description is given.

[0061] S302. Obtain the second sample input to the large model in real time.

[0062] In the embodiments of the present application, step S302 can be implemented in any one of the embodiments of the present application, and no limitation is made here and no further description is given.

[0063] S303. For any one layer of the feature space in each layer of the feature space, extract the feature representation of the second sample in any one layer of the feature space.

[0064] S304. Determine the feature distribution corresponding to any one layer of the feature space, and calculate the Mahalanobis distance between the feature representation and the feature distribution.

[0065] In some embodiments, the feature representation of the second sample in the feature space of any layer of the large model can be extracted, the feature distribution corresponding to the feature space of any layer can be determined, and the Mahalanobis distance between the feature representation and the feature distribution in the feature space of any layer can be calculated.

[0066] In some embodiments, the formula for calculating the Mahalanobis distance is as follows:

[0067]

[0068] Where M represents the Mahalanobis distance, represents the mean vector of the j-th class in the i-th layer of the large model, and N j represents the number of the first samples of the j-th class, and f i (x) represents the mean vector of a sample, and ∑ represents the covariance matrix.

[0069] S305. Normalize the Mahalanobis distance, and determine the maximum Mahalanobis distance from the normalized Mahalanobis distance as the global anomaly score.

[0070] In some embodiments, since the large model has multiple layers of feature spaces, and each layer of feature space corresponds to a Mahalanobis distance, the Mahalanobis distance can be normalized to obtain the normalized Mahalanobis distance. The formula for the normalized Mahalanobis distance is as follows:

[0071]

[0072] Where Norm(M i (x)) represents the normalized Mahalanobis distance, M i (x) represents the Mahalanobis distance, μ i represents the mean, and σ i represents the variance.

[0073] Furthermore, the normalized Mahalanobis distances can be compared in magnitude to determine the maximum Mahalanobis distance, and the maximum Mahalanobis distance can be used as the global anomaly score.

[0074] Optionally, the formula for determining the global anomaly score, that is, the maximum Mahalanobis distance, is as follows:

[0075] Score(x) = max{Norm(M i (x)), 1 ≤ i ≤ L}(5)

[0076] Where Score(x) represents the global anomaly score, and Norm(M i (x)) represents the normalized Mahalanobis distance.

[0077] S306. Based on the global anomaly score, identify whether there are contaminated samples in the second sample to monitor the security of the large model.

[0078] In some embodiments, whether there are contaminated samples in the second sample can be identified according to the magnitude of the anomaly score threshold and the global anomaly score. By obtaining the preset anomaly score threshold and determining whether the global anomaly score is greater than the anomaly score threshold.

[0079] In some embodiments, in response to the global anomaly score being less than the anomaly score threshold, it is determined that there are no contaminated samples in the second sample; in response to the global anomaly score being greater than the anomaly score threshold, it is determined that there are contaminated samples in the second sample.

[0080] In some embodiments, when it is determined that there are no contaminated samples in the second sample, the large model uses the second sample for data processing and can output the output result corresponding to the second sample. That is to say, the output result corresponding to the second sample can be determined based on the large model.

[0081] Furthermore, based on the output result, it can be verified whether the large model has been attacked to monitor the security of the large model in real time.

[0082] In some embodiments, an output rule can be obtained, and the output result can be verified based on the output rule to verify whether the large model has been attacked. The output result can also be sent to the client for manual verification of the output result to verify whether the large model has been attacked.

[0083] In the method for monitoring the security of the large model provided by the embodiments of the present application, by determining the feature representation of the second sample in the feature space of each layer of the large model, the Mahalanobis distance can be calculated according to the feature representation and the feature distribution. Further, according to the Mahalanobis distance, it is judged whether there are contaminated samples in the second sample, so as to monitor the security of the large model. Thus, potential attack behaviors can be detected and defended in real time, providing a strong guarantee for the security and reliability of the large model.

[0084] On the basis of the above embodiments, the embodiments of the present application can explain the training process of the large model, such as Figure 4 shown, the training process of the large model includes but is not limited to the following steps:

[0085] S401, determine the training set of the large model, and the training set includes multiple training samples.

[0086] In some embodiments, the training set of the large model includes multiple training samples, and the training samples can be contaminated samples or non - contaminated samples. That is to say, the training set may contain contaminated samples and non - contaminated samples.

[0087] S402, determine the average confidence of the training samples, and identify candidate contaminated samples from the training samples based on the average confidence.

[0088] In some embodiments, candidate contaminated samples are identified from the training samples by determining the average confidence of each training sample and taking the training samples with an average confidence greater than the confidence threshold as the candidate contaminated samples.

[0089] In some embodiments, the confidence of the training samples can be determined based on the difference information between the clean samples and the contaminated samples. By obtaining a contaminated first reference sample and an uncontaminated second reference sample, and determining the difference information between the first reference sample and the second reference sample for model training.

[0090] In some embodiments, the main differences between clean samples and contaminated samples during the machine learning model training process may be reflected in aspects such as the prediction accuracy of the model, the loss value, and the gradient change. By monitoring the change in the loss value and the gradient change of the large model on the first reference sample and the second reference sample as the difference information for model training.

[0091] Furthermore, based on the difference information, the average confidence of the training samples can be determined, and the training samples with an average confidence greater than the confidence threshold are taken as candidate contaminated samples. That is, the average confidence of the training samples can be calculated according to the difference information.

[0092] Optionally, for the training sample x i , the formula for calculating its average confidence is as follows:

[0093]

[0094] where μ i represents the average confidence, 1 ≤ e ≤ E represents the training cycle, b e,i represents the confidence level, the larger this value, the greater the bias, and the greater the probability that the training sample x i is a contaminated sample.

[0095] Optionally, the formula for determining the confidence level is as follows:

[0096]

[0097] where θe represents the parameter, and y i represents the difference information.

[0098] S403. Perform a nearest neighbor search on the candidate contaminated samples, and determine the contaminated samples from the candidate contaminated samples according to the results of the nearest neighbor search.

[0099] In some embodiments, a nearest neighbor search is performed on each candidate contaminated sample, and the nearest neighbor samples of the candidate contaminated samples are obtained as the results of the nearest neighbor search. Furthermore, the contaminated samples can be determined from the candidate contaminated samples according to the nearest neighbor samples.

[0100] In some embodiments, for any sample in the candidate contaminated samples, a nearest neighbor search is performed on any sample to determine K nearest neighbor samples corresponding to any sample, where K is a natural number greater than 1. By determining whether the K nearest neighbor samples include the candidate contaminated samples, it can be determined whether any sample is a contaminated sample.

[0101] In some embodiments, in response to the K nearest neighbor samples including the candidate contaminated samples, it is determined that any sample is a contaminated sample; in response to the K nearest neighbor samples not including the candidate contaminated samples, it is determined that any sample is a clean sample that has not been contaminated. A nearest neighbor search is performed on each candidate contaminated sample until the nearest neighbor search for the candidate contaminated samples ends, obtaining the determined contaminated samples and the clean samples that have not been contaminated.

[0102] S404, generate a contaminated sample set based on the contaminated samples, and remove the contaminated sample set from the training set to obtain a target training set, so that the large model is trained using the target training set.

[0103] In some embodiments, a contaminated sample set can be generated according to the determined contaminated samples. The contaminated sample set is a subset of the training set. That is to say, in order to enable the large model to be trained on clean samples, by removing the contaminated sample set from the training set, a clean target training set can be obtained, and thus the large model can be trained using the target training set.

[0104] In the large model security monitoring method provided by the embodiments of the present application, by determining contaminated samples in the training samples, the identified contaminated samples can be removed from the training set, ensuring that the model is trained on clean data, thereby significantly reducing the success rate of backdoor attacks and being able to efficiently identify and eliminate backdoor contamination in the training data.

[0105] Corresponding to the large model security monitoring methods proposed in the above several embodiments, an embodiment of the present application also proposes a large model security monitoring device. Since the large model security monitoring device proposed in the embodiments of the present application corresponds to the large model security monitoring methods proposed in the above several embodiments, the implementation manners of the above large model security monitoring methods are also applicable to the large model security monitoring device proposed in the embodiments of the present application and will not be described in detail in the following embodiments.

[0106] To implement the above embodiments, the present application also proposes a large model security monitoring device.

[0107] Figure 5 It is a schematic structural diagram of a large model security monitoring device provided by the embodiments of the present application.

[0108] As Figure 5 shown, the large model security monitoring device 500 includes:

[0109] The first determination module 501 is configured to input an uncontaminated first sample into a trained large model, and determine the feature distribution of the first sample in the feature space of each layer of the large model;

[0110] The second determination module 502 is configured to obtain a second sample that is input into the large model in real time, determine the feature representation of the second sample in the feature space of each layer, and determine the Mahalanobis distance between the feature representation and the feature distribution;

[0111] The identification module 503 is configured to determine the global anomaly score of the large model based on the Mahalanobis distance, and identify whether there is a contaminated sample in the second sample based on the global anomaly score, so as to monitor the security of the large model.

[0112] In a possible implementation manner of the embodiment of the present application, the first determination module 501 is further configured to: construct a multivariate Gaussian distribution model in the feature space of each layer of the large model; calculate the mean vector and covariance matrix of the first sample based on the multivariate Gaussian distribution model, and use the mean vector and covariance matrix as the feature distribution.

[0113] In a possible implementation manner of the embodiment of the present application, the second determination module 502 is further configured to: for any layer of the feature space in each layer of the feature space, extract the feature representation of the second sample in any layer of the feature space; determine the feature distribution corresponding to any layer of the feature space, and calculate the Mahalanobis distance between the feature representation and the feature distribution.

[0114] In a possible implementation manner of the embodiment of the present application, the identification module 503 is further configured to: perform normalization processing on the Mahalanobis distance, and determine the maximum Mahalanobis distance from the normalized Mahalanobis distance as the global anomaly score.

[0115] In a possible implementation manner of the embodiment of the present application, the identification module 503 is further configured to: obtain a preset anomaly score threshold, and determine whether the global anomaly score is greater than the anomaly score threshold; in response to the global anomaly score being less than the anomaly score threshold, determine that there is no contaminated sample in the second sample; in response to the global anomaly score being greater than the anomaly score threshold, determine that there is a contaminated sample in the second sample.

[0116] In a possible implementation manner of the embodiment of the present application, the identification module 503 is further configured to: determine the output result corresponding to the second sample based on the large model; verify whether the large model is attacked based on the output result, so as to monitor the security of the large model in real time.

[0117] In a possible implementation manner of the embodiment of the present application, the first determination module 501 is further configured to: determine a training set of the large model, where the training set includes a plurality of training samples; determine the average confidence of the training samples, and identify candidate contaminated samples from the training samples based on the average confidence; perform a nearest neighbor search on the candidate contaminated samples, and determine contaminated samples from the candidate contaminated samples according to the result of the nearest neighbor search; generate a contaminated sample set based on the contaminated samples, and remove the contaminated sample set from the training set to obtain a target training set, so that the large model is trained using the target training set.

[0118] In a possible implementation manner of the embodiment of the present application, the first determination module 501 is further configured to: obtain a contaminated first reference sample and an uncontaminated second reference sample, and determine the difference information between the first reference sample and the second reference sample for model training; based on the difference information, determine the average confidence of the training samples, and use the training samples with an average confidence greater than the confidence threshold as candidate contaminated samples.

[0119] In a possible implementation manner of the embodiment of the present application, the first determination module 501 is further configured to: for any sample in the candidate contaminated samples, perform a nearest neighbor search on the any sample to determine K nearest neighbor samples corresponding to the any sample, where K is a natural number greater than 1; in response to the K nearest neighbor samples including candidate contaminated samples, determine the any sample as a contaminated sample.

[0120] In the large model security monitoring device provided by the embodiment of the present application, by determining the feature distribution of the first sample in the feature space of each layer of the large model and determining the feature representation of the second sample in the feature space of each layer of the large model, the Mahalanobis distance can be calculated according to the feature representation and the feature distribution. Further, it is determined whether the second sample is a contaminated sample according to the Mahalanobis distance, so as to monitor the security of the large model. Thus, potential attack behaviors can be detected and defended in real time, providing a strong guarantee for the security and reliability of the large model.

[0121] It should be noted that the foregoing explanation of the embodiment of the large model security monitoring method is also applicable to the large model security monitoring device of this embodiment, and will not be repeated here.

[0122] To implement the above embodiment, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiment.

[0123] To implement the above embodiment, the present application also proposes a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the method provided in the foregoing embodiment.

[0124] To implement the above embodiments, the present application further provides a computer program product, including a computer program, which when executed by a processor, implements the method provided by the foregoing embodiments.

[0125] The collection, storage, use, processing, transmission, provision, and application, etc. of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0126] It should be noted that personal information from users should be collected for legal and reasonable purposes and not shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and signing an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps need to be taken to defend and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0127] The present application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present application is expected to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.

[0128] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0129] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0130] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions can be performed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0131] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0132] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0133] Those of ordinary skill in the art can understand that all or part of the steps carried out in the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0134] In addition, each functional unit in the various embodiments of the present application may be integrated into a processing module, may exist physically separately for each unit, or two or more units may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0135] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A monitoring method for large model security, characterized in that, The method includes: Inputting an uncontaminated first sample into a trained large model, and determining the feature distribution of the first sample in the feature space of each layer of the large model; Obtaining a second sample input into the large model in real time, determining the feature representation of the second sample in the feature space of each layer, and determining the Mahalanobis distance between the feature representation and the feature distribution; Based on the Mahalanobis distance, determining the global anomaly score of the large model, and based on the global anomaly score, identifying whether there is a contaminated sample in the second sample to monitor the security of the large model.

2. The method according to claim 1, wherein The determining the feature distribution of the first sample in the feature space of each layer of the large model includes: Constructing a multivariate Gaussian distribution model in the feature space of each layer of the large model; Based on the multivariate Gaussian distribution model, calculating the mean vector and covariance matrix of the first sample, and using the mean vector and covariance matrix as the feature distribution.

3. The method according to claim 1, wherein The determining the feature representation of the second sample in the feature space of each layer and determining the Mahalanobis distance between the feature representation and the feature distribution includes: For any layer of the feature space in each layer of the feature space, extracting the feature representation of the second sample in the any layer of the feature space; Determining the feature distribution corresponding to the any layer of the feature space, and calculating the Mahalanobis distance between the feature representation and the feature distribution.

4. The method according to claim 3, wherein The determining the global anomaly score of the large model based on the Mahalanobis distance includes: Performing normalization processing on the Mahalanobis distance, and determining the maximum Mahalanobis distance from the normalized Mahalanobis distance as the global anomaly score.

5. The method according to claim 4, characterized in that The identifying whether there is a contaminated sample in the second sample based on the global anomaly score includes: Obtaining a preset anomaly score threshold, and determining whether the global anomaly score is greater than the anomaly score threshold; In response to the global anomaly score being less than the anomaly score threshold, determining that there is no contaminated sample in the second sample; In response to the global anomaly score being greater than the anomaly score threshold, determining that there is a contaminated sample in the second sample.

6. The method according to claim 5, characterized in that, After the determining that there is no contaminated sample in the second sample in response to the global anomaly score being less than the anomaly score threshold, it further includes: Determining an output result corresponding to the second sample based on the large model; Based on the output result, verifying whether the large model is attacked to monitor the security of the large model in real time.

7. The method according to claim 1, characterized in that, The training process of the large model includes: Determining a training set of the large model, where the training set includes a plurality of training samples; Determining the average confidence of the training samples, and identifying candidate contaminated samples from the training samples based on the average confidence; Performing a nearest neighbor search on the candidate contaminated samples, and determining contaminated samples from the candidate contaminated samples according to the result of the nearest neighbor search; Generating a contaminated sample set based on the contaminated samples, and removing the contaminated sample set from the training set to obtain a target training set, so that the large model is trained using the target training set.

8. The method according to claim 7, wherein Determining the average confidence of the training samples and identifying candidate contaminated samples based on the average confidence includes: Obtaining a first reference sample that is contaminated and a second reference sample that is not contaminated, and determining the difference information between the first reference sample and the second reference sample for model training; Based on the difference information, determining the average confidence of the training samples, and taking the training samples with an average confidence greater than the confidence threshold as the candidate contaminated samples.

9. The method according to claim 7, wherein Performing a nearest neighbor search on the candidate contaminated samples and determining contaminated samples from the candidate contaminated samples according to the results of the nearest neighbor search includes: For any sample in the candidate contaminated samples, performing a nearest neighbor search on the any sample to determine K nearest neighbor samples corresponding to the any sample, where K is a natural number greater than 1; In response to the K nearest neighbor samples including the candidate contaminated samples, determining the any sample as a contaminated sample.

10. A monitoring device for the security of large models, characterized in that, The device includes: A first determination module, configured to input an uncontaminated first sample into a trained large model and determine the feature distribution of the first sample in the feature space of each layer of the large model; A second determination module, configured to obtain a second sample that is input into the large model in real time, determine the feature representation of the second sample in the feature space of each layer, and determine the Mahalanobis distance between the feature representation and the feature distribution; An identification module, configured to determine the global anomaly score of the large model based on the Mahalanobis distance, and identify whether there are contaminated samples in the second sample based on the global anomaly score to monitor the security of the large model.

11. An electronic device, characterized in that, Includes: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1-9.

13. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Cited By

  • Large language model-oriented backdoor detection and purification method, system and equipment

    CN121881354A