Intelligent Grid Deep Learning Model Copyright Protection Method, Device and System Based on Model Confidence Region

By applying model confidence domain technology in the smart grid deep learning model, generating feature identification sets and training copyright detection models, the problem of low copyright protection detection accuracy in the smart grid is solved, and effective defense and data security guarantees are achieved for copyright infringement.

CN119249376BActive Publication Date: 2025-07-18ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411475941.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-07-18
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

In the prior art, the copyright protection method of the smart grid deep learning model has low accuracy when detecting model copyright infringement attacks, and cannot effectively defend against attackers obtaining user privacy data and grid operation data through model extraction attacks, resulting in data leakage and security risks.

Method used

By obtaining the smart grid depth model set, using model confidence domain technology to search for feature data samples in the proprietary data set, generate a confidence domain feature point set, and reduce the dimensionality of gradient vectors, generate a model feature identification set, train a copyright detection model, and improve the accuracy of copyright infringement detection.

Benefits of technology

Effectively defend against attackers forged identities to replace users to participate in model training, improving the accuracy and robustness of feature identification matching, ensuring the copyright security of data sets deployed in the smart grid, and preventing data leakage and unauthorized utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119249376B_ABST
    Figure CN119249376B_ABST
Patent Text Reader

Abstract

The present application relates to an intelligent power grid deep learning model copyright protection method, device, system, computer-readable storage medium and computer program product based on a model confidence region. The method includes: obtaining an intelligent power grid deep model set; for each model in the intelligent power grid deep model set, searching for characteristic data samples in a proprietary dataset where the model prediction value is close to the model discrimination boundary to obtain a confidence region characteristic point set; performing dimensionality reduction on the gradient vectors of the confidence region characteristic point set on the model discrimination boundary according to the method of linear discriminant analysis to obtain a perturbation vector of the confidence region characteristic point set; generating a model characteristic identification set corresponding to the intelligent power grid deep model set according to the change in the prediction labels of each model before and after combining the confidence region characteristic point set and the perturbation vector; training a copyright detection model to be trained according to the model characteristic identification set to obtain a pre-trained copyright detection model. Using this method can improve the detection accuracy of model copyright infringement attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to an intelligent grid deep learning model copyright protection method, device, system, computer-readable storage medium, and computer program product based on a model confidence region. Background Art

[0002] With the rapid development of smart grid technology, machine learning models are increasingly widely used in power systems. In particular, the progress of deep learning technology has promoted the development of various advanced analysis algorithms, including power grid flow prediction, user electricity consumption behavior analysis, equipment anomaly detection, etc. However, these machine learning-based systems rely on high-quality data sets, which makes them potential targets for intellectual property theft.

[0003] For example, machine learning models in smart grids usually need to access and analyze a large amount of user data, such as household electricity consumption habits, electricity consumption, etc. Once these data are obtained by an attacker through a model extraction attack, it will not only violate the privacy of users, but may also be used for malicious purposes, such as fraud or stealing user identity information. At the same time, the smart grid also contains a large amount of power grid operation data, such as power load, equipment status, etc. These data are crucial for the safe operation of the power grid. If an attacker obtains these data through a model extraction attack, it may lead to the leakage of power grid operation data, thus endangering the safe and stable operation of the power grid. The attacker can use these data for power market manipulation or cyber attacks against the power grid. Therefore, how to protect the copyright of neural networks in this situation has become an urgent problem to be solved.

[0004] In related technologies, some studies have been dedicated to defending against model extraction attacks and protecting model privacy. However, these methods still have some limitations in practical applications, resulting in a low detection accuracy for model copyright infringement attacks. Summary of the Invention

[0005] Based on this, it is necessary to provide an intelligent grid deep learning model copyright protection method, device, system, computer-readable storage medium, and computer program product based on a model confidence region that can improve the detection accuracy of model copyright infringement attacks for the above technical problems.

[0006] In a first aspect, the present application provides an intelligent grid deep learning model copyright protection method based on a model confidence region, including:

[0007] Obtain a set of smart grid deep models; the set of smart grid deep models includes a target copyright protection model, as well as a target risk model and a target homomorphic model corresponding to the target copyright protection model; the target risk model is a model with an infringement problem of the smart grid dataset with respect to the target copyright protection model; the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain;

[0008] For each model in the set of smart grid deep models, search for characteristic data samples in the smart grid proprietary dataset whose distance between the corresponding model prediction value and the corresponding model discrimination boundary satisfies a preset condition, and obtain a set of confidence domain characteristic points; the model discrimination boundary represents the region that can be predicted as other types within a preset displacement in the gradient direction of the model's inference on the characteristic data sample;

[0009] Determine the gradient vector of the set of confidence domain characteristic points on the model discrimination boundary, and perform dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, and use the gradient vector after dimensionality reduction as the perturbation vector of the set of confidence domain characteristic points;

[0010] Generate a set of model characteristic identifiers corresponding to the set of smart grid deep models according to the change in the prediction labels of each model before and after combining the set of confidence domain characteristic points and the perturbation vector;

[0011] Train the copyright detection model to be trained according to the set of model characteristic identifiers to obtain a pre-trained copyright detection model; the pre-trained copyright detection model is used to detect whether the model to be detected infringes the copyright of the target copyright protection model.

[0012] In one embodiment, the determining the gradient vector of the set of confidence domain characteristic points on the model discrimination boundary and performing dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis includes:

[0013] Group the gradient vectors according to the class labels, and determine the mean vector of each class and the overall mean vector;

[0014] Determine the within-class scatter matrix and the between-class scatter matrix according to the mean vector of each class and the overall mean vector;

[0015] By solving the eigenvalues and eigenvectors of the product of the inverse matrix of the within-class scatter matrix and the between-class scatter matrix, select the eigenvectors whose cumulative contribution rate satisfies the preset condition for dimensionality reduction processing to obtain the gradient vector after dimensionality reduction.

[0016] In one embodiment, the training the copyright detection model to be trained according to the set of model characteristic identifiers includes:

[0017] Expand the model feature identification set by searching for homologous data subsets to obtain an expanded model feature identification set;

[0018] Train the copyright detection model to be trained through the expanded model feature identification set.

[0019] In one embodiment, the expanding the model feature identification set by searching for homologous data subsets to obtain an expanded model feature identification set includes:

[0020] Divide the model feature identification set into k subsets, and regard each model feature identification in the model feature identification set as a data point; each subset includes a core point; the core point is represented as randomly selected from the model feature identification set;

[0021] Determine the distance from each data point to each core point, and assign each data point to the subset to which the nearest core point belongs;

[0022] Update the core point of each subset to the mean value of all data points within the subset, and return to the step of determining the distance from each data point to each core point until the core point no longer changes or reaches the maximum number of iterations, and perform data augmentation processing on the data points within each subset to obtain the expanded model feature identification set.

[0023] In one embodiment, the obtaining the intelligent grid depth model set includes:

[0024] Obtain the training data set used by the target copyright protection model during training, and train the target homomorphic model on the training data set; the target homomorphic model includes a fully homomorphic model, an architecture homomorphic model, and a problem domain homomorphic model;

[0025] The fully homomorphic model is obtained by performing multiple trainings on the premise that all experimental settings are the same as those of the target copyright protection model;

[0026] The architecture homomorphic model is obtained by fine-tuning the hyperparameters during training and performing multiple trainings on the premise that the architecture is the same as that of the target copyright protection model;

[0027] The problem domain homomorphic model is obtained by performing multiple trainings under the premise of using the same source data set as the target copyright protection model and with a preset combination of similar model structures and randomly selected hyperparameters.

[0028] The copyright detection model to be trained includes an autoencoder recognition network and a mapping network. Training the copyright detection model to be trained according to the model feature identification set includes:

[0029] Training the autoencoder recognition network and the mapping network using self-supervised learning methods to map the model feature identification set to a low-dimensional representation space;

[0030] If the model feature identification is , the positive pair of the model feature identification is , and the negative pair of the model feature identification is , the loss function of self-supervised learning can be expressed as:

[0031]

[0032] where represents the degree of similarity of , and represents the degree of similarity of .

[0033] In a second aspect, the present application also provides an intelligent grid deep learning model copyright protection device based on a model confidence region, including:

[0034] An acquisition module, configured to acquire an intelligent grid deep model set; the intelligent grid deep model set includes a target copyright protection model, as well as a target risk model and a target homomorphic model corresponding to the target copyright protection model; the target risk model is a model with an intelligent grid dataset infringement problem with respect to the target copyright protection model; the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain;

[0035] A search module, configured to search, for each model in the intelligent grid deep model set, for a feature data sample whose distance between the corresponding model prediction value and the corresponding model discrimination boundary in the intelligent grid proprietary dataset meets a preset condition, to obtain a confidence region feature point set; the model discrimination boundary represents a region that can be predicted as other types within a preset displacement in the gradient direction of the model's inference on the feature data sample;

[0036] A determination module, configured to determine the gradient vector of the confidence region feature point set on the model discrimination boundary, and perform dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, and use the dimensionality-reduced gradient vector as the perturbation vector of the confidence region feature point set;

[0037] ​A generation module, configured to generate a model feature identification set corresponding to the intelligent grid deep model set according to the prediction label changes before and after the combination of the confidence domain feature point set and the perturbation vector for each of the models;

[0038] A training module, configured to train a copyright detection model to be trained according to the model feature identification set, so as to obtain a pre-trained copyright detection model; the pre-trained copyright detection model is used to detect whether a model to be detected infringes the copyright of the target copyright protection model.

[0039] In a third aspect, the present application further provides an intelligent grid deep learning model copyright protection system based on a model confidence domain. The system includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0040] In a fourth aspect, the present application further provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0041] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0042] The above intelligent grid deep learning model copyright protection method, device, system, computer-readable storage medium, and computer program product based on the model confidence region obtain a target copyright protection model, a target risk model corresponding to the target copyright protection model, and a target homomorphic model. The target risk model is a model with an intelligent grid dataset infringement problem with respect to the target copyright protection model; the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain. For each model in the intelligent grid deep model set, feature data samples whose distance between the corresponding model prediction value and the corresponding model discrimination boundary in the intelligent grid proprietary dataset satisfies a preset condition are searched to obtain a confidence region feature point set, so that the obtained confidence region feature point set can accurately describe the confidence region information of the deep learning model. Then, the gradient vector of the confidence region feature point set on the model discrimination boundary is dimensionally reduced according to the method of linear discriminant analysis to obtain a perturbation vector of the confidence region feature point set. Thus, according to the change in the prediction labels before and after combining the confidence region feature point set and the perturbation vector, a model feature identification set corresponding to the intelligent grid deep model set can be accurately generated. The pre-trained copyright detection model is obtained by training the copyright detection model to be trained with the model feature identification set. The pre-trained copyright detection model can be used to detect whether the model to be detected infringes the copyright of the target copyright protection model, and thus can effectively defend against attackers who forge identities to replace users in the iteration after the start of training. At the same time, by extracting the confidence region features of the model to generate feature identifications, the risk model and the homomorphic model can be effectively distinguished, improving the accuracy and robustness of feature identification matching, protecting the dataset copyright security of model deployment in the intelligent grid, and preventing data leakage and unauthorized use. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 It is a schematic flowchart of a method for protecting the copyright of an intelligent grid deep learning model based on the model confidence region in one embodiment;

[0045] Figure 2 It is a schematic flowchart of a method for protecting the copyright of an intelligent grid deep learning model based on the model confidence region in another embodiment;

[0046] Figure 3 It is a schematic flowchart of a method for protecting the copyright of an intelligent grid deep learning model based on the model confidence region in yet another embodiment;

[0047] Figure 4 It is a structural block diagram of an intelligent grid deep learning model copyright protection device based on model confidence domain in an embodiment;

[0048] Figure 5 It is an internal structure diagram of a computer device in an embodiment. Specific implementation manners

[0049] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0051] In one embodiment, as Figure 1 shown, a method for copyright protection of an intelligent grid deep learning model based on model confidence domain is provided. In this embodiment, this method is exemplified by being applied to a computer device. It can be understood that the computer device can be a terminal, a server, or a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0052] Step S110, obtaining an intelligent grid deep model set.

[0053] Among them, the intelligent grid deep model set includes a target copyright protection model, as well as a target risk model and a target homomorphic model corresponding to the target copyright protection model.

[0054] Among them, the target risk model is a model with an intelligent grid data set infringement problem with the target copyright protection model.

[0055] Among them, the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain.

[0056] Among them, the target copyright protection model can be a machine learning model deployed in the intelligent grid and requiring intelligent grid data copyright protection.

[0057] In practical applications, the target copyright protection model can be named as the smart grid copyright protection model.

[0058] In specific implementation, a computer device can obtain a smart grid deep model set; the smart grid deep model set includes the target copyright protection model, as well as the target risk model and the target homomorphic model corresponding to the target copyright protection model.

[0059] Among them, the computer device can generate the target risk model through a model extraction algorithm and generate the target homomorphic model using the white-box permission of the defender.

[0060] Furthermore, the computer device can train the homomorphic model by fine-tuning the data set and parameters to simulate different models in the same problem domain. The homomorphic model can be trained on the training data set used during the training process of the target copyright protection model, but with different hyperparameters or model structures. By simulating the attacker, the target risk model is generated through the model extraction algorithm, and the (Application Programming Interface, API) of the target copyright protection model is queried multiple times, and the target risk model with smart grid data set infringement problems is simulated and trained using the returned prediction results.

[0061] Specifically, the detailed steps of the model extraction algorithm are as follows. The simulated attacker repeatedly queries the target copyright protection model through the API In a problem domain where there may be smart grid data set leakage and copyright infringement, define this data set as The target copyright protection model The training data set used during the training process is

[0062] is The corresponding label. The simulated attacker obtains A part of the data set, randomly extracts a model extraction data set with a size of from the data set where represents the proportion of the model extraction data set in the training data set, and 1.5 can be used to achieve the effect of balancing the learning model features and simulating the actual attack scenario. By querying to obtain the prediction result The model extraction data set can be expressed as and use the result to train the risk model, with the optimization objective being so as to realize the feature learning of the risk model for the target copyright protection model and obtain the target risk model . Among them, ​and respectively represent the neural network parameters of the target risk model and the target copyright protection model. After multiple extractions and trainings, a risk model set is formed .

[0063] In some embodiments, during the process of the computer device obtaining the intelligent power grid depth model set, the computer device can obtain the training data set used by the target copyright protection model during the training process, and train a target homomorphic model on the training data set; the target homomorphic model includes a fully homomorphic model, an architecture homomorphic model, and a problem domain homomorphic model.

[0064] Among them, the fully homomorphic model is obtained through multiple trainings on the premise that all experimental settings are the same as those of the target copyright protection model.

[0065] Among them, the architecture homomorphic model is obtained through multiple trainings by fine-tuning the hyperparameters during training on the premise that the architecture is the same as that of the target copyright protection model.

[0066] Among them, the problem domain homomorphic model is obtained through multiple trainings under the premise of using the same source data set as the target copyright protection model, with a preset similar model structure and a randomly selected combination of hyperparameters.

[0067] Specifically, during the training process of the homomorphic model, considering that it is necessary to distinguish the target copyright protection model and the homomorphic model as much as possible during the subsequent recognition network training stage, the entire training data set is used for training. According to the proximity between the target homomorphic model and the target copyright protection model, the target homomorphic model set is divided into the following three subsets: the fully homomorphic model , the architecture homomorphic model , and the problem domain homomorphic model . Among them, the fully homomorphic model represents the result of multiple trainings on the premise that all experimental settings are the same as those of the target copyright protection model, the architecture homomorphic model represents the result of multiple trainings by fine-tuning the hyperparameters during training, such as the learning rate, batch size, and optimizer selection, while keeping the architecture the same as that of the target copyright protection model, and the problem domain homomorphic model represents the result of training under the premise of only using the same source data set, with a preset similar model structure and a randomly selected combination of hyperparameters. The combination , and can be used to build the intelligent power grid depth model set .

[0068] Step S120: For each model in the smart grid deep model set, search for feature data samples in the smart grid proprietary dataset where the distance between the corresponding model prediction value and the corresponding model discrimination boundary meets a preset condition, to obtain a confidence region feature point set.

[0069] Among them, the model discrimination boundary represents the region that can be predicted as other types within a preset displacement in the gradient direction of the model's inference on the feature data sample.

[0070] Furthermore, before obtaining the confidence region feature point set, perform data augmentation on the smart grid proprietary dataset to generate an augmented smart grid proprietary dataset. The methods for data augmentation of the image dataset in the smart grid proprietary dataset include rotation, flipping, scaling, slicing, etc., to increase data diversity; the methods for data augmentation of the digital dataset in the smart grid proprietary dataset include adding noise, slicing data, etc., to enhance the robustness of the model.

[0071] In this way, for each model in the smart grid deep model set, based on the augmented smart grid proprietary dataset, extract the confidence region features of the model set. Specifically, the computer device can, through the white-box permission of the smart grid proprietary dataset, use an iterative method to search for feature data samples (data points) in the augmented smart grid proprietary dataset where the distance between the corresponding model prediction value and the corresponding model discrimination boundary meets a preset condition (such as being less than a preset distance threshold), and fuse the feature data samples belonging to different labels in combination with the problem domain. To ensure that the obtained confidence region feature point set can accurately describe the confidence region information of the deep learning model, use the method of attitude matching to find the feature data sample with the highest attitude from different label classes, and at the same time obtain feature data samples whose similarity with the decision sample meets a preset threshold through a heuristic iterative algorithm, and set the iterative algorithm to stop within the model discrimination boundary (discrimination boundary region) or when the number of iterations ends. Among them, the model discrimination boundary represents the region that can be predicted as other types within a displacement in the gradient direction of the model's inference on the feature data sample. Specifically, let the model ( ) have a gradient of for the input feature data sample , then the discrimination boundary region can be expressed as .

[0072] Step S130: Determine the gradient vector of the confidence region feature point set on the model discrimination boundary, and perform dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, and use the dimensionality-reduced gradient vector as the perturbation vector of the confidence region feature point set.

[0073] In a specific implementation, after obtaining the confidence region feature point set, the computer device can determine the gradient vector of the confidence region feature point set on the model discrimination boundary, and perform dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, and use the gradient vector after dimensionality reduction as the perturbation vector of the confidence region feature point set.

[0074] Specifically, the computer device can group the gradient vectors according to the class labels and determine the mean vector of each class and the overall mean vector . Then, according to the mean vector of each class and the overall mean vector , determine the within-class scatter matrix and the between-class scatter matrix . Then, by solving the eigenvalues and eigenvectors of the product of the inverse matrix of the within-class scatter matrix and the between-class scatter matrix , select the eigenvectors whose cumulative contribution rate meets the preset conditions (for example, the first 95% cumulative contribution rate), and use these eigenvectors to project the original gradient vector into a low-dimensional space to complete the dimensionality reduction process and obtain the gradient vector after dimensionality reduction. Use the gradient vector after dimensionality reduction as the perturbation vector of the confidence region feature point set.

[0075] Step S140, generate a model feature identification set corresponding to the smart grid deep model set according to the prediction label changes of each model before and after combining the confidence region feature point set and the perturbation vector.

[0076] In a specific implementation, the computer device can record the prediction label changes of each model before and after combining the confidence region feature point set and the perturbation vector. The generation function of the model feature identification can be expressed as: . Among them, is the confidence region feature point set, is the linearly dimension-reduced gradient vector (i.e., the gradient vector after dimensionality reduction).

[0077] Step S150, train the copyright detection model to be trained according to the model feature identification set to obtain a pre-trained copyright detection model.

[0078] Among them, the pre-trained copyright detection model is used to detect whether the model to be detected infringes the copyright of the target copyright protection model.

[0079] In a specific implementation, the computer device can train the copyright detection model to be trained according to the model feature identification set to obtain a pre-trained copyright detection model, and the pre-trained copyright detection model is used to detect whether the model to be detected infringes the copyright of the target copyright protection model.

[0080] Further, the computer device can expand the model feature identification set by using the method of searching for subsets of homologous data to increase the diversity of training data, obtain the expanded model feature identification set, and train the copyright detection model to be trained based on the expanded model feature identification set.

[0081] Among them, the method of searching for subsets of homologous data is as follows: The computer device can divide the model feature identification set into k subsets, and regard each model feature identification in the model feature identification set as a data point; each subset includes a core point; the core point is randomly selected from the model feature identification set; determine the distance from each data point to each core point, and assign each data point to the subset to which the nearest core point belongs; update the core point of each subset to the mean value of all data points in the subset, and return to the step of determining the distance from each data point to each core point until the core point no longer changes or reaches the maximum number of iterations, and perform data augmentation processing on the data points in each subset to obtain the expanded model feature identification set 。

[0082] In some embodiments, the copyright detection model to be trained includes an autoencoder recognition network and a mapping network. The computer device can use the method of self-supervised learning to train the recognition network and the mapping network, and map the model feature identifications in the expanded model feature identification set to a low-dimensional representation space. The goal of self-supervised learning is to make the distance between similar data points in the representation space as close as possible, and the distance between different data points as far as possible. Specifically, divide the expanded model feature identification set into a training set and a test set, and use the method of self-supervised learning to train the recognition network so that the distance between similar data points in the representation space is as close as possible, and the distance between different data points is as far as possible, and verify the performance of the recognition network to ensure that it can accurately distinguish the feature identifications of different models. First, perform feature identification data preparation, and divide the expanded model feature identification set into a training set and a test set . The training set is used to train the recognition network and the mapping network, and the test set is used to verify the performance of the network. Suppose the expanded model feature identification set contains samples, and each sample is composed of the label of the feature point superimposed with the label of the dimensionality-reduced feature vector, that is , where 。

[0083] Further, the recognition network adopts an autoencoder structure, including two parts: an encoder and a decoder. The encoder inputs the high-dimensional feature identification data The representation vector mapped to a low dimension , the decoder restores the low-dimensional representation vector to high-dimensional data .

[0084] The encoder part consists of fully connected layers, with a ReLU activation function connected after each layer. Suppose the encoder contains layers, then the output of the th layer is expressed as: , where and are the weight matrix and bias vector of the th layer respectively, is the input data. The decoder part is symmetric to the encoder and also consists of fully connected layers, with a ReLU activation function connected after each layer. The output of the decoder is the restored high-dimensional data.

[0085] The mapping network consists of three identical sub-networks, which share weights and structures. Each sub-network receives a low-dimensional representation vector as input and outputs a new low-dimensional representation vector. These three sub-networks process anchor points, positive samples, and negative sample data respectively. Each sub-network consists of several fully connected layers, with a ReLU activation function connected after each layer. The self-supervised learning method is used to map the low-dimensional representation vector output by the recognition network to the final representation vector. The goal of self-supervised learning is to make the distance between data points of the same category in the representation space as close as possible, and the distance between data points of different categories as far as possible. Let the model feature identifier be , its positive pair is , and the negative pair is . The loss function of self-supervised learning can be expressed as:

[0086] ;

[0087] where represents the similarity degree of , and represents the similarity degree of .

[0088] can be represented by the cosine similarity degree:

[0089] , .

[0090] In the above intelligent grid deep learning model copyright protection method based on the model confidence region, by obtaining the target copyright protection model, as well as the target risk model and the target homomorphic model corresponding to the target copyright protection model, the target risk model is a model with an intelligent grid dataset infringement problem with respect to the target copyright protection model; the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain; and for each model in the intelligent grid deep model set, search for characteristic data samples in the intelligent grid proprietary dataset whose distance between the corresponding model prediction value and the corresponding model discrimination boundary satisfies a preset condition, to obtain a confidence region feature point set, so that the obtained confidence region feature point set can accurately describe the confidence region information of the deep learning model, and then reduce the dimension of the gradient vector of the confidence region feature point set on the model discrimination boundary according to the method of linear discriminant analysis, to obtain the perturbation vector of the confidence region feature point set, so that the prediction label changes before and after the combination of the confidence region feature point set and the perturbation vector can be used to accurately generate the model feature identification set corresponding to the intelligent grid deep model set, and train the copyright detection model to be trained through the model feature identification set to obtain a pre-trained copyright detection model; the pre-trained copyright detection model can be used to detect whether the model to be detected infringes the copyright of the target copyright protection model, and thus can effectively defend against attackers forging identities to replace users to participate in model training during the iteration after the start of training; at the same time, by extracting the confidence region features of the model to generate feature identifications, the risk model and the homomorphic model can be effectively distinguished, improving the accuracy and robustness of feature identification matching, ensuring the copyright security of the dataset deployed in the intelligent grid, and preventing data leakage and unauthorized use.

[0091] In some embodiments, the computer device can use the pre-trained copyright detection model to detect whether the model to be detected infringes the copyright of the target copyright protection model. First, obtain the model feature identifications corresponding to all models in the intelligent grid deep model set. For the model to be detected, generate the to-be-detected model feature identification corresponding to the model to be detected according to the above method of extracting model feature identifications. Match the generated to-be-detected model feature identification with the target copyright protection model feature identification, and calculate the similarity between the model feature identifications. The homomorphism degree can be expressed as the cosine similarity between the to-be-detected model feature identification and the model feature identification in the feature space. According to the cosine similarity between the to-be-detected model feature identification and the model feature identification of the target copyright protection model, judge whether the model to be detected infringes the data copyright of the target copyright protection model. If the similarity exceeds the threshold, it is judged that the model to be detected infringes the copyright of the target copyright protection model, realizing the protection of the copyright security of the dataset deployed in the intelligent grid and preventing data leakage and unauthorized use.

[0092] Furthermore, based on the degree of similarity between the model feature identifier of the target risk model and the model feature identifier of the target homomorphic model respectively and the model feature identifier to be detected, it can be determined whether the model to be detected belongs to the risk model or the homomorphic model, effectively distinguishing the risk model and the homomorphic model, and improving the accuracy and robustness of feature identifier matching.

[0093] In another embodiment, as Figure 2 shown, a method for copyright protection of deep learning models in smart grids based on model confidence regions is provided. Taking the application of this method to a computer device as an example, the method includes the following steps:

[0094] Step S202: Obtain a set of smart grid deep models. For each model in the set of smart grid deep models, search for characteristic data samples in the smart grid proprietary dataset whose distance between the corresponding model prediction value and the corresponding model discrimination boundary meets a preset condition, and obtain a set of confidence region characteristic points.

[0095] Step S204: Determine the gradient vector of the set of confidence region characteristic points on the model discrimination boundary, group the gradient vectors according to the class labels, and determine the mean vector of each class and the overall mean vector.

[0096] Step S206: Determine the within-class scatter matrix and the between-class scatter matrix according to the mean vector of each class and the overall mean vector.

[0097] Step S208: By solving the eigenvalues and eigenvectors of the product of the inverse matrix of the within-class scatter matrix and the between-class scatter matrix, select the eigenvectors whose cumulative contribution rate meets the preset condition for dimensionality reduction processing, obtain the dimensionality-reduced gradient vector, and use the dimensionality-reduced gradient vector as the perturbation vector of the set of confidence region characteristic points.

[0098] Step S210: Generate a set of model feature identifiers corresponding to the set of smart grid deep models according to the change in the predicted labels of each model before and after combining the set of confidence region characteristic points and the perturbation vector.

[0099] Step S212: Expand the set of model feature identifiers by means of searching for a subset of homologous data to obtain an expanded set of model feature identifiers.

[0100] Step S214: Train the copyright detection model to be trained through the expanded set of model feature identifiers.

[0101] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a method for copyright protection of deep learning models in smart grids based on model confidence regions described above.

[0102] In another embodiment, for the convenience of those skilled in the art to understand, Figure 3It provides a flowchart of yet another method for copyright protection of deep learning models in smart grids based on model confidence regions. As Figure 3 shown, the method for copyright protection of deep learning models in smart grids based on model confidence regions in this embodiment can be divided into four parts: S1, the preparation stage of the deep learning model set for smart grids. Generate a target risk model through a model extraction algorithm, and generate a target homomorphic model by fine-tuning the dataset and parameters, thereby building a deep learning model set for smart grids; S2, the generation stage of the confidence region feature identification set. Perform data augmentation on the proprietary dataset of smart grids, extract the model discrimination boundary to determine the confidence region feature point set, and obtain the reduced-dimensional gradient vector through linear discriminant analysis; S3, the expansion of the model feature identification set and the training stage of the recognition network. Expand the model feature identification set by searching for subsets of homologous data, and train the autoencoder recognition network and the contrastive learning mapping network; S4, the copyright detection stage of the deep learning model in smart grids. Generate the model feature identifications corresponding to all models in the deep learning model set for smart grids, and the recognition network calculates the spatial similarity of the model feature identifications to determine the privacy infringement threshold.

[0103] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a method for copyright protection of deep learning models in smart grids based on model confidence regions described above.

[0104] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0105] Based on the same inventive concept, the embodiments of the present application also provide a device for copyright protection of deep learning models in smart grids based on model confidence regions for implementing the above-mentioned method for copyright protection of deep learning models in smart grids based on model confidence regions. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the device for copyright protection of deep learning models in smart grids based on model confidence regions provided below can refer to the limitations of the method for copyright protection of deep learning models in smart grids based on model confidence regions described above, and will not be repeated here.

[0106] In an exemplary embodiment, as Figure 4 shown, there is provided an intelligent grid deep learning model copyright protection device based on a model confidence region, including: an acquisition module 410, a search module 420, a determination module 430, a generation module 440, and a training module 450, where:

[0107] The acquisition module 410 is configured to acquire a set of intelligent grid deep models; the set of intelligent grid deep models includes a target copyright protection model, as well as a target risk model and a target homomorphic model corresponding to the target copyright protection model; the target risk model is a model with an intelligent grid dataset infringement problem with respect to the target copyright protection model; the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain.

[0108] The search module 420 is configured to search, for each model in the set of intelligent grid deep models, in the intelligent grid proprietary dataset for characteristic data samples whose distance between the corresponding model prediction value and the corresponding model discrimination boundary satisfies a preset condition, to obtain a confidence region feature point set; the model discrimination boundary represents a region that can be predicted as other types within a preset displacement in the gradient direction of the model's inference on the characteristic data sample.

[0109] The determination module 430 is configured to determine the gradient vector of the confidence region feature point set on the model discrimination boundary, and perform dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, and use the dimensionality-reduced gradient vector as the perturbation vector of the confidence region feature point set.

[0110] The generation module 440 is configured to generate a model feature identification set corresponding to the set of intelligent grid deep models according to the change in the prediction labels of each model before and after combining the confidence region feature point set and the perturbation vector.

[0111] The training module 450 is configured to train a copyright detection model to be trained according to the model feature identification set to obtain a pre-trained copyright detection model; the pre-trained copyright detection model is used to detect whether a model to be detected infringes the copyright of the target copyright protection model.

[0112] In one of the embodiments, the determination module 430 is specifically configured to group the gradient vectors according to category labels, determine the mean vector of each category and the overall mean vector; determine the within-class scatter matrix and the between-class scatter matrix according to the mean vector of each category and the overall mean vector; by solving the eigenvalues and eigenvectors of the product of the inverse matrix of the within-class scatter matrix and the between-class scatter matrix, select the eigenvectors whose cumulative contribution rate satisfies a preset condition for dimensionality reduction processing to obtain the dimensionality-reduced gradient vector.

[0113] In one embodiment, the training module 450 is specifically configured to expand the model feature identification set by searching for a subset of homologous data, so as to obtain an expanded model feature identification set; and train the copyright detection model to be trained through the expanded model feature identification set.

[0114] In one embodiment, the training module 450 is specifically configured to divide the model feature identification set into k subsets, and regard each model feature identification in the model feature identification set as a data point; each of the subsets includes a core point; the core point is represented as being randomly selected from the model feature identification set; determine the distance from each data point to each core point, and assign each data point to the subset to which the nearest core point belongs; update the core point of each subset to the mean value of all data points within the subset, and return to the step of determining the distance from each data point to each core point until the core point no longer changes or reaches the maximum number of iterations, and perform data augmentation processing on the data points within each subset to obtain the expanded model feature identification set.

[0115] In one embodiment, the acquisition module 410 is specifically configured to acquire the training data set used in the training process of the target copyright protection model, and train the target homomorphic model on the training data set; the target homomorphic model includes a fully homomorphic model, an architecture homomorphic model, and a problem domain homomorphic model; the fully homomorphic model is obtained by performing multiple trainings on the premise that all experimental settings are the same as those of the target copyright protection model; the architecture homomorphic model is obtained by fine-tuning the hyperparameters during training and performing multiple trainings on the premise that the architecture is the same as that of the target copyright protection model; the problem domain homomorphic model is obtained by performing multiple trainings under the premise of using the same source data set as the target copyright protection model and under the combination of a preset similar model structure and randomly selected hyperparameters.

[0116] In one embodiment, the copyright detection model to be trained includes an autoencoder recognition network and a mapping network. The training module 450 is specifically configured to train the autoencoder recognition network and the mapping network by using a self-supervised learning method, and map the model feature identification set to a low-dimensional representation space.

[0117] If the model feature identification is , the positive pair of the model feature identification is , the negative pair of the model feature identification is , the loss function of self-supervised learning can be expressed as:

[0118]

[0119] Among them, represents the same attitude of represents the same attitude of

[0120] Each module in the above intelligent grid deep learning model copyright protection device based on the model confidence region can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0121] In an exemplary embodiment, a computer device is provided. The computer device can be an intelligent grid deep learning model copyright protection system based on the model confidence region, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store model feature identification set data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for protecting the copyright of an intelligent grid deep learning model based on the model confidence region.

[0122] Those skilled in the art can understand that Figure 5 the structure shown in

[0123] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0124] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0125] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0126] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0127] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, Artificial Intelligence (AI) processors, etc., without limitation.

[0128] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.

Claims

1. An intelligent grid deep learning model copyright protection method based on model confidence regions, characterized in that The method includes: Obtaining a set of deep models for the smart grid; the set of deep models for the smart grid includes a target copyright protection model, as well as a target risk model and a target homomorphic model corresponding to the target copyright protection model; the target risk model is a model with an infringement problem of the smart grid dataset with respect to the target copyright protection model; the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain; For each model in the set of deep models for the smart grid, searching in the smart grid proprietary dataset for characteristic data samples whose distance between the corresponding model prediction value and the corresponding model discrimination boundary satisfies a preset condition, to obtain a set of confidence domain characteristic points; the model discrimination boundary represents the area that can be predicted as other types within a preset displacement in the gradient direction of the model's inference on the characteristic data sample; Determining the gradient vector of the set of confidence domain characteristic points on the model discrimination boundary, and performing dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, and taking the dimensionality-reduced gradient vector as the perturbation vector of the set of confidence domain characteristic points; Generating a set of model characteristic identifiers corresponding to the set of deep models for the smart grid according to the change in the prediction labels of each model before and after combining the set of confidence domain characteristic points and the perturbation vector; Training a copyright detection model to be trained according to the set of model characteristic identifiers to obtain a pre-trained copyright detection model; the pre-trained copyright detection model is used to detect whether a model to be detected infringes the copyright of the target copyright protection model.

2. The method according to claim 1, wherein The determining the gradient vector of the set of confidence domain characteristic points on the model discrimination boundary, and performing dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, includes: Grouping the gradient vectors according to the class labels, and determining the mean vector of each class and the overall mean vector; Determining the within-class scatter matrix and the between-class scatter matrix according to the mean vector of each class and the overall mean vector; By solving the eigenvalues and eigenvectors of the product of the inverse matrix of the within-class scatter matrix and the between-class scatter matrix, selecting the eigenvectors whose cumulative contribution rate satisfies a preset condition for dimensionality reduction processing, to obtain the dimensionality-reduced gradient vector.

3. The method according to claim 1, wherein The training the copyright detection model to be trained according to the set of model characteristic identifiers includes: Expanding the set of model characteristic identifiers by means of searching for a subset of homologous data to obtain an expanded set of model characteristic identifiers; Training the copyright detection model to be trained through the expanded set of model characteristic identifiers.

4. The method according to claim 3, characterized in that, The expanding the set of model characteristic identifiers by means of searching for a subset of homologous data to obtain an expanded set of model characteristic identifiers includes: Dividing the set of model characteristic identifiers into k subsets, and taking each model characteristic identifier in the set of model characteristic identifiers as a data point; each subset includes a core point; the core point is randomly selected from the set of model characteristic identifiers; Determining the distance from each data point to each core point, and assigning each data point to the subset to which the nearest core point belongs; Update the core point of each of the subsets to the mean value of all the data points within that subset, and return the step of determining the distance from each of the data points to each of the core points until the core points no longer change or the maximum number of iterations is reached. Perform data augmentation processing on the data points within each of the subsets to obtain the extended model feature identification set.

5. The method according to claim 1, wherein The obtaining of the intelligent grid depth model set includes: Obtain the training data set used during the training of the target copyright protection model, and train the target homomorphic model on the training data set; the target homomorphic model includes a fully homomorphic model, an architecture homomorphic model, and a problem domain homomorphic model; The fully homomorphic model is obtained by performing multiple trainings on the premise that all experimental settings are the same as those of the target copyright protection model; The architecture homomorphic model is obtained by fine-tuning the hyperparameters during training and performing multiple trainings on the premise that the architecture is the same as that of the target copyright protection model; The problem domain homomorphic model is obtained by performing multiple trainings under a combination of a preset similar model structure and randomly selected hyperparameters on the premise of using the same source data set as the target copyright protection model.

6. The method according to claim 1, characterized in that, The copyright detection model to be trained includes an autoencoder recognition network and a mapping network. The training of the copyright detection model to be trained according to the model feature identification set includes: Train the autoencoder recognition network and the mapping network using a self-supervised learning method to map the model feature identification set to a low-dimensional feature space; If the model feature identifier is , the positive pair of the model feature identifier is , the negative pair of the model feature identifier is , and the loss function of self-supervised learning can be expressed as: ; Among them, represents 's degree of agreement, represents 's degree of agreement.

7. An intelligent grid deep learning model copyright protection device based on a model confidence region, characterized in that, The device includes: An obtaining module, configured to obtain an intelligent grid depth model set; the intelligent grid depth model set includes a target copyright protection model, as well as a target risk model and a target homomorphic model corresponding to the target copyright protection model; the target risk model is a model with an intelligent grid data set infringement problem with respect to the target copyright protection model; the target homomorphic model is a model obtained by simulating the target copyright protection model in the same problem domain; A searching module, configured to search, for each model in the intelligent grid depth model set, for a feature data sample in the intelligent grid proprietary data set whose distance between the corresponding model prediction value and the corresponding model discrimination boundary satisfies a preset condition, to obtain a confidence domain feature point set; the model discrimination boundary represents a region that can be predicted as another type within a preset displacement in the gradient direction of the model's inference on the feature data sample; A determining module, configured to determine the gradient vector of the confidence domain feature point set on the model discrimination boundary, and perform dimensionality reduction on the obtained gradient vector according to the method of linear discriminant analysis, and use the gradient vector after dimensionality reduction as the perturbation vector of the confidence domain feature point set; A generating module, configured to generate a model feature identification set corresponding to the intelligent grid depth model set according to the change in the prediction labels of each model before and after combining the confidence domain feature point set and the perturbation vector; A training module, configured to train a copyright detection model to be trained according to the model feature identification set, so as to obtain a pre-trained copyright detection model; the pre-trained copyright detection model is used to detect whether a model to be detected infringes on the copyright of the target copyright protection model.

8. An intelligent power grid deep learning model copyright protection system based on a model confidence region, including a memory and a processor, where the memory stores a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method for protecting deep learning model based on confidential computing

    US11886554B1

  • Video detection method and apparatus, and device, storage medium and program product

    WO2023221634A1