Model determination method and device, equipment, medium and product
By obtaining the multimodal training set and generating the trigger set, and training the deep neural network with a preset encryption algorithm, the watermark integrity problem in the multimodal network is solved and the model ownership verification ability is enhanced.
Patent Information
- Application Number
- CN202510414958.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
In multimodal deep neural networks, the prior art is difficult to ensure the integrity of watermarks in key integration layers, resulting in weakening of model ownership verification capabilities and watermarks are prone to tampering.
By obtaining the multimodal training set and generating the multimodal trigger set, training the training model with a preset encryption algorithm, generating a watermark model, and evaluating it to determine the ownership verification results, ensuring that the watermark maintains integrity under different modal types.
Maintaining the integrity of the watermark under different types of data prevents unauthorized tampering, enhancing the model's ability to verify ownership, ensuring that the watermark remains intact and is detected throughout the entire life of the model.
Smart Images

Figure CN120296709A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model security protection, and in particular, to a method, device, equipment, medium and product for determining a model. Background Art
[0002] Deep neural networks (DNNs) play a crucial role in connected and autonomous driving systems, enabling vehicles to interpret and respond to complex environments. DNNs are a class of machine learning algorithms designed to identify patterns and learn from data. They consist of multiple layers of interconnected neurons that can capture intricate features and relationships in the data. DNNs can process large and complex datasets and make accurate predictions and decisions, thus revolutionizing various fields.
[0003] The prior art usually uses the method of embedding watermarks in the model to verify the ownership of the model and avoid the abuse of DNN models and related malicious activities.
[0004] However, as the complexity of data increases, multi-modal data has emerged compared to single-modal data, which contains various types of data, such as text, images, audio, and video at the same time. In a multi-modal network, a specific layer is responsible for integrating information from different modalities. At this time, it is impossible to ensure the integrity of the watermark in these key integration layers, thus weakening the ability to verify ownership. DNN models often use datasets for fine-tuning to improve their performance in specific activities or environments, but this process may inadvertently damage or weaken any embedded watermark, thus weakening the ability to verify ownership, and the embedded watermark may be tampered with, resulting in the invalidation of the watermark and affecting the ownership of the DNN model. Summary of the Invention
[0005] The present invention provides a method, device, equipment, medium and product for determining a model to ensure the integrity of the watermark embedded in the model.
[0006] According to the first aspect of the present invention, a method for determining a model is provided, including:
[0007] Obtain a multi-modal training set and generate a multi-modal trigger set;
[0008] Train the model to be trained according to the multi-modal training set, the multi-modal trigger set and a preset encryption algorithm to obtain a watermark model;
[0009] Evaluate the watermark model, determine the ownership verification result of the watermark model and obtain a final model.
[0010] According to the second aspect of the present invention, a device for determining a model is provided, including:
[0011] A data acquisition module, configured to acquire a multi-modal training set and generate a multi-modal trigger set;
[0012] A model training module, configured to train a model to be trained according to the multi-modal training set, the multi-modal trigger set and a preset encryption algorithm to obtain a watermark model;
[0013] A model determination module, configured to evaluate the watermark model, determine the ownership verification result of the watermark model and obtain a final model.
[0014] According to a third aspect of the present invention, there is provided an electronic device, including:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the model determination method according to any embodiment of the present invention.
[0018] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for enabling a processor to implement the model determination method according to any embodiment of the present invention when executed.
[0019] According to a fifth aspect of the present invention, an embodiment of the present invention further provides a computer program product including a computer program, and when the computer program is executed by a processor, it implements the model determination method according to any embodiment of the present invention.
[0020] The technical solution of the embodiment of the present invention is to acquire a multi-modal training set and generate a multi-modal trigger set; train a model to be trained according to the multi-modal training set, the multi-modal trigger set and a preset encryption algorithm to obtain a watermark model; evaluate the watermark model, determine the ownership verification result of the watermark model and obtain a final model. By combining the multi-modal training set and the trigger set with a preset encryption algorithm for model training, the model training requirements for different types of data are met, the integrity of the watermark can still be maintained under different modal types, and unauthorized watermark tampering is prevented, thereby enhancing the ability of the model to verify ownership.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0023] Figure 1 is a flowchart of a method for determining a model provided in Embodiment 1 of the present invention;
[0024] Figure 2 is a flowchart of a method for determining a model provided in Embodiment 2 of the present invention;
[0025] Figure 3 is a schematic structural diagram of a device for determining a model provided in Embodiment 3 of the present invention;
[0026] Figure 4 is a schematic structural diagram of an electronic device implementing the embodiments of the present invention. Detailed Embodiments
[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0029] Embodiment 1
[0030] Figure 1The following is a flowchart of a method for determining a model provided by Embodiment 1 of the present invention. This embodiment is applicable to the training of a model for determining ownership function. This method can be executed by a model determination device, which can be implemented in the form of hardware and / or software, and the model determination device can be configured in an electronic device. As Figure 1 shown, the method includes:
[0031] S110. Obtain a multimodal training set and generate a multimodal trigger set.
[0032] In this embodiment, the multimodal training set can be understood as a set containing multiple training samples, and each training sample contains data of at least one modality type. The multimodal trigger set can be understood as the trigger set corresponding to different modality types, where the trigger set can be understood as the part for adding a watermark to the model.
[0033] Specifically, the processor can obtain the multimodal training set, and generate a corresponding trigger set for the data under each modality type through the multimodal training set to form a multimodal trigger set to meet the requirements of different modality types.
[0034] Exemplarily, a standard dataset can be used, such as CIFAR-10 for image classification, or a multimodal dataset combining text and images, such as MNIST for the image modality and Free Spoken Digits for the audio modality, etc.
[0035] S120. Train the model to be trained according to the multimodal training set, the multimodal trigger set, and a preset encryption algorithm to obtain a watermark model.
[0036] In this embodiment, the preset encryption algorithm can be understood as an algorithm for model encryption, such as the Advanced Encryption Standard (AES), etc. The watermark model can be understood as the model obtained after embedding the watermark.
[0037] Specifically, the processor can mix the training samples in the multimodal training set with the samples in the multimodal trigger set, input them into the model to be trained for training to obtain an initial model that meets the training end condition, and can also fine-tune the watermark model with additional data to obtain a fine-tuned model, and encrypt a specific layer of the fine-tuned model through the preset encryption algorithm to obtain a watermark model to ensure the security of the watermark and at the same time ensure the function of the model.
[0038] S130. Evaluate the watermark model, determine the ownership verification result of the watermark model, and obtain the final model.
[0039] In this embodiment, the ownership verification result can be understood as the conclusion or state obtained from the process of determining the ownership of the watermark model. The final model can be understood as the model after passing the verification.
[0040] Specifically, when the training ends, the processor can evaluate the watermark model, unlock the model with the key, and trigger the trigger set included therein to verify the trigger accuracy of the watermark, thereby obtaining the ownership verification result of the watermark model, so as to determine the performance of the watermark model and obtain a high-precision watermark model as the final model.
[0041] The technical solution of the embodiment of the present invention is to obtain a multi-modal training set and generate a multi-modal trigger set; train the model to be trained according to the multi-modal training set, the multi-modal trigger set and the preset encryption algorithm to obtain a watermark model; evaluate the watermark model, determine the ownership verification result of the watermark model and obtain the final model. By combining the multi-modal training set and the trigger set with the preset encryption algorithm for model training, the model training requirements of different types of data are met, the integrity of the watermark can still be maintained under different modal types, and unauthorized watermark tampering is prevented, thereby enhancing the ability of the model to verify ownership.
[0042] Embodiment 2
[0043] Figure 2 The following is a flowchart of a method for determining a model provided in Embodiment 2 of the present invention. This embodiment is a further refinement between the above embodiments. As Figure 2 shown, the method includes:
[0044] S201. Obtain a multi-modal training set and randomly select training samples from the multi-modal training set.
[0045] In this embodiment, the training sample can be understood as the data used to train the model, which may include data under at least one modal type. For example, a training sample may include x images, y audios, and z text data, or only include data under one modal type.
[0046] Specifically, the processor can obtain the multi-modal training set and randomly select training samples from the multi-modal training set to generate subsequent trigger sets.
[0047] S202. For the modal type of the training sample, generate trigger samples according to the specific embedding content, training sample, and noise under the modal type.
[0048] In this embodiment, the modality type can be understood as the type of data, such as images, text, audio, video, etc. The specific embedded content can be understood as the relevant content for adding watermarks, such as identification information, copyright information, and source information, etc. Noise can be understood as the added random or pseudo-random data used to interfere with potential piracy, illegal copying, or unauthorized use, and can also ensure the authenticity and integrity of the content. Among them, the noise is also related to the modality type, such as adding extra bytes in text, changing pixel values in images, or adding background noise in audio, etc. The trigger sample can be understood as the key data, events, or conditions for activating or starting the watermark embedding, detection, extraction, or verification process under different modalities.
[0049] Specifically, for different modality types corresponding to the training samples, the processor can add the specific embedded content and noise corresponding to the modality type to the original data of the training samples to generate the trigger sample corresponding to this training sample.
[0050] S203. Generate a trigger set based on the trigger samples under each modality type.
[0051] Specifically, the processor can synthesize the trigger samples under different modality types to generate a trigger set containing multiple modality types, or each modality type can separately correspond to a trigger set.
[0052] S204. Mix the training samples in the multi-modal training set with the trigger samples in the multi-modal trigger set to obtain a mixed training sample set.
[0053] In this embodiment, the mixed training sample set can be understood as a set containing training samples and trigger samples.
[0054] Specifically, the processor can randomly mix the trigger samples into the training samples to obtain a mixed training sample set.
[0055] S205. Train the model to be trained through the mixed sample set to obtain an initial watermark model with embedded watermarks.
[0056] In this embodiment, the model to be trained can be understood as the model that needs to be trained, and can include models in the following scenarios: for example, the initial model when the model is initially constructed, the model to be verified when the model owner suspects that the model has been stolen, and the model when the existing model is modified or adjusted later. The initial watermark model can be understood as the initial watermark model obtained after integrating the watermark into the model parameters.
[0057] Specifically, the processor can input the samples of the mixed sample set into the model to be trained for training until the training end requirement is met, and obtain an initial watermark model embedded with a watermark. This enables the model to learn and recognize two types of data, and this process ensures that the model integrates the watermark into its parameters, thereby effectively embedding the trigger set into the learned representations.
[0058] S206. Encrypt the layer to be encrypted in the initial watermark model according to a preset encryption algorithm to obtain an intermediate watermark model.
[0059] In this embodiment, the selection of the layer to be encrypted involves selecting a specific layer from a trained watermark model, which has a significant impact on the learning process but is not important for the main performance indicators of the model. Usually, layers such as convolutional layers are selected because of their influence on the model behavior.
[0060] Exemplarily, a DNN for image classification usually consists of a multi-layer structure, and these layers gradually transform the input data. The architecture can vary according to the specific problem, but common structures include the following layers: Input layer: This layer receives image data, and each input image is represented as a matrix of pixel values. Convolutional layers: These layers apply convolutional filters to the input image to detect features such as edges, textures, and patterns. Activation layer: After convolution or linear transformation, the activation layer introduces non-linearity, enabling the network to learn more complex patterns. Pooling layers: These layers reduce the spatial dimensions (height and width) of the feature map, reducing the number of parameters and computational complexity while retaining key features. Fully connected layers: After feature extraction through convolution and pooling, these layers are responsible for classifying the image based on the learned features. Output layer: The last layer that generates the predicted class label for the input image.
[0061] Among them, the layer to be encrypted is any at least one layer in the convolutional layer. Although encryption brings some complexity, it is limited to a single layer, so it is controllable, and it enhances the security and robustness of the model, making the solution both practical and effective. There are several important reasons for applying watermark and encryption technologies to the convolutional layer of a neural network: Embedding the watermark into the model itself helps protect the ownership of the trained model. Applying encryption to the data processed in the convolutional layer can ensure privacy. Encrypting the data in the convolutional layer can ensure that the information is protected during transmission. Encryption can prevent the leakage of sensitive model parameters. It helps with integrity verification.
[0062] In this embodiment, the intermediate watermark model can be understood as an encrypted watermark model.
[0063] Specifically, traditional watermarks often cannot maintain their integrity during the fine-tuning and incremental training processes. Therefore, the processor can encrypt the layer to be encrypted in the initial watermark model according to a preset encryption algorithm to ensure that the watermark can remain complete and be detected even after the model is modified, thereby enhancing the robustness.
[0064] S207. Train the intermediate watermark model with the obtained incremental dataset to obtain a watermark model.
[0065] In this embodiment, the incremental dataset can be understood as a set of updated data samples.
[0066] Specifically, the processor can use the data in the incremental dataset as the input of the intermediate watermark model and continuously train and learn with this latest data to obtain a watermark model.
[0067] Among them, based on the above embodiment, the step of training the intermediate watermark model with the obtained incremental dataset to obtain a watermark model can be refined as:
[0068] Train the unencrypted layers in the intermediate watermark model with the obtained incremental dataset to obtain a watermark model.
[0069] Specifically, since one of the convolutional layers in the intermediate watermark model has been encrypted, during the incremental training process, the encrypted layer will retain the watermark in all modalities. The processor inputs the incremental data in the incremental dataset into the intermediate watermark model and continues to learn the incremental data through the unencrypted layers. The watermark embedded in the encrypted layer remains secure and unchanged, ensuring the integrity and ownership of the model even when the model learns from new data.
[0070] S208. Unlock the watermark model with a key to obtain an unlocked model.
[0071] In this embodiment, the key is used for decryption. The unlocked model can be understood as the model after unlocking.
[0072] Specifically, when encrypting, a key for decryption will be sent to the personnel related to the model to perform decryption through the key to realize the original function of the model. The processor can decrypt the watermark model with the key to enable the model to run normally and obtain an unlocked model.
[0073] S209. Extract the trigger samples in the unlocked model and check the status of the trigger samples to obtain an ownership verification result.
[0074] In this embodiment, the status of the trigger samples is used to reflect that the watermark is complete and valid.
[0075] Specifically, the processor can extract the trigger samples in the unlocked model. By evaluating the response of the unlocked model to the trigger samples, it can be confirmed whether the watermark is in an active state, that is, to determine the existence and integrity of the watermark, so as to obtain an ownership verification result.
[0076] S210. When the ownership verification result is valid, obtain the final model.
[0077] Specifically, when the ownership verification result is valid, obtain the final model that has not been stolen and has ownership. When the ownership verification result is invalid, it indicates that the unlocking model has been stolen.
[0078] Exemplarily, an example is given in a specific scenario. When an adversary steals a model, fine-tunes it with their own custom dataset, and deploys it as a public service. For simplicity, assume that the adversary's goal is to fine-tune the stolen model without changing its layers and use this modified model as a service. The present application serves as an owner verification process: when the original model owner suspects that the model has been stolen, they can use the processor to access the suspicious model through the API. Then, the owner can retrain the model using the original training data and evaluate its performance on the encrypted trigger dataset by unlocking it with a key to determine the ownership verification result and perform watermark verification. This retraining step ensures a fair evaluation because only if the model was initially watermarked will its trigger accuracy increase and the verification result be accurate.
[0079] The technical solution of the embodiment of the present invention ensures that the watermark can be effectively embedded in all types of data used by the model by developing different trigger sets for each modality in the multimodal network. By training the model to be trained with a mixed sample set, enabling the model to be trained to learn and identify two types of data, it can ensure that the watermark is integrated into the model parameters and effectively inserted into the representations it has learned. This integration can retain the presence of the watermark and maintain its integrity throughout the life cycle of the model (including incremental training and subsequent decryption processes). By presetting an encryption algorithm to encrypt the layer to be encrypted in the initial watermark model and encrypting a layer that integrates all mode information, unauthorized tampering is prevented, ensuring that the watermark remains intact and detectable even after the model is modified, thereby enhancing the robustness. By decrypting the encrypted layer of the model with a key, it can be ensured that the decrypted layer does not damage the accuracy of the model, and thus the embedded watermark can not only provide strong protection for the model but also not affect the main functions and test accuracy of the model, enhancing the model's ability to verify ownership.
[0080] Exemplarily, a specific example is presented. The model architecture of the model to be trained selects common deep neural network architectures, such as ResNet for image classification, or multimodal architectures that can handle text and images. The model is trained by combining conventional multimodal training data and multimodal trigger sets embedded with watermarks. After initial training, additional data is used for fine-tuning and incremental training to evaluate the robustness of the watermark. The preset encryption algorithm uses the Advanced Encryption Standard (AES) to encrypt specific layers of the trained watermark model to ensure the security of the watermark while maintaining the functionality of the model. The encrypted layer is decrypted with a key to obtain an unlocked model to evaluate the performance of the model on test data and trigger sets. The experimental results are shown in the following table:
[0081] Table 1 Validation Results Table under Different Incremental Datasets and Different Encryption Layers
[0082]
[0083]
[0084] Among them, the third column represents the watermark / trigger accuracy (WA), which is used to measure the performance of the model on the watermark trigger set and indicates the degree of preservation and detection of the watermark. Baseline: Before any modification, the model should be able to accurately identify a high-precision trigger set. For example, if the watermark accuracy (WA) is 95%, it means the model can reliably identify the watermark. After fine-tuning and incremental training: The model maintains a high WA value, showing robustness. If the WA after fine-tuning is 92%, there is only a slight decrease, indicating that the watermark remains basically intact. If the WA value after incremental training is 90%, it means that the watermark can still be retained even after further training. As can be seen from Table 1, the watermark accuracy is very high and can be maintained in different incremental data ratios and different encryption layers.
[0085] Table 2 Results under Different Trigger Set Types and Encryption Layers
[0086]
[0087] Among them, the fourth column represents the test accuracy (TA), which is used to evaluate the performance of the model when completing the main task using the test dataset and reflects the overall functionality and effectiveness of the model. Baseline: The initial test accuracy of the model on its main task may be 85%, which is a reference point. After fine-tuning / enhanced training / encryption: The test accuracy remains unchanged, such as 83%, to ensure that the practicality of the model is not greatly reduced due to the encryption and watermark processes. The test accuracy is still similar to the baseline, such as 84%, indicating that the encryption process has little or no impact on the model performance. As can be seen from Table 2, under different trigger types, the watermark and test accuracy can be maintained unchanged.
[0088] Example 3
[0089] Figure 3 This is a schematic structural diagram of a model determination device provided in Embodiment 3 of the present invention. As Figure 3 shown, the device includes:
[0090] A data acquisition module 31, configured to acquire a multi-modal training set and generate a multi-modal trigger set;
[0091] A model training module 32, configured to train a model to be trained according to the multi-modal training set, the multi-modal trigger set, and a preset encryption algorithm to obtain a watermark model;
[0092] A model determination module 33, configured to evaluate the watermark model, determine the ownership verification result of the watermark model, and obtain a final model.
[0093] The technical solution of the embodiment of the present invention is to acquire a multi-modal training set and generate a multi-modal trigger set; train a model to be trained according to the multi-modal training set, the multi-modal trigger set, and a preset encryption algorithm to obtain a watermark model; evaluate the watermark model, determine the ownership verification result of the watermark model, and obtain a final model. By combining the multi-modal training set and the trigger set with a preset encryption algorithm for model training, the model training requirements for different types of data are met, the integrity of the watermark can still be maintained under different modal types, and unauthorized watermark tampering is prevented, thereby enhancing the ability of the model to verify ownership.
[0094] Further, the model training module 32 includes:
[0095] A first determination unit, configured to mix the training samples in the multi-modal training set with the trigger samples in the multi-modal trigger set to obtain a mixed training sample set;
[0096] A second determination unit, configured to train the model to be trained through the mixed sample set to obtain an initial watermark model with an embedded watermark;
[0097] A third determination unit, configured to encrypt the layer to be encrypted in the initial watermark model according to a preset encryption algorithm to obtain an intermediate watermark model;
[0098] A fourth determination unit, configured to train the intermediate watermark model through the obtained incremental data set to obtain a watermark model.
[0099] Wherein, the layer to be encrypted is any at least one layer in the convolutional layer.
[0100] Wherein, the fourth determination unit is specifically configured to:
[0101] Train the unencrypted layer in the intermediate watermark model through the obtained incremental data set to obtain a watermark model.
[0102] Further, the model determination module 33 is specifically configured to:
[0103] Unlock the watermark model with a key to obtain an unlocked model;
[0104] Extract trigger samples from the unlocked model and check the status of the trigger samples to obtain an ownership verification result;
[0105] When the ownership verification result is valid, obtain a final model.
[0106] Further, the data acquisition module 31 is specifically configured to:
[0107] Randomly select training samples from the multimodal training set;
[0108] For the modal type of the training samples, generate trigger samples according to the specific embedding content, the training samples, and noise under the modal type;
[0109] Generate a trigger set based on the trigger samples under each modal type.
[0110] The model determination device provided by the embodiments of the present invention can execute the model determination method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0111] Embodiment 4
[0112] Figure 4 FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0113] As Figure 4As shown, the electronic device 40 includes at least one processor 41 and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0114] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disc, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0115] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the method for determining a model.
[0116] In some embodiments, the method for determining a model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the method for determining a model described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the method for determining a model in any other appropriate manner (e.g., by means of firmware).
[0117] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0118] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0119] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0121] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0122] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0123] In one embodiment, the embodiment of the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the model determination method of any embodiment of the present invention.
[0124] In the process of implementing the computer program product, computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0125] It should be understood that the various forms of the flow shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0126] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining a model, characterized in that, Including: Obtain a multi-modal training set and generate a multi-modal trigger set; Train the model to be trained according to the multi-modal training set, the multi-modal trigger set and a preset encryption algorithm to obtain a watermark model; Evaluate the watermark model, determine the ownership verification result of the watermark model and obtain a final model.
2. The method according to claim 1, characterized in that The step of training the model to be trained according to the multi-modal training set, the multi-modal trigger set and a preset encryption algorithm to obtain a watermark model includes: Mix the training samples in the multi-modal training set with the trigger samples in the multi-modal trigger set to form a mixed training sample set; Train the model to be trained through the mixed sample set to obtain an initial watermark model with embedded watermarks; Encrypt the layer to be encrypted in the initial watermark model according to a preset encryption algorithm to obtain an intermediate watermark model; Train the intermediate watermark model through the obtained incremental data set to obtain a watermark model.
3. The method according to claim 2, wherein The layer to be encrypted is any at least one layer in the convolutional layer.
4. The method according to claim 2, wherein The step of training the intermediate watermark model through the obtained incremental data set to obtain a watermark model includes: Train the unencrypted layer in the intermediate watermark model through the obtained incremental data set to obtain a watermark model.
5. The method according to claim 1, characterized in that The step of evaluating the watermark model, determining the ownership verification result of the watermark model and obtaining a final model includes: Unlock the watermark model with a key to obtain an unlocked model; Extract the trigger samples in the unlocked model and check the status of the trigger samples to obtain an ownership verification result; When the ownership verification result is valid, obtain a final model.
6. The method according to claim 1, wherein The step of obtaining a multi-modal training set and generating a multi-modal trigger set includes: Randomly select training samples from the multi-modal training set; For the modal type of the training samples, generate trigger samples according to the specific embedded content, the training samples and noise under the modal type; Generate a trigger set based on the trigger samples under each modal type.
7. An apparatus for determining a model, characterized in that, Including: A data acquisition module for obtaining a multi-modal training set and generating a multi-modal trigger set; A model training module for training the model to be trained according to the multi-modal training set, the multi-modal trigger set and a preset encryption algorithm to obtain a watermark model; A model determination module for evaluating the watermark model, determining the ownership verification result of the watermark model and obtaining a final model.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the model determination method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to execute the model determination method according to any one of claims 1-6 when executed.
10. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the method for determining the model according to any one of claims 1-6.