Data processing, model training, detection, task processing methods, equipment and media

By alternately training the first deep learning model and the second deep learning model, a pre-trained protein model carrying watermark information is constructed, which solves the problem of insufficient security of black-box watermarking technology in protein pre-trained models and achieves the security of the model and the effectiveness of the encoding vector.

CN116386730BActive Publication Date: 2026-06-02BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
Filing Date
2023-02-23
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing black-box watermarking techniques cannot effectively protect the security of protein pre-trained models, especially in the absence of target categories in self-supervised learning, making the models vulnerable to theft or attack.

Method used

By alternately training the first and second deep learning models, a pre-trained model of watermarked proteins carrying watermark information is constructed using the sample protein dataset. This ensures that the training process does not affect each other, thereby achieving the security and effectiveness of the model.

Benefits of technology

This improves the security of protein pre-trained models and the effectiveness of encoding vectors, ensuring that the model's copyright and data privacy are not violated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386730B_ABST
    Figure CN116386730B_ABST
Patent Text Reader

Abstract

This disclosure provides a data processing, model training, detection, and task processing method, device, and medium, relating to the field of artificial intelligence technology, and particularly to the field of deep learning technology. The specific implementation scheme is as follows: acquiring first target protein data; processing the first target protein data using a watermarked protein pre-trained model to obtain a first target encoding vector; the watermarked protein pre-trained model is a trained first deep learning model, and the trained first deep learning model and the trained second deep learning model are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset, the training processes of the first deep learning model and the second deep learning model influencing each other; the sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the field of deep learning technology. Specifically, it relates to a data processing, model training, detection, and task processing method, apparatus, and medium. Background Technology

[0002] With the development of computer technology, artificial intelligence (AI) technology has also developed. AI technology can include computer vision, speech recognition, natural language processing, machine learning, deep learning, big data processing, and knowledge graph technology.

[0003] Artificial intelligence technology has been widely applied in various fields. For example, it can be used to process protein data. Summary of the Invention

[0004] This disclosure provides a data processing, model training, detection, and task processing method, device, and medium.

[0005] According to one aspect of this disclosure, a data processing method is provided, comprising: acquiring first target protein data; and processing the first target protein data using a watermarked protein pre-trained model to obtain a first target encoding vector; wherein the watermarked protein pre-trained model is a first deep learning model that has been trained, and a second deep learning model that has been trained is used to detect the relationship between a model to be detected and the watermarked protein pre-trained model; the first deep learning model and the second deep learning model that have been trained are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset, and the training processes of the first deep learning model and the second deep learning model influence each other; wherein the sample protein dataset includes a first sample protein dataset, and the first sample protein dataset is a predetermined privacy dataset.

[0006] According to another aspect of this disclosure, a model detection method is provided, comprising: processing a detection protein dataset using a model to be detected to obtain a sixth target encoding vector set, wherein the detection protein dataset is obtained based on a first sample protein dataset, predetermined mask data, and a key tuple; processing the sixth target encoding vector set using the key tuple to obtain a target decoding vector set; and determining detection information of the model to be detected based on a reference vector and the target decoding vector set, wherein the detection information characterizes the relationship between the model to be detected and a watermarked protein pre-trained model; wherein the key tuple is a trained second deep learning model, the watermarked protein pre-trained model is a trained first deep learning model, the trained second deep learning model and the trained first deep learning model are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset, and the training processes of the first deep learning model and the second deep learning model influence each other; wherein the sample protein dataset includes the first sample protein dataset, and the first sample protein dataset is a predetermined privacy dataset.

[0007] According to one aspect of this disclosure, a method for training a pre-trained model is provided, comprising: alternately training a first deep learning model and a second deep learning model using a sample protein dataset, such that the training processes of the first deep learning model and the second deep learning model influence each other; and determining the trained first deep learning model as a watermarked protein pre-trained model; wherein the trained second deep learning model is used to detect the relationship between a model to be detected and the watermarked protein pre-trained model; wherein the sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset.

[0008] According to another aspect of this disclosure, a training method for a protein task model is provided, comprising: training a watermarked protein pre-training model using a fourth sample protein dataset to obtain a protein task model; wherein the watermarked protein pre-training model is trained using a pre-training model training method.

[0009] According to another aspect of this disclosure, a protein task processing method is provided, comprising: inputting second target protein data into a protein task model to obtain output information; wherein the protein task model is trained using a protein task model training method.

[0010] According to another aspect of this disclosure, a data processing apparatus is provided, comprising: an acquisition module for acquiring first target protein data; and a first obtaining module for processing the first target protein data using a watermarked protein pre-trained model to obtain a first target encoding vector; wherein the watermarked protein pre-trained model is a trained first deep learning model, and a trained second deep learning model is used to detect the relationship between a model to be detected and the watermarked protein pre-trained model; the trained first deep learning model and the trained second deep learning model are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset, and the training processes of the first deep learning model and the second deep learning model influence each other; wherein the sample protein dataset includes a first sample protein dataset, and the first sample protein dataset is a predetermined privacy dataset.

[0011] According to another aspect of this disclosure, a model detection apparatus is provided, comprising: a second obtaining module, configured to process a detection protein dataset using a model to be detected to obtain a sixth target encoding vector set, wherein the detection protein dataset is obtained based on a first sample protein dataset, predetermined mask data, and a key tuple; a third obtaining module, configured to process the sixth target encoding vector set using the key tuple to obtain a target decoding vector set; and a determining module, configured to determine detection information of the model to be detected based on a reference vector and the target decoding vector set, wherein the detection information characterizes the relationship between the model to be detected and a watermarked protein pre-trained model; wherein the key tuple is a trained second deep learning model, the watermarked protein pre-trained model is a trained first deep learning model, the trained second deep learning model and the trained first deep learning model are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset, and the training processes of the first deep learning model and the second deep learning model influence each other; wherein the sample protein dataset includes the first sample protein dataset, and the first sample protein dataset is a predetermined privacy dataset.

[0012] According to another aspect of this disclosure, a training apparatus for a pre-trained model is provided, comprising: a first training module for alternately training a first deep learning model and a second deep learning model using a sample protein dataset, such that the training processes of the first deep learning model and the second deep learning model influence each other; and a second determining module for determining the trained first deep learning model as a watermarked protein pre-trained model; wherein the trained second deep learning model is used to detect the relationship between a model to be detected and the watermarked protein pre-trained model; wherein the sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset.

[0013] According to another aspect of this disclosure, a training apparatus for a protein task model is provided, comprising: a second training module for training a watermarked protein pre-training model using a fourth sample protein dataset to obtain a protein task model; wherein the watermarked protein pre-training model is trained using a training apparatus for a pre-training model.

[0014] According to another aspect of this disclosure, a protein task processing apparatus is provided, comprising: an input module for inputting second target protein data into a protein task model to obtain output information; wherein the protein task model is trained using a protein task model training device.

[0015] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described in this disclosure.

[0016] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods described in this disclosure.

[0017] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.

[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0019] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0020] Figure 1 The illustration schematically shows an exemplary system architecture in which data processing methods, model detection methods, pre-trained model training methods, protein task model training methods, protein task processing methods and apparatus can be applied according to embodiments of the present disclosure.

[0021] Figure 2 A flowchart illustrating a data processing method according to an embodiment of the present disclosure is shown schematically.

[0022] Figure 3 A flowchart illustrating a model detection method according to an embodiment of the present disclosure is shown schematically.

[0023] Figure 4 A flowchart illustrating a training method for a pre-trained model according to an embodiment of the present disclosure is shown schematically.

[0024] Figure 5 This illustration schematically shows an example of a method for obtaining a second sample protein dataset for the i-th round based on a first sample protein dataset, predetermined mask data, and a second encoder for the (i-1)-th round, according to an embodiment of the present disclosure.

[0025] Figure 6 The illustration shows an example of a method according to an embodiment of the present disclosure for training a second encoder and decoder for the (i-1)th round using a reference vector, a third sample protein dataset, a first encoder for the (i-1)th round, a decoder for the (i-1)th round, and a second sample protein dataset for the i-th round to obtain a second encoder and decoder for the i-th round.

[0026] Figure 7A The illustration schematically shows an example of a method for adjusting the model parameters of the second encoder and decoder in the (i-1)th round according to an embodiment of the present disclosure, based on a reference vector, a first decoding vector set in the i-th round, a third decoding vector set in the i-th round, a second decoding vector set in the i-th round, and a fourth decoding vector set in the i-th round, to obtain the second encoder and decoder in the i-th round.

[0027] Figure 7B The illustration shows an example diagram of a method according to an embodiment of the present disclosure for obtaining the value of the second loss function in the i-th round based on a second loss function, a reference vector, a second decoding vector set in the i-th round, and a fourth decoding vector set in the i-th round;

[0028] Figure 7C The illustration shows an example diagram of a method according to an embodiment of the present disclosure for obtaining the value of the first loss function in the i-th round based on a first loss function, a reference vector, a first decoding vector set in the i-th round, and a third decoding vector set in the i-th round.

[0029] Figure 8The illustration shows an example of a method according to an embodiment of the present disclosure for training the first encoder in the (i-1)th round using a first sample protein dataset, a third sample protein dataset, a reference vector, a first encoder in the (i-1)th round, and a decoder in the i-th round to obtain the first encoder in the i-th round.

[0030] Figure 9A The illustration shows an example of a method according to an embodiment of the present disclosure, which adjusts the model parameters of the first encoder in the (i-1)th round based on a reference vector, a first encoding vector set corresponding to the third sample protein dataset in the i-th round, a third encoding vector set corresponding to the third sample protein dataset, a fifth decoding vector set corresponding to the first encoding vector set, and a sixth decoding vector set corresponding to the fifth encoding vector set, to obtain the first encoder method in the i-th round.

[0031] Figure 9B The illustration shows an example of a method for obtaining the value of the fourth loss function in the i-th round based on the fourth loss function according to an embodiment of the present disclosure, using a reference vector, a fifth decoding vector set corresponding to the first encoding vector set in the i-th round, and a sixth decoding vector set corresponding to the fifth encoding vector set.

[0032] Figure 9C The illustration shows an example of a method for obtaining the value of the third loss function in the i-th round based on the third loss function according to an embodiment of the present disclosure, which is based on the first encoding vector set corresponding to the third sample protein dataset in the i-th round and the third encoding vector set corresponding to the third sample protein dataset.

[0033] Figure 10 The illustration shows an example diagram of a method for determining detection information of a model to be detected based on a reference vector and a target decoding vector set according to an embodiment of the present disclosure;

[0034] Figure 11 A flowchart illustrating a method for training a protein task model according to an embodiment of the present disclosure is shown schematically.

[0035] Figure 12 A flowchart illustrating a protein task processing method according to an embodiment of the present disclosure is shown schematically.

[0036] Figure 13 A block diagram of a data processing apparatus according to an embodiment of the present disclosure is shown schematically;

[0037] Figure 14 A block diagram of a model detection apparatus according to an embodiment of the present disclosure is shown schematically;

[0038] Figure 15 A block diagram of a training apparatus for a pre-trained model according to an embodiment of the present disclosure is shown schematically.

[0039] Figure 16 A block diagram of a training apparatus for a protein task model according to an embodiment of the present disclosure is shown schematically.

[0040] Figure 17 A block diagram of a protein task processing apparatus according to embodiments of the present disclosure is schematically shown; and

[0041] Figure 18 The diagram illustrates a block diagram of an electronic device suitable for implementing a data processing method, a model detection method, a pre-trained model training method, a protein task model training method, and a protein task processing method according to embodiments of the present disclosure. Detailed Implementation

[0042] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0043] Proteins are essential components of all cells and tissues in the human body. They are fundamental to life and the main agents of life activities. Pre-trained language models and methods can be applied to train protein task models. For example, self-supervised methods can be used for pre-training to obtain a pre-trained protein model, which can then be fine-tuned using downstream task data to obtain the final protein task model.

[0044] However, since the label information in self-supervised methods is automatically generated, pre-trained models can learn from unlabeled sample datasets, but this also requires significant resources. These resources can include data resources and computing power, which leads to pre-training models typically being performed by well-resourced AI companies and then shared via cloud platforms.

[0045] To address the risk of protein pre-trained models being stolen or attacked in the aforementioned methods, neural network watermarking technology has emerged. Neural network watermarking technology includes at least one of the following: white-box watermarking, black-box watermarking, box-free watermarking, and fragile neural network watermarking. White-box watermarking allows verifiers to access the network's internal structure and information such as weights when verifying copyright. Box-free watermarking is suitable for copyright authentication of generative networks. Fragile watermarking can detect malicious tampering with the network's functionality based on the extent of watermark corruption. Black-box watermarking is suitable for verifiers who cannot access the network internally but can only call the network remotely via an API (Application Programming Interface).

[0046] Regarding black-box watermarking technology, since it requires setting up a classifier for a specific task, i.e., specifying the target category, and protein pre-training models use self-supervised learning, which does not have such a target category, black-box watermarking technology cannot be applied to protein pre-training models based on amino acid sequences, thus making it impossible to guarantee the security of protein pre-training models.

[0047] Therefore, this disclosure proposes a data processing scheme. For example, first target protein data is acquired. The first target protein data is processed using a watermarked protein pre-trained model to obtain a first target encoding vector. The watermarked protein pre-trained model is a first deep learning model that has been trained. A second deep learning model that has been trained is used to detect the relationship between the model to be detected and the watermarked protein pre-trained model. The first and second deep learning models are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training processes of the first and second deep learning models influence each other. The sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset.

[0048] According to embodiments of this disclosure, since the first sample protein dataset is a predetermined privacy dataset, the trained first deep learning model is obtained by alternately training a second deep learning model and the first deep learning model using the sample protein dataset. The training processes of the first and second deep learning models influence each other. Therefore, the trained first deep learning model is thus determined as a watermarked protein pre-training model, realizing the construction of a watermarked protein pre-training model carrying watermark information. Furthermore, since the first target encoding vector is obtained by processing the first target protein data using the watermarked protein pre-training model, the security of the first target encoding vector is guaranteed, thereby improving the effectiveness of the first target encoding vector.

[0049] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0050] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0051] Figure 1 The illustration schematically depicts an exemplary system architecture from which data processing methods, model detection methods, pre-trained model training methods, protein task model training methods, protein task processing methods, and apparatus can be applied according to embodiments of the present disclosure.

[0052] It is important to note that Figure 1 The examples shown are merely examples of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. However, they do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture for applying data processing methods, model detection methods, pre-trained model training methods, protein task model training methods, protein task processing methods, and apparatus may include a terminal device. However, the terminal device can implement the data processing methods, model detection methods, pre-trained model training methods, protein task model training methods, protein task processing methods, and apparatus provided by embodiments of this disclosure without interacting with a server.

[0053] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as at least one of wired and wireless communication links. The terminal devices may include at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0054] Users can interact with server 105 via network 104 using at least one of the first terminal device 101, second terminal device 102, and third terminal device 103 to receive or send messages, etc. At least one of the first terminal device 101, second terminal device 102, and third terminal device 103 may have various communication client applications installed. For example, at least one of knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and social media platform software, etc.

[0055] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and supporting web browsing. For example, the electronic devices can include at least one of smartphones, tablets, laptops, and desktop computers.

[0056] Server 105 can be a server that provides various services. For example, Server 105 can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system. It solves the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.

[0057] It should be noted that the data processing method, model detection method, and protein task processing method provided in the embodiments of this disclosure can generally be executed by one of the first terminal device 101, the second terminal device 102, and the third terminal device 103. Accordingly, the dataset processing device, model detection device, and protein task processing device provided in the embodiments of this disclosure can also be disposed in one of the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0058] Alternatively, the data processing method, model detection method, and protein task processing method provided in the embodiments of this disclosure can generally also be executed by server 105. Correspondingly, the data processing apparatus, model detection apparatus, and protein task processing apparatus provided in the embodiments of this disclosure can generally be located in server 105. The data processing method, model detection method, and protein task processing method provided in the embodiments of this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with at least one of the first terminal device 101, second terminal device 102, third terminal device 103, and server 105. Correspondingly, the data processing apparatus, model detection apparatus, and protein task processing apparatus provided in the embodiments of this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with at least one of the first terminal device 101, second terminal device 102, third terminal device 103, and server 105.

[0059] It should be noted that the training methods for the pre-trained model and the protein task model provided in this embodiment can generally be executed by server 105. Correspondingly, the training apparatus for the pre-trained model and the training apparatus for the protein task model provided in this embodiment can generally be located in server 105. The training methods for the pre-trained model and the training methods for the protein task model provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with at least one of the first terminal device 101, the second terminal device 102, the third terminal device 103, and server 105. Correspondingly, the training apparatus for the pre-trained model and the training apparatus for the protein task model provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with at least one of the first terminal device 101, the second terminal device 102, the third terminal device 103, and server 105.

[0060] Alternatively, the training methods for the pre-trained model and the protein task model provided in this embodiment of the present disclosure can generally be executed by one of the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the training apparatus for the pre-trained model and the training apparatus for the protein task model provided in this embodiment of the present disclosure can also be provided in one of the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0061] It should be understood that Figure 1 The number of first terminal devices, second terminal devices, third terminal devices, networks, and servers shown in the diagram is merely illustrative. Depending on implementation needs, any number of first terminal devices, second terminal devices, third terminal devices, networks, and servers can be included.

[0062] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0063] Figure 2 A flowchart illustrating a data processing method according to an embodiment of the present disclosure is shown schematically.

[0064] like Figure 2 As shown, the method 200 includes operations S210 to S220.

[0065] In operation S210, acquire the data for the first target protein.

[0066] In operation S220, the watermarked protein pre-trained model is used to process the first target protein data to obtain the first target encoding vector.

[0067] According to embodiments of this disclosure, the watermark protein pre-training model can be a first deep learning model that has been trained. A second deep learning model that has been trained can be used to detect the relationship between the model to be detected and the watermark protein pre-training model. The first and second deep learning models can be obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training processes of the first and second deep learning models can influence each other.

[0068] According to embodiments of this disclosure, a sample protein dataset may include a first sample protein dataset, which is a predetermined privacy dataset.

[0069] According to embodiments of this disclosure, protein data is polymeric compound data formed by amino acid data combining in a specific order through dehydration condensation to form a polypeptide chain, and then combining one or more polypeptide chains according to a specific spatial structure. Protein data may contain elemental data such as carbon, hydrogen, oxygen, and nitrogen.

[0070] According to embodiments of this disclosure, protein data may include protein sequence data. Protein sequence data may refer to the amino-terminal composition of the protein data and the sequence of amino acids. A database of protein sequence data may refer to a database that uses computer functions to analyze biological information. By applying computer algorithms, DNA (Deoxyribonucleic Acid) and protein sequence data are compared to detect evolutionary relationships between structure, function, and sequence.

[0071] According to embodiments of this disclosure, the first sample protein dataset can be a predetermined privacy dataset, which is not publicly disclosed. The predetermined privacy dataset can be constructed by the copyright owner and is only accessible to the copyright owner. The construction method of the predetermined privacy dataset can be configured according to actual business needs and is not limited here. For example, at least one protein sequence data within a predetermined sequence data length range can be selected from a protein sequence data database. The at least one protein sequence data is then converted into sequence data in a target sequence data format. The predetermined sequence data length range and the target sequence data format can be configured according to actual business needs and are not limited here. For example, the predetermined sequence data length range can be 1050–1200. The sequence data length of the target sequence data format can be 1024.

[0072] According to embodiments of this disclosure, a watermarked protein pre-training model can treat a protein sequence as a sentence and the amino acids in the protein sequence as tags (i.e., words in natural language processing), so that the concepts of the pre-trained language model can be copied into the watermarked protein pre-training model.

[0073] According to embodiments of this disclosure, the watermarked protein pre-training model can be obtained by fine-tuning the model parameters of a clean protein pre-training model. The clean protein pre-training model can refer to an initial first deep learning model. The initial first deep learning model can refer to a first deep learning model that has not been trained using a fifth sample protein dataset. The fifth sample protein dataset can be a predetermined publicly available dataset. The absolute value of the difference between the output information of the watermarked protein pre-training model and the output information of the clean protein pre-training model is less than or equal to a third predetermined threshold, thereby making the watermarked protein pre-training model functionally similar to the clean protein pre-training model.

[0074] According to embodiments of this disclosure, the model structures of the first deep learning model and the second deep learning model can be configured according to actual business needs, and are not limited herein. For example, the first deep learning model may include at least one of the following: a first convolutional neural network (CNN), a first recurrent neural network (RNN), a first gated recurrent unit (GRU), a first long short-term memory (LSTM) network, and an encoder in a first Transformer. The first deep learning model may include a first encoder. The first encoder may include at least one of the following: a first convolutional neural network, a first recurrent neural network, a first gated recurrent unit, a first long short-term memory network, and a first Transformer. The second deep learning model may include a second encoder and a decoder. The second encoder may include at least one of the following: a second convolutional neural network, a second recurrent neural network, a second gated recurrent unit, a second long short-term memory network, and an encoder in a second Transformer. The decoder may include at least one of the following: a fully connected network and a decoder in a second Transformer.

[0075] According to embodiments of this disclosure, the trained second deep learning model can be used to detect the relationship between the model to be detected and the watermarked protein pre-trained model, that is, the trained second deep learning model can be used to determine the copyright ownership information of the model to be detected. The trained first and second deep learning models can be obtained by alternately training the second and first deep learning models using a sample protein dataset. This can include: firstly, training the previous round of the second deep learning model using at least a portion of the sample protein dataset to obtain the current round of the second deep learning model. While keeping the model parameters of the current round of the second deep learning model unchanged, training the previous round of the first deep learning model using at least a portion of the sample protein dataset to obtain the current round of the first deep learning model, thereby allowing the training processes of the first and second deep learning models to influence each other. The sample protein datasets used to train the first and second deep learning models may contain at least a common portion.

[0076] According to embodiments of this disclosure, the first target encoding vector can be obtained by processing the first target protein data using a watermarked protein pre-training model. For example, the first target encoding vector can be obtained by using a watermarked protein pre-training model to extract features from the first target protein data.

[0077] According to embodiments of this disclosure, since the first sample protein dataset is a predetermined privacy dataset, the trained first deep learning model is obtained by alternately training a second deep learning model and the first deep learning model using the sample protein dataset. The training processes of the first and second deep learning models influence each other. Therefore, the trained first deep learning model is thus determined as a watermarked protein pre-training model, realizing the construction of a watermarked protein pre-training model carrying watermark information. Furthermore, since the first target encoding vector is obtained by processing the first target protein data using the watermarked protein pre-training model, the security of the first target encoding vector is guaranteed, thereby improving the effectiveness of the first target encoding vector.

[0078] According to embodiments of this disclosure, operation S210 may include the following operations.

[0079] Based on the first target protein data, the second target encoding vector is obtained. An attention strategy is then used to process the second target encoding vector to obtain the first target encoding vector.

[0080] According to embodiments of this disclosure, an attention strategy can be used to focus on important information with high weight and ignore unimportant information with low weight, and can exchange information with other information by sharing important information, thereby realizing the transmission of important information. In embodiments of this disclosure, the attention strategy can extract information about each dimension vector of the second target encoding vector itself and the information between each dimension vector, so as to better complete the feature extraction of the first target protein data.

[0081] According to embodiments of this disclosure, feature extraction can be performed on the first target protein data to obtain a second target encoding vector. The second target encoding vector can then be processed at least one level based on an attention strategy to obtain the first target encoding vector.

[0082] According to embodiments of this disclosure, the key matrix, value matrix, and query matrix can be matrices in an attention mechanism. A first target encoding vector can be obtained by processing the second target encoding vector used as the key matrix, value matrix, and query matrix based on an attention strategy. For example, an attention unit can be determined according to the attention strategy. The first target encoding vector is obtained by processing the second target encoding vector used as the key matrix, value matrix, and query matrix using the attention unit.

[0083] According to embodiments of this disclosure, an attention strategy is used to process the second target encoding vector to obtain the first target encoding vector. Since the attention strategy can extract information about each dimension vector of the second target encoding vector itself and the information between each dimension vector, the accuracy of the first target encoding vector is improved.

[0084] According to embodiments of this disclosure, processing a second target encoding vector using an attention strategy to obtain a first target encoding vector may include the following operations.

[0085] The second target encoding vector is processed using a self-attention strategy to obtain the third target encoding vector. Based on the second and third target encoding vectors, the fourth target encoding vector is obtained. Based on the fourth target encoding vector, the first target encoding vector is obtained.

[0086] According to embodiments of this disclosure, the self-attention strategy may include a multi-head self-attention strategy. A second target encoding vector, used as a key matrix, value matrix, and query matrix, can be subjected to self-attention processing to obtain a third target encoding vector. The second and third target encoding vectors can be fused to obtain a seventh target encoding vector. The seventh target encoding vector is then processed to obtain a fourth target encoding vector. For example, the seventh target encoding vector can be subjected to residual concatenation and normalization to obtain the fourth target encoding vector. Normalization may include one of the following: batch normalization (BN) and layer normalization (LN). The fourth target encoding vector can be processed to obtain a first target encoding vector.

[0087] According to embodiments of this disclosure, obtaining a first target encoding vector based on a fourth target encoding vector may include the following operations.

[0088] The fourth target encoding vector is processed using a multilayer perceptron strategy to obtain the fifth target encoding vector. Based on the fourth and fifth target encoding vectors, the first target encoding vector is then obtained.

[0089] According to embodiments of this disclosure, a multilayer perceptron strategy can be used to extract features from a fourth target encoding vector. A multilayer perceptron can be determined based on the strategy. The fourth target encoding vector is processed using the multilayer perceptron to obtain a fifth target encoding vector. The fourth and fifth target encoding vectors can be fused to obtain an eighth target encoding vector. The eighth target encoding vector is then processed to obtain a first target encoding vector. For example, residual concatenation and normalization can be performed on the eighth target encoding vector to obtain the first target encoding vector.

[0090] According to embodiments of this disclosure, operation S210 may include the following operations.

[0091] The first target protein data is processed using a long-term dependency information learning strategy to obtain the first target encoding vector.

[0092] According to embodiments of this disclosure, a long-term dependency information learning strategy may include one of a unidirectional long-term dependency information learning strategy and a bidirectional long-term dependency information learning strategy.

[0093] According to embodiments of this disclosure, a long-term dependency learning strategy can be used to determine the sequence features of a first target protein data, enabling the sequence features to reflect information carried by long-term memory during the determination process. A unidirectional long-term dependency learning strategy can be used to obtain a first target encoding vector by processing the first target protein data in a forward direction. A bidirectional long-term dependency learning strategy can be used to obtain a first target encoding vector by processing the first target protein data in both forward and reverse directions.

[0094] According to embodiments of this disclosure, a first target protein data is processed using a long-term dependent information learning strategy to obtain a first target encoding vector. The long-term dependent information learning strategy can achieve long-term memory, thus improving the accuracy of the first target encoding vector.

[0095] According to embodiments of this disclosure, when the long-term dependent information learning strategy includes a unidirectional long-term dependent information learning strategy, processing the first target protein data based on the long-term dependent information learning strategy to obtain the first target encoding vector may include the following operations.

[0096] The first target protein data is subjected to positive long-term dependency learning to obtain the first target encoding vector.

[0097] According to embodiments of this disclosure, when the long-term dependency information learning strategy includes a bidirectional long-term dependency information learning strategy, processing the first target protein data based on the long-term dependency information learning strategy to obtain the first target encoding vector may include the following operations.

[0098] The first target protein data is subjected to forward long-term dependency learning and reverse long-term dependency learning to obtain the first target encoding vector.

[0099] According to embodiments of this disclosure, forward long-term dependency learning can refer to long-term dependency learning using the input information at the current position and the hidden state information at the previous position. Reverse long-term dependency learning can refer to long-term dependency learning using the hidden state information at the next position and the input information at the current position.

[0100] Figure 3 A flowchart illustrating a model detection method according to an embodiment of the present disclosure is shown schematically.

[0101] like Figure 3 As shown, the method 300 includes steps S310 to S330.

[0102] In operation S310, the detection protein dataset is processed using the model to be detected to obtain the sixth target encoding vector set.

[0103] In operation S320, the sixth target encoded vector set is processed using the key tuple to obtain the target decoded vector set.

[0104] In operation S330, the detection information of the model to be detected is determined based on the reference vector and the target decoding vector set.

[0105] According to embodiments of this disclosure, the protein detection dataset can be obtained from a first sample protein dataset, predetermined mask data, and a key tuple. Detection information can characterize the relationship between the model to be detected and the watermarked protein pre-trained model. The key tuple can be a trained second deep learning model. The watermarked protein pre-trained model can be a trained first deep learning model. The trained second deep learning model and the trained first deep learning model can be obtained by alternately training the second deep learning model and the first deep learning model using the sample protein dataset. The training processes of the first deep learning model and the second deep learning model can influence each other.

[0106] According to embodiments of this disclosure, the sample protein dataset may include a first sample protein dataset. The first sample protein dataset may be a predetermined privacy dataset.

[0107] According to embodiments of this disclosure, the second deep learning model may include a second encoder and a decoder. The key tuple may include the trained second encoder and the trained decoder. The protein detection dataset may be obtained based on a first sample protein dataset, predetermined mask data, and the second encoder included in the key tuple. The detection information may characterize the relationship between the model to be detected and the watermarked protein pre-trained model; that is, the detection information may characterize the copyright ownership information of the model to be detected.

[0108] According to embodiments of this disclosure, the model to be detected can be configured according to actual business needs, and is not limited thereto. For example, the model to be detected may refer to a suspicious pre-trained model. A suspicious pre-trained model may refer to a watermark protein pre-trained model obtained through the training method of a pre-trained model. Alternatively, a suspicious pre-trained model may refer to other suspected pre-trained models.

[0109] According to embodiments of this disclosure, the protein detection dataset may include at least one protein detection dataset. The sixth target encoding vector set may include a sixth target encoding vector corresponding to the at least one protein detection dataset. The target decoding vector set may include target decoding vectors corresponding to the at least one sixth target encoding vector.

[0110] According to embodiments of this disclosure, a protein detection dataset can be input into the model to be detected, outputting a sixth target encoding vector set. The sixth target encoding vector set is processed using a decoder included in the key tuple to obtain a target decoding vector set. After obtaining the target decoding vector set, watermark detection can be performed on the model to be detected based on the baseline vector and the target decoding vector set.

[0111] According to embodiments of this disclosure, the target decoding vector set is obtained by processing the sixth target encoding vector set using a key tuple. The detection protein dataset is obtained based on the first sample protein dataset, predetermined mask data, and the key tuple. The sixth target encoding vector set is obtained by processing the detection protein dataset using the model to be detected. Based on this, since the detection information of the model to be detected is determined based on the reference vector and the seventh decoding vector set, the detection information can characterize the watermark detection result of the model to be detected, thereby verifying the effectiveness of the model watermark.

[0112] According to embodiments of this disclosure, operation S330 may include the following operations.

[0113] The similarity between the baseline vector and the target decoded vector set is determined to obtain the target similarity set. Based on the target similarity set, the detection information of the model to be detected is determined.

[0114] According to embodiments of this disclosure, the target similarity set may include at least one target similarity. Target similarity can characterize the degree of similarity between a reference vector and a target decoded vector. The relationship between the numerical value of the target similarity and the degree of similarity can be configured according to actual business needs, and is not limited herein.

[0115] According to embodiments of this disclosure, a predetermined similarity determination method can be used to determine the similarity between a reference vector and at least one target decoded vector included in the target decoded vector set, thereby obtaining at least one target similarity. The predetermined similarity determination method can be configured according to actual business needs and is not limited thereto. For example, the predetermined similarity determination method may include at least one of the following: a distance-based similarity determination method, a cosine similarity determination method, a correlation coefficient-based similarity determination method, and a similarity determination method based on a set perspective.

[0116] For example, distance-based similarity determination methods may include at least one of the following: Euclidean distance (ED), Manhattan distance (MD), Chebyshev distance (CD), Minkowski distance (MD), Standardized Euclidean distance (SED), Mahalanobis distance (MD), and Lance Williams distance (LWD). Cosine-based similarity determination methods may include at least one of the following: cosine similarity and the Tanimoto coefficient. Correlation-based similarity determination methods may include the Pearson Correlation Coefficient (PCC). Similarity determination methods based on a set perspective may include the Jaccard Similarity Coefficient (JSC).

[0117] According to embodiments of this disclosure, determining the similarity between a reference vector and at least one target decoding vector included in the target decoding vector set to obtain at least one target similarity may include: determining the similarity between the reference vector and at least one target decoding vector respectively to obtain at least one target similarity.

[0118] According to embodiments of this disclosure, detection information of a model to be detected can be determined based on at least one target similarity and a predetermined threshold. For example, for a target similarity among at least one target similarity, detection information corresponding to that target similarity can be determined based on the predetermined threshold and the target similarity. Detection information of the model to be detected is obtained based on the detection information corresponding to each of the at least one target similarity.

[0119] According to embodiments of this disclosure, determining the detection information corresponding to the target similarity based on a first predetermined threshold and the target similarity may include: if the target similarity is greater than the first predetermined threshold, the detection information corresponding to the target similarity indicates that the model to be detected belongs to a clean protein pre-trained model, i.e., belongs to the initial first encoder. If the target similarity is less than or equal to the first predetermined threshold, the detection information corresponding to the target similarity indicates that the model to be detected does not belong to a clean protein pre-trained model, i.e., does not belong to the initial first encoder. The first predetermined threshold can be configured according to actual business needs and is not limited thereto.

[0120] According to embodiments of this disclosure, obtaining detection information for a model to be detected based on detection information corresponding to at least one target similarity can include: determining that the detection information for the model to be detected is information indicating that the model to be detected does not belong to a clean protein pre-training model when the number of target detection information is greater than a second predetermined threshold; and determining that the detection information for the model to be detected is information indicating that the model to be detected belongs to a clean protein pre-training model when the number of target detection information is less than or equal to the second predetermined threshold. The second predetermined threshold may be greater than or equal to 0 and less than or equal to the number of detection protein data included in the detection protein dataset. Target detection information may refer to detection information indicating that the model to be detected does not belong to a clean protein pre-training model.

[0121] Figure 4 A flowchart illustrating a training method for a pre-trained model according to an embodiment of the present disclosure is shown schematically.

[0122] like Figure 4 As shown, the method 400 includes operations S410 to S420.

[0123] In operation S410, the first deep learning model and the second deep learning model are trained alternately using the sample protein dataset, so that the training process of the first deep learning model and the training process of the second deep learning model influence each other.

[0124] When operating the S420, the first deep learning model that has been trained is identified as the watermark protein pre-training model.

[0125] According to embodiments of this disclosure, the trained second deep learning model can be used to detect the relationship between the model to be detected and the watermark protein pre-trained model.

[0126] According to embodiments of this disclosure, the sample protein dataset may include a first sample protein dataset. The first sample protein dataset may be a predetermined privacy dataset.

[0127] According to embodiments of this disclosure, the training process of the watermark protein pre-training model can be found in the description of the relevant sections above, and will not be repeated here.

[0128] According to an embodiment of this disclosure, since the first sample protein dataset is a predetermined privacy dataset, the first deep learning model that has been trained is obtained by alternately training the second deep learning model and the first deep learning model using the sample protein dataset. The training process of the first deep learning model and the training process of the second deep learning model influence each other. Therefore, the first deep learning model that has been trained is determined as the watermark protein pre-training model, thus realizing the construction of a watermark protein pre-training model carrying watermark information.

[0129] The above are merely exemplary embodiments, but are not limited thereto. Other training methods for pre-trained models known in the art may also be included, as long as they can construct a watermark protein pre-trained model carrying watermark information.

[0130] According to embodiments of this disclosure, the sample protein dataset may further include a second sample protein dataset and a third sample protein dataset. The third sample protein dataset may be a pre-determined publicly available dataset.

[0131] According to embodiments of this disclosure, the second sample protein dataset may be obtained based on the first sample protein dataset, predetermined mask data, and a second deep learning model.

[0132] According to embodiments of this disclosure, the third sample protein dataset can be a pre-defined public dataset. The pre-defined public dataset can provide reliable protein sequences associated with annotations. Annotations can be derived from research findings in the literature and E-value-verified computational analysis results. Annotations can include at least one of the following: protein and gene name annotations, protein function annotations, enzyme-specific information annotations, subcellular localization annotations, protein-protein interaction annotations, expression mode annotations, location and role annotations of important domains and sites, and annotations of protein variants resulting from protein hydrolysis and post-translational modifications.

[0133] According to embodiments of this disclosure, a second sample protein dataset can be obtained from a first sample protein dataset, predetermined mask data, and a second deep learning model. This can include: the second sample protein dataset being obtained from a first intermediate dataset and a second intermediate dataset. The first intermediate dataset can be obtained by processing the first sample protein dataset using predetermined mask data. The second intermediate dataset can be obtained by processing the predetermined mask data using a second deep learning model. For example, the second intermediate dataset can be obtained by processing the predetermined mask data using a second encoder in the second deep learning model.

[0134] For example, the first intermediate dataset may be obtained by processing the first sample protein dataset using predetermined mask data. This may include: for at least one first sample protein dataset, the first intermediate dataset may be obtained by replacing predetermined characters in the first sample protein dataset using predetermined mask data. The predetermined characters in the first sample protein dataset may be randomly determined from the first sample protein dataset.

[0135] According to embodiments of this disclosure, a first deep learning model may include a first encoder. A second deep learning model may include a second encoder and a decoder.

[0136] According to embodiments of this disclosure, operation S410 may include the following operations.

[0137] The training operations of training the second encoder and decoder using the second sample protein dataset, the third sample protein dataset, the reference vector, the first encoder and the decoder are executed alternately, as well as the training operations of training the first encoder using the first sample protein dataset, the third sample protein dataset, the reference vector, the first encoder and the decoder, are executed until a predetermined termination condition is met.

[0138] According to embodiments of this disclosure, the first encoder, decoder, and second encoder can be configured according to actual business needs, and are not limited herein. For example, the first encoder may include at least one of the following: a first convolutional neural network, a first recurrent neural network, a first gated recurrent unit, a first long short-term memory network, and an encoder in a first Transformer. The second encoder may include at least one of the following: a second convolutional neural network, a second recurrent neural network, a second gated recurrent unit, a second long short-term memory network, and an encoder in a second Transformer. The decoder may include at least one of the following: a fully connected network and a decoder in a second Transformer.

[0139] According to embodiments of this disclosure, the predetermined termination condition may include at least one of the following: the maximum number of training epochs has been reached and the model has converged. The baseline vector may be a randomly initialized vector that satisfies the target sequence data format. The second sample protein dataset may be obtained based on the first sample protein dataset, predetermined mask data, and the second encoder.

[0140] According to embodiments of this disclosure, training operations that alternately execute training operations using a second sample protein dataset, a third sample protein dataset, a reference vector, a first encoder, and a decoder to train a second encoder and a decoder, and training operations that alternately execute training operations using a first sample protein dataset, a third sample protein dataset, a reference vector, a first encoder, and a decoder to train a first encoder, until a predetermined termination condition is met, may include the following operations.

[0141] Based on the first sample protein dataset, the predetermined mask data, and the second encoder of round i-1, the second sample protein dataset for round i is obtained. Using the baseline vector, the third sample protein dataset, the first encoder of round i-1, the decoder of round i-1, and the second sample protein dataset for round i, the second encoder and decoder for round i are trained, resulting in the second encoder and decoder for round i. Using the first sample protein dataset, the third sample protein dataset, the baseline vector, the first encoder of round i-1, and the decoder for round i, the first encoder for round i-1 is trained, resulting in the first encoder for round i.

[0142] According to embodiments of this disclosure, i can be an integer greater than 1.

[0143] According to embodiments of this disclosure, multiple rounds of training are performed on the first encoder, decoder, and second encoder until a predetermined termination condition is met. The first encoder, after training, is determined as the watermark protein pre-trained model.

[0144] According to embodiments of this disclosure, in response to I > 1, when 1 < i ≤ I, the second encoder and decoder for the (i-1)th round can be trained first to obtain the second encoder and decoder for the i-th round. While keeping the model parameters of the trained second encoder and decoder for the i-th round unchanged, the first encoder for the (i-1)th round is trained. When i = 1, the initial second encoder and decoder can be trained first to obtain the second encoder and decoder for the 1st round. While keeping the model parameters of the second encoder and decoder for the 1st round unchanged, the initial first encoder is trained to obtain the first encoder for the 1st round. In response to I = 1, the initial second encoder and decoder can be trained first to obtain the trained second encoder and decoder. While keeping the model parameters of the trained second encoder and decoder unchanged, the initial first encoder is trained to obtain the trained first encoder.

[0145] According to embodiments of this disclosure, a first encoder may include a first input layer and a first hidden layer. The first encoder in the (i-1)th round can be used to encode a third sample protein dataset and a second sample protein dataset in the i-th round. The third sample protein dataset may be a predetermined public dataset. The third sample protein dataset may include at least one third sample protein data. For example, the first input layer of the first encoder in the (i-1)th round can be used to encode the third sample protein dataset to obtain a first intermediate encoded dataset. The first hidden layer of the first encoder is then used to process the first intermediate encoded dataset to obtain a first encoded vector set. Similarly, the first input layer of the first encoder in the (i-1)th round can be used to encode the second sample protein dataset in the i-th round to obtain a second intermediate encoded dataset. The first hidden layer of the first encoder is then used to process the second intermediate encoded dataset to obtain a second encoded vector set.

[0146] According to embodiments of this disclosure, the decoder may include a second hidden layer and a first output layer. The decoder in the (i-1)th round can be used to reconstruct the first and second encoded vector sets in the i-th round. For example, the second hidden layer of the decoder in the (i-1)th round can be used to decode the first encoded vector set in the i-th round to obtain a first auxiliary decoded dataset. The first output layer of the decoder is then used to process the first auxiliary decoded dataset to obtain a first decoded vector set. The second hidden layer of the decoder in the (i-1)th round can be used to decode the second encoded vector set to obtain a second auxiliary decoded dataset. The first output layer of the decoder is then used to process the second auxiliary decoded dataset to obtain a second decoded vector set.

[0147] According to embodiments of this disclosure, the second encoder may include a second input layer and a third hidden layer. The (i-1)th round of the second encoder can be used to encode the first sample protein dataset to obtain the second sample protein dataset for the i-th round. The first sample protein dataset may be a predetermined privacy dataset. The first sample protein dataset may include at least one first sample protein data. For example, the second input layer of the (i-1)th round of the second encoder can be used to encode predetermined mask data to obtain the first intermediate dataset for the i-th round. The third hidden layer of the second encoder is used to process the first intermediate dataset for the i-th round to obtain the second sample protein dataset for the i-th round.

[0148] According to embodiments of this disclosure, since the first sample protein dataset is a predetermined privacy dataset, and the second sample protein dataset of the i-th round is obtained based on the first sample protein dataset, predetermined mask data, and the second encoder of the (i-1)-th round, the second encoder and decoder of the (i-1)-th round are trained using the predetermined disclosed third sample protein dataset, reference vector, first encoder of the (i-1)-th round, decoder of the (i-1)-th round, and second sample protein dataset of the i-th round. The first encoder of the (i-1)-th round is trained using the third sample protein dataset, first sample protein dataset, reference vector, first encoder of the (i-1)-th round, and decoder of the i-th round. Thus, the trained first encoder is determined as the watermark protein pre-training model, thereby realizing the construction of a watermark protein pre-training model carrying watermark information.

[0149] According to embodiments of this disclosure, obtaining the second sample protein dataset for the i-th round based on the first sample protein dataset, predetermined mask data, and the second encoder for the (i-1)-th round may include the following operations.

[0150] Based on the predetermined mask data and the second encoder of round i-1, the first intermediate dataset of round i is obtained. Based on the second intermediate dataset and the first intermediate dataset of round i, the second sample protein dataset of round i is obtained.

[0151] According to embodiments of this disclosure, the second intermediate dataset may be obtained based on the first sample protein dataset and predetermined mask data.

[0152] According to embodiments of this disclosure, the second encoder may include a third input layer and a fourth hidden layer. The (i-1)th round of the second encoder can be used to encode the first sample protein dataset to obtain the second sample protein dataset of the i-th round. A second intermediate dataset can be obtained based on the first sample protein dataset and predetermined mask data. The predetermined mask data can be encoded using the third input layer of the (i-1)th round of the second encoder to obtain the first intermediate dataset of the i-th round. The second intermediate dataset and the first intermediate dataset of the i-th round are processed using the fourth hidden layer of the second encoder to obtain the second sample protein dataset of the i-th round.

[0153] According to embodiments of this disclosure, since the second intermediate dataset is obtained based on the first sample protein dataset and the predetermined mask data, and the first intermediate dataset of the i-th round is obtained based on the predetermined mask data and the second encoder of the (i-1)-th round, the second sample protein dataset of the i-th round can be obtained based on the second intermediate dataset and the first intermediate dataset of the i-th round.

[0154] According to embodiments of this disclosure, a second sample protein dataset can be determined according to the following formula (1).

[0155]

[0156] According to embodiments of this disclosure, x k,j ∈D k x p,j ∈D p D k This can characterize the second sample protein dataset. p This can characterize the first sample protein dataset. k,j It can characterize the j-th second-sample protein data x in the second-sample protein dataset. k x p,j It can characterize the j-th first-sample protein data x in the first-sample protein dataset. p Mask can represent predefined mask data. T can represent the second encoder. It can represent element-wise multiplication. It can characterize the second intermediate dataset. It can characterize the first intermediate dataset.

[0157] According to embodiments of this disclosure, after obtaining the first sample protein dataset D... pNext, a predefined mask data Mask and a second encoder T can be constructed. This is done by injecting the predefined mask data Mask and the second encoder T into the first sample protein dataset D. p In this way, the second sample protein dataset D is obtained. k The injection method can be configured according to actual business needs, and is not limited here.

[0158] For example, for the first sample protein dataset D p The j-th first sample protein data x p The first intermediate dataset for the i-th round can be obtained based on the predetermined mask data Mask and the second encoder T for the (i-1)-th round. Based on the j-th first sample protein data x in the first sample protein dataset p Together with the pre-defined mask data, a second intermediate dataset is obtained. According to the second intermediate dataset and the first intermediate dataset of the i-th round We obtain the second sample protein dataset D from the i-th round. k .

[0159] The following is for reference. Figure 5 , Figure 6 , Figure 7A , Figure 7B , Figure 7C , Figure 8 , Figure 9A , Figure 9B , Figure 9C and Figure 10 The training method of the pre-trained model according to the embodiments of this disclosure will be further explained in conjunction with specific embodiments.

[0160] Figure 5 The illustration shows an example diagram of a method for obtaining a second sample protein dataset for the i-th round based on a first sample protein dataset, predetermined mask data, and a second encoder for the (i-1)-th round, according to an embodiment of the present disclosure.

[0161] like Figure 5 As shown, in 500, the first intermediate dataset 503 of the i-th round can be obtained based on the predetermined mask data 501 and the second encoder 502 of the (i-1)-th round. The second sample protein dataset 505 of the i-th round is obtained based on the second intermediate dataset 504 and the first intermediate dataset 503 of the i-th round.

[0162] According to embodiments of this disclosure, training a second encoder and decoder for the (i-1)th round using the reference vector, the third sample protein dataset, the first encoder for the (i-1)th round, the decoder for the (i-1)th round, and the second sample protein dataset for the i-th round to obtain the second encoder and decoder for the i-th round may include the following operations.

[0163] The first encoder in round i-1 processes the third protein dataset and the second protein dataset in round i, respectively, to obtain the first encoded vector set corresponding to the third protein dataset and the second encoded vector set corresponding to the second protein dataset in round i. The decoder in round i-1 processes the first encoded vector set corresponding to the third protein dataset and the second encoded vector set corresponding to the second protein dataset in round i, respectively, to obtain the first decoded vector set corresponding to the first encoded vector set and the second decoded vector set corresponding to the second encoded vector set in round i. The decoder in round i-1 processes the third encoded vector set corresponding to the third protein dataset and the fourth encoded vector set corresponding to the second protein dataset in round i, respectively, to obtain the third decoded vector set corresponding to the third encoded vector set and the fourth decoded vector set corresponding to the fourth encoded vector set in round i. Based on the baseline vector, the first decoded vector set, the third decoded vector set, the second decoded vector set, and the fourth decoded vector set in round i, the model parameters of the second encoder and decoder in round i-1 are adjusted to obtain the second encoder and decoder in round i.

[0164] According to embodiments of this disclosure, the third encoding vector set corresponding to the third sample protein dataset and the fourth encoding vector set corresponding to the second sample protein dataset in the i-th round can be obtained by processing the third sample protein dataset and the second sample protein dataset in the i-th round respectively using the initial first encoder.

[0165] According to embodiments of this disclosure, the first encoder in the (i-1)th round may include a fourth input layer and a fifth hidden layer. The first encoder in the (i-1)th round can be used to encode the third sample protein dataset and the second sample protein dataset in the i-th round. For example, the fourth input layer of the first encoder in the (i-1)th round can be used to encode the third sample protein dataset to obtain a third intermediate encoded dataset. The fifth hidden layer of the first encoder is then used to process the third intermediate encoded dataset to obtain a first encoded vector set. Similarly, the fourth input layer of the first encoder in the (i-1)th round can be used to encode the second sample protein dataset in the i-th round to obtain a fourth intermediate encoded dataset. The fifth hidden layer of the first encoder is then used to process the fourth intermediate encoded dataset to obtain a second encoded vector set.

[0166] According to embodiments of this disclosure, the decoder in the (i-1)th round may include a sixth hidden layer and a second output layer. The decoder in the (i-1)th round can be used to reconstruct the first and second encoded vector sets of the i-th round. For example, the sixth hidden layer of the decoder in the (i-1)th round can be used to decode the first encoded vector set of the i-th round to obtain a third auxiliary decoded dataset. The second output layer of the decoder is then used to process the third auxiliary decoded dataset to obtain a first decoded vector set. The sixth hidden layer of the decoder in the (i-1)th round can be used to decode the second encoded vector set to obtain a fourth auxiliary decoded dataset. The second output layer of the decoder is then used to process the fourth auxiliary decoded dataset to obtain a second decoded vector set.

[0167] According to embodiments of this disclosure, the watermarked protein pre-training model can be obtained by fine-tuning the model parameters of a clean protein pre-training model. The clean protein pre-training model can refer to an initial first encoder. The initial first encoder can include a fifth input layer and a seventh hidden layer. The initial first encoder can be used to encode a third sample protein dataset and a second sample protein dataset from the i-th round. For example, the fifth input layer of the (i-1)-th round first encoder can be used to encode the third sample protein dataset to obtain a fifth intermediate encoded dataset. The seventh hidden layer of the first encoder is used to process the first intermediate encoded dataset to obtain the third encoded vector set from the i-th round. The first input layer of the (i-1)-th round first encoder can be used to encode the second sample protein dataset from the i-th round to obtain a sixth intermediate encoded dataset. The seventh hidden layer of the first encoder is used to process the sixth intermediate encoded dataset to obtain a fourth encoded vector set.

[0168] Alternatively, the decoder in the (i-1)th round can be used to reconstruct the third and fourth encoded vector sets from the i-th round. For example, the sixth hidden layer of the decoder in the (i-1)th round can be used to decode the third encoded vector set from the i-th round to obtain the fifth auxiliary decoded dataset. The second output layer of the decoder is then used to process the fifth auxiliary decoded dataset to obtain the third decoded vector set. Similarly, the sixth hidden layer of the decoder in the (i-1)th round can be used to decode the fourth encoded vector set to obtain the sixth auxiliary decoded dataset. The second output layer of the decoder is then used to process the sixth auxiliary decoded dataset to obtain the fourth decoded vector set.

[0169] According to embodiments of this disclosure, after obtaining the first decoding vector set, the third decoding vector set, the second decoding vector set, and the fourth decoding vector set of the i-th round, the model parameters of the second encoder and decoder of the (i-1)-th round can be adjusted based on the reference vector, the first decoding vector set, the third decoding vector set, the second decoding vector set, and the fourth decoding vector set of the i-th round to obtain the second encoder and decoder of the i-th round.

[0170] Figure 6 The illustration shows an example of a method according to an embodiment of the present disclosure for training a second encoder and decoder for the (i-1)th round using a reference vector, a third sample protein dataset, a first encoder for the (i-1)th round, a decoder for the (i-1)th round, and a second sample protein dataset for the i-th round, to obtain a second encoder and decoder for the i-th round.

[0171] like Figure 6 As shown in 600, the first encoder 603_1 of the (i-1)th round can be used to process the third sample protein dataset 601 and the second sample protein dataset 602 of the i-th round respectively, to obtain the first encoding vector set 604 corresponding to the third sample protein dataset 601 and the second encoding vector set 605 corresponding to the second sample protein dataset 602 of the i-th round.

[0172] The decoder 603_3 of the (i-1)th round processes the first encoding vector set 604 corresponding to the third sample protein dataset 601 and the second encoding vector set 605 corresponding to the second sample protein dataset 602 in the i-th round, respectively, to obtain the first decoding vector set 606 corresponding to the first encoding vector set 604 and the second decoding vector set 607 corresponding to the second encoding vector set 605 in the i-th round.

[0173] The decoder 603_3 of the (i-1)th round processes the third encoding vector set 608 corresponding to the third sample protein dataset 601 and the fourth encoding vector set 609 corresponding to the second sample protein dataset 602 in the i-th round, respectively, to obtain the third decoding vector set 610 corresponding to the third encoding vector set 608 and the fourth decoding vector set 611 corresponding to the fourth encoding vector set 609 in the i-th round.

[0174] Based on the reference vector 612, the first decoding vector set 606 of the i-th round, the third decoding vector set 610 of the i-th round, the second decoding vector set 607 of the i-th round, and the fourth decoding vector set 611 of the i-th round, the model parameters of the second encoder 603_2 and the decoder 603_3 of the (i-1)-th round are adjusted to obtain the second encoder and decoder of the i-th round.

[0175] According to embodiments of this disclosure, adjusting the model parameters of the second encoder and decoder in the (i-1)th round based on the reference vector, the first decoding vector set in the i-th round, the third decoding vector set in the i-th round, the second decoding vector set in the i-th round, and the fourth decoding vector set in the i-th round to obtain the second encoder and decoder in the i-th round may include the following operations.

[0176] Based on the first loss function, the value of the first loss function for round i is obtained according to the baseline vector, the first decoded vector set for round i, and the third decoded vector set for round i. Based on the second loss function, the value of the second loss function for round i is obtained according to the baseline vector, the second decoded vector set for round i, and the fourth decoded vector set for round i. Based on the first and second loss function values ​​for round i, the model parameters of the second encoder and decoder for round i-1 are adjusted to obtain the second encoder and decoder for round i.

[0177] According to embodiments of this disclosure, the first loss function and the second loss function can be configured according to actual business needs, and are not limited herein. For example, the first loss function and the second loss function may include at least one of the following: regression loss function, exponential loss function, squared error loss function, absolute error loss function, cross-entropy loss function, Hinge loss function and Huber loss function, binary classification loss function, binary cross-entropy loss function, multi-class loss function and multi-class cross-entropy loss function.

[0178] According to embodiments of this disclosure, the first loss function value for round i can be obtained based on a first loss function, using a reference vector, a first decoding vector set for round i, and a third decoding vector set for round i. The model parameters of the second encoder and decoder for round i-1 are adjusted based on the first loss function value for round i until a predetermined termination condition is met. For example, the model parameters of the predetermined model can be adjusted using a backpropagation algorithm or a stochastic gradient descent algorithm until the predetermined termination condition is met. The second encoder and decoder for round i-1 obtained under the condition of meeting the predetermined termination condition are determined as candidate second encoders and decoders for round i.

[0179] According to embodiments of this disclosure, the second loss function value for the i-th round can be obtained based on the second loss function, using a reference vector, the second decoding vector set for the i-th round, and the fourth decoding vector set for the i-th round. The model parameters of the candidate second encoder and decoder for the i-th round are adjusted based on the second loss function value for the i-th round until a predetermined termination condition is met. For example, the model parameters of the predetermined model can be adjusted using a backpropagation algorithm or a stochastic gradient descent algorithm until the predetermined termination condition is met. The candidate second encoder and decoder for the i-th round obtained under the condition of meeting the predetermined termination condition are determined as the second encoder and decoder for the i-th round.

[0180] Alternatively, based on the first loss function, the value of the first loss function in round i can be obtained using the baseline vector, the first decoding vector set in round i, and the third decoding vector set in round i. Based on the second loss function, the value of the second loss function in round i can be obtained using the baseline vector, the second decoding vector set in round i, and the fourth decoding vector set in round i. Based on the value of the first loss function in round i and the value of the second loss function in round i, the value of the first target loss function in round i can be determined.

[0181] According to embodiments of this disclosure, after obtaining the first objective loss function value for the i-th round, the model parameters of the second encoder and decoder for the (i-1)-th round can be adjusted based on the first objective loss function value for the i-th round until a predetermined termination condition is met. For example, the model parameters of the predetermined model can be adjusted according to the backpropagation algorithm or the stochastic gradient descent algorithm until the predetermined termination condition is met. The second encoder and decoder for the (i-1)-th round obtained under the condition of meeting the predetermined termination condition are determined as candidate second encoders and decoders for the i-th round. The predetermined termination condition may include at least one of output value convergence and the number of training rounds reaching the maximum number of training rounds.

[0182] Figure 7A The illustration schematically shows an example of a method for adjusting the model parameters of the second encoder and decoder in the (i-1)th round according to an embodiment of the present disclosure, based on a reference vector, a first decoding vector set in the i-th round, a third decoding vector set in the i-th round, a second decoding vector set in the i-th round, and a fourth decoding vector set in the i-th round, to obtain the second encoder and decoder in the i-th round.

[0183] like Figure 7A As shown, in 700A, the reference vector 701, the first decoded vector set 702 of the i-th round, and the third decoded vector set 703 of the i-th round can be input into the first loss function 704, and the first loss function value 705 of the i-th round can be output. The reference vector 701, the second decoded vector set 706 of the i-th round, and the fourth decoded vector set 707 of the i-th round can be input into the second loss function 708, and the second loss function value 709 of the i-th round can be output.

[0184] Based on the first loss function value 705 and the second loss function value 709 of the i-th round, adjust the model parameters of the second encoder 710_1 and the decoder 710_2 of the (i-1)-th round to obtain the second encoder and decoder of the i-th round.

[0185] According to embodiments of this disclosure, the second loss function value for the i-th round is obtained based on the second loss function, the reference vector, the second decoding vector set for the i-th round, and the fourth decoding vector set for the i-th round. This may include the following operations.

[0186] Determine the similarity between the baseline vector and the second decoded vector set of round i to obtain the first similarity set of round i. Determine the similarity between the baseline vector and the fourth decoded vector set of round i to obtain the second similarity set of round i. Based on the second loss function, and according to the first and second similarity sets of round i, obtain the value of the second loss function for round i.

[0187] According to embodiments of this disclosure, the similarity between a reference vector and each second decoded vector in the second decoded vector set of the i-th round can be determined based on a predetermined similarity determination method, resulting in multiple first similarities for the i-th round, i.e., a first similarity set for the i-th round. Similarly, the similarity between a reference vector and each fourth decoded vector in the fourth decoded vector set of the i-th round can be determined based on the predetermined similarity determination method, resulting in multiple second similarities for the i-th round, i.e., a second similarity set for the i-th round.

[0188] According to embodiments of this disclosure, a first similarity can characterize the degree of similarity between the reference vector and the second decoding vector in the i-th round. A second similarity can characterize the degree of similarity between the reference vector and the fourth decoding vector in the i-th round. The values ​​of the first and second similarities and the relationship between the degree of similarity can be configured according to actual business needs, and are not limited herein.

[0189] For example, a higher first similarity value indicates a greater similarity between the base vector and the second decoded vector in round i. Conversely, a lower first similarity value indicates a lower similarity between the base vector and the second decoded vector in round i. Alternatively, a lower first similarity value indicates a greater similarity between the base vector and the second decoded vector in round i. Conversely, a higher first similarity value indicates a lower similarity between the base vector and the second decoded vector in round i. The first similarity value can be configured according to actual business needs and is not limited here.

[0190] For example, a higher second similarity value indicates a greater similarity between the baseline vector and the fourth decoded vector in the i-th round. Conversely, a lower second similarity value indicates a lower similarity between the baseline vector and the fourth decoded vector in the i-th round. Alternatively, a lower second similarity value indicates a greater similarity between the baseline vector and the fourth decoded vector in the i-th round. Conversely, a higher second similarity value indicates a lower similarity between the baseline vector and the fourth decoded vector in the i-th round. The second similarity value can be configured according to actual business needs and is not limited here.

[0191] According to embodiments of this disclosure, during the training of the second encoder and decoder, the similarity between the reference vector and the second decoder vector in the i-th round can be made to increase, while the similarity between the reference vector and the fourth decoder vector in the i-th round can be made to decrease, and the similarity between the reference vector and the second decoder vector in the i-th round is greater than the similarity between the reference vector and the fourth decoder vector in the i-th round.

[0192] According to embodiments of this disclosure, the second loss function can be determined according to the following formula (2).

[0193]

[0194] According to embodiments of this disclosure, s k It can represent the reference vector. P * P can represent the first encoder. P can represent the initial first encoder. G can represent the decoder. L2 can represent the second loss function.

[0195] According to embodiments of this disclosure, D k This can characterize the second sample protein dataset. |D k | can represent the number of second-sample protein data points included in the second-sample protein dataset. x k,j It can characterize the j-th second-sample protein data x in the second-sample protein dataset. k P * (x k,j ) can characterize x k,j The second encoded vector. G(P) * (x k,j )) can characterize P * (x k,j The second decoding vector of ). P(x) k,j ) can characterize x k,j The fourth encoded vector. G(P(x) k,j )) can characterize P(x k,j The fourth decoding vector of ). sim(G(P) * (x k,j )), s k ) can characterize s k With G(P) * (x k,j The first similarity between sim(G(P(x)) k,j )), s k ) can characterize s k With G(P(x) k,j The second similarity between )) |D k | can be an integer greater than or equal to 1.

[0196] According to embodiments of this disclosure, the first encoder P of the (i-1)th round can be utilized. * Processing with the second sample protein dataset D k The j-th second sample protein data x k This yields protein data x for each second sample. k The corresponding second encoding vector P* (x k,j The initial first encoder P can be used to process the second sample protein dataset D in the i-th round. k The j-th second sample protein data x k This yields protein data x for each second sample. k Each corresponding fourth encoding vector P(x) k,j ).

[0197] According to embodiments of this disclosure, the decoder G of the (i-1)th round can be used to process the j-th second encoded vector P in the second encoded vector set of the i-th round. * (x k,j ), to obtain the second encoding vector P of the i-th round and the j-th second encoding vector. * (x k,j The second decoding vector G(P) corresponding to ) * (x k,j The decoder G of the (i-1)th round can be used to process the j-th fourth encoded vector P(x) in the fourth encoded vector set of the i-th round. k,j ), to obtain the fourth encoding vector P(x) of the i-th round and the j-th encoding vector. k,j The fourth decoding vector G(P(x) corresponds to k,j )).

[0198] According to embodiments of this disclosure, a reference vector s can be determined. k and the j-th second decoding vector G(P) in the i-th round * (x k,j The similarity between )) is used to obtain the j-th first similarity sim(G(P) in the i-th round. * (x k,j )), s k The reference vector s can be determined. k and the j-th fourth decoding vector G(P(x) in the i-th round k,j The similarity between )) is used to obtain the j-th second similarity sim(G(P(x)) in the i-th round. k,j )), s k It can be based on the second loss function L2, according to the j-th first similarity sim(G(P) in the i-th round). * (x k,j )), s k ) and the j-th second similarity in the i-th round sim(G(P(x) k,j )), s k ), thus obtaining the second loss function value for the i-th round.

[0199] Figure 7BThe illustration shows an example diagram of a method for obtaining the value of the second loss function in the i-th round based on a second loss function according to an embodiment of the present disclosure, using a reference vector, a second decoding vector set in the i-th round, and a fourth decoding vector set in the i-th round.

[0200] like Figure 7B As shown in Figure 700B, the similarity between the baseline vector 711 and the second decoded vector set 712 of the i-th round can be determined, resulting in the first similarity set 713 of the i-th round. The similarity between the baseline vector 711 and the fourth decoded vector set 714 of the i-th round is determined, resulting in the second similarity set 715 of the i-th round. The first similarity set 713 and the second similarity set 715 of the i-th round can be input into the second loss function 716, and the second loss function value 717 of the i-th round is output.

[0201] According to embodiments of this disclosure, the first loss function value for the i-th round is obtained based on the first loss function, the reference vector, the first decoding vector set for the i-th round, and the third decoding vector set for the i-th round. This may include the following operations.

[0202] Determine the similarity between the baseline vector and the first decoded vector set of round i to obtain the third similarity set of round i. Determine the similarity between the baseline vector and the third decoded vector set of round i to obtain the fourth similarity set of round i. Based on the first loss function, and according to the third and fourth similarity sets of round i, obtain the value of the first loss function for round i.

[0203] According to embodiments of this disclosure, the similarity between a reference vector and each first decoded vector in the first decoded vector set of the i-th round can be determined based on a predetermined similarity determination method, resulting in multiple third similarities for the i-th round, i.e., a set of third similarities for the i-th round. Based on the predetermined similarity determination method, the similarity between the reference vector and each third decoded vector in the third decoded vector set of the i-th round can be determined, resulting in multiple fourth similarities for the i-th round, i.e., a set of fourth similarities for the i-th round.

[0204] According to embodiments of this disclosure, the third similarity can characterize the degree of similarity between the reference vector and the first decoded vector in the i-th round. The fourth similarity can characterize the degree of similarity between the reference vector and the third decoded vector in the i-th round. The values ​​of the third and fourth similarities and the relationship between the degree of similarity can be configured according to actual business needs, and are not limited herein.

[0205] For example, a higher third similarity value indicates a greater similarity between the reference vector and the first decoded vector in round i. Conversely, a lower third similarity value indicates a lower similarity between the reference vector and the first decoded vector in round i. Alternatively, a lower third similarity value indicates a greater similarity between the reference vector and the first decoded vector in round i. Conversely, a higher third similarity value indicates a lower similarity between the reference vector and the first decoded vector in round i. The third similarity value can be configured according to actual business needs and is not limited here.

[0206] For example, a higher fourth similarity value indicates a greater similarity between the baseline vector and the third decoded vector in round i. Conversely, a lower fourth similarity value indicates a lower similarity between the baseline vector and the third decoded vector in round i. Alternatively, a lower fourth similarity value indicates a greater similarity between the baseline vector and the third decoded vector in round i. Conversely, a higher fourth similarity value indicates a lower similarity between the baseline vector and the third decoded vector in round i. The fourth similarity value can be configured according to actual business needs and is not limited here.

[0207] According to embodiments of this disclosure, during the training of the second encoder and decoder, the similarity between the reference vector and the first decoded vector in the i-th round can be made to decrease, and the similarity between the reference vector and the third decoded vector in the i-th round can be made to decrease.

[0208] According to embodiments of this disclosure, the first loss function can be determined according to the following formula (3).

[0209]

[0210] According to embodiments of this disclosure, s k It can represent the reference vector. P * P can represent the first encoder. P can represent the initial first encoder. G can represent the decoder. L1 can represent the first loss function.

[0211] According to embodiments of this disclosure, D t It can characterize the third sample protein dataset. |D t | can represent the number of third-sample protein data points included in the third-sample protein dataset. x t,m It can characterize the m-th third-sample protein data x in the third-sample protein dataset. t P * (x t,m ) can characterize x t,m The first encoded vector. G(P) * (x t,m )) can characterize P * (x t,m The first decoded vector of ). P(x)t,m ) can characterize x t,m The third encoded vector. G(P(x) t,m )) can characterize P(x t,m The third decoding vector of ). sim(G(P) * (x t,m )), s k ) can characterize s k With G(P) * (x t,m The third similarity between sim(G(P(x)) t,m )), s k ) can characterize s k With G(P(x) t,m The fourth similarity between them. |D t | can be an integer greater than or equal to 1.

[0212] According to embodiments of this disclosure, the first encoder P of the (i-1)th round can be utilized. * Processing with the third sample protein dataset D t The m-th third sample protein data x t This yields protein data x for each m-th third sample. t The corresponding first encoding vector P * (x t,m The initial first encoder P can be used to process the third sample protein dataset D in the i-th round. t The m-th third sample protein data x t Obtain protein data x for each third sample t Each corresponding third encoding vector P(x) t,m ).

[0213] According to embodiments of this disclosure, the decoder G of the (i-1)th round can be used to process the m-th first encoded vector P in the first encoded vector set of the i-th round. * (x t,m ), to obtain the first encoding vector P of the i-th round and the m-th first encoding vector. * (x t,m The first decoded vector G(P) corresponding to ) * (x t,m The decoder G of the (i-1)th round can be used to process the m-th third-coded vector P(x) in the third-coded vector set of the i-th round. t,m ), to obtain the third encoding vector P(x) of the i-th round and the m-th third encoding vector. t,m The third decoding vector G(P(x)) corresponds to t,m )).

[0214] According to embodiments of this disclosure, a reference vector s can be determined.k and the m-th first decoding vector G(P) in the i-th round * (x t,m The similarity between )) is used to obtain the m-th third similarity sim(G(P) in the i-th round. * (x t,m )), s k The reference vector s can be determined. k and the m-th third decoding vector G(P(x) in the i-th round t,m The similarity between )) is used to obtain the m-th fourth similarity sim(G(P(x)) in the i-th round. t,m )), s k Based on the first loss function L1, and according to the m-th third similarity sim(G(P) in the i-th round), * (x t,m )), s k ) and the m-th fourth similarity in the i-th round sim(G(P(x) t,m )), s k ), thus obtaining the first loss function value for the i-th round.

[0215] Figure 7C The illustration shows an example diagram of a method according to an embodiment of the present disclosure for obtaining the value of the first loss function in the i-th round based on a first loss function, a reference vector, a first decoding vector set in the i-th round, and a third decoding vector set in the i-th round.

[0216] like Figure 7C As shown, in 700C, the similarity between the baseline vector 718 and the first decoded vector set 719 of the i-th round can be determined, resulting in the third similarity set 720 of the i-th round. The similarity between the baseline vector 718 and the third decoded vector set 721 of the i-th round is then determined, resulting in the fourth similarity set 722 of the i-th round. The third similarity set 720 and the fourth similarity set 722 of the i-th round can then be input into the first loss function 723, outputting the first loss function value 724 of the i-th round.

[0217] According to embodiments of this disclosure, training the first encoder for the (i-1)th round using a first sample protein dataset, a third sample protein dataset, a reference vector, a first encoder for the (i-1)th round, and a decoder for the i-th round to obtain the first encoder for the i-th round may include the following operations.

[0218] The first encoder in round i-1 processes the first sample protein dataset to obtain the fifth encoded vector set corresponding to the first sample protein dataset in round i. The decoder in round i processes the first encoded vector set corresponding to the third sample protein dataset and the fifth encoded vector set corresponding to the first sample protein dataset in round i, respectively, to obtain the fifth decoded vector set corresponding to the first encoded vector set and the sixth decoded vector set corresponding to the fifth encoded vector set in round i. Based on the baseline vector, the first encoded vector set corresponding to the third sample protein dataset in round i, the third encoded vector set corresponding to the third sample protein dataset, the fifth decoded vector set corresponding to the first encoded vector set, and the sixth decoded vector set corresponding to the fifth encoded vector set, the model parameters of the first encoder in round i-1 are adjusted to obtain the first encoder in round i.

[0219] According to embodiments of this disclosure, the first encoder in the (i-1)th round may include a sixth input layer and an eighth hidden layer. The first encoder in the (i-1)th round can be used to encode a first sample protein dataset. For example, the sixth input layer of the first encoder in the (i-1)th round can be used to encode the first sample protein dataset to obtain a seventh intermediate encoded dataset. The eighth hidden layer of the first encoder is then used to process the seventh intermediate encoded dataset to obtain a fifth encoded vector set.

[0220] According to embodiments of this disclosure, the decoder in the i-th round may include a ninth hidden layer and a third output layer. The decoder in the i-th round can be used to reconstruct the first and fifth encoded vector sets of the i-th round. For example, the ninth hidden layer of the decoder in the i-th round can be used to decode the first encoded vector set of the i-th round to obtain a seventh auxiliary decoded dataset. The third output layer of the decoder is used to process the seventh auxiliary decoded dataset to obtain a fifth decoded vector set. The ninth hidden layer of the decoder in the i-th round can be used to decode the fifth encoded vector set to obtain an eighth auxiliary decoded dataset. The third output layer of the decoder is used to process the eighth auxiliary decoded dataset to obtain a sixth decoded vector set.

[0221] According to embodiments of this disclosure, after obtaining the first encoding vector set, the third encoding vector set, the fifth decoding vector set, and the sixth decoding vector set of the i-th round, the model parameters of the first encoder in the (i-1)-th round can be adjusted based on the reference vector, the first encoding vector set, the third encoding vector set, the fifth decoding vector set, and the sixth decoding vector set of the i-th round to obtain the first encoder of the i-th round.

[0222] Figure 8The illustration shows an example of a method according to an embodiment of the present disclosure for training the first encoder in the (i-1)th round using a first sample protein dataset, a third sample protein dataset, a reference vector, a first encoder in the (i-1)th round, and a decoder in the i-th round to obtain the first encoder in the i-th round.

[0223] like Figure 8 As shown, in 800, the first encoder 802_1 of the (i-1)th round can be used to process the first sample protein dataset 801 to obtain the fifth encoding vector set 803 corresponding to the first sample protein dataset 801 in the i-th round.

[0224] The decoder 802_2 of the i-th round processes the first encoding vector set 805 corresponding to the third sample protein dataset and the fifth encoding vector set 803 corresponding to the first sample protein dataset 801 of the i-th round, respectively, to obtain the fifth decoding vector set 806 corresponding to the first encoding vector set 805 and the sixth decoding vector set 804 corresponding to the fifth encoding vector set 803 of the i-th round.

[0225] Based on the baseline vector 807, the first encoding vector set 805 corresponding to the third sample protein dataset in the i-th round, the third encoding vector set 808 corresponding to the third sample protein dataset, the fifth decoding vector set 806 corresponding to the first encoding vector set 805, and the sixth decoding vector set 804 corresponding to the fifth encoding vector set 803, the model parameters of the first encoder 802_1 in the (i-1)-th round are adjusted to obtain the first encoder in the i-th round.

[0226] According to embodiments of this disclosure, adjusting the model parameters of the first encoder in the (i-1)th round based on the baseline vector, the first encoding vector set corresponding to the third sample protein dataset in the i-th round, the third encoding vector set corresponding to the third sample protein dataset, the fifth decoding vector set corresponding to the first encoding vector set, and the sixth decoding vector set corresponding to the fifth encoding vector set, to obtain the first encoder in the i-th round, may include the following operations.

[0227] Based on the third loss function, the value of the third loss function for round i is obtained according to the first encoding vector set corresponding to the third sample protein dataset and the third encoding vector set corresponding to the third sample protein dataset. Based on the fourth loss function, the value of the fourth loss function for round i is obtained according to the baseline vector, the fifth decoding vector set corresponding to the first encoding vector set and the sixth decoding vector set corresponding to the fifth encoding vector set. Based on the values ​​of the third and fourth loss functions for round i, the model parameters of the first encoder for round i-1 are adjusted to obtain the first encoder for round i.

[0228] According to embodiments of this disclosure, the third and fourth loss functions can be configured according to actual business needs, and are not limited herein. For example, the third and fourth loss functions may include at least one of the following: regression loss function, exponential loss function, squared error loss function, absolute error loss function, cross-entropy loss function, Hinge loss function and Huber loss function, binary classification loss function, binary cross-entropy loss function, multi-class loss function and multi-class cross-entropy loss function.

[0229] According to embodiments of this disclosure, the third loss function value for the i-th round can be obtained based on the third loss function and using the first and third encoding vector sets of the i-th round. The fourth loss function value for the i-th round is obtained based on the second loss function and using the reference vector, the fifth and sixth decoding vector sets of the i-th round. The second target loss function value for the i-th round is determined based on the third and fourth loss function values ​​of the i-th round.

[0230] According to embodiments of this disclosure, after obtaining the second objective loss function value for the i-th round, the model parameters of the first encoder for the (i-1)-th round can be adjusted based on the second objective loss function value for the i-th round until a predetermined termination condition is met. For example, the model parameters of the predetermined model can be adjusted according to the backpropagation algorithm or the stochastic gradient descent algorithm until the predetermined termination condition is met. The first encoder for the (i-1)-th round obtained under the condition of meeting the predetermined termination condition is determined as the candidate first encoder for the i-th round. The predetermined termination condition may include at least one of output value convergence and the number of training rounds reaching the maximum number of training rounds.

[0231] Figure 9A The illustration shows an example of a method according to an embodiment of the present disclosure, which adjusts the model parameters of the first encoder in the (i-1)th round to obtain the first encoder in the i-th round based on a reference vector, a first encoding vector set corresponding to the third sample protein dataset in the i-th round, a third encoding vector set corresponding to the third sample protein dataset, a fifth decoding vector set corresponding to the first encoding vector set, and a sixth decoding vector set corresponding to the fifth encoding vector set.

[0232] like Figure 9AAs shown in Figure 900A, the first encoding vector set 902 and the third encoding vector set 903 corresponding to the third sample protein dataset in the i-th round can be input into the third loss function 904, outputting the third loss function value 905 for the i-th round. The baseline vector 901, the fifth decoding vector set 906 corresponding to the first encoding vector set 902 in the i-th round, and the sixth decoding vector set 907 corresponding to the fifth encoding vector set can be input into the fourth loss function 908, outputting the fourth loss function value 909 for the i-th round. Based on the third loss function value 905 and the fourth loss function value 909 for the i-th round, the model parameters of the first encoder 910_1 in the (i-1)-th round can be adjusted to obtain the first encoder for the i-th round.

[0233] According to embodiments of this disclosure, the fourth loss function value for the i-th round is obtained based on the reference vector, the fifth decoding vector set corresponding to the first encoding vector set in the i-th round, and the sixth decoding vector set corresponding to the fifth encoding vector set. This can include the following operations.

[0234] Determine the similarity between the baseline vector and the fifth decoded vector set corresponding to the first encoded vector set in the i-th round, thus obtaining the fifth similarity set for the i-th round. Determine the similarity between the baseline vector and the sixth decoded vector set corresponding to the fifth encoded vector set in the i-th round, thus obtaining the sixth similarity set for the i-th round. Based on the fourth loss function, and according to the fifth and sixth similarity sets in the i-th round, obtain the value of the fourth loss function for the i-th round.

[0235] According to embodiments of this disclosure, the similarity between a reference vector and each fifth decoded vector in the fifth decoded vector set of the i-th round can be determined based on a predetermined similarity determination method, resulting in multiple fifth similarities for the i-th round, i.e., a set of fifth similarities for the i-th round. Similarly, the similarity between a reference vector and each sixth decoded vector in the sixth decoded vector set of the i-th round can be determined based on the predetermined similarity determination method, resulting in multiple sixth similarities for the i-th round, i.e., a set of sixth similarities for the i-th round.

[0236] According to embodiments of this disclosure, the fifth similarity can characterize the degree of similarity between the reference vector and the fifth decoded vector in the i-th round. The sixth similarity can characterize the degree of similarity between the reference vector and the sixth decoded vector in the i-th round. The numerical values ​​of the fifth and sixth similarities and the relationship between their similarity degrees can be configured according to actual business needs, and are not limited herein.

[0237] According to embodiments of this disclosure, during the training of the first encoder, the similarity between the reference vector and the sixth decoder vector in the i-th round can be increased, the similarity between the reference vector and the fifth decoder vector in the i-th round can be decreased, and the similarity between the reference vector and the sixth decoder vector in the i-th round is greater than the similarity between the reference vector and the fifth decoder vector in the i-th round.

[0238] According to embodiments of this disclosure, the fourth loss function can be determined according to the following formula (4).

[0239]

[0240] According to embodiments of this disclosure, s k It can represent the reference vector. P * G can represent the first encoder. G can represent the decoder. L4 can represent the fourth loss function.

[0241] According to embodiments of this disclosure, D t It can characterize the third sample protein dataset. |D t | can represent the number of third-sample protein data points included in the third-sample protein dataset. x t,m It can characterize the m-th third-sample protein data x in the third-sample protein dataset. t P * (x t,m ) can characterize x t,m The first encoded vector. G(P) * (x t,m )) can characterize P * (x t,m The fifth decoding vector of ). sim(G(P) * (x t,m )), s k ) can characterize s k With G(P) * (x t,m The fifth similarity between )) is |D t | can be an integer greater than or equal to 1.

[0242] According to embodiments of this disclosure, D p This can characterize the first sample protein dataset. |D p | can represent the number of first-sample protein data points included in the first-sample protein dataset. x p,j It can characterize the j-th first-sample protein data x in the first-sample protein dataset. p P * (x p,j ) can characterize x p,j The fifth encoded vector. G(P)* (x p,j )) can characterize P * (x p,j The sixth decoded vector of ). sim(G(P) * (x p,j )), s k ) can characterize s k With G(P) * (x p,j The sixth similarity between )) |D p | can be an integer greater than or equal to 1.

[0243] According to embodiments of this disclosure, the first encoder P of the i-th round can be utilized. * Processing with the third sample protein dataset D t The m-th third sample protein data x t Obtain protein data x for each third sample t The corresponding first encoding vector P * (x t,m The decoder G of the i-th round can be used to process the m-th first encoded vector P in the first encoded vector set of the i-th round. * (x t,m ), to obtain the first encoding vector P of the i-th round and the m-th first encoding vector. * (x t,m The fifth decoding vector G(P) corresponds to * (x t,m )).

[0244] According to embodiments of this disclosure, a reference vector s can be determined. k and the m-th fifth decoding vector G(P) in the i-th round * (x t,m The similarity between )) is used to obtain the m-th fifth similarity sim(G(P) in the i-th round. * (x t,m )), s k The reference vector s can be determined. k and the m-th sixth decoding vector G(P) in the i-th round * (x p,j The similarity between )) is used to obtain the m-th sixth similarity sim(G(P) in the i-th round. * (x p,j )), s k It can be based on the second loss function L4, according to the m-th fifth similarity sim(G(P) in the i-th round). * (x t,m )), s k The sixth similarity between the i-th round and the m-th round is sim(G(P)). * (xp,j )), s k ), thus obtaining the fourth loss function value for the i-th round.

[0245] Figure 9B The illustration shows an example of a method according to an embodiment of the present disclosure for obtaining the value of the fourth loss function in the i-th round based on the fourth loss function, a reference vector, a fifth decoding vector set corresponding to the first encoding vector set in the i-th round, and a sixth decoding vector set corresponding to the fifth encoding vector set.

[0246] like Figure 9B As shown, in 900B, the similarity between the baseline vector 911 and the fifth decoded vector set 912 corresponding to the first encoded vector set in the i-th round can be determined, resulting in the fifth similarity set 913 for the i-th round. The similarity between the baseline vector 911 and the sixth decoded vector set 914 corresponding to the fifth encoded vector set in the i-th round can be determined, resulting in the sixth similarity set 915 for the i-th round. The fifth similarity set 913 and the sixth similarity set 915 for the i-th round can be input into the fourth loss function 916, outputting the fourth loss function value 917 for the i-th round.

[0247] According to embodiments of this disclosure, the value of the third loss function in the i-th round is obtained based on the third loss function, according to the first encoding vector set corresponding to the third sample protein dataset in the i-th round and the third encoding vector set corresponding to the third sample protein dataset. This may include the following operations.

[0248] Determine the similarity between the first encoding vector set corresponding to the third protein dataset in round i and the third encoding vector set corresponding to the third protein dataset in round i, thus obtaining the seventh similarity set in round i. Based on the third loss function, obtain the value of the third loss function in round i according to the seventh similarity set in round i.

[0249] According to embodiments of this disclosure, the similarity between a reference vector and each first encoded vector in the first encoded vector set of the i-th round can be determined based on a predetermined similarity determination method, resulting in multiple seventh similarities for the i-th round, i.e., a set of seventh similarities for the i-th round. The seventh similarity can characterize the degree of similarity between the reference vector and the first encoded vectors of the i-th round. The relationship between the value of the seventh similarity and the degree of similarity can be configured according to actual business needs and is not limited herein.

[0250] According to embodiments of this disclosure, during the training of the first encoder, the similarity between the first encoding vector in the i-th round and the third encoding vector in the i-th round can be made to increase.

[0251] According to embodiments of this disclosure, the third loss function can be determined according to the following formula (5).

[0252]

[0253] According to embodiments of this disclosure, P * P can represent the first encoder. P can represent the initial first encoder, and L3 can represent the third loss function.

[0254] According to embodiments of this disclosure, D t It can characterize the third sample protein dataset. |D t | can represent the number of third-sample protein data points included in the third-sample protein dataset. x t,m It can characterize the m-th third-sample protein data x in the third-sample protein dataset. t P * (x t,m ) can characterize x t , m The first encoded vector. P(x) t,m ) can characterize x t , m The third encoded vector. sim(P) * (x t,m ), P(x t,m )) can characterize P * (x t,m ) and P(x t,m The seventh similarity between them. |D t | can be an integer greater than or equal to 1.

[0255] According to embodiments of this disclosure, the first encoder P of the i-th round can be utilized. * Processing with the third sample protein dataset D t The m-th third sample protein data x t This yields protein data x for each m-th third sample. t The corresponding first encoding vector P * (x t,m ) and the third encoding vector P(x t,m It is possible to determine the m-th first encoding vector and the m-th third encoding vector P(x) in the i-th round. t,m The similarity between ) is used to obtain the m-th seventh similarity sim(P) in the i-th round. * (x t,m ), P(x t,m It can be based on the third loss function L3, according to the m-th seventh similarity sim(P) in the i-th round. * (x t,m ), P(x t,m )), thus obtaining the third loss function value for the i-th round.

[0256] Figure 9C The illustration shows an example of a method for obtaining the value of the third loss function in the i-th round based on the third loss function according to an embodiment of the present disclosure, using the first encoding vector set corresponding to the third sample protein dataset in the i-th round and the third encoding vector set corresponding to the third sample protein dataset.

[0257] like Figure 9C As shown in Figure 900C, the similarity between the first encoding vector set 918 and the third encoding vector set 919 corresponding to the third sample protein dataset in the i-th round can be determined, resulting in the seventh similarity set 920 for the i-th round. The seventh similarity set 920 for the i-th round can then be input into the third loss function 921, outputting the third loss function value 922 for the i-th round.

[0258] Figure 10 The illustration shows an example diagram of a method for determining detection information of a model to be detected based on a reference vector and a target decoding vector set according to an embodiment of the present disclosure.

[0259] like Figure 10 As shown, in step 1000, the first sample protein dataset 1001 and the predetermined mask data 1002 are input to the second encoder 1003_1 included in the key tuple 1003, and the output is the detected protein dataset 1004. The detected protein dataset 1004 is processed by the model to be detected 1005 to obtain the sixth target encoding vector set 1006. The sixth target encoding vector set 1006 is processed by the decoder 1003_2 included in the key tuple 1003 to obtain the target decoding vector set 1007. Based on the reference vector 1008 and the target decoding vector set 1007, the detection information 1009 of the model to be detected is determined.

[0260] According to embodiments of this disclosure, the first deep learning model may include a pre-trained language model.

[0261] According to embodiments of this disclosure, the pre-trained language model can be configured according to actual business needs, and is not limited herein. For example, the pre-trained language model may include at least one of the following: BERT (Bidirectional Encoder Representations from Transformers), ERNIE (Enhanced Language Representation with Informative Entities), RoBERTa (A Robustly Optimized BERT) model, and MASS (MAsked Sequence to Sequence Pretraining) model.

[0262] Figure 11 A flowchart illustrating a method for training a protein task model according to an embodiment of the present disclosure is shown.

[0263] like Figure 11 As shown, the method 1100 includes operation S1110.

[0264] In operation S1110, the watermark protein pre-training model is trained using the fourth sample protein dataset to obtain the protein task model.

[0265] According to embodiments of this disclosure, the watermark protein pre-training model can be obtained by training using the training method of the pre-training model described in embodiments of this disclosure.

[0266] According to embodiments of this disclosure, a pre-trained model for watermarked proteins can be trained using a fourth sample protein dataset to obtain a protein task model. The fourth sample protein dataset can be a dataset with downstream specific task labels. The protein task model can be used to perform a specific prediction task. The prediction task can include at least one of the following: a secondary structure prediction task and a transmembrane protein localization task.

[0267] According to embodiments of this disclosure, since the watermarked protein pre-training model is obtained using the training method of the pre-training model, by training the watermarked protein pre-training model using the fourth sample protein dataset, a protein task model injected with watermark information can be obtained, thereby ensuring the security of the protein task model.

[0268] The above are merely exemplary embodiments, but are not limited thereto. Other training methods for protein task models known in the art may also be included, as long as the safety of the protein task model can be guaranteed.

[0269] Figure 12A flowchart illustrating a protein task processing method according to an embodiment of the present disclosure is shown schematically.

[0270] like Figure 12 As shown, the method 1200 includes operation S1210.

[0271] In operation S1210, the data of the second target protein is input into the protein task model to obtain the output information.

[0272] According to embodiments of this disclosure, a protein task model can be trained using the protein task model training method described in embodiments of this disclosure.

[0273] According to embodiments of this disclosure, a protein task model can be trained using a fourth sample protein dataset to obtain a pre-trained watermark protein model. The protein task model is then used to process the second target protein data to obtain output information.

[0274] According to embodiments of this disclosure, since the protein task model is obtained using a protein task model training method, by inputting the second target protein data into the protein task model, prediction of the second target protein data can be achieved, thereby improving the efficiency and accuracy of protein task processing.

[0275] The above are merely exemplary embodiments, but are not limited thereto. Other protein task processing methods known in the art may also be included, as long as they can improve the efficiency and accuracy of protein task processing.

[0276] Figure 13 A block diagram of a data processing apparatus according to an embodiment of the present disclosure is shown schematically.

[0277] like Figure 13 As shown, the data processing device 1300 may include an acquisition module 1310 and a first acquisition module 1320.

[0278] The acquisition module 1310 is used to acquire the data of the first target protein.

[0279] The first acquisition module 1320 is used to process the first target protein data using a watermarked protein pre-trained model to obtain the first target encoding vector.

[0280] According to embodiments of this disclosure, the watermark protein pre-training model is a first deep learning model that has been trained. A second deep learning model, also trained, is used to detect the relationship between the model to be detected and the watermark protein pre-training model. The first and second deep learning models are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training processes of the first and second deep learning models influence each other.

[0281] According to embodiments of this disclosure, the sample protein dataset includes a first sample protein dataset. The first sample protein dataset is a predetermined privacy dataset.

[0282] According to embodiments of this disclosure, the first obtaining module 1310 may include a first obtaining submodule and a second obtaining submodule.

[0283] The first acquisition submodule is used to obtain the second target coding vector based on the first target protein data.

[0284] The second acquisition submodule is used to process the second target encoding vector using an attention strategy to obtain the first target encoding vector.

[0285] According to embodiments of this disclosure, the second obtaining submodule may include a first obtaining unit, a second obtaining unit, and a third obtaining unit.

[0286] The first acquisition unit is used to process the second target encoding vector based on a self-attention strategy to obtain the third target encoding vector.

[0287] The second obtaining unit is used to obtain the fourth target encoding vector based on the second target encoding vector and the third target encoding vector.

[0288] The third obtaining unit is used to obtain the first target encoding vector based on the fourth target encoding vector.

[0289] According to embodiments of this disclosure, the third obtaining unit may include a first obtaining subunit and a second obtaining subunit.

[0290] The first acquisition subunit is used to process the fourth target encoding vector based on the multilayer perceptron strategy to obtain the fifth target encoding vector.

[0291] The second obtaining subunit is used to obtain the first target encoding vector based on the fourth target encoding vector and the fifth target encoding vector.

[0292] According to embodiments of this disclosure, the first obtaining module 1310 may include a third obtaining submodule.

[0293] The third submodule is used to process the first target protein data based on a long-term dependency information learning strategy to obtain the first target encoding vector. The long-term dependency information learning strategy includes one of a unidirectional long-term dependency information learning strategy or a bidirectional long-term dependency information learning strategy.

[0294] According to embodiments of this disclosure, when the long-term dependent information learning strategy includes a unidirectional long-term dependent information learning strategy, the third acquisition submodule may include a fourth acquisition unit.

[0295] The fourth acquisition unit is used to learn the long-term dependency information of the first target protein data to obtain the first target encoding vector.

[0296] According to embodiments of this disclosure, when the long-term dependency information learning strategy includes a bidirectional long-term dependency information learning strategy, the third acquisition submodule may include a fifth acquisition unit.

[0297] The fifth acquisition unit is used to learn forward long-term dependency information and reverse long-term dependency information from the first target protein data to obtain the first target encoding vector.

[0298] Figure 14 A block diagram of a model detection apparatus according to an embodiment of the present disclosure is shown schematically.

[0299] like Figure 14 As shown, the model detection device 1400 may include a second acquisition module 1410, a third acquisition module 1420, and a first determination module 1430.

[0300] The second acquisition module 1410 is used to process the detection protein dataset using the model to be detected, and obtain the sixth target encoding vector set. The detection protein dataset is obtained based on the first sample protein dataset, predetermined mask data, and key tuples.

[0301] The third acquisition module 1420 is used to process the sixth target encoding vector set using the key tuple to obtain the target decoding vector set.

[0302] The first determining module 1430 is used to determine the detection information of the model to be detected based on the reference vector and the target decoding vector set. The detection information characterizes the relationship between the model to be detected and the watermark protein pre-trained model.

[0303] According to embodiments of this disclosure, the key tuple is a trained second deep learning model. The watermark protein pre-trained model is a trained first deep learning model. The trained second deep learning model and the trained first deep learning model are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training processes of the first deep learning model and the second deep learning model influence each other.

[0304] According to embodiments of this disclosure, the sample protein dataset includes a first sample protein dataset. The first sample protein dataset is a predetermined privacy dataset.

[0305] According to embodiments of this disclosure, the first determining module 1430 may include a fourth obtaining submodule and a first determining submodule.

[0306] The fourth submodule is used to determine the similarity between the baseline vector and the target decoded vector set, thereby obtaining the target similarity set.

[0307] The first determination submodule is used to determine the detection information of the model to be detected based on the target similarity set.

[0308] Figure 15 A block diagram of a training apparatus for a pre-trained model according to an embodiment of the present disclosure is shown schematically.

[0309] like Figure 15 As shown, the training device 1500 for the pre-trained model may include a first training module 1510 and a second determination module 1520.

[0310] The first training module 1510 is used to alternately train the first deep learning model and the second deep learning model using the sample protein dataset, so that the training process of the first deep learning model and the training process of the second deep learning model influence each other.

[0311] The second determination module 1520 is used to determine the first deep learning model that has been trained as the watermark protein pre-training model.

[0312] According to embodiments of this disclosure, the trained second deep learning model is used to detect the relationship between the model to be detected and the watermark protein pre-trained model.

[0313] According to embodiments of this disclosure, the sample protein dataset includes a first sample protein dataset. The first sample protein dataset is a predetermined privacy dataset.

[0314] According to embodiments of this disclosure, the sample protein dataset further includes a second sample protein dataset and a third sample protein dataset. The third sample protein dataset is a pre-determined publicly available dataset.

[0315] According to embodiments of this disclosure, the second sample protein dataset is obtained based on the first sample protein dataset, predetermined mask data, and a second deep learning model.

[0316] According to embodiments of this disclosure, a first deep learning model includes a first encoder. A second deep learning model includes a second encoder and a decoder.

[0317] According to embodiments of this disclosure, the first training module 1510 may include a training submodule.

[0318] The training submodule is used to alternately execute training operations to train the second encoder and decoder using the second sample protein dataset, the third sample protein dataset, the benchmark vector, the first encoder and the decoder, and training operations to train the first encoder using the first sample protein dataset, the third sample protein dataset, the benchmark vector, the first encoder and the decoder, until a predetermined termination condition is met.

[0319] According to embodiments of this disclosure, the training submodule may include a sixth acquisition unit, a seventh acquisition unit, and an eighth acquisition unit.

[0320] The sixth obtaining unit is used to obtain the second sample protein dataset of the i-th round based on the first sample protein dataset, the predetermined mask data and the second encoder of the (i-1)-th round.

[0321] The seventh acquisition unit is used to train the second encoder and decoder of the (i-1)th round using the benchmark vector, the third sample protein dataset, the first encoder of the (i-1)th round, the decoder of the (i-1)th round, and the second sample protein dataset of the i-th round, so as to obtain the second encoder and decoder of the i-th round.

[0322] The eighth acquisition unit is used to train the first encoder of the (i-1)th round using the first sample protein dataset, the third sample protein dataset, the benchmark vector, the first encoder of the (i-1)th round, and the decoder of the i-th round, to obtain the first encoder of the i-th round.

[0323] According to embodiments of this disclosure, i is an integer greater than 1.

[0324] According to embodiments of this disclosure, the sixth obtaining unit may include a third obtaining subunit and a fourth obtaining subunit.

[0325] The third obtaining subunit is used to obtain the first intermediate dataset for the i-th round based on the predetermined mask data and the second encoder of the (i-1)-th round.

[0326] The fourth sub-unit is used to obtain the second sample protein dataset for the i-th round based on the second intermediate dataset and the first intermediate dataset for the i-th round. The second intermediate dataset is obtained based on the first sample protein dataset and predetermined mask data.

[0327] According to embodiments of this disclosure, the seventh obtaining unit may include a fifth obtaining subunit, a sixth obtaining subunit, a seventh obtaining subunit, and a first adjusting subunit.

[0328] The fifth subunit is used to process the third sample protein dataset and the second sample protein dataset of the i-1th round using the first encoder, to obtain the first encoding vector set corresponding to the third sample protein dataset and the second encoding vector set corresponding to the second sample protein dataset of the i-1th round.

[0329] The sixth subunit is used to process the first encoding vector set corresponding to the third sample protein dataset and the second encoding vector set corresponding to the second sample protein dataset in the i-1th round using the decoder, to obtain the first decoding vector set corresponding to the first encoding vector set and the second decoding vector set corresponding to the second encoding vector set in the i-1th round.

[0330] The seventh sub-unit is used to process the third encoded vector set corresponding to the third sample protein dataset and the fourth encoded vector set corresponding to the second sample protein dataset in the i-th round using the decoder of the (i-1)th round, respectively, to obtain the third decoded vector set corresponding to the third encoded vector set and the fourth decoded vector set corresponding to the fourth encoded vector set in the i-th round. The third encoded vector set corresponding to the third sample protein dataset and the fourth encoded vector set corresponding to the second sample protein dataset in the i-th round are obtained by processing the third sample protein dataset and the second sample protein dataset in the i-th round using the initial first encoder.

[0331] The first adjustment subunit is used to adjust the model parameters of the second encoder and decoder in the (i-1)th round based on the reference vector, the first decoding vector set in the i-th round, the third decoding vector set in the i-th round, the second decoding vector set in the i-th round, and the fourth decoding vector set in the i-th round, so as to obtain the second encoder and decoder in the i-th round.

[0332] According to embodiments of this disclosure, the first adjustment subunit can be used for:

[0333] Based on the first loss function, the value of the first loss function for round i is obtained according to the baseline vector, the first decoded vector set for round i, and the third decoded vector set for round i. Based on the second loss function, the value of the second loss function for round i is obtained according to the baseline vector, the second decoded vector set for round i, and the fourth decoded vector set for round i. Based on the first and second loss function values ​​for round i, the model parameters of the second encoder and decoder for round i-1 are adjusted to obtain the second encoder and decoder for round i.

[0334] According to embodiments of this disclosure, the second loss function value for the i-th round is obtained based on the second loss function, the reference vector, the second decoding vector set for the i-th round, and the fourth decoding vector set for the i-th round. This may include the following operations.

[0335] Determine the similarity between the baseline vector and the second decoded vector set of round i to obtain the first similarity set of round i. Determine the similarity between the baseline vector and the fourth decoded vector set of round i to obtain the second similarity set of round i. Based on the second loss function, and according to the first and second similarity sets of round i, obtain the value of the second loss function for round i.

[0336] According to embodiments of this disclosure, the first loss function value for the i-th round is obtained based on the first loss function, the reference vector, the first decoding vector set for the i-th round, and the third decoding vector set for the i-th round. This may include the following operations.

[0337] Determine the similarity between the baseline vector and the first decoded vector set of round i to obtain the third similarity set of round i. Determine the similarity between the baseline vector and the third decoded vector set of round i to obtain the fourth similarity set of round i. Based on the first loss function, and according to the third and fourth similarity sets of round i, obtain the value of the first loss function for round i.

[0338] According to embodiments of this disclosure, the eighth obtaining unit may include an eighth obtaining subunit, a ninth obtaining subunit, and a second adjusting subunit.

[0339] The eighth subunit is used to process the first sample protein dataset using the first encoder in the (i-1)th round to obtain the fifth encoding vector set corresponding to the first sample protein dataset in the i-th round.

[0340] The ninth subunit is used to process the first encoding vector set corresponding to the third sample protein dataset and the fifth encoding vector set corresponding to the first sample protein dataset in the i-th round using the decoder in the i-th round, to obtain the fifth decoding vector set corresponding to the first encoding vector set and the sixth decoding vector set corresponding to the fifth encoding vector set in the i-th round.

[0341] The second adjustment subunit is used to adjust the model parameters of the first encoder in the (i-1)th round based on the baseline vector, the first encoding vector set corresponding to the third sample protein dataset in the i-th round, the third encoding vector set corresponding to the third sample protein dataset, the fifth decoding vector set corresponding to the first encoding vector set, and the sixth decoding vector set corresponding to the fifth encoding vector set, so as to obtain the first encoder in the i-th round.

[0342] According to embodiments of this disclosure, the second adjustment subunit can be used for:

[0343] Based on the third loss function, the value of the third loss function for round i is obtained according to the first encoding vector set corresponding to the third sample protein dataset and the third encoding vector set corresponding to the third sample protein dataset. Based on the fourth loss function, the value of the fourth loss function for round i is obtained according to the baseline vector, the fifth decoding vector set corresponding to the first encoding vector set and the sixth decoding vector set corresponding to the fifth encoding vector set. Based on the values ​​of the third and fourth loss functions for round i, the model parameters of the first encoder for round i-1 are adjusted to obtain the first encoder for round i.

[0344] According to embodiments of this disclosure, the fourth loss function value for the i-th round is obtained based on the reference vector, the fifth decoding vector set corresponding to the first encoding vector set in the i-th round, and the sixth decoding vector set corresponding to the fifth encoding vector set. This can include the following operations.

[0345] Determine the similarity between the baseline vector and the fifth decoded vector set corresponding to the first encoded vector set in the i-th round, thus obtaining the fifth similarity set for the i-th round. Determine the similarity between the baseline vector and the sixth decoded vector set corresponding to the fifth encoded vector set in the i-th round, thus obtaining the sixth similarity set for the i-th round. Based on the fourth loss function, and according to the fifth and sixth similarity sets in the i-th round, obtain the value of the fourth loss function for the i-th round.

[0346] According to embodiments of this disclosure, the value of the third loss function in the i-th round is obtained based on the third loss function, according to the first encoding vector set corresponding to the third sample protein dataset in the i-th round and the third encoding vector set corresponding to the third sample protein dataset. This may include the following operations.

[0347] Determine the similarity between the first encoding vector set corresponding to the third protein dataset in round i and the third encoding vector set corresponding to the third protein dataset in round i, thus obtaining the seventh similarity set in round i. Based on the third loss function, obtain the value of the third loss function in round i according to the seventh similarity set in round i.

[0348] According to embodiments of this disclosure, the first deep learning model includes a pre-trained language model.

[0349] Figure 16 A block diagram of a training apparatus for a protein task model according to an embodiment of the present disclosure is shown schematically.

[0350] like Figure 16 As shown, the training device 1600 for the protein task model may include a second training module 1610.

[0351] The second training module 1610 is used to train the watermark protein pre-training model using the fourth sample protein dataset to obtain the protein task model.

[0352] According to embodiments of this disclosure, the watermark protein pre-training model is obtained by training using the training device of the pre-training model of embodiments of this disclosure.

[0353] Figure 17 A block diagram of a protein task processing apparatus according to an embodiment of the present disclosure is shown schematically.

[0354] like Figure 17 As shown, the protein task processing apparatus 1700 may include a fourth acquisition module 1710.

[0355] The fourth acquisition module 1710 is used to input the second target protein data into the protein task model and obtain the output information.

[0356] According to embodiments of this disclosure, the protein task model is trained using the training apparatus of the protein task model of embodiments of this disclosure.

[0357] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0358] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described in the present disclosure.

[0359] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the present disclosure.

[0360] According to embodiments of the present disclosure, a computer program product includes a computer program that, when executed by a processor, implements the methods described in the present disclosure.

[0361] Figure 18The diagram schematically illustrates an electronic device suitable for implementing data processing methods, model detection methods, pre-trained model training methods, protein task model training methods, and protein task processing methods according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0362] like Figure 18 As shown, the electronic device 1800 includes a computing unit 1801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1802 or a computer program loaded from a storage unit 1808 into a random access memory (RAM) 1803. The RAM 1803 may also store various programs and data required for the operation of the electronic device 1800. The computing unit 1801, ROM 1802, and RAM 1803 are interconnected via a bus 1804. An input / output (I / O) interface 1805 is also connected to the bus 1804.

[0363] Multiple components in electronic device 1800 are connected to I / O interface 1805, including: input unit 1806, such as keyboard, mouse, etc.; output unit 1807, such as various types of displays, speakers, etc.; storage unit 1808, such as disk, optical disk, etc.; and communication unit 1809, such as network card, modem, wireless transceiver, etc. Communication unit 1809 allows electronic device 1800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0364] The computing unit 1801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1801 performs the various methods and processes described above, such as data processing methods, model detection methods, pre-trained model training methods, protein task model training methods, and protein task processing methods. For example, in some embodiments, the data processing methods, model detection methods, pre-trained model training methods, protein task model training methods, and protein task processing methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1800 via ROM 1802 and / or communication unit 1809. When the computer program is loaded into RAM 1803 and executed by computing unit 1801, one or more steps of the data processing method, model detection method, pre-trained model training method, protein task model training method, and protein task processing method described above can be performed. Alternatively, in other embodiments, computing unit 1801 can be configured to perform the data processing method, model detection method, pre-trained model training method, protein task model training method, and protein task processing method by any other suitable means (e.g., by means of firmware).

[0365] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0366] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0367] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0368] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0369] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0370] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0371] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0372] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A data processing method, comprising: Acquire the data for the first target protein; as well as The first target protein data is processed using a watermarked protein pre-trained model to obtain the first target encoding vector; The watermark protein pre-training model is a first deep learning model that has been trained. The second deep learning model that has been trained is used to detect the relationship between the model to be detected and the watermark protein pre-training model. The first deep learning model and the second deep learning model that have been trained are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training process of the first deep learning model and the training process of the second deep learning model influence each other. The sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset. The sample protein dataset further includes a second sample protein dataset and a third sample protein dataset, wherein the third sample protein dataset is a pre-determined publicly available dataset; The second sample protein dataset is obtained based on the first sample protein dataset, the predetermined mask data, and the second deep learning model. The first deep learning model includes a first encoder, and the second deep learning model includes a second encoder and a decoder. The method of alternately training the second deep learning model and the first deep learning model using a sample protein dataset includes: The training operations of training the second encoder and the decoder using the second sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder are performed alternately, as well as the training operation of training the first encoder using the first sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder, until a predetermined termination condition is met.

2. The method according to claim 1, wherein, The process of using a watermarked protein pre-trained model to process the first target protein data to obtain a first target encoding vector includes: Based on the first target protein data, a second target coding vector is obtained; and The second target encoding vector is processed using an attention strategy to obtain the first target encoding vector.

3. The method according to claim 2, wherein, The step of processing the second target encoding vector using an attention strategy to obtain the first target encoding vector includes: The second target encoding vector is processed using a self-attention strategy to obtain the third target encoding vector; Based on the second target encoding vector and the third target encoding vector, a fourth target encoding vector is obtained; and The first target encoding vector is obtained based on the fourth target encoding vector.

4. The method according to claim 3, wherein, The step of obtaining the first target encoding vector based on the fourth target encoding vector includes: The fourth target encoding vector is processed based on a multilayer perceptron strategy to obtain the fifth target encoding vector; and The first target encoding vector is obtained based on the fourth target encoding vector and the fifth target encoding vector.

5. The method according to claim 1, wherein, The process of using a watermarked protein pre-trained model to process the first target protein data to obtain a first target encoding vector includes: The first target protein data is processed based on a long-term dependency information learning strategy to obtain the first target encoding vector, wherein the long-term dependency information learning strategy includes one of a unidirectional long-term dependency information learning strategy and a bidirectional long-term dependency information learning strategy.

6. The method according to claim 5, wherein, When the long-term dependency information learning strategy includes the unidirectional long-term dependency information learning strategy, the step of processing the first target protein data based on the long-term dependency information learning strategy to obtain the first target encoding vector includes: The first target protein data is subjected to positive long-term dependency information learning to obtain the first target encoding vector; Wherein, when the long-term dependency information learning strategy includes the bidirectional long-term dependency information learning strategy, the step of processing the first target protein data based on the long-term dependency information learning strategy to obtain the first target encoding vector includes: The first target protein data is subjected to forward long-term dependency learning and reverse long-term dependency learning to obtain the first target encoding vector.

7. A model detection method, comprising: The detection protein dataset is processed using the model to be detected to obtain a sixth target encoding vector set, wherein the detection protein dataset is obtained based on the first sample protein dataset, predetermined mask data and key tuples; The sixth target encoded vector set is processed using the key tuple to obtain the target decoded vector set; and Based on the baseline vector and the target decoding vector set, the detection information of the model to be detected is determined, wherein the detection information characterizes the relationship between the model to be detected and the watermark protein pre-training model. The key tuple is a trained second deep learning model, the watermark protein pre-trained model is a trained first deep learning model, and the trained second deep learning model and the trained first deep learning model are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training process of the first deep learning model and the training process of the second deep learning model influence each other. The sample protein dataset includes the first sample protein dataset, which is a predetermined privacy dataset.

8. The method according to claim 7, wherein, The step of determining the detection information of the model to be detected based on the reference vector and the target decoding vector set includes: Determine the similarity between the baseline vector and the target decoded vector set to obtain the target similarity set; and Based on the target similarity set, the detection information of the model to be detected is determined.

9. A training method for a pre-trained model, comprising: The first deep learning model and the second deep learning model are trained alternately using a sample protein dataset, so that the training processes of the first deep learning model and the second deep learning model influence each other. as well as The first deep learning model that has been trained is identified as the watermark protein pre-training model. The trained second deep learning model is used to detect the relationship between the model to be detected and the watermark protein pre-trained model. The sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset. The sample protein dataset further includes a second sample protein dataset and a third sample protein dataset, wherein the third sample protein dataset is a pre-determined publicly available dataset; The second sample protein dataset is obtained based on the first sample protein dataset, the predetermined mask data, and the second deep learning model. The first deep learning model includes a first encoder, and the second deep learning model includes a second encoder and a decoder. The step of alternately training a first deep learning model and a second deep learning model using a sample protein dataset, such that the training processes of the first and second deep learning models influence each other, includes: The training operations of training the second encoder and the decoder using the second sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder are performed alternately, as well as the training operation of training the first encoder using the first sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder, until a predetermined termination condition is met.

10. The method according to claim 9, wherein, The alternating execution of training operations using the second sample protein dataset, the third sample protein dataset, the benchmark vector, the first encoder, and the decoder to train the second encoder and the decoder, and training operations using the first sample protein dataset, the third sample protein dataset, the benchmark vector, the first encoder, and the decoder to train the first encoder, until the predetermined termination condition is met, includes: Based on the first sample protein dataset, the predetermined mask data, and the second encoder of the (i-1)th round, the second sample protein dataset of the i-th round is obtained; Using the baseline vector, the third sample protein dataset, the first encoder of round i-1, the decoder of round i-1, and the second sample protein dataset of round i, train the second encoder and decoder of round i-1 to obtain the second encoder and decoder of round i; and The first encoder of the (i-1)th round is trained using the first sample protein dataset, the third sample protein dataset, the baseline vector, the first encoder of the (i-1)th round, and the decoder of the i-th round to obtain the first encoder of the i-th round; Where i is an integer greater than 1.

11. The method according to claim 10, wherein, The step of obtaining the second sample protein dataset for the i-th round based on the first sample protein dataset, the predetermined mask data, and the second encoder for the (i-1)-th round includes: Based on the predetermined mask data and the second encoder of the (i-1)th round, the first intermediate dataset of the i-th round is obtained; and The second sample protein dataset for the i-th round is obtained based on the second intermediate dataset and the first intermediate dataset for the i-th round, wherein the second intermediate dataset is obtained based on the first sample protein dataset and the predetermined mask data.

12. The method according to claim 11, wherein, The process of training the second encoder and decoder for the (i-1)th round using the baseline vector, the third sample protein dataset, the first encoder for the (i-1)th round, the decoder for the (i-1)th round, and the second sample protein dataset for the i-th round, to obtain the second encoder and decoder for the i-th round, includes: The first encoder of the (i-1)th round is used to process the third sample protein dataset and the second sample protein dataset of the i-th round respectively, to obtain the first encoding vector set corresponding to the third sample protein dataset and the second encoding vector set corresponding to the second sample protein dataset of the i-th round; The decoder of the (i-1)th round processes the first encoding vector set corresponding to the third sample protein dataset and the second encoding vector set corresponding to the second sample protein dataset of the i-th round, respectively, to obtain the first decoding vector set corresponding to the first encoding vector set and the second decoding vector set corresponding to the second encoding vector set of the i-th round; The decoder in the (i-1)th round processes the third encoded vector set corresponding to the third sample protein dataset and the fourth encoded vector set corresponding to the second sample protein dataset in the i-th round, respectively, to obtain the third decoded vector set corresponding to the third encoded vector set and the fourth decoded vector set corresponding to the fourth encoded vector set in the i-th round. The third encoded vector set corresponding to the third sample protein dataset and the fourth encoded vector set corresponding to the second sample protein dataset in the i-th round are obtained by processing the third sample protein dataset and the second sample protein dataset in the i-th round using the initial first encoder. Based on the reference vector, the first decoding vector set of the i-th round, the third decoding vector set of the i-th round, the second decoding vector set of the i-th round, and the fourth decoding vector set of the i-th round, the model parameters of the second encoder and decoder of the (i-1)-th round are adjusted to obtain the second encoder and decoder of the i-th round.

13. The method according to claim 12, wherein, The step of adjusting the model parameters of the second encoder and decoder in the (i-1)th round based on the reference vector, the first decoding vector set of the i-th round, the third decoding vector set of the i-th round, the second decoding vector set of the i-th round, and the fourth decoding vector set of the i-th round to obtain the second encoder and decoder in the i-th round includes: Based on the first loss function, the value of the first loss function in the i-th round is obtained according to the reference vector, the first decoding vector set in the i-th round, and the third decoding vector set in the i-th round; Based on the second loss function, and according to the baseline vector, the second decoding vector set of the i-th round, and the fourth decoding vector set of the i-th round, the value of the second loss function for the i-th round is obtained; and Based on the first and second loss function values ​​of the i-th round, the model parameters of the second encoder and decoder in the (i-1)-th round are adjusted to obtain the second encoder and decoder in the i-th round.

14. The method according to claim 13, wherein, The step of obtaining the second loss function value for the i-th round based on the second loss function, according to the baseline vector, the second decoding vector set of the i-th round, and the fourth decoding vector set of the i-th round, includes: Determine the similarity between the baseline vector and the second decoding vector set of the i-th round to obtain the first similarity set of the i-th round; Determine the similarity between the baseline vector and the fourth decoded vector set of the i-th round to obtain the second similarity set of the i-th round; and Based on the second loss function, the value of the second loss function for the i-th round is obtained according to the first similarity set and the second similarity set for the i-th round.

15. The method according to claim 13 or 14, wherein, The step of obtaining the first loss function value for the i-th round based on the first loss function, according to the baseline vector, the first decoding vector set of the i-th round, and the third decoding vector set of the i-th round, includes: Determine the similarity between the baseline vector and the first decoding vector set of the i-th round to obtain the third similarity set of the i-th round; Determine the similarity between the baseline vector and the third decoding vector set of the i-th round to obtain the fourth similarity set of the i-th round; and Based on the first loss function, the value of the first loss function for the i-th round is obtained according to the third and fourth similarity sets of the i-th round.

16. The method according to any one of claims 12 to 14, wherein, The step of training the first encoder for the (i-1)th round using the first sample protein dataset, the third sample protein dataset, the baseline vector, the first encoder for the (i-1)th round, and the decoder for the i-th round to obtain the first encoder for the i-th round includes: The first sample protein dataset is processed using the first encoder in the (i-1)th round to obtain the fifth encoding vector set corresponding to the first sample protein dataset in the i-th round; The decoder of the i-th round processes the first encoding vector set corresponding to the third sample protein dataset and the fifth encoding vector set corresponding to the first sample protein dataset in the i-th round, respectively, to obtain the fifth decoding vector set corresponding to the first encoding vector set and the sixth decoding vector set corresponding to the fifth encoding vector set in the i-th round; and Based on the baseline vector, the first encoding vector set corresponding to the third sample protein dataset in the i-th round, the third encoding vector set corresponding to the third sample protein dataset, the fifth decoding vector set corresponding to the first encoding vector set, and the sixth decoding vector set corresponding to the fifth encoding vector set, the model parameters of the first encoder in the (i-1)-th round are adjusted to obtain the first encoder in the i-th round.

17. The method according to claim 16, wherein, The step of adjusting the model parameters of the first encoder in the (i-1)th round, based on the baseline vector, the first encoding vector set corresponding to the third sample protein dataset in the i-th round, the third encoding vector set corresponding to the third sample protein dataset, the fifth decoding vector set corresponding to the first encoding vector set, and the sixth decoding vector set corresponding to the fifth encoding vector set, to obtain the first encoder in the i-th round, includes: Based on the third loss function, the value of the third loss function for the i-th round is obtained according to the first encoding vector set corresponding to the third sample protein dataset and the third encoding vector set corresponding to the third sample protein dataset in the i-th round. Based on the fourth loss function, the value of the fourth loss function for the i-th round is obtained according to the reference vector, the fifth decoding vector set corresponding to the first encoding vector set in the i-th round, and the sixth decoding vector set corresponding to the fifth encoding vector set; and Based on the third and fourth loss function values ​​of the i-th round, the model parameters of the first encoder in the (i-1)-th round are adjusted to obtain the first encoder of the i-th round.

18. The method according to claim 17, wherein, The step of obtaining the fourth loss function value for the i-th round based on the fourth loss function, according to the reference vector, the fifth decoding vector set corresponding to the first encoding vector set in the i-th round, and the sixth decoding vector set corresponding to the fifth encoding vector set, includes: Determine the similarity between the baseline vector and the fifth decoding vector set corresponding to the first encoding vector set in the i-th round to obtain the fifth similarity set in the i-th round; Determine the similarity between the baseline vector and the sixth decoded vector set corresponding to the fifth encoded vector set in the i-th round, to obtain the sixth similarity set for the i-th round; and Based on the fourth loss function, the value of the fourth loss function in the i-th round is obtained according to the fifth and sixth similarity sets in the i-th round.

19. The method according to claim 17 or 18, wherein, The step of obtaining the third loss function value for the i-th round based on the third loss function, according to the first encoding vector set corresponding to the third sample protein dataset in the i-th round and the third encoding vector set corresponding to the third sample protein dataset, includes: Determine the similarity between the first encoding vector set corresponding to the third sample protein dataset in the i-th round and the third encoding vector set corresponding to the third sample protein dataset, to obtain the seventh similarity set in the i-th round; and Based on the third loss function, the value of the third loss function in the i-th round is obtained according to the seventh similarity set in the i-th round.

20. The method according to any one of claims 9-14 and 17-18, wherein, The first deep learning model includes a pre-trained language model.

21. A training method for a protein task model, comprising: A pre-trained model for watermarking proteins was trained using the fourth sample protein dataset to obtain the protein task model. The watermark protein pre-training model is obtained by training using the method according to any one of claims 9 to 20.

22. A protein task processing method, comprising: Input the data of the second target protein into the protein task model to obtain the output information; The protein task model is trained using the method described in claim 21.

23. A data processing apparatus, comprising: The acquisition module is used to acquire the data of the first target protein; as well as The first acquisition module is used to process the first target protein data using a watermarked protein pre-trained model to obtain a first target encoding vector. The watermark protein pre-training model is a first deep learning model that has been trained. The second deep learning model that has been trained is used to detect the relationship between the model to be detected and the watermark protein pre-training model. The first deep learning model and the second deep learning model that have been trained are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training process of the first deep learning model and the training process of the second deep learning model influence each other. The sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset. The sample protein dataset further includes a second sample protein dataset and a third sample protein dataset, wherein the third sample protein dataset is a pre-determined publicly available dataset; The second sample protein dataset is obtained based on the first sample protein dataset, the predetermined mask data, and the second deep learning model. The first deep learning model includes a first encoder, and the second deep learning model includes a second encoder and a decoder. The step of alternately training the first deep learning model and the second deep learning model using a sample protein dataset includes: The training operations of training the second encoder and the decoder using the second sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder are performed alternately, as well as the training operation of training the first encoder using the first sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder, until a predetermined termination condition is met.

24. A model testing device, comprising: The second acquisition module is used to process the detection protein dataset using the model to be detected to obtain the sixth target encoding vector set, wherein the detection protein dataset is obtained based on the first sample protein dataset, the predetermined mask data and the key tuple; The third obtaining module is used to process the sixth target encoded vector set using the key tuple to obtain the target decoded vector set; and The first determining module is used to determine the detection information of the model to be detected based on the reference vector and the target decoding vector set, wherein the detection information characterizes the relationship between the model to be detected and the watermark protein pre-training model. The key tuple is a trained second deep learning model, the watermark protein pre-trained model is a trained first deep learning model, and the trained second deep learning model and the trained first deep learning model are obtained by alternately training the second deep learning model and the first deep learning model using a sample protein dataset. The training process of the first deep learning model and the training process of the second deep learning model influence each other. The sample protein dataset includes the first sample protein dataset, which is a predetermined privacy dataset.

25. A training device for a pre-trained model, comprising: The first training module is used to alternately train the first deep learning model and the second deep learning model using the sample protein dataset, so that the training process of the first deep learning model and the training process of the second deep learning model influence each other. as well as The second determination module is used to determine the first deep learning model that has been trained as the watermark protein pre-training model. The trained second deep learning model is used to detect the relationship between the model to be detected and the watermark protein pre-trained model. The sample protein dataset includes a first sample protein dataset, which is a predetermined privacy dataset. The sample protein dataset further includes a second sample protein dataset and a third sample protein dataset, wherein the third sample protein dataset is a pre-determined publicly available dataset; The second sample protein dataset is obtained based on the first sample protein dataset, the predetermined mask data, and the second deep learning model. The first deep learning model includes a first encoder, and the second deep learning model includes a second encoder and a decoder. The step of alternately training a first deep learning model and a second deep learning model using a sample protein dataset, such that the training processes of the first and second deep learning models influence each other, includes: The training operations of training the second encoder and the decoder using the second sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder are performed alternately, as well as the training operation of training the first encoder using the first sample protein dataset, the third sample protein dataset, the reference vector, the first encoder, and the decoder, until a predetermined termination condition is met.

26. A training device for a protein task model, comprising: The second training module is used to train the watermark protein pre-training model using the fourth sample protein dataset to obtain the protein task model. The watermark protein pre-training model is obtained by training the device according to claim 25.

27. A protein task processing apparatus, comprising: The fourth acquisition module is used to input the second target protein data into the protein task model and obtain the output information; The protein task model is trained using the apparatus according to claim 26.

28. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 22.

29. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 22.

30. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 22.