Decoder training method, model detection method, device and equipment

By training a molecular graph generation model and a decoder, and combining this with target digital signature matching, the problem of protecting the ownership of molecular graph datasets is solved. This enables the legality detection of the molecular graph generation model, ensuring the legitimate use of the dataset.

CN116959615BActive Publication Date: 2026-05-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-10-14
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, obtaining training datasets for molecular graph generation models is costly and requires protecting their ownership, making it difficult to effectively detect whether unauthorized molecular graph generation models are trained on protected datasets.

Method used

By training a first molecular graph generation model and decoder based on the original molecular graph dataset, and iteratively adjusting the model parameters to generate and decode molecular graphs carrying ownership information, the training source of the molecular graph generation model is detected by matching latent variables with target digital signatures.

Benefits of technology

It achieves effective protection of molecular graph datasets, and can detect whether molecular graph generation models are trained on datasets that carry ownership information, ensuring the legitimate use of dataset ownership.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116959615B_ABST
    Figure CN116959615B_ABST
Patent Text Reader

Abstract

The application provides a training method and a model detection method of a decoder, a device and equipment, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing model training based on an original molecular graph data set to obtain a first molecular graph generation model; and based on an initial molecular graph, iteratively performing the following steps to train the first molecular graph generation model and the decoder to obtain a target molecular graph generation model and a target decoder: in each iteration process, inputting the initial molecular graph into the first molecular graph generation model, obtaining a predicted molecular graph of the current iteration process through the first molecular graph generation model; inputting the predicted molecular graph into the decoder to obtain a first hidden variable of the predicted molecular graph through the decoder; and matching the first hidden variable with a target digital signature, and adjusting the model parameters of the first molecular graph generation model and the decoder based on the matching result. The method can effectively protect the molecular graph data set.
Need to check novelty before this filing date? Find Prior Art

Description

Decoder training methods, model detection methods, devices and equipment Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a decoder training method, model detection method, apparatus and device. Background Technology

[0002] Because chemical space is discrete and vast, manually designing new molecular structures is often time-consuming and labor-intensive. Therefore, molecular graph generation models are used to generate new molecular structures. However, these models need to be trained on molecular graph datasets, which are acquired at considerable cost by their owners. Therefore, it is necessary to protect the ownership of these datasets. This necessitates training a decoder to detect whether unauthorized molecular graph generation models are trained on these protected datasets. Summary of the Invention

[0003] This application provides a decoder training method, model detection method, apparatus, and device, which can effectively protect molecular graph datasets. The technical solution is as follows:

[0004] On the one hand, a method for training a decoder is provided, the method comprising:

[0005] The model is trained based on the original molecular graph dataset to obtain the first molecular graph generation model. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The first molecular graph generation model is used to generate molecular graphs, which are used to represent the molecular structure of any molecule.

[0006] Based on the initial molecular graph, the following steps are iteratively performed to train the first molecular graph generation model and decoder, resulting in the target molecular graph generation model and target decoder:

[0007] In any iteration, the initial molecular graph is input into the first molecular graph generation model, and the predicted molecular graph for this iteration is obtained through the first molecular graph generation model.

[0008] The predicted molecular graph is input into the decoder, and the first latent variable of the predicted molecular graph is obtained through the decoder. The first latent variable is used to describe the ownership information of the predicted molecular graph.

[0009] The first latent variable is matched with the target digital signature. Based on the matching result, the model parameters of the first molecular graph generation model and the decoder are adjusted. The target digital signature is generated based on the ownership information of the original molecular graph dataset.

[0010] On the other hand, a model detection method is provided, the method comprising:

[0011] Multiple molecular diagrams are obtained, which are generated based on a second molecular diagram generation model, and the molecular diagrams are used to represent the molecular structure of any molecule;

[0012] The multiple molecular graphs are input into a target decoder, which maps each molecular graph to a latent variable space to obtain a second latent variable for each molecular graph. The second latent variable is used to describe the ownership information of the molecular graph. The target decoder is trained based on the original molecular graph dataset and the target digital signature. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The target digital signature is generated based on the ownership information of the target molecular graph dataset. The target molecular graph dataset is obtained by embedding ownership information into the original molecular graph dataset.

[0013] The second hidden variables of the plurality of molecular graphs are matched with the target digital signature respectively;

[0014] In the case where a target molecular graph exists among the multiple molecular graphs, it is determined that the second molecular graph generation model is trained based on the target molecular graph dataset, and the second latent variable of the target molecular graph matches the target digital signature.

[0015] On the other hand, a training device for a decoder is provided, the device comprising:

[0016] The training module is used to train the model based on the original molecular graph dataset to obtain the first molecular graph generation model. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The first molecular graph generation model is used to generate molecular graphs, which are used to represent the molecular structure of any molecule.

[0017] The iterative execution module is used to iteratively execute the following steps based on the initial molecular graph to train the first molecular graph generation model and decoder, thereby obtaining the target molecular graph generation model and target decoder:

[0018] In any iteration, the initial molecular graph is input into the first molecular graph generation model, and the predicted molecular graph for this iteration is obtained through the first molecular graph generation model.

[0019] The predicted molecular graph is input into the decoder, and the first latent variable of the predicted molecular graph is obtained through the decoder. The first latent variable is used to describe the ownership information of the predicted molecular graph.

[0020] The first latent variable is matched with the target digital signature. Based on the matching result, the model parameters of the first molecular graph generation model and the decoder are adjusted. The target digital signature is generated based on the ownership information of the original molecular graph dataset.

[0021] In some embodiments, the matching result is a matching parameter, the first hidden variable includes a first number of real numbers with positive and negative signs, the target digital signature includes a first number of binary numbers representing positive and negative signs, the matching parameter refers to the proportion of the target real number among the first number of real numbers, and the positive and negative signs of the target real number and the binary numbers compared with the target real number are matched.

[0022] In any iteration, training is performed based on a second number of initial molecular graphs. The iteration execution module is used to:

[0023] The mean value of the matching parameters of the second number of predicted molecular maps is determined to obtain the target matching parameters;

[0024] Determine the target molecular mass parameters of the second number of predicted molecular maps, wherein the target molecular mass parameters are used to describe the quality of the predicted molecular maps generated in this iteration;

[0025] The target matching parameters and the target molecular mass parameters are weighted and summed to obtain a first reward value, which is used to describe the reward obtained in generating the predicted molecular map during this iteration.

[0026] Based on the first reward value, adjust the model parameters of the first molecular graph generation model and the decoder.

[0027] In some embodiments, the predicted molecular graph includes predicted intermediate subgraphs at multiple stages, each stage being obtained by adding atoms and chemical bonds to the predicted intermediate subgraph of the previous stage.

[0028] The iterative execution module is used for:

[0029] For each stage, the target matching parameter and the target molecular mass parameter of the stage are weighted and summed to obtain the second reward value of the stage.

[0030] The first reward value is obtained by weighted summing of the second reward values ​​of the multiple stages.

[0031] In some embodiments, the target matching parameter for each stage is the mean of the matching parameters of a second number of predicted intermediate subgraphs for that stage.

[0032] In some embodiments, the iterative execution module is configured to:

[0033] For each stage, the mean of the correctness parameters of the second number of predicted intermediate subgraphs of the stage is used as the target correctness parameter of the stage, which is used to describe the chemical correctness of the molecular graph.

[0034] The ratio of the number of first intermediate subgraphs to the number of multiple molecular graph samples is used as the target novelty parameter of the stage. The first intermediate subgraph is the intermediate subgraph that is different from the multiple molecular graph samples in the predicted intermediate subgraphs of the second number.

[0035] The ratio of the number of the second intermediate subgraphs to the second number is used as the target uniqueness parameter of the stage, wherein the second intermediate subgraph is an intermediate subgraph that is not repeated in the first intermediate subgraph.

[0036] The target molecular mass parameter of the stage is obtained by weighted summation of the target correctness parameter, the target novelty parameter, and the target uniqueness parameter.

[0037] In some embodiments, the iterative execution module is configured to:

[0038] Gaussian noise is added to the predicted molecular graph, and the predicted molecular graph with added Gaussian noise is input into the decoder.

[0039] In some embodiments, the apparatus further includes:

[0040] The input module is used to input multiple molecular map samples from the original molecular map dataset into the target molecular map generation model respectively;

[0041] The acquisition module is used to obtain a target molecular graph dataset through the target molecular graph generation model. The target molecular graph dataset includes multiple molecular graph samples carrying ownership information.

[0042] On the other hand, a model detection device is provided, the device comprising:

[0043] An acquisition module is used to acquire multiple molecular diagrams, which are generated based on a second molecular diagram generation model, and the molecular diagrams are used to represent the molecular structure of any molecule.

[0044] An input module is used to input the multiple molecular graphs into a target decoder, which maps the multiple molecular graphs to a latent variable space to obtain second latent variables for the multiple molecular graphs. The second latent variables are used to describe the ownership information of the molecular graphs. The target decoder is trained based on an original molecular graph dataset and a target digital signature. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The target digital signature is generated based on the ownership information of the target molecular graph dataset. The target molecular graph dataset is obtained by embedding ownership information into the original molecular graph dataset.

[0045] The matching module is used to match the second hidden variables of the plurality of molecular graphs with the target digital signature respectively;

[0046] The determination module is used to determine, when a target molecular graph exists in the plurality of molecular graphs, that the second molecular graph generation model is trained based on the target molecular graph dataset, and that the second latent variable of the target molecular graph matches the target digital signature.

[0047] In some embodiments, the second hidden variable includes a first number of real numbers with positive and negative signs, and the target digital signature includes a first number of binary numbers representing positive and negative signs;

[0048] The matching module is used for:

[0049] For the second latent variable of each molecular graph, the first number of real numbers is compared with the first number of binary numbers. If the proportion of the target real number is greater than a preset proportion, it is determined that the second latent variable of the molecular graph matches the target digital signature. The sign of the target real number matches the sign represented by the binary number compared with the target real number.

[0050] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the decoder training method or model detection method in the embodiments of this application.

[0051] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to implement the decoder training method or model detection method as described in the embodiments of this application.

[0052] On the other hand, a computer program product is provided, the computer program product including computer program code stored in a computer-readable storage medium, a processor of a computer device reading the computer program code from the computer-readable storage medium, the processor executing the computer program code, causing the computer device to execute the decoder training method or model detection method described in any of the above implementations.

[0053] In this embodiment, a predicted molecular graph is obtained through a first molecular graph generation model, and a first latent variable of the predicted molecular graph is obtained based on a decoder. Then, based on the matching result of the first latent variable and the target digital signature, the model parameters of the first molecular graph generation model and the decoder are adjusted. Since the first latent variable can describe the ownership information of the predicted molecular graph, and the target digital signature is generated based on the ownership information of the original molecular graph dataset, the trained target molecular graph generation model can generate a molecular graph carrying the ownership information of the original molecular graph dataset. The trained target decoder can decode the molecular graph carrying the ownership information to obtain the ownership information of the molecular graph. Furthermore, by decoding the molecular graph of any molecular graph generation model through the target decoder, it is possible to detect whether the molecular graph generation model is trained based on the molecular graph dataset carrying the ownership information, thereby achieving effective protection of the molecular graph dataset. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0056] Figure 2 is a flowchart of a decoder training method provided in an embodiment of this application;

[0057] Figure 3 is a flowchart of another decoder training method provided in an embodiment of this application;

[0058] Figure 4 is a flowchart of a model detection method provided in an embodiment of this application;

[0059] Figure 5 is a schematic diagram of an overall framework provided in an embodiment of this application;

[0060] Figure 6 is a block diagram of a decoder training device provided in an embodiment of this application;

[0061] Figure 7 is a block diagram of a model detection device provided in an embodiment of this application;

[0062] Figure 8 is a block diagram of a terminal provided in an embodiment of this application;

[0063] Figure 9 is a block diagram of a server provided in an embodiment of this application. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0065] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0066] In this application, the term "at least one" means one or more, and "multiple" means two or more.

[0067] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the molecular map datasets involved in this application were all obtained with full authorization.

[0068] The following is an explanation of the terms used in this application.

[0069] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0070] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0071] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0072] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learn-by-doing.

[0073] The following describes the implementation environment involved in this application:

[0074] The decoder training method provided in this application can be executed by a computer device. In some embodiments, the computer device is at least one of a terminal and a server. The following is a schematic diagram of the implementation environment of the decoder training method provided in this application. Referring to Figure 1, the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited herein. In some embodiments, the terminal 101 is used to acquire an original molecular graph dataset and send it to the server 102. The server 102 is used to train a target molecular graph generation model and a target decoder based on the original molecular graph dataset.

[0075] In some embodiments, terminal 101 may be a smartphone, tablet, laptop, desktop computer, smart voice interaction device, smart home appliance, vehicle terminal, aircraft, etc., but is not limited thereto. In some embodiments, server 102 may be an independent server, a server cluster or distributed system composed of multiple servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.

[0076] Figure 2 is a flowchart of a decoder training method according to an embodiment of this application. Referring to Figure 2, this embodiment of the application illustrates the method as being executed by a server, and the method includes the following steps:

[0077] 201. The server trains the model based on the original molecular graph dataset to obtain the first molecular graph generation model. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The first molecular graph generation model is used to generate molecular graphs, which are used to represent the molecular structure of any molecule.

[0078] In this embodiment of the application, the server can train the first molecular graph generation model based on the original molecular graph dataset using model algorithms such as GCPN (Graph Convolutional Policy Network), MoFlow (flow-based molecular graph generation model), or GraphAF (autoregressive flow-based molecular graph generation model).

[0079] In this embodiment of the application, the ownership information is used to describe the organization or individual to which the multiple molecular map samples included in the original molecular map dataset belong; in this embodiment of the application, the example of multiple molecular map samples included in the original molecular map dataset belonging to the same organization or individual is used for illustration.

[0080] In this embodiment, the server learns the distribution patterns of the original molecular graph dataset based on the original molecular graph dataset. These distribution patterns include the atomic distribution patterns and chemical bond distribution patterns within the molecular graph, and then generate a first molecular graph generation model that reflects these distribution patterns. The distribution patterns of the original molecular graph dataset can be represented as follows:

[0081] 202. Based on the initial molecular graph, the server iteratively trains the first molecular graph generation model and decoder to obtain the target molecular graph generation model and target decoder.

[0082] In this embodiment, during any iteration, the server inputs the initial molecular graph into a first molecular graph generation model, and obtains the predicted molecular graph for the current iteration through the first molecular graph generation model; the predicted molecular graph is input into a decoder, and the first latent variable of the predicted molecular graph is obtained through the decoder. The first latent variable is used to describe the ownership information of the predicted molecular graph; the first latent variable is matched with a target digital signature, and the model parameters of the first molecular graph generation model and the decoder are adjusted based on the matching result. The target digital signature is generated based on the ownership information of the original molecular graph dataset.

[0083] In this embodiment, the initial molecular graph is the starting graph for generating the predicted molecular graph in each iteration. It can be any molecular graph containing atoms and chemical bonds, such as a molecular graph sample from the original molecular dataset, or a blank graph without atoms and chemical bonds. The server adds new atoms and chemical bonds to the initial molecular graph using the first molecular graph generation model to obtain the predicted molecular graph. In this embodiment, the initial molecular graph in any iteration is described as a blank graph.

[0084] In this embodiment, during the iterative training of the first molecular graph generation model and decoder, if the iteration stopping condition is met in any iteration, the server outputs the first molecular graph generation model and decoder used in that iteration as the target molecular graph generation model and target decoder. The iteration stopping condition can be that the number of iterations reaches a preset number, the loss value reaches a preset value, or the loss value reaches a convergent state. The decoder is used to map the predicted molecular graph to the latent variable space to obtain the first latent variable of the predicted molecular graph, which includes a first number of real numbers.

[0085] In this embodiment, the target molecular graph generation model can map molecular graph samples in the original molecular graph dataset that do not carry ownership information to molecular graphs that carry ownership information. The target decoder can then obtain the first latent variable of the molecular graph, thereby obtaining the ownership information of the molecular graph.

[0086] In this embodiment, a predicted molecular graph is obtained through a first molecular graph generation model, and a first latent variable of the predicted molecular graph is obtained based on a decoder. Then, based on the matching result of the first latent variable and the target digital signature, the model parameters of the first molecular graph generation model and the decoder are adjusted. Since the first latent variable can describe the ownership information of the predicted molecular graph, and the target digital signature is generated based on the ownership information of the original molecular graph dataset, the trained target molecular graph generation model can generate a molecular graph carrying the ownership information of the original molecular graph dataset. The trained target decoder can decode the molecular graph carrying the ownership information to obtain the ownership information of the molecular graph. Furthermore, by decoding the molecular graph of any molecular graph generation model through the target decoder, it is possible to detect whether the molecular graph generation model is trained based on the molecular graph dataset carrying the ownership information, thereby achieving effective protection of the molecular graph dataset.

[0087] Figure 2 above illustrates the basic process of training the decoder. The training process of the decoder will be further described below based on Figure 3. Figure 3 is a flowchart of a decoder training method according to an embodiment of this application. Referring to Figure 3, this embodiment uses server execution as an example for illustration. The method includes the following steps:

[0088] 301. The server trains the model based on the original molecular graph dataset to obtain the first molecular graph generation model. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The first molecular graph generation model is used to generate molecular graphs, which are used to represent the molecular structure of any molecule.

[0089] This step is the same as step 201, and will not be repeated here.

[0090] In this embodiment, the server iteratively executes steps 302-308 to iteratively train the first molecular graph generation model and decoder to obtain the target molecular graph generation model and target decoder. In any iteration, the initial molecular graph is used as the starting graph to generate the predicted molecular graph. The initial molecular graph used in the iteration process can be a blank graph.

[0091] In this embodiment of the application, during any iteration, the server trains based on a second number of initial molecular maps. This second number can be set and changed as needed, such as being 50.

[0092] 302. In any iteration, the server inputs the second number of initial molecular graphs into the first molecular graph generation model. Through the first molecular graph generation model, the server obtains the predicted molecular graphs corresponding to the second number of initial molecular graphs in this iteration.

[0093] In this embodiment, the server processes a second number of initial molecular graphs in two ways. In one implementation, the server sequentially inputs the second number of initial molecular graphs into the first molecular graph generation model to obtain a second number of predicted molecular graphs. This reduces training error by performing one iteration of training using multiple initial molecular graphs. In another implementation, the server can set a second number of first molecular graph generation models with identical model parameters. This allows the second number of predicted molecular graphs to be obtained synchronously through these models. This synchronous processing of multiple initial molecular graphs using multiple first molecular graph generation models improves processing efficiency. It should be noted that since the first molecular graph generation model uses probabilistic random sampling, it may generate different predicted molecular graphs for the same initial molecular graph.

[0094] In this embodiment, after inputting an initial molecular graph into a first molecular graph generation model, the model maps the initial molecular graph to a latent variable space to obtain the latent variables of the initial molecular graph. Then, the latent variables are mapped back to the molecular graph space to obtain a predicted molecular graph. The latent variables of the initial molecular graph describe its molecular structure and ownership information. The latent variable space can be represented as follows: R represents a real number, and k represents the dimension of the latent variables in the latent variable space.

[0095] In this embodiment, the first molecular graph generation model selects atoms and chemical bonds from the molecular graph space to generate a predicted molecular graph. This molecular graph space includes multiple types of molecules and multiple types of chemical bonds, and each molecule has at most n atoms. Therefore, this molecular graph space can be represented as follows: 'b' represents the set of chemical bonds, corresponding to the edges representing chemical bonds in the molecular diagram, and 'b' represents the number of different types of chemical bonds. 'a' represents the set of atoms, 'a' represents the number of categories of molecules, and 'a' represents the nodes representing atoms in the molecular diagram.

[0096] In this embodiment of the application, the first molecular graph generation model maps the latent variable space to the molecular graph space, which can be represented as follows: θ represents the model parameters of the first molecular graph generation model, h θ This represents the first molecular graph generation model.

[0097] 303. The server inputs the second number of predicted molecular graphs into the decoder. The decoder obtains the first latent variable for each of the second number of predicted molecular graphs. The first latent variable is used to describe the ownership information of the predicted molecular graph.

[0098] In this embodiment, the first latent variable includes a first number of real numbers with positive and negative signs. This first number can be set and changed as needed, such as 128 or 256. The first latent variable can be a vector, in which the first number of real numbers are arranged sequentially.

[0099] In this embodiment, the server processes a second number of predicted molecular graphs in two ways. In one implementation, the server sequentially inputs the second number of predicted molecular graphs into a decoder to obtain a second number of first latent variables. This reduces training error by performing one iteration of training using multiple predicted molecular graphs. In another implementation, the server sets up a second number of decoders with identical model parameters, meaning the second number of first latent variables can be obtained synchronously through these decoders. This synchronous processing of multiple predicted molecular graphs using multiple decoders improves processing efficiency.

[0100] In some embodiments, the process by which the server inputs each predicted molecular graph into the decoder includes the following steps: the server adds Gaussian noise to the predicted molecular graph and inputs the Gaussian-noise-added predicted molecular graph into the decoder. Since unauthorized users typically modify the molecular graphs in the dataset before using it, in this embodiment, by adding random perturbation (Gaussian noise) to the predicted molecular graph before inputting it into the decoder, the decoder is trained based on the randomly perturbated molecular graph. This allows the decoder to effectively handle modified molecular graphs, thereby improving the decoder's applicability.

[0101] The above embodiments illustrate the addition of Gaussian noise to the predicted molecular graph. In other embodiments, other types of random perturbations, such as salt-and-pepper noise, can be added to the predicted molecular graph. Furthermore, various random perturbations can be added to the predicted molecular graph, without any specific limitations.

[0102] 304. The server matches the first hidden variable of the second number with the target digital signature to obtain the matching parameters of the second number.

[0103] In this embodiment, the target digital signature includes a first number of binary digits representing positive and negative. These binary digits include 0 and 1, and can be set to 1 for positive and 0 for negative, or vice versa, so that the two types of digits in the binary signature represent positive and negative respectively. In this embodiment, the target digital signature can be a vector, with the first number of binary digits arranged sequentially. The target digital signature can be represented as {0,1,...,1,0}. ss represents the first number. This allows matching the target digital signature with the real numbers with positive or negative signs in the first hidden variable to obtain matching parameters. In this embodiment, the matching parameters refer to the proportion of the target real number among the first number of real numbers, and the match between the positive or negative sign of the target real number and the positive or negative sign represented by the binary number compared with the target real number.

[0104] In this embodiment, the server obtains the target digital signature based on the ownership information using a public-key cryptography algorithm. In this embodiment, the ownership information can be ASCII (American Standard Code for Information Interchange) code of a legal statement text, used to represent the ownership of the molecular graph dataset. The public-key cryptography algorithm can be selected as needed, such as RSA (an encryption algorithm), where the target digital signature is generated by the private key and verified by the public key. The public and private keys form a key pair belonging to the same organization or individual. Successful verification of the target digital signature using the public key indicates that the ownership information was set by the organization or individual to which the private and public keys belong. Thus, even if the public key is made public and the ownership information is exposed, multiple verifications of ownership can still be completed using the public key, and the verification results indicate that the ownership information was set by the organization or individual to which the private and public keys belong, meaning that the ownership of the molecular graph dataset belongs to that organization or individual, thereby improving the effectiveness of the verification.

[0105] In this embodiment, the server performs synchronization processing on the first hidden variable of the second number, that is, it synchronizes the first hidden variable of the second number with the target digital signature to obtain the matching result of the second number; this improves the matching efficiency.

[0106] 305. The server determines the mean of the second number of matching parameters to obtain the target matching parameters.

[0107] In this embodiment, the predicted molecular graph includes intermediate predicted subgraphs at multiple stages, each intermediate predicted subgraph being obtained by adding atoms and chemical bonds to the intermediate predicted subgraph of the previous stage. This generates a second number of intermediate predicted subgraphs at each stage. Accordingly, the target matching parameter for each stage is the mean of the matching parameters of the second number of intermediate predicted subgraphs for that stage. Using the mean as the target matching parameter for that stage makes the target matching parameter more representative and accurate.

[0108] 306. The server determines the target molecular mass parameter for a second number of predicted molecular maps, which is used to describe the quality of the predicted molecular maps generated in this iteration.

[0109] In this embodiment of the application, the target molecular mass parameter is a parameter used to describe the overall quality of the second number of predicted molecular maps; the quality of the molecular map includes the chemical correctness, novelty, and uniqueness of the molecular map. Accordingly, for each stage, the process by which the server determines the target molecular mass parameter of the second number of predicted molecular maps includes the following steps (1)-(4):

[0110] (1) The server takes the mean of the correctness parameters of the second number of predicted intermediate subgraphs in this stage as the target correctness parameter for this stage, which is used to describe the chemical correctness of the molecular graph.

[0111] In this embodiment, the server can determine the correctness parameters of the predicted intermediate subgraph based on a third-party sub-library. For example, the third-party sub-library could be the RDKit (Chemical Molecular Computational Retrieval) library.

[0112] (2) The server uses the ratio of the number of the first intermediate subgraphs to the number of the multiple molecular graph samples as the target novelty parameter for this stage. The first intermediate subgraph is the intermediate subgraph that is different from the multiple molecular graph samples in the predicted intermediate subgraphs of the second number.

[0113] In this embodiment of the application, the target novelty parameter is used to describe the proportion of newly emerging intermediate subgraphs in the second number of predicted intermediate subgraphs, that is, it can effectively represent the novelty of the predicted intermediate subgraphs generated in this stage.

[0114] (3) The server uses the ratio of the number of the second intermediate subgraphs to the second number as the target uniqueness parameter for this stage. The second intermediate subgraph is an intermediate subgraph that is not repeated in the first intermediate subgraph.

[0115] In this embodiment, the server can remove duplicate intermediate subgraphs from the first intermediate subgraph to obtain the second intermediate subgraph. The target uniqueness parameter describes the proportion of non-repeating intermediate subgraphs in the first intermediate subgraph, effectively representing the uniqueness of the predicted intermediate subgraph generated at this stage.

[0116] (4) The server performs a weighted sum of the target's correctness parameter, novelty parameter, and uniqueness parameter to obtain the target's molecular mass parameter for this stage.

[0117] In the embodiments of this application, the weights of the target correctness parameter, the target novelty parameter, and the target uniqueness parameter can be set and changed as needed, and no specific limitations are made here.

[0118] In the embodiments of this application, since the target correctness parameter, the target novelty parameter, and the target uniqueness parameter can respectively represent the chemical correctness, novelty, and uniqueness of the generated prediction intermediate subgraph, the target molecular mass parameter is comprehensive and representative.

[0119] It should be noted that the target correctness parameter, target novelty parameter, and target uniqueness parameter are the basic parameters for determining the target molecular mass parameter. In some embodiments, the target molecular mass parameter also needs to be determined based on the business objective. For example, if the business objective is to generate a molecular map of proteins, then the target molecular mass parameter also needs to be determined based on the parameters used to describe the protein quality of the predicted molecular map, thereby improving the specificity and accuracy of the determined target molecular mass parameter.

[0120] In this embodiment, the numbering of steps (1)-(3) is only for ease of description. The server can execute steps (1)-(3) sequentially or execute steps (1)-(3) synchronously. In this embodiment, the synchronous execution of steps (1)-(3) is used as an example to improve the efficiency of determining the target molecular mass parameter.

[0121] In this embodiment, during any iteration, if the intermediate prediction subgraph generated at any stage violates the target constraints, the intermediate prediction subgraph generated at that stage is invalid, meaning the current iteration is invalid and the process terminates. These target constraints include chemical property restrictions, such as the generation of isolated nodes in the intermediate prediction subgraph at that stage, or the generation of He in the intermediate prediction subgraph, which only considers elements C, H, and O. This indicates that the intermediate prediction subgraph violates the target constraints. This improves the effectiveness of model training.

[0122] In this embodiment, during any iteration, if training is performed based on an initial molecular map, then at any stage, the target matching parameter is the matching parameter of the predicted intermediate subgraph for that stage. The target correctness parameter is the correctness parameter of the predicted intermediate subgraph for that stage. If the predicted intermediate subgraph is different from multiple molecular map samples, the target novelty parameter for that stage can be set to 1; if the predicted intermediate subgraph is the same as multiple molecular map samples, the target novelty parameter for that stage can be set to 0. The server performs a weighted sum of the target correctness parameter and the target novelty parameter to obtain the target molecular mass parameter for that stage. This approach, performing one iteration of training based on an initial molecular map, improves training efficiency.

[0123] 307. The server performs a weighted summation of the target matching parameters and the target molecular mass parameters to obtain a first reward value, which is used to describe the reward obtained in generating the predicted molecular map during this iteration.

[0124] In this embodiment of the application, since the predicted molecular graph includes intermediate subgraphs in multiple stages, the process by which the server weights and sums the target matching parameters and the target molecular mass parameters to obtain the first reward value includes the following steps: for each stage, the server weights and sums the target matching parameters and the target molecular mass parameters of that stage to obtain the second reward value for that stage; the server weights and sums the second reward values ​​of multiple stages to obtain the first reward value.

[0125] In the embodiments of this application, the weights of the target matching parameter and the target molecular mass parameter can be set and changed as needed, and no specific limitation is made here.

[0126] In this embodiment, the reward value is determined by fusing the target matching parameters and the target molecular mass parameters. This ensures that the subsequently generated target decoder can not only effectively detect ownership information, but also guarantees the quality of the molecular graph generated by the target molecular graph generation model. In other words, it protects the molecular graph dataset while avoiding adverse effects on the molecular graph generated by the molecular graph generation model.

[0127] In this embodiment, the weights of the second reward values ​​at each stage can be set and changed as needed. For example, the weights of multiple stages can be equal or the weights of multiple stages can increase sequentially. In this embodiment, the sequential increase of the weights of multiple stages is used as an example. Since the predicted intermediate subgraphs of later stages are closer to the final predicted subgraphs, their second reward values ​​are more accurate, thus the accuracy of the weighted first reward value is higher.

[0128] In this embodiment, a first reward value is obtained by fusing second reward values ​​from multiple stages, so that the first reward value can cover the entire process of generating the predicted molecular map. Subsequently, the model parameters are adjusted based on the first reward value, which can improve the effectiveness and comprehensiveness of the adjustment.

[0129] 308. Based on the first reward value, the server adjusts the model parameters of the first molecular graph generation model and the decoder.

[0130] In this embodiment, model training is based on a reinforcement learning algorithm. This reinforcement learning algorithm can be set and modified as needed. In this embodiment, a Markov decision process is used. Taking the reinforcement learning algorithm as an example, steps 302-308 above are used to train the first molecular graph generation model and decoder. The state space represents the state space, where at each stage t, the predicted intermediate subgraph is used as the state, and the initial stage state is an empty graph. represents the action space, which involves adding new atoms and their corresponding edges to the predicted intermediate subgraph from the previous stage using the first molecular graph generation model. P represents the state transition function, used to determine whether the predicted intermediate subgraph violates the target constraints when sufficient observations are made of it. r represents the reward value. γ represents the discount factor, γ∈[0,1], used to adjust the reward value.

[0131] In this embodiment, if the reward value obtained by the server based on the execution of the target action is greater than the reward value before the execution of the target action, the server adjusts the model parameters to increase the probability of executing the target action, which is the action corresponding to generating the prediction molecular graph; if the reward value obtained based on the execution of the target action is less than the reward value before the execution of the target action, the server adjusts the model parameters to decrease the probability of executing the target action, thereby adjusting the model parameters towards a higher reward value.

[0132] In some embodiments, the weights of the second reward values ​​in multiple stages are equal, and the second reward values ​​in multiple stages are summed to obtain the first reward value; then for a given initial state-action pair The first reward value can be obtained by the following formula (1).

[0133]

[0134] Where, π θ The term represents the strategy, or action, performed to generate the predicted molecular map based on (s,a); t represents the stage, and r represents the phase. t γ represents the second reward value in stage t, and γ represents the discount factor. This represents the first reward value.

[0135] In this embodiment, the objective of the first reward value is to maximize the reward value, that is, the objective of reinforcement learning is to maximize the expected return. In some embodiments, this maximization of expected return is achieved through a proximal policy optimization algorithm, which is shown in the following formula (2).

[0136]

[0137] Where θ represents the model parameters, π θ This represents the strategy after the model parameters are updated. This represents the strategy before the model parameters were updated. This represents the ratio of two strategies, with the strategy clipped to the interval (1-∈, 1+∈) to prevent the difference between the two strategies from becoming too large. In other words, it controls the gradient adjustment of the model parameters to avoid excessively large gradients. Represents the maximum expected value of the model parameters; ∈ represents the hyperparameter, which can be set and changed as needed, such as ∈ being 0.2; This represents the advantage function estimated at stage t, used to describe the relevant reward value.

[0138] In this embodiment of the application, a molecular graph dataset carrying ownership information is generated based on a trained target molecular graph generation model. The process includes the following steps: the server inputs multiple molecular graph samples from the original molecular graph dataset into the target molecular graph generation model; the server obtains the target molecular graph dataset through the target molecular graph generation model, and the target molecular graph dataset includes multiple molecular graph samples carrying ownership information.

[0139] In this embodiment, for any molecular graph sample in the original molecular graph dataset, the server inputs the molecular graph sample into the target molecular graph generation model. The target molecular graph generation model then maps the molecular graph sample to a latent variable space to obtain the latent variables of the molecular graph sample. These latent variables describe the molecular structure and ownership information of the molecular graph sample. Then, the latent variables of the molecular graph sample are mapped back to the molecular graph space to obtain a molecular graph sample carrying ownership information. In some embodiments, the server stores all molecular graph samples carrying ownership information in a database. This database stores molecular graph samples carrying ownership information, and the target molecular graph dataset is then obtained based on the molecular graph samples in the database.

[0140] In this embodiment of the application, a molecular graph carrying ownership information is obtained through a target molecular graph generation model, which can be represented as follows: This represents the target molecular graph generation model. This represents a molecular diagram that does not carry ownership information. A molecular diagram representing ownership information.

[0141] In this embodiment of the application, ownership information is embedded into the original molecular graph dataset through the target molecular graph generation model, making the molecular graph generation model trained on the target molecular graph dataset traceable, thereby effectively protecting the ownership of the molecular graph dataset.

[0142] It's important to note that designing new molecular structures, a fundamental problem in drug discovery, is a complex process. On one hand, the designed molecular structure must possess certain chemical properties to be applicable to different drugs; on the other hand, due to the discrete and vast nature of chemical space, manually designing new molecular structures is often time-consuming and labor-intensive. Deep learning-based molecular graph generation models offer an opportunity to discover new molecular structures. By training a molecular graph generation model on a molecular graph dataset, new molecular structures can be derived from this model, significantly reducing the time and manpower costs of drug discovery. Meanwhile, unlike traditional image or text data, molecular graphs require greater protection, primarily for the following reasons: First, acquiring molecular graph datasets is more costly than acquiring image or text data. For example, image or text data can be obtained directly from the internet, while molecular graph datasets often require theoretical knowledge or chemical experiments, resulting in a relative scarcity of truly effective molecular graph datasets. Second, the providers of molecular graph datasets (such as pharmaceutical companies) and the users (such as artificial intelligence companies) are often not the same person or organization, necessitating measures to prevent the misuse of molecular graph datasets. For example, providers want to ensure that molecular graph datasets are used by authorized users and prevent unauthorized users from accessing them; and providers also need to prevent departing employees from taking molecular graph datasets with them and leaking trade secrets.

[0143] In this embodiment, a predicted molecular graph is obtained through a first molecular graph generation model, and a first latent variable of the predicted molecular graph is obtained based on a decoder. Then, based on the matching result of the first latent variable and the target digital signature, the model parameters of the first molecular graph generation model and the decoder are adjusted. Since the first latent variable can describe the ownership information of the predicted molecular graph, and the target digital signature is generated based on the ownership information of the original molecular graph dataset, the trained target molecular graph generation model can generate a molecular graph carrying the ownership information of the original molecular graph dataset. The trained target decoder can decode the molecular graph carrying the ownership information to obtain the ownership information of the molecular graph. Furthermore, by decoding the molecular graph of any molecular graph generation model through the target decoder, it is possible to detect whether the molecular graph generation model is trained based on the molecular graph dataset carrying the ownership information, thereby achieving effective protection of the molecular graph dataset.

[0144] Figures 2-3 above illustrate the decoder training process, and Figure 4 is a flowchart of a model detection method provided according to an embodiment of this application, implementing the target decoder trained based on any of the above embodiments. Referring to Figure 4, this embodiment will illustrate the method using server execution as an example. The method includes the following steps:

[0145] 401. The server obtains multiple molecular graphs, which are generated based on the second molecular graph generation model. The molecular graphs are used to represent the molecular structure of any molecule.

[0146] In this embodiment of the application, the second molecular graph generation model is a molecular graph generation model that is not authorized to use the target molecular graph dataset, that is, whether the molecular graph generation model to be detected is trained using the target molecular graph dataset.

[0147] 402. The server inputs the multiple molecular graphs into the target decoder, which maps the multiple molecular graphs to the latent variable space to obtain the second latent variables of the multiple molecular graphs. The second latent variables are used to describe the ownership information of the molecular graphs. The target decoder is trained based on the original molecular graph dataset and the target digital signature. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The target digital signature is generated based on the ownership information of the target molecular graph dataset. The target molecular graph dataset is obtained by embedding the ownership information into the original molecular graph dataset.

[0148] In this embodiment, the second latent variable is analogous to the first latent variable and will not be described again. The process of generating the target digital signature based on the ownership information of the target molecular graph dataset is similar to step 304 and will not be described again. The target molecular graph dataset includes multiple molecular graphs carrying ownership information.

[0149] 403. The server matches the second hidden variables of the multiple molecular graphs with the target digital signature respectively.

[0150] In this embodiment, the second latent variable includes a first number of real numbers with positive and negative signs, and the target digital signature includes a first number of binary numbers representing positive and negative signs. Accordingly, the server matches the second latent variables of multiple molecular graphs with the target digital signature, including the following steps: For each molecular graph's second latent variable, the server compares the first number of real numbers with the first number of binary numbers. If the proportion of the target real number is greater than a preset proportion, it determines that the second latent variable of the molecular graph matches the target digital signature, and the positive and negative signs of the target real number and the binary numbers compared with the target real number match.

[0151] In this embodiment, the preset ratio can be set and changed as needed, such as a preset ratio of 90%. The proportion of the target real number can be the difference between 1 and the bit error rate, where the bit error rate refers to the proportion of real numbers in the first number of real numbers whose positive or negative sign does not match the positive or negative sign represented by the binary digits they are compared with.

[0152] 404. If the target molecular graph exists in the multiple molecular graphs, the server determines that the second molecular graph generation model is trained based on the target molecular graph dataset, and the second latent variable of the target molecular graph matches the target digital signature.

[0153] In this embodiment, if the second molecular graph generation model is trained based on the target molecular graph dataset, random perturbations may be added to the molecular graphs in the target molecular graph dataset before training to modify the target molecular graph dataset and remove the ownership information of the molecular graph dataset. Then, training is performed based on the modified molecular graph dataset, so that not all molecular graphs generated by the second molecular graph generation model can detect ownership information. However, in this embodiment, if the target molecular graph exists in multiple molecular graphs, it is determined that the second molecular graph generation model is trained based on the target molecular graph dataset, thereby ensuring the effectiveness of detection.

[0154] In other embodiments, if the number of target molecular maps in multiple molecular maps exceeds a preset number, the server determines that the second molecular map generation model is trained based on the target molecular map dataset, further improving the effectiveness and accuracy of model detection.

[0155] Referring to Figure 5, which is a schematic diagram of an overall framework provided in an embodiment of this application, the overall framework includes a codec training framework 501, a target molecular graph dataset generation framework 502, and a model detection framework 503. Based on the codec training framework, a target molecular graph generation model and a target decoder can be trained. Based on the target molecular graph dataset generation framework, a target molecular graph dataset carrying ownership information can be generated. Based on the model detection framework, it is possible to detect whether the molecular graph generation model is trained on the target molecular graph dataset carrying ownership information.

[0156] In this embodiment of the application, the second latent variable of the molecular graph is determined by the target decoder. Since the second latent variable is used to describe the ownership information of the molecular graph, the target digital signature is generated based on the ownership information of the target molecular graph dataset. Therefore, if the second latent variable matches the target digital signature, it can be said that the second molecular graph generation model is trained based on the target molecular graph dataset, thereby improving the accuracy of model detection.

[0157] Figure 6 is a block diagram of a decoder training apparatus according to an embodiment of this application. The apparatus is used to execute the steps of the above-described decoder training method. Referring to Figure 6, the apparatus includes:

[0158] Training module 601 is used to train the model based on the original molecular graph dataset to obtain the first molecular graph generation model. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The first molecular graph generation model is used to generate molecular graphs, which are used to represent the molecular structure of any molecule.

[0159] The iterative execution module 602 is used to iteratively execute the following steps based on the initial molecular graph to train the first molecular graph generation model and decoder, thereby obtaining the target molecular graph generation model and target decoder: In any iteration, the initial molecular graph is input into the first molecular graph generation model, and the predicted molecular graph for this iteration is obtained through the first molecular graph generation model; the predicted molecular graph is input into the decoder, and the first latent variable of the predicted molecular graph is obtained through the decoder. The first latent variable is used to describe the ownership information of the predicted molecular graph; the first latent variable is matched with the target digital signature, and the model parameters of the first molecular graph generation model and decoder are adjusted based on the matching result. The target digital signature is generated based on the ownership information of the original molecular graph dataset.

[0160] In some embodiments, the matching result is a matching parameter, the first hidden variable includes a first number of real numbers with positive and negative signs, the target digital signature includes a first number of binary numbers representing positive and negative signs, the matching parameter refers to the proportion of the target real number among the first number of real numbers, and the positive and negative signs of the target real number and the positive and negative signs represented by the binary numbers compared with the target real number.

[0161] During any iteration, training is performed based on the second number of initial molecular maps. The iteration execution module 602 is used for:

[0162] The mean value of the matching parameters of the second number of predicted molecular maps is determined to obtain the target matching parameters;

[0163] Determine the target molecular mass parameters for the second number of predicted molecular maps. The target molecular mass parameters are used to describe the quality of the predicted molecular maps generated in this iteration.

[0164] The target matching parameters and target molecular mass parameters are weighted and summed to obtain the first reward value, which describes the reward obtained in generating the predicted molecular map during this iteration.

[0165] Based on the first reward value, adjust the model parameters of the first molecular graph generation model and the decoder.

[0166] In some embodiments, the predicted molecular graph includes predicted intermediate subgraphs at multiple stages, each stage of which is obtained by adding atoms and chemical bonds to the predicted intermediate subgraph of the previous stage.

[0167] Iterative execution module 602 is used for:

[0168] For each stage, the target matching parameter and the target molecular mass parameter of the stage are weighted and summed to obtain the second reward value of the stage.

[0169] The first reward value is obtained by weighted summation of the second reward values ​​from multiple stages.

[0170] In some embodiments, the target matching parameter for each stage is the mean of the matching parameters of a second number of predicted intermediate subgraphs for that stage.

[0171] In some embodiments, the iteration execution module 602 is configured to:

[0172] For each stage, the mean of the correctness parameters of the second number of predicted intermediate subgraphs of the stage is used as the target correctness parameter of the stage. The correctness parameter is used to describe the chemical correctness of the molecular graph.

[0173] The ratio of the number of first intermediate subgraphs to the number of multiple molecular graph samples is used as the target novelty parameter for the stage. The first intermediate subgraph is the intermediate subgraph that is different from the multiple molecular graph samples in the second number of predicted intermediate subgraphs.

[0174] The ratio of the number of the second intermediate subgraphs to the number of the second intermediate subgraphs is used as the target uniqueness parameter of the stage. The second intermediate subgraph is an intermediate subgraph that is not repeated in the first intermediate subgraph.

[0175] The target molecular mass parameter for the stage is obtained by weighted summation of the target correctness parameter, target novelty parameter, and target uniqueness parameter.

[0176] In some embodiments, the iteration execution module 602 is configured to:

[0177] Gaussian noise is added to the predicted molecular graph, and the predicted molecular graph with added Gaussian noise is then input into the decoder.

[0178] In some embodiments, the apparatus further includes:

[0179] The input module is used to input multiple molecular graph samples from the original molecular graph dataset into the target molecular graph generation model.

[0180] The acquisition module is used to generate a target molecular graph dataset from the target molecular graph generation model. The target molecular graph dataset includes multiple molecular graph samples carrying ownership information.

[0181] In this embodiment, a predicted molecular graph is obtained through a first molecular graph generation model, and a first latent variable of the predicted molecular graph is obtained based on a decoder. Then, based on the matching result of the first latent variable and the target digital signature, the model parameters of the first molecular graph generation model and the decoder are adjusted. Since the first latent variable can describe the ownership information of the predicted molecular graph, and the target digital signature is generated based on the ownership information of the original molecular graph dataset, the trained target molecular graph generation model can generate a molecular graph carrying the ownership information of the original molecular graph dataset. The trained target decoder can decode the molecular graph carrying the ownership information to obtain the ownership information of the molecular graph. Furthermore, by decoding the molecular graph of any molecular graph generation model through the target decoder, it is possible to detect whether the molecular graph generation model is trained based on the molecular graph dataset carrying the ownership information, thereby achieving effective protection of the molecular graph dataset.

[0182] Figure 7 is a block diagram of a model detection apparatus according to an embodiment of this application. This apparatus is used to perform the steps of the above-described model detection method. Referring to Figure 7, the apparatus includes:

[0183] The acquisition module 701 is used to acquire multiple molecular diagrams, which are generated based on the second molecular diagram generation model. The molecular diagrams are used to represent the molecular structure of any molecule.

[0184] The input module 702 is used to input multiple molecular graphs into the target decoder. The target decoder maps the multiple molecular graphs to the latent variable space to obtain the second latent variables of the multiple molecular graphs. The second latent variables are used to describe the ownership information of the molecular graphs. The target decoder is trained based on the original molecular graph dataset and the target digital signature. The original molecular graph dataset includes multiple molecular graph samples without ownership information. The target digital signature is generated based on the ownership information of the target molecular graph dataset. The target molecular graph dataset is obtained by embedding the ownership information into the original molecular graph dataset.

[0185] Matching module 703 is used to match the second hidden variables of multiple molecular graphs with the target digital signature respectively;

[0186] The determination module 704 is used to determine, in the case where a target molecular graph exists in multiple molecular graphs, the second molecular graph generation model is trained based on the target molecular graph dataset, and the second latent variable of the target molecular graph matches the target digital signature.

[0187] In some embodiments, the second hidden variable includes a first number of real numbers with positive and negative signs, and the target digital signature includes a first number of binary numbers representing positive and negative signs.

[0188] Matching module 703 is used for:

[0189] For the second latent variable of each molecular graph, the first number of real numbers is compared with the first number of binary numbers. If the proportion of the target real number is greater than the preset proportion, it is determined that the second latent variable of the molecular graph matches the target digital signature, and the sign of the target real number matches the sign represented by the binary number compared with the target real number.

[0190] In this embodiment of the application, the second latent variable of the molecular graph is determined by the target decoder. Since the second latent variable is used to describe the ownership information of the molecular graph, the target digital signature is generated based on the ownership information of the target molecular graph dataset. Therefore, if the second latent variable matches the target digital signature, it can be said that the second molecular graph generation model is trained based on the target molecular graph dataset, thereby improving the accuracy of model detection.

[0191] In the embodiments of this application, the computer device can be a terminal or a server. When the computer device is a terminal, the terminal acts as the execution subject to implement the technical solution provided in the embodiments of this application; when the computer device is a server, the server acts as the execution subject to implement the technical solution provided in the embodiments of this application; or, the technical solution provided in this application can be implemented through the interaction between the terminal and the server. The embodiments of this application do not limit this.

[0192] Figure 8 shows a structural block diagram of a terminal 800 provided in an exemplary embodiment of this application. The terminal 800 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 800 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0193] Typically, terminal 800 includes a processor 801 and a memory 802.

[0194] Processor 801 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0195] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one program code, which is executed by the processor 801 to implement the decoder training method or model detection method provided in the method embodiments of this application.

[0196] In some embodiments, the terminal 800 may also optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.

[0197] Peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 801 and memory 802. In some embodiments, processor 801, memory 802 and peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 801, memory 802 and peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0198] The radio frequency (RF) circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 804 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 804 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0199] Display screen 805 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 801 for processing. In this case, display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 805, disposed on the front panel of terminal 800; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 800 or in a folded design; in other embodiments, display screen 805 may be a flexible display screen, disposed on a curved or folded surface of terminal 800. Furthermore, display screen 805 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 805 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0200] The camera assembly 806 is used to acquire images or videos. Optionally, the camera assembly 806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0201] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 801 for processing, or input to the radio frequency circuit 804 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 800. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 807 may also include a headphone jack.

[0202] Power supply 808 is used to supply power to the various components in terminal 800. Power supply 808 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 808 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0203] In some embodiments, the terminal 800 further includes one or more sensors 809. The one or more sensors 809 include, but are not limited to, an accelerometer 810, a gyroscope 811, a pressure sensor 812, an optical sensor 813, and a proximity sensor 814.

[0204] Accelerometer 810 can detect the magnitude of acceleration on the three coordinate axes of a coordinate system established by terminal 800. For example, accelerometer 810 can be used to detect the components of gravitational acceleration on the three coordinate axes. Processor 801 can control display screen 805 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 810. Accelerometer 810 can also be used for games or for acquiring user motion data.

[0205] The gyroscope sensor 811 can detect the orientation and rotation angle of the terminal 800. The gyroscope sensor 811, in conjunction with the accelerometer sensor 810, can collect 3D motion data from the user on the terminal 800. Based on the data collected by the gyroscope sensor 811, the processor 801 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0206] The pressure sensor 812 can be disposed on the side bezel of the terminal 800 and / or the lower layer of the display screen 805. When the pressure sensor 812 is disposed on the side bezel of the terminal 800, it can detect the user's grip signal on the terminal 800, and the processor 801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 812. When the pressure sensor 812 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0207] An optical sensor 813 is used to collect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity collected by the optical sensor 813. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity collected by the optical sensor 813.

[0208] The proximity sensor 814, also known as a distance sensor, is typically located on the front panel of the terminal 800. The proximity sensor 814 is used to detect the distance between the user and the front of the terminal 800. In one embodiment, when the proximity sensor 814 detects that the distance between the user and the front of the terminal 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from a screen-on state to a screen-off state; when the proximity sensor 814 detects that the distance between the user and the front of the terminal 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from a screen-off state to a screen-on state.

[0209] Those skilled in the art will understand that the structure shown in FIG8 does not constitute a limitation on the terminal 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0210] Figure 9 is a schematic diagram of a server structure according to an embodiment of this application. The server 900 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 901 and one or more memories 902. The memories 902 are used to store executable program code, and the processors 901 are configured to execute the executable program code to implement the decoder training method or model detection method provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0211] This application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the training method or model detection method of the decoder in any of the above implementations.

[0212] This application also provides a computer program product, which includes computer program code. The computer program code is stored in a computer-readable storage medium. The processor of the computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to execute the training method or model detection method of the decoder in any of the above implementations.

[0213] In some embodiments, the computer program product involved in the present application can be deployed and executed on a computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.

[0214] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for training a decoder, characterized in that, The method includes: training a model based on an original molecular graph dataset to obtain a first molecular graph generation model, wherein the original molecular graph dataset includes multiple molecular graph samples without ownership information, the first molecular graph generation model is used to generate molecular graphs, and the molecular graphs are used to represent the molecular structure of any molecule; based on the initial molecular graph, iteratively executing the following steps to train the first molecular graph generation model and the decoder to obtain a target molecular graph generation model and a target decoder: in any iteration, the initial molecular graph is input into the first molecular graph generation model, and the predicted molecular graph for the current iteration is obtained through the first molecular graph generation model; the predicted molecular graph is input into the decoder, and the first latent variable of the predicted molecular graph is obtained through the decoder, the first latent variable being used to describe the ownership information of the predicted molecular graph; the first latent variable is matched with a target digital signature, and based on the matching result, the model parameters of the first molecular graph generation model and the decoder are adjusted, wherein the target digital signature is generated based on the ownership information of the original molecular graph dataset.

2. The method according to claim 1, characterized in that, The matching result is a matching parameter. The first latent variable includes a first number of real numbers with positive and negative signs. The target digital signature includes a first number of binary numbers representing positive and negative signs. The matching parameter refers to the proportion of the target real number among the first number of real numbers, and the positive and negative signs of the target real number and the positive and negative signs represented by the binary numbers compared with the target real number. In any iteration, training is performed based on a second number of initial molecular graphs. The adjustment of the model parameters of the first molecular graph generation model and the decoder based on the matching result includes: determining the mean of the matching parameters of the second number of predicted molecular graphs to obtain the target matching parameter; determining the target molecular mass parameter of the second number of predicted molecular graphs, the target molecular mass parameter being used to describe the quality of the predicted molecular graphs generated in this iteration; weighted summing of the target matching parameter and the target molecular mass parameter to obtain a first reward value, the first reward value being used to describe the reward obtained for generating the predicted molecular graph in this iteration; and adjusting the model parameters of the first molecular graph generation model and the decoder based on the first reward value.

3. The method according to claim 2, characterized in that, The predicted molecular graph includes intermediate predicted subgraphs at multiple stages, each intermediate predicted subgraph being obtained by adding atoms and chemical bonds to the intermediate predicted subgraph of the previous stage; the weighted summation of the target matching parameter and the target molecular mass parameter to obtain the first reward value includes: for each stage, weighted summation of the target matching parameter and the target molecular mass parameter of the stage to obtain the second reward value of the stage; and weighted summation of the second reward values ​​of the multiple stages to obtain the first reward value.

4. The method according to claim 3, characterized in that, The target matching parameter for each stage is the mean of the matching parameters of the second number of predicted intermediate subgraphs for that stage.

5. The method according to claim 3, characterized in that, The determination of the target molecular mass parameter for the second number of predicted molecular maps includes: for each stage, taking the mean of the correctness parameters of the second number of predicted intermediate submaps for that stage as the target correctness parameter for that stage, the correctness parameter being used to describe the chemical correctness of the molecular map; taking the ratio of the number of first intermediate submaps to the number of the plurality of molecular map samples as the target novelty parameter for that stage, the first intermediate submap being an intermediate submap in the second number of predicted intermediate submaps that is different from the plurality of molecular map samples; taking the ratio of the number of second intermediate submaps to the second number as the target uniqueness parameter for that stage, the second intermediate submap being an intermediate submap that is not repeated in the first intermediate submap; and weighted summing the target correctness parameter, the target novelty parameter, and the target uniqueness parameter to obtain the target molecular mass parameter for that stage.

6. The method according to claim 1, characterized in that, The step of inputting the predicted molecular map into the decoder includes: adding Gaussian noise to the predicted molecular map and inputting the predicted molecular map with added Gaussian noise into the decoder.

7. The method according to claim 1, characterized in that, The method further includes: inputting multiple molecular graph samples from the original molecular graph dataset into the target molecular graph generation model; and obtaining a target molecular graph dataset through the target molecular graph generation model, wherein the target molecular graph dataset includes multiple molecular graph samples carrying ownership information.

8. A model detection method, characterized in that, The method includes: acquiring multiple molecular graphs, which are generated based on a second molecular graph generation model, and the molecular graphs are used to represent the molecular structure of any molecule; inputting the multiple molecular graphs into a target decoder, and mapping the multiple molecular graphs to a latent variable space through the target decoder to obtain second latent variables of the multiple molecular graphs, the second latent variables being used to describe the ownership information of the molecular graphs; the target decoder being trained based on an original molecular graph dataset and a target digital signature; the original molecular graph dataset including multiple molecular graph samples without ownership information; the target digital signature being generated based on the ownership information of the target molecular graph dataset; the target molecular graph dataset being obtained by embedding ownership information into the original molecular graph dataset; and the target decoder being obtained through the training method of any one of claims 1-7; matching the second latent variables of the multiple molecular graphs with the target digital signature; and, if a target molecular graph exists among the multiple molecular graphs, determining that the second molecular graph generation model is trained based on the target molecular graph dataset, and that the second latent variable of the target molecular graph matches the target digital signature.

9. The method according to claim 8, characterized in that, The second hidden variable includes a first number of real numbers with positive and negative signs, and the target digital signature includes a first number of binary numbers representing positive and negative signs; The step of matching the second hidden variables of the plurality of molecular graphs with the target digital signature includes: for the second hidden variable of each molecular graph, comparing the first number of real numbers with the first number of binary numbers; if the proportion of the target real number is greater than a preset proportion, determining that the second hidden variable of the molecular graph matches the target digital signature, wherein the sign of the target real number matches the sign represented by the binary number compared with the target real number.

10. A training device for a decoder, characterized in that, The apparatus includes: a training module for training a model based on an original molecular graph dataset to obtain a first molecular graph generation model, wherein the original molecular graph dataset includes multiple molecular graph samples without ownership information, the first molecular graph generation model is used to generate molecular graphs, and the molecular graphs are used to represent the molecular structure of any molecule; and an iterative execution module for iteratively executing the following steps based on an initial molecular graph to train the first molecular graph generation model and a decoder to obtain a target molecular graph generation model and a target decoder: in any iteration, the initial molecular graph is input into the first molecular graph generation model, and a predicted molecular graph for the current iteration is obtained through the first molecular graph generation model; the predicted molecular graph is input into the decoder, and a first latent variable of the predicted molecular graph is obtained through the decoder, the first latent variable being used to describe the ownership information of the predicted molecular graph; the first latent variable is matched with a target digital signature, and based on the matching result, the model parameters of the first molecular graph generation model and the decoder are adjusted, wherein the target digital signature is generated based on the ownership information of the original molecular graph dataset.

11. The apparatus according to claim 10, characterized in that, The matching result is the matching parameter. The first latent variable includes a first number of real numbers with positive and negative signs. The target digital signature includes a first number of binary numbers representing positive and negative signs. The matching parameter refers to the proportion of the target real number among the first number of real numbers, and the positive and negative signs of the target real number and the positive and negative signs represented by the binary numbers compared with the target real number. In any iteration, training is performed based on a second number of initial molecular graphs. The iteration execution module is used to: determine the mean of the matching parameters of the second number of predicted molecular graphs to obtain the target matching parameter; determine the target molecular mass parameter of the second number of predicted molecular graphs, the target molecular mass parameter being used to describe the quality of the predicted molecular graph generated in this iteration; and perform a weighted summation of the target matching parameter and the target molecular mass parameter to obtain a first reward value, the first reward value being used to describe the reward obtained for generating the predicted molecular graph in this iteration. Based on the first reward value, adjust the model parameters of the first molecular graph generation model and the decoder.

12. The apparatus according to claim 11, characterized in that, The predicted molecular graph includes intermediate predicted subgraphs at multiple stages, each intermediate predicted subgraph being obtained by adding atoms and chemical bonds to the intermediate predicted subgraph of the previous stage; the iterative execution module is used to: for each stage, perform a weighted summation of the target matching parameter and the target molecular mass parameter of the stage to obtain a second reward value for the stage; The first reward value is obtained by weighted summing of the second reward values ​​of the multiple stages.

13. The apparatus according to claim 12, characterized in that, The target matching parameter for each stage is the mean of the matching parameters of the second number of predicted intermediate subgraphs for that stage.

14. The apparatus according to claim 12, characterized in that, The iterative execution module is configured to: for each stage, take the mean of the correctness parameters of the second number of predicted intermediate subgraphs of the stage as the target correctness parameter of the stage, the correctness parameter being used to describe the chemical correctness of the molecular graph; take the ratio of the number of first intermediate subgraphs to the number of the plurality of molecular graph samples as the target novelty parameter of the stage, the first intermediate subgraph being an intermediate subgraph different from the plurality of molecular graph samples in the second number of predicted intermediate subgraphs; take the ratio of the number of second intermediate subgraphs to the second number as the target uniqueness parameter of the stage, the second intermediate subgraph being an intermediate subgraph that is not repeated in the first intermediate subgraph; and perform a weighted summation of the target correctness parameter, the target novelty parameter, and the target uniqueness parameter to obtain the target molecular mass parameter of the stage.

15. The apparatus according to claim 10, characterized in that, The iterative execution module is used to: add Gaussian noise to the predicted molecular graph, and input the predicted molecular graph with added Gaussian noise into the decoder.

16. The apparatus according to claim 10, characterized in that, The device further includes: an input module for inputting multiple molecular graph samples from the original molecular graph dataset into the target molecular graph generation model; and an acquisition module for obtaining a target molecular graph dataset through the target molecular graph generation model, wherein the target molecular graph dataset includes multiple molecular graph samples carrying ownership information.

17. A model testing device, characterized in that, The apparatus includes: an acquisition module for acquiring multiple molecular graphs, the multiple molecular graphs being generated based on a second molecular graph generation model, the molecular graphs representing the molecular structure of any molecule; an input module for inputting the multiple molecular graphs into a target decoder, the target decoder mapping the multiple molecular graphs to a latent variable space to obtain second latent variables of the multiple molecular graphs, the second latent variables describing the ownership information of the molecular graphs, the target decoder being trained based on an original molecular graph dataset and a target digital signature, the original molecular graph dataset including multiple molecular graph samples without ownership information, the target digital signature being generated based on the ownership information of the target molecular graph dataset, the target molecular graph dataset being obtained by embedding ownership information into the original molecular graph dataset, the target decoder being obtained by the training method of any one of claims 1-7; a matching module for matching the second latent variables of the multiple molecular graphs with the target digital signature; and a determination module for determining, when a target molecular graph exists among the multiple molecular graphs, that the second molecular graph generation model is trained based on the target molecular graph dataset, and the second latent variable of the target molecular graph matches the target digital signature.

18. The apparatus according to claim 17, characterized in that, The second latent variable includes a first number of real numbers with positive and negative signs, and the target digital signature includes a first number of binary numbers representing positive and negative signs; the matching module is used to: for each molecular graph's second latent variable, compare the first number of real numbers with the first number of binary numbers, and if the proportion of the target real number is greater than a preset proportion, determine that the second latent variable of the molecular graph matches the target digital signature, wherein the positive and negative signs of the target real number and the positive and negative signs represented by the binary numbers compared with the target real number match.

19. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as the training method of the decoder according to any one of claims 1 to 7 or the model detection method according to any one of claims 8 to 9.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one computer program, which is used to execute the training method of the decoder according to any one of claims 1 to 7 or the model detection method according to any one of claims 8 to 9.

21. A computer program product, characterized in that, The computer program product includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform a training method for a decoder as claimed in any one of claims 1 to 7 or a model detection method as claimed in any one of claims 8 to 9.