Data processing method and molecular processing model training method

By receiving the pending molecular data and molecular domain information in the molecular processing task, and using the molecular processing model for cross-domain feature extraction and downstream task processing, the problem of lack of universality in the existing technology is solved, and more efficient molecular processing and model multiplexing is achieved.

WO2025141487A1PCT designated stage expired Publication Date: 2025-07-03CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/063168
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-25
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Most of the methods in the prior art are tailored to a single domain data domain, limiting the pre-trained model to capture molecular interactions between different chemical domains, resulting in poor learning performance and lack of universality in downstream tasks.

Method used

Provide a data processing method, by receiving the pending molecular data and molecular domain information in the molecular processing task, using the molecular processing model to determine the pending sub-data and data type information based on the molecular domain information, generate molecular processing results, and use a cross-domain molecular processing model for feature extraction and downstream task processing.

Benefits of technology

Molecular processing spanning multiple data formats is realized, model multiplexing rate and downstream task processing efficiency are improved, and single-field molecular characterization performance is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024063168_03072025_PF_FP_ABST
    Figure IB2024063168_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a data processing method and a molecular processing model training method. The data processing method comprises: receiving a molecular processing task, wherein the molecular processing task comprises molecular data to be processed and molecular domain information corresponding to said molecular data; and inputting said molecular data and the molecular domain information into a molecular processing model, so as to obtain a molecular processing result that is output by the molecular processing model, wherein the molecular processing model determines, on the basis of the molecular domain information, at least one piece of sub-data to be processed that corresponds to said molecular data and data type information of each piece of sub-data to be processed, determines, on the basis of each piece of sub-data to be processed and the data type information of each piece of sub-data to be processed, feature information of each piece of sub-data to be processed, and generates a molecular processing result on the basis of the feature information of each piece of sub-data to be processed. By means of the molecular processing model provided in the embodiments of the present disclosure, a universal molecular processing model is provided, and a plurality of data formats are spanned, thereby improving the model reusability of the molecular processing model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Data Processing Method, Method for Training a Molecular Processing Model This disclosure claims priority to Chinese patent application number 202311864584.2, filed with the China Patent Office on December 29, 2023, the entire contents of which are incorporated herein by reference. Technical Field: The embodiments of the present disclosure relate to the field of computer technology, and more particularly to a data processing method. Background: For the traditional pharmaceutical industry, the journey from standard biological research to drug launch takes many years, with a relatively low success rate. Overall, drug development is a costly, time-consuming, and high-risk undertaking. In the rapidly developing digital market, AI technology has been incorporated into traditional pharmaceutical manufacturing, and pre-training of biomolecules has attracted increasing attention. Fine-tuning large-scale pre-trained models can significantly improve the performance of various downstream biological tasks, such as molecular property prediction models, virtual screening models, affinity estimation models, and molecule generation models. Consequently, significant effort has been invested in pre-training biomolecules. However, most existing methods are tailored to a single domain data domain, focusing on small molecules or proteins. This limits the ability of pre-trained models to capture molecular interactions between different chemical domains, thereby limiting learning performance in downstream tasks that rely heavily on this information, such as structure-based binding affinity prediction and virtual screening. Furthermore, these methods lack versatility. Therefore, obtaining a model that can simultaneously address multiple downstream problems and achieve better performance has become a pressing challenge for those skilled in the art. In light of this, embodiments of the present disclosure provide a data processing method. One or more embodiments of the present disclosure also relate to a data processing apparatus, a method for training a molecular processing model, a computing device, a computer-readable storage medium, and a computer program to address the technical deficiencies of the prior art. According to a first aspect of an embodiment of the present disclosure, a data processing method is provided, comprising: receiving a molecular processing task, wherein the molecular processing task includes molecular data to be processed and molecular domain information corresponding to the molecular data to be processed; inputting the molecular data to be processed and the molecular domain information into a molecular processing model, and obtaining a molecular processing result output by the molecular processing model, wherein the molecular processing model determines, based on the molecular domain information, at least one sub-data to be processed corresponding to the molecular data to be processed and data type information of each sub-data to be processed; determining feature information of each sub-data to be processed based on each sub-data to be processed and the data type information of each sub-data to be processed; and generating a molecular processing result based on the feature information of each sub-data to be processed.According to a second aspect of an embodiment of the present disclosure, a data processing method is provided, comprising: receiving a small molecule property prediction task, wherein the small molecule property prediction task includes small molecule data to be processed and molecular domain information corresponding to the small molecule data to be processed; inputting the small molecule data to be processed and the molecular domain information into a molecular property prediction model, and obtaining a molecular property prediction result output by the molecular property prediction model, wherein the molecular property prediction model determines data type information corresponding to the small molecule data to be processed based on the molecular domain information, determines feature information of sub-data to be processed based on the data type information and the small molecule data to be processed, and generates a molecular property prediction result based on the features of the sub-data to be processed. According to a third aspect of an embodiment of the present disclosure, a data processing method is provided, comprising: receiving a virtual screening task, wherein the virtual screening task includes complex data to be processed and molecular domain information corresponding to the complex data to be processed; inputting the complex data to be processed and the molecular domain information into a molecular virtual screening model, and obtaining a molecular virtual screening result output by the molecular virtual screening model, wherein the molecular virtual screening model determines small molecule data to be processed and protein data to be processed corresponding to the complex data to be processed based on the molecular domain information, determines feature information of sub-data to be processed based on the small molecule data to be processed and the protein data to be processed, and generates a molecular virtual screening result based on the feature data to be processed. According to a fourth aspect of an embodiment of the present disclosure, a method for training a molecular processing model is provided, applied to a cloud-side device, comprising: obtaining a first training sample, wherein the first training sample includes first sample molecular data and sample molecular domain information corresponding to the first sample molecular data; training a molecular processing pre-trained model based on the first sample molecular data and the sample molecular domain information corresponding to the first sample molecular data, wherein the molecular processing pre-trained model includes a classification layer, an embedding layer, and an encoding layer; obtaining a second training sample, wherein the second training sample includes second sample molecular data and sample label data corresponding to the second sample molecular data; training a molecular processing model based on the second sample molecular data and the sample label data corresponding to the second sample molecular data, and obtaining model parameters of the molecular processing model, wherein the molecular processing model includes the molecular processing pre-trained model and a downstream task layer; and sending the model parameters of the molecular processing model to an end-side device. According to a fifth aspect of an embodiment of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is configured to store computer-executable instructions, and the processor is configured to execute the computer-executable instructions. When executed by the processor, the computer-executable instructions implement the steps of the aforementioned data processing method or molecular processing model training method.According to a sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, storing computer-executable instructions. When executed by a processor, the instructions implement the steps of the aforementioned data processing method or molecular processing model training method. According to a seventh aspect of the embodiments of the present disclosure, a computer program is provided. When executed on a computer, the computer is instructed to execute the steps of the aforementioned data processing method or molecular processing model training method. A data processing method provided in one embodiment of the present disclosure includes receiving a molecular processing task, wherein the molecular processing task includes molecular data to be processed and molecular domain information corresponding to the molecular data to be processed; inputting the molecular data to be processed and the molecular domain information into a molecular processing model; obtaining a molecular processing result output by the molecular processing model; wherein the molecular processing model determines at least one sub-data to be processed corresponding to the molecular data to be processed and data type information of each sub-data to be processed based on the molecular domain information; determines feature information of each sub-data to be processed based on each sub-data to be processed and the data type information of each sub-data to be processed; and generating a molecular processing result based on the feature information of each sub-data to be processed. Through the methods provided in the embodiments of the present disclosure, a molecular processing model can split the molecular data to be processed based on its molecular domain information, obtaining at least one sub-data to be processed and data type information corresponding to each sub-data. Feature information for each sub-data to be processed is further determined based on the sub-data and its data type information, so that the feature information for each sub-data to be processed is determined based on the specificity and relationship between the data type information. A molecular processing result is then generated based on the feature information for each sub-data to be processed. The embodiments of the present disclosure provide a universal molecular processing model that can encode molecules across a wide range of biochemical fields, including small molecules, proteins, protein-small molecule complexes, and more. These molecules, spanning multiple data formats, can be processed within the same molecular processing model, eliminating the need to train separate models based on molecule type. This improves the model reuse rate of the molecular processing model. This allows the molecular processing model to handle different molecular processing tasks based on different downstream tasks, improving model efficiency. Furthermore, cross-domain molecular representation learning can also improve molecular representation performance within a single domain.BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 is an architectural diagram of a data processing system provided by one embodiment of the present disclosure; Figure 2 is a flow chart of a data processing method provided by one embodiment of the present disclosure; Figure 3 is a schematic diagram of the structure of the encoding layer provided by one embodiment of the present disclosure; Figure 4 is a schematic diagram of the structure of the embedding layer in the model training phase provided by one embodiment of the present disclosure; Figure 5 is a flow chart of a data processing method for a small molecule property prediction task provided by one embodiment of the present disclosure; Figure 6 is a flow chart of a data processing method for a virtual screening task provided by one embodiment of the present disclosure; Figure 7 is a flow chart of a molecular processing model training method provided by one embodiment of the present disclosure; Figure 8 is a schematic diagram of the structure of a data processing apparatus provided by one embodiment of the present disclosure; and Figure 9 is a block diagram of a computing device provided by one embodiment of the present disclosure. The following description sets forth numerous specific details to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art may make similar generalizations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific embodiments disclosed below. The terminology used in one or more embodiments of the present disclosure is for the purpose of describing specific embodiments only and is not intended to limit the present disclosure. As used in one or more embodiments of the present disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be understood that although the terms first, second, etc. may be employed in one or more embodiments of the present disclosure to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, the first could be referred to as the second, and similarly, the second could be referred to as the first, without departing from the scope of one or more embodiments of the present disclosure. Depending on the context, the word "if," as used herein, could be interpreted as "when..." or "when..." or "in response to determining." In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant region, and corresponding operation portals are provided for users to choose to authorize or refuse.The large model in one or more embodiments of the present disclosure specifically refers to a deep learning model with large-scale model parameters, typically including hundreds of millions, tens of billions, or even hundreds of billions of model parameters. Large models, also known as foundation models (Foundation Mode I), are pre-trained using large-scale unlabeled corpora to produce pre-trained models with over 100 million parameters. Such models are adaptable to a wide range of downstream tasks and exhibit good generalization capabilities. Examples include Large Language Models (LLM) and multi-modal pre-training models (multi-modal pre-training mode I). In practical applications, large models only require a small number of samples to fine-tune the pre-trained model and can be applied to various tasks. Large models can be widely used in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. Key application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. First, the terms used in one or more embodiments of this disclosure are explained.

[0002] MODE: Mixture-of-Domain-Experts (MIXED-DOMAIN-EXPERTS), a method derived from multimodality, aims to capture multi-domain specificity and inter-domain relationships. Goodness-of-Goal Denoising: Atomic coordinate denoising, which randomly adds noise to coordinates and forces the model to predict this noise, is often used as a pre-training task for molecules.

[0003] Masked Token Denoising: Masked denoising randomly masks token types and allows the model to recover the masked token types. It is often used as a pre-training task in natural language processing, vision, and molecular biology. Traditionally, the pharmaceutical industry's pipeline, from standard biological research to drug approval and marketing, takes years, with a low success rate. Overall, drug development is a costly, time-consuming, and high-risk undertaking. In today's rapidly developing digital landscape, artificial intelligence has made significant progress and has begun to be integrated into traditional pharmaceutical processes. In recent years, pre-training of biomolecules has attracted increasing attention. Fine-tuning large pre-trained models can significantly improve the performance of various downstream biological tasks, such as molecular property prediction, virtual screening, affinity estimation, molecule generation, and protein structure prediction. Consequently, researchers have invested significant effort in pre-training biomolecules to leverage the inherent potential of large-scale unlabeled molecular corpora. However, existing methods are all based on single-domain data domains, focusing on small molecules or proteins. This limits the ability of pre-trained models to capture molecular interactions across different chemical domains. Consequently, this limits learning performance in downstream tasks that rely heavily on this information, such as structure-based affinity prediction and virtual screening, and also lacks general applicability. To address this, the present disclosure provides a data processing method. One or more embodiments of the present disclosure also involve a data processing apparatus, a molecular processing model training method, a computing device, and a computer-readable storage medium, each of which is described in detail in the following embodiments.1 , which shows an architecture diagram of a data processing system provided by one embodiment of the present disclosure. The data processing system may include a client 100 and a server 200. The client 100 is configured to send molecular data to be processed and molecular domain information corresponding to the molecular data to be processed to the server 200. The server 200 is configured to input the molecular data to be processed and the molecular domain information into a molecular processing model to obtain a molecular processing result output by the molecular processing model, wherein the molecular processing model determines at least one sub-data to be processed corresponding to the molecular data to be processed and data type information of each sub-data to be processed based on the molecular domain information, determines feature information of each sub-data to be processed based on each sub-data to be processed and the data type information of each sub-data to be processed, and generates a molecular processing result based on the feature information of each sub-data to be processed. The molecular processing result is then sent to the client 100. The client 100 is further configured to receive the molecular processing result sent by the server 200. A data processing system may include multiple clients 100 and a server 200. The clients 100 can be referred to as client-side devices, and the server 200 can be referred to as cloud-side devices. Multiple clients 100 can establish a communication connection through the server 200. In a molecular processing scenario, the server 200 provides molecular processing services to multiple clients 100. Multiple clients 100 can act as senders or receivers, communicating through the server 200. Users can interact with the server 200 through the clients 100 to receive data from other clients 100 or send data to other clients 100. In a molecular processing scenario, users can publish data streams to the server 200 through the clients 100. The server 200 generates molecular processing results based on the data streams and pushes the molecular processing results to other clients with which they have established communication. The connection between the clients 100 and the server 200 is established via a network. The network provides the medium for the communication link between the clients 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.The data transmitted by the client 100 may need to undergo encoding, transcoding, compression, and other processing before being published to the server 200. The client 100 may be a browser, an app (Application Program), a web application such as an H5 (HyperText Markup Languages, Version 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 may be developed based on a software development kit (SDK) for a corresponding service provided by the server 200, such as a real-time communication (RTC) SDK. The client 100 may be deployed in an electronic device and may rely on the device or certain apps in the device to operate. For example, the electronic device may have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, or a personal computer. Electronic devices can also typically be configured with various other types of applications, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, and the like. The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that provide backend training support for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers or as a single server. The server can also be a server in a distributed system or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data and artificial intelligence platforms, or intelligent cloud computing servers or intelligent cloud hosts with artificial intelligence technology. It is worth noting that the data processing methods provided in the embodiments of the present disclosure are generally executed by the server. However, in other embodiments of the present disclosure, the client may also have similar functions to the server and thereby execute the data processing methods provided in the embodiments of the present disclosure. In other embodiments, the data processing methods provided in the embodiments of the present disclosure may also be jointly executed by the client and the server.Referring to FIG. 2 , FIG. 2 shows a flow chart of a data processing method provided by an embodiment of the present disclosure, specifically comprising the following steps: Step 202: Receive a molecular processing task, wherein the molecular processing task includes molecular data to be processed and molecular domain information corresponding to the molecular data to be processed. A molecular processing task specifically refers to a task for processing molecular data in the embodiments provided by the present disclosure. In practical applications, molecular processing tasks may include molecular property prediction tasks, virtual screening tasks, affinity estimation tasks, molecular generation tasks, protein structure prediction tasks, and the like. A molecular processing task typically carries molecular data to be processed for the molecular processing task. The molecular domain information corresponding to the molecular data to be processed specifically refers to the type of the molecular data to be processed. In practical applications, the molecular data to be processed may be small molecule data to be processed, protein data to be processed, or data of a small molecule-protein complex to be processed. The molecular domain information specifically indicates the type of the molecular data to be processed. For example, the molecular domain information corresponding to small molecule data may be "small molecule," the molecular domain information corresponding to protein data may be "protein," and the molecular domain information corresponding to small molecule-protein complex data may be "complex." In one specific embodiment provided by the present disclosure, a molecular processing task is received. The molecular processing task carries molecular data 1 to be processed, and the molecular domain information corresponding to the molecular data 1 is "small molecule." For another example, the molecular processing task carries molecular data 2 to be processed, and the molecular domain information corresponding to the molecular data 2 is "protein." For another example, the molecular processing task carries molecular data 3 to be processed, and the molecular domain information corresponding to the molecular data 3 is "complex." It should be noted that in the embodiments provided by the present disclosure, the molecular data to be processed are all digitized structures of physical molecules. The molecular data to be processed can be either 2D or 3D structures. A 2D structure specifically means that the molecular data to be processed includes nodes and edges, where nodes are atoms and edges are chemical bonds between atoms. A 3D structure specifically means that, in addition to the chemical bonds between atoms, it also includes the three-dimensional coordinates of each atom in the molecular data to be processed, i.e., the spatial positional relationships between atoms.Step 204: Input the molecular data to be processed and the molecular domain information into a molecular processing model to obtain a molecular processing result output by the molecular processing model. The molecular processing model determines, based on the molecular domain information, at least one sub-data to be processed corresponding to the molecular data to be processed and the data type information of each sub-data to be processed. Based on each sub-data to be processed and the data type information of each sub-data to be processed, the model determines feature information of each sub-data to be processed, and generates a molecular processing result based on the feature information of each sub-data to be processed. After obtaining the molecular data to be processed and the molecular domain information, the molecular data to be processed and the molecular domain information are input together into the molecular processing model for processing, obtaining a molecular processing result output by the molecular processing model. A molecular processing model specifically refers to a model used to perform molecular processing tasks. In the method provided in the embodiments of the present disclosure, the molecular processing model first determines, based on the molecular domain information, the sub-data to be processed corresponding to the molecular data to be processed and the data type information corresponding to each sub-data to be processed. Specifically, the sub-data to be processed refers to the elements that make up the molecular data to be processed. For example, when the molecular data to be processed is small molecule data, the data type information of the sub-data to be processed is small molecule; when the molecular data to be processed is protein data, the data type information of the sub-data to be processed is protein; and when the molecular data to be processed is a complex, the data type of the sub-data to be processed is both small molecule and protein. After determining the data type information of each sub-data to be processed, the molecular processing model performs feature extraction based on each sub-data to be processed and the corresponding data type information, obtaining feature information corresponding to each sub-data to be processed. It then performs downstream task processing based on the feature information of each sub-data to be processed to obtain the final molecular processing result. The molecular processing model provided in the embodiments of the present disclosure can encode molecules across a wide range of biochemical fields, including small molecules, proteins, protein-small molecule complexes, and so on. This avoids the need to train different molecular processing models for each molecular processing task. In an embodiment provided herein, a molecule processing model includes a feature extraction module and a downstream task module. The feature extraction model can be trained through model pre-training so that it can process various types of molecules and output feature information of sub-data to be processed. The downstream task module generates different results based on the feature information of each sub-data to be processed, depending on the downstream task.In a specific embodiment provided by the present disclosure, the molecular processing model includes a classification layer, an embedding layer, an encoding layer, and a downstream task layer; the molecular data to be processed includes at least one sub-data to be processed; inputting the molecular data to be processed and the molecular domain information into the molecular processing model to obtain a molecular processing result output by the molecular processing model includes S2042-S2048:

[0004] S2042: Input the molecular data to be processed and the molecular domain information into the classification layer, and obtain at least one sub-data to be processed and data type information corresponding to each sub-data to be processed, determined by the classification layer based on the molecular domain information. In practical applications, the classification layer, embedding layer, and encoding layer can all be understood as feature extraction modules of the molecular processing model, and the downstream task layer can be understood as the downstream task module of the molecular processing model. The classification layer specifically refers to the process of segmenting the molecular data to be processed based on the molecular domain information of the molecular data to be processed. In one specific embodiment provided herein, taking the molecular domain information of molecular data 1 to be processed as "small molecule" as an example, the classification layer determines that molecular data 1 to be processed includes: sub-data to be processed 1, and data type information of sub-data 1 to be processed as "small molecule." In another specific embodiment provided herein, taking the molecular domain information of molecular data 2 to be processed as "protein" as an example, the classification layer determines that molecular data 2 to be processed includes: sub-data to be processed 1, and data type information of sub-data 1 to be processed as "protein." In another specific embodiment provided by the present disclosure, taking the molecular domain information of the molecular data to be processed 3 as "complex" as an example, the classification layer determines that the molecular data to be processed 3 includes: sub-data to be processed 1 and sub-data to be processed 2. The data type information of sub-data to be processed 1 is "small molecule", and the data type information of sub-data to be processed 2 is "protein". The purpose of the classification layer is to split the molecular data to be processed input into the molecular processing model to obtain at least one sub-data to be processed, facilitating subsequent processing by the molecular processing model based on the data type information of each sub-data to be processed. Through the classification layer, the molecular processing model can process molecules across a wide range of biochemical fields, eliminating the need to train different models for different types of molecules.

[0005] S2044: Input each sub-data to be processed and the data type information corresponding to each sub-data to be processed into the embedding layer, and obtain embedded feature information corresponding to each sub-data to be processed, which is output by the embedding layer based on each sub-data to be processed and the data type information corresponding to each sub-data to be processed. The embedding layer specifically encodes data to generate feature information recognizable by the terminal. In the embodiments provided herein, each sub-data to be processed and the data type corresponding to each sub-data to be processed are input into the embedding layer for embedding processing, thereby generating embedded feature information corresponding to each sub-data to be processed. Furthermore, after the sub-data to be processed is input into the embedding layer, embedding processing is first performed on the sub-data to be processed to obtain initial embedded feature information of the sub-data to be processed; then, structural position information of the sub-data to be processed is embedded to obtain structural position embedded feature information; and finally, data type information of the sub-data to be processed is embedded to obtain data type embedded feature information. Finally, the initial sub-data to be processed embedding feature information, the structural position embedding feature information and the data type embedding feature information are fused to obtain the sub-data to be processed embedding feature information corresponding to the sub-data to be processed.

[0006] S2046: Input the embedded feature information of each sub-data to be processed into the encoding layer to obtain the molecular coding feature information to be processed output by the encoding layer, wherein the molecular coding feature information to be processed is obtained based on the data type information. After obtaining the embedded feature information of the sub-data to be processed, it is input into the encoding layer for feature extraction. After processing by the encoding layer, the molecular coding feature information to be processed corresponding to the embedded feature information of the sub-data to be processed is obtained. In actual applications, a domain expert mixing module is provided in the encoding layer to perform feature enhancement on the embedded feature information of the sub-data to be processed, thereby capturing the specificities and relationships between multiple data types. The ultimately generated molecular coding feature information to be processed is obtained by utilizing the specificities and relationships between the multiple data types captured by the type mixing module. In one embodiment of the present disclosure, the encoding layer includes a shared self-attention module and a type mixing module. The encoding layer embeds feature information of each sub-data to be processed and outputs the encoded feature information of the molecule to be processed. This includes: inputting each sub-data to be processed into the shared self-attention module to obtain initial molecular feature information corresponding to each sub-data to be processed; inputting each initial molecular feature information into the type mixing module to obtain type mixing feature information corresponding to each initial molecular feature information; and generating the encoded feature information of the molecule to be processed based on each type mixing feature information. Furthermore, the encoding layer includes both a shared self-attention module and a type mixing module. The shared self-attention module is used to facilitate alignment between data of different data types and to learn cross-domain interaction information. The type mixing module, also known as a domain expert mixture module, is used to capture the specificity of multiple data types and the relationship between data type information, ultimately mixing to generate the encoded feature information of the molecule to be processed. When the embedded feature information of the sub-data to be processed is input into the encoding layer, it is first processed by the shared self-attention module to obtain the initial molecular feature information corresponding to the embedded feature information of the sub-data to be processed. The initial molecular feature information specifically refers to the information output by the shared self-attention module. This initial molecular feature information is then input into the type mixing module, which performs feature processing based on the type information corresponding to each molecule's initial feature information to obtain type mixing feature information corresponding to each molecule's initial feature information. The type mixing feature information corresponding to each molecule's initial feature information is then fused to generate the final molecular encoding feature information to be processed.In practical applications, the encoding layer includes multiple identical encoding stacks, meaning the aforementioned processing within the encoding layer is repeated multiple times. For example, the encoding layer includes six encoding stacks, each of which includes a shared self-attention module and a type mixing module. The output of the first encoding stack serves as the input to the second encoding stack, the output of the second encoding stack serves as the input to the third encoding stack, and so on, until the output of the last encoding stack is reached. The output of the last encoding stack can then be used as the output of the encoding layer. Alternatively, the outputs of each encoding stack can be fused based on their weights to form the output of the encoding layer. In the embodiments provided herein, this is not limited to this and will only be determined based on practical applications. In one specific embodiment provided herein, the type mixing module includes at least one type feature submodule. Inputting each molecule's initial feature information into the type mixing module to obtain type mixing feature information corresponding to each molecule's initial feature information includes: determining the data type information corresponding to each molecule's initial feature information and the type feature submodule corresponding to each data type; and inputting each molecule's initial feature information into the corresponding type feature submodule to obtain type mixing feature information corresponding to each molecule's initial feature information. Referring to FIG3 , FIG3 shows a schematic diagram of the structure of the encoding layer provided by an embodiment of the present disclosure. As shown in FIG3 , the encoding layer includes a shared self-attention module and a type mixing module. The type mixing module includes two type feature submodules, each of which processes its corresponding initial molecular feature information. Specifically, the two type feature submodules in the type mixing module are a small molecule submodule and a protein submodule. As shown in FIG3 , the initial molecular feature information is input into the encoding layer and first processed by the shared self-attention module to obtain the initial molecular feature information corresponding to the embedded feature information of the sub-data to be processed. The initial molecular feature information is then input into the type mixing module, which outputs the encoded feature information of the molecular to be processed. The type mixing module includes two type feature submodules, a small molecule submodule and a protein submodule. The data types corresponding to the sub-data to be processed include small molecules and proteins. Specifically, when the data type corresponding to the initial molecular feature information is small molecule, the initial molecular feature information is input into the small molecule submodule; when the data type corresponding to the initial molecular feature information is protein, the initial molecular feature information is input into the protein submodule. In a specific embodiment provided by the present disclosure, when the molecular data 1 to be processed is a small molecule, the data type information of the sub-data 1 to be processed is a small molecule. Then, the initial molecular feature information corresponding to the sub-data 1 to be processed is input into the small molecule sub-module for feature extraction, and finally type mixed feature information is obtained.When the molecular data to be processed 2 is a protein, the data type information of the sub-data to be processed 1 is included as protein. The initial molecular feature information corresponding to the sub-data to be processed 1 is then input into the protein sub-module for feature extraction, ultimately obtaining mixed-type feature information. When the molecular data to be processed 3 is a complex, the data type information of the sub-data to be processed 1 is included as small molecule, and the data type information of the sub-data to be processed 2 is protein. The initial molecular feature information corresponding to the sub-data to be processed 1 is then input into the small molecule sub-module for feature extraction, and the initial molecular feature information corresponding to the sub-data to be processed 2 is input into the protein sub-module for feature extraction. The two types of information are then mixed to obtain mixed-type feature information.

[0007] S2048: Input the coding feature information of the molecules to be processed into the downstream task layer to obtain a molecular processing result output by the downstream task layer. After feature extraction by multiple coding stacks in the encoder, the final coding feature information of the molecules to be processed can be obtained. The coding feature information of the molecules to be processed is then input into the downstream task layer for processing to obtain a molecular processing result corresponding to the molecular processing task. In practical applications, the downstream task layer corresponds to the molecular processing task. If the molecular processing task is a molecular property prediction task, the downstream task layer is the molecular property prediction layer; if the molecular processing task is a virtual screening task, the downstream task layer is the virtual screening layer; if the molecular processing task is an affinity estimation task, the downstream task layer is the affinity estimation layer; if the molecular processing task is a molecule generation task, the downstream task layer is the molecule generation layer; if the molecular processing task is a protein structure prediction task, the downstream task layer is the protein structure prediction layer. The method provided by the embodiments of the present disclosure provides a molecular processing model. In this model, molecular data to be processed can be split based on its molecular domain information to obtain at least one sub-data to be processed and the data type information corresponding to each sub-data. The encoding layer of the feature extraction module of the molecular processing model includes a shared self-attention module and a type mixing module. The shared self-attention module is used to promote data alignment between different types of data and learn similar information between different data types. The type mixing module includes multiple type feature sub-modules. The corresponding type feature sub-module is determined based on the data type information corresponding to each sub-data to be processed, thereby capturing the specificity and relationship between multiple data types, ultimately generating encoding feature information for the molecules to be processed. Finally, through the downstream task layer, the molecular processing results of the molecular processing task are output. The molecular processing model provided by the embodiments of the present disclosure provides a universal molecular processing model that can encode molecules across a wide range of biochemical fields, including small molecules, proteins, protein-small molecule complexes, and so on. These molecules can be processed within the same molecular processing model across multiple data formats, eliminating the need to train separate models based on molecule type, thereby improving the model reuse rate of the molecular processing model. This enables the molecular processing model to handle different molecular processing tasks according to different downstream tasks, improving the efficiency of model use. At the same time, cross-domain molecular representation learning can also improve the molecular representation performance in a single domain.In a specific embodiment provided by the present disclosure, the molecular processing model is trained by the following steps: obtaining a first training sample, wherein the first training sample includes first sample molecular data and sample molecular domain information corresponding to the first sample molecular data; training a molecular processing pre-trained model based on the first sample molecular data and the sample molecular domain information corresponding to the first sample molecular data, wherein the molecular processing pre-trained model includes a classification layer, an embedding layer, and an encoding layer; obtaining a second training sample, wherein the second training sample includes second sample molecular data and sample label data corresponding to the second sample molecular data; training a molecular processing model based on the second sample molecular data and the sample label data corresponding to the second sample molecular data, wherein the molecular processing model includes the molecular processing pre-trained model and a downstream task layer. As described in the above embodiment, the molecular processing model includes a feature extraction module and a downstream task module. The feature extraction module is shared between different molecular processing models, while different molecular processing models have different downstream task modules. Therefore, the first training sample specifically refers to a sample used to train the feature extraction module. In the embodiment provided by the present disclosure, the first training sample includes the first sample molecular data and the sample molecular domain information corresponding to the first sample molecular data. In practical applications, the first sample molecular data in the first training sample includes a data format of the sample molecular data, which may be a 2D format or a 3D format. Specifically, the 2D structure refers to the molecular data to be processed comprising nodes and edges, where nodes are atoms and edges are chemical bonds between atoms. Specifically, the 3D structure refers to the molecular data to be processed comprising, in addition to the chemical bonds between atoms, the three-dimensional coordinates of each atom in the molecular data to be processed, i.e., the spatial positional relationships between atoms. The first sample molecular data is used to pre-train the feature extraction module in the molecular processing model. Specifically, the classification layer, embedding layer, and encoding layer in the feature extraction module are pre-trained. The process of training the feature extraction module of the molecular processing model using the first sample molecular data employs a self-supervised training approach. After obtaining the feature extraction module, corresponding second sample molecular data is obtained based on different downstream tasks. The feature extraction module and downstream task layers are further trained using the second sample molecular data. The feature extraction module can be understood as a pre-trained molecular processing model. In the method provided herein, the feature extraction module (pre-trained molecular processing model) is first pre-trained. The feature extraction module and downstream task layers are then combined to form a molecular processing model. The molecular processing model is further trained using the second sample molecular data, thereby obtaining a molecular processing model.In the process of further training the molecular processing model using the second sample molecular data, a supervised approach is used for training. In another specific embodiment provided by the present disclosure, the molecular processing pre-training model includes a classification layer, an embedding layer, and an encoding layer; training the molecular processing pre-training model based on first sample molecular data and sample molecular domain information corresponding to the first sample molecular data includes: inputting the first sample molecular data and the sample molecular domain information into the classification layer to obtain at least one sample sub-data determined by the classification layer based on the sample molecular domain information and sample data type information corresponding to each sample sub-data; inputting each sample sub-data and the sample data type information corresponding to each sample sub-data into the embedding layer to obtain sample sub-data embedding feature information corresponding to each sample sub-data; inputting each sample sub-data embedding feature information into the encoding layer to obtain sample molecular encoding feature information output by the encoding layer, wherein the sample molecular encoding feature information is obtained based on each sample data type information; inputting the sample molecular encoding feature information into the output layer to obtain a molecular pre-processing result output by the output layer; performing multi-task learning of continuous atomic coordinates and categorical atomic types based on the molecular pre-processing result to obtain a first model loss value; and adjusting model parameters of the molecular processing pre-training model based on the first model loss value until a model training stop condition is met. During the model training phase, the first sample molecule data and sample molecule domain information are input into the classification layer. After classification processing by the sample classification layer, at least one sample sub-data corresponding to the first sample molecule data and sample data type information corresponding to each sample sub-data are obtained. Each sample sub-data and the sample data type information corresponding to each sample sub-data are then input into the embedding layer. In the embedding layer, embedded feature information corresponding to each sample sub-data is generated based on the sample sub-data's feature information, position encoding information, and sample data type information. The embedded feature information of the sample sub-data is input into the encoding layer, which includes a shared self-attention module and a type mixing module. The type mixing module includes at least one type feature sub-module. The embedded feature information of the sample sub-data is input into the shared self-attention module for feature extraction. The output feature information is input into the type mixing module. Based on the sample data type information corresponding to each sample sub-data, the type feature sub-module is input into the corresponding type feature sub-module. This allows the type feature sub-module to learn the specificity and relationship between multiple data types, ultimately outputting the sample molecule encoded feature information.Finally, the sample molecule encoding feature information is input into the output layer to obtain a molecule preprocessing result output by the output layer. After obtaining the molecule preprocessing result, a model loss value is calculated using multi-task learning of continuous atomic coordinates and classified atomic types. The feature extraction module (molecule processing pretraining model) is jointly pretrained using a multi-task learning method. After a preset number of steps of model training, a pretrained feature extraction module (molecule processing pretraining model) is obtained. In another specific embodiment provided by the present disclosure, each sample sub-data and the sample data type information corresponding to each sample sub-data are input into the embedding layer to obtain the sample sub-data embedding feature information corresponding to each sample sub-data, including: embedding the sample sub-data and the sample data type information to obtain sample sub-data feature information and sample data type embedding information; determining a structure position embedding strategy for the sample sub-data according to a preset structure control strategy, and determining the structure position embedding information corresponding to the sample sub-data according to the structure embedding strategy, wherein the structure position embedding strategy includes a 2D strategy or a 3D strategy; The sample data type embedding information and the structure position embedding information generate sample sub-data embedding feature information corresponding to the sample sub-data. In practical applications, the position encoding information of the sample sub-data is determined based on the position information of the sample sub-data. In practical applications, the position encoding information of the sample sub-data can be 2D encoding information or 3D encoding information. To improve the processing versatility of the model, a structure control unit can be provided to determine the structure position embedding strategy of the sample sub-data using a preset structure control strategy. Referring to FIG4 , FIG4 shows a schematic diagram of the structure of the embedding layer during the model training phase according to an embodiment of the present disclosure. As shown in FIG4 , during the embedding process of the first sample molecular data, the sample sub-data needs to be embedded to obtain sample sub-data features; the sample data type information needs to be embedded to obtain sample data type embedding information; and the structure information of the sample sub-data is inputted via the structure control unit. When the 2D strategy is selected, the two-dimensional structure information of the sample sub-data is inputted; when the 3D strategy is selected, the three-dimensional structure information of the sample sub-data is inputted. The structure position embedding information is obtained based on the structure information of the sample sub-data. Finally, the sample sub-data embedding feature information is obtained according to the sample sub-data feature, the sample data type embedding information and the structure position embedding information.Using 2D and 3D structure control units, the model's learning capabilities are enhanced by controlling whether 2D or 3D features are used to obtain structural position embedding information for molecular data during each training session. For complexes, proteins within a 5-angstrom region of the small molecule domain can be input into the network to obtain structural information. During pre-training of the feature extraction module, the structure control unit can determine the structural type of the sample sub-data input into the embedding layer, allowing the model to learn a variety of molecular structures. This allows the feature extraction module to process both 2D and 3D sample molecular data, enabling subsequent molecular processing models to identify molecules from a wide range of biochemical domains, thereby improving the model's generalization. In another specific embodiment provided by the present disclosure, the molecular processing model includes a molecular processing pre-training model and a downstream task layer. Training the molecular processing model based on second sample molecular data and sample label data corresponding to the second sample molecular data includes: inputting the second sample molecular data into the molecular processing pre-training model to obtain molecular encoding feature information output by the molecular processing pre-training model; inputting the molecular encoding feature information into the downstream task layer to obtain predicted label data output by the downstream task layer; calculating a second model loss value based on the predicted label data and the sample label data; adjusting model parameters of the molecular processing pre-training model and the model parameters of the downstream task layer based on the second model loss value, and continuing to train the molecular processing model until a model training stop condition is met. After obtaining the pre-trained feature extraction module (molecular processing pre-training model), the molecular processing model that requires further training is formed by combining it with the downstream task layer corresponding to a specific downstream task. In this process, the second sample molecular data serves as a training sample for supervised training. The second sample molecular data is input into the molecular processing pre-training model for processing to obtain molecular encoding feature information output by the molecular processing pre-training model. The molecular encoding feature information is input into the downstream task layer to obtain predicted label data output by the downstream task layer. The predicted label data is then compared with the sample label data corresponding to the second sample molecular data to calculate the model loss value. At this point, the molecular processing model is not yet trained. After obtaining the predicted label data, the model loss value can be calculated based on the memory comparison of the predicted label data and the sample label data. In the method provided in the present disclosure, there are many methods for calculating the model loss value, such as the cross-beam loss function, the maximum loss function, the average loss function, etc. In the present disclosure, the specific method of the loss function is not limited and is subject to actual application.After obtaining the second model loss value, the model parameters of the molecular processing model can be adjusted based on the second model loss value. Specifically, the second model loss value can be back-propagated to sequentially adjust the model parameters in the downstream task layer and the molecular processing pre-trained model. It should be noted that in the method provided herein, the molecular processing pre-trained model is pre-trained using the first sample molecular data. After obtaining the molecular processing pre-trained model, the model parameters of the molecular processing pre-trained model are not fixed. Instead, the model parameters are further fine-tuned during further training of the molecular processing model using the second sample molecular data, thereby obtaining the final molecular processing model. During the model training process, a shared self-attention module approach is employed, combined with a type mixing module. Cross-domain knowledge is learned to correlate in the shared self-attention module, and specificity between types is established in the type mixing module. This results in a unified multi-domain feature extraction module (molecular processing pre-trained model). This ensures that when deploying the molecular processing model, the deployment and usage costs do not increase with the increase in downstream tasks. Referring to FIG5 , FIG5 is a flowchart of a data processing method provided by an embodiment of the present disclosure, applied to a molecular property prediction task, comprising: Step 502: Receiving a small molecule property prediction task, wherein the small molecule property prediction task includes small molecule data to be processed and molecular domain information corresponding to the small molecule data to be processed. Step 504: Inputting the small molecule data to be processed and the molecular domain information into a molecular property prediction model to obtain a molecular property prediction result output by the molecular property prediction model, wherein the molecular property prediction model determines data type information corresponding to the small molecule data to be processed based on the molecular domain information, determines feature information of sub-data to be processed based on the data type information and the small molecule data to be processed, and generates a molecular property prediction result based on the features of the sub-data to be processed. Referring to FIG6 , FIG6 is a flowchart of a data processing method provided by an embodiment of the present disclosure, applied to a virtual screening task, comprising: Step 602: Receiving a virtual screening task, wherein the virtual screening task includes complex data to be processed and molecular domain information corresponding to the complex data to be processed. Step 604: Input the complex data to be processed and the molecular domain information into a molecular virtual screening model to obtain a molecular virtual screening result output by the molecular virtual screening model, wherein the molecular virtual screening model determines the small molecule data to be processed and the protein data to be processed corresponding to the complex data to be processed based on the molecular domain information, determines feature information of the sub-data to be processed based on the small molecule data to be processed and the protein data to be processed, and generates a molecular virtual screening result based on the feature data to be processed.Referring to Figure 7, a flowchart of a molecular processing model training method provided in accordance with an embodiment of the present disclosure is shown. The method, applied to a cloud-side device, specifically includes the following steps: Step 702: Obtain a first training sample, wherein the first training sample includes first sample molecular data and sample molecular domain information corresponding to the first sample molecular data. Step 704: Train a molecular processing pre-trained model based on the first sample molecular data and the sample molecular domain information corresponding to the first sample molecular data. The molecular processing pre-trained model includes a classification layer, an embedding layer, and an encoding layer. Step 706: Obtain a second training sample, wherein the second training sample includes second sample molecular data and sample label data corresponding to the second sample molecular data. Step 708: Train a molecular processing model based on the second sample molecular data and the sample label data corresponding to the second sample molecular data to obtain model parameters for the molecular processing model. The molecular processing model includes the molecular processing pre-trained model and a downstream task layer. Step 710: Send the model parameters of the molecular processing model to the end-side device. It should be noted that steps 702 and 708 are implemented in the same manner as the above-described molecular processing model training method and are not further described in detail in this embodiment of the present disclosure. In practical applications, since model training requires a large amount of data and sufficient computing resources, the end-side device may not have the corresponding processing capabilities. Therefore, the model training process can be implemented on the cloud-side device. After obtaining the model parameters of the molecular processing model, the cloud-side device can also send the model parameters to the end-side device. The end-side device can locally construct a question-and-answer model based on the model parameters of the molecular processing model and further utilize the molecular processing model to perform question-and-answer tasks. Corresponding to the above-described method embodiment, the present disclosure also provides an embodiment of a data processing device. Figure 8 shows a schematic structural diagram of a data processing device provided in one embodiment of the present disclosure. As shown in FIG8 , the apparatus includes: a receiving module 802 configured to receive a molecular processing task, wherein the molecular processing task includes molecular data to be processed and molecular domain information corresponding to the molecular data to be processed; and a processing module 804 configured to input the molecular data to be processed and the molecular domain information into a molecular processing model to obtain a molecular processing result output by the molecular processing model, wherein the molecular processing model determines at least one sub-data to be processed corresponding to the molecular data to be processed and data type information of each sub-data to be processed based on the molecular domain information, determines feature information of each sub-data to be processed based on each sub-data to be processed and the data type information of each sub-data to be processed, and generates a molecular processing result based on the feature information of each sub-data to be processed.Optionally, the molecular processing model includes a classification layer, an embedding layer, an encoding layer, and a downstream task layer, and the molecular data to be processed includes at least one sub-data to be processed; the processing module 804 is further configured to: input the molecular data to be processed and the molecular domain information into the classification layer, to obtain at least one sub-data to be processed and data type information corresponding to each sub-data to be processed determined by the classification layer based on the molecular domain information; input each sub-data to be processed and the data type information corresponding to each sub-data to be processed into the embedding layer, to obtain embedded feature information of each sub-data to be processed output by the embedding layer based on each sub-data to be processed and the data type information corresponding to each sub-data to be processed; input the embedded feature information of each sub-data to be processed into the encoding layer, to obtain coding feature information of the molecule to be processed output by the encoding layer, wherein the coding feature information of the molecule to be processed is obtained based on each data type information; and input the coding feature information of the molecule to be processed into the downstream task layer, to obtain a molecular processing result output by the downstream task layer. Optionally, the encoding layer includes a shared self-attention module and a type mixing module; the processing module 804 is further configured to: input each sub-data to be processed into the shared self-attention module to obtain initial molecular feature information corresponding to each sub-data to be processed; input each initial molecular feature information into the type mixing module to obtain type mixing feature information corresponding to each initial molecular feature information; and generate encoding feature information of the molecule to be processed based on each type mixing feature information. Optionally, the type mixing module includes at least one type feature sub-module; the processing module 804 is further configured to: determine the data type information corresponding to each initial molecular feature information and the type feature sub-module corresponding to each data type information; input each initial molecular feature information into the corresponding type feature sub-module to obtain type mixing feature information corresponding to each initial molecular feature information. Optionally, the at least one type feature sub-module includes a small molecule sub-module and a protein sub-module; and the sub-data to be processed include small molecule data to be processed, protein data to be processed, or complex data to be processed.Optionally, the apparatus further includes a training module configured to: obtain a first training sample, wherein the first training sample includes first sample molecular data and sample molecular domain information corresponding to the first sample molecular data; train a molecular processing pre-training model based on the first sample molecular data and the sample molecular domain information corresponding to the first sample molecular data, wherein the molecular processing pre-training model includes a classification layer, an embedding layer, and an encoding layer; obtain a second training sample, wherein the second training sample includes second sample molecular data and sample label data corresponding to the second sample molecular data; train a molecular processing model based on the second sample molecular data and the sample label data corresponding to the second sample molecular data, wherein the molecular processing model includes a molecular processing pre-training model and a downstream task layer. Optionally, the molecular processing pre-training model includes a classification layer, an embedding layer, and an encoding layer; the training module is further configured to: input the first sample molecule data and the sample molecule domain information into the classification layer, obtain at least one sample sub-data determined by the classification layer according to the sample molecule domain information and sample data type information corresponding to each sample sub-data; input each sample sub-data and the sample data type information corresponding to each sample sub-data into the embedding layer, obtain sample sub-data embedding feature information corresponding to each sample sub-data; input each sample sub-data embedding feature information into the encoding layer, obtain sample molecule encoding feature information output by the encoding layer, wherein the sample molecule encoding feature information is obtained according to each sample data type information; input the sample molecule encoding feature information into the output layer, obtain a molecular preprocessing result output by the output layer; perform multi-task learning of continuous atomic coordinates and classified atomic types based on the molecular preprocessing result to obtain a first model loss value; and adjust model parameters of the molecular processing pre-training model based on the first model loss value until a model training stop condition is met. Optionally, the training module is further configured to: perform embedding processing on the sample sub-data and the sample data type information to obtain sample sub-data feature information and sample data type embedding information; determine a structural position embedding strategy for the sample sub-data according to a preset structural control strategy, and determine structural position embedding information corresponding to the sample sub-data according to the structural embedding strategy, wherein the structural position embedding strategy includes a 2D strategy or a 3D strategy; and generate sample sub-data embedding feature information corresponding to the sample sub-data according to the sample sub-data feature information, the sample data type embedding information, and the structural position embedding information.Optionally, the molecular processing model includes a molecular processing pre-training model and a downstream task layer; the training module is further configured to: input the second sample molecular data into the molecular processing pre-training model to obtain molecular encoding feature information output by the molecular processing pre-training model; input the molecular encoding feature information into the downstream task layer to obtain predicted label data output by the downstream task layer; calculate a second model loss value based on the predicted label data and the sample label data; adjust model parameters of the molecular processing pre-training model and the model parameters of the downstream task layer based on the second model loss value, and continue training the molecular processing model until a model training stop condition is met. The apparatus provided by the embodiments of the present disclosure provides a molecular processing model in which molecular data to be processed can be split based on molecular domain information of the molecular data to be processed to obtain at least one sub-data to be processed and data type information corresponding to each sub-data to be processed. The encoding layer of the feature extraction module of the molecular processing model includes a shared self-attention module and a type mixing module. The shared self-attention module is used to promote data alignment between different types of data and learn similar information between different data types. The type mixing module includes multiple type feature submodules. It determines the corresponding type feature submodule based on the data type information corresponding to the different sub-data to be processed, thereby capturing the specificity and relationships between multiple data types, ultimately generating the encoding feature information of the molecules to be processed. Finally, through the downstream task layer, the molecular processing results of the molecular processing task are output. The molecular processing model provided by the disclosed embodiments provides a universal molecular processing model that can encode molecules across a wide range of biochemical fields, including small molecules, proteins, protein-small molecule complexes, and so on. This model can be processed within the same molecular processing model across multiple data formats, eliminating the need to train separate models based on molecule type and improving the model reuse rate of the molecular processing model. This allows the molecular processing model to handle different molecular processing tasks based on different downstream tasks, improving model utilization efficiency. Furthermore, cross-domain molecular representation learning can also improve molecular representation performance within a single domain. The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the aforementioned data processing method share the same concept. For details not described in detail in the technical solution of the data processing device, reference can be made to the description of the technical solution of the aforementioned data processing method. FIG. 9 shows a block diagram of a computing device 900 according to an embodiment of the present disclosure. Components of computing device 900 include, but are not limited to, a memory 910 and a processor 920.oThe processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data. The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface. In one embodiment of the present disclosure, the aforementioned components of the computing device 900 and other components not shown in FIG. 9 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG. 9 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed. Computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). Computing device 900 may also be a mobile or stationary server.The processor 920 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned data processing method or molecular processing model training method. The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the aforementioned data processing method or molecular processing model training method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned data processing method or molecular processing model training method. An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned data processing method or molecular processing model training method. The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the aforementioned data processing method or molecular processing model training method are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned data processing method or molecular processing model training method. One embodiment of the present disclosure further provides a computer program, wherein, when executed on a computer, the computer executes the steps of the aforementioned data processing method or molecular processing model training method. The above is an illustrative embodiment of a computer program according to this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the aforementioned data processing method or molecular processing model training method are based on the same concept. For details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the aforementioned data processing method or molecular processing model training method. The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous. The computer instructions comprise computer program code, which may be in source code form, object code form, executable file, or some intermediate form.The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. It should be noted that, for ease of description, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present disclosure are not limited by the order of the actions described, as, depending on the embodiments of the present disclosure, certain steps may be performed in a different order or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily required for the embodiments of the present disclosure. In the above embodiments, the description of each embodiment has its own emphasis. For portions not described in detail in a particular embodiment, reference should be made to the relevant description of the other embodiments. The preferred embodiments of the present disclosure disclosed above are merely provided to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to the specific implementations described. Obviously, many modifications and variations are possible based on the content of the embodiments of this disclosure. This disclosure selects and describes these embodiments in detail to better explain the principles and practical applications of the embodiments of this disclosure, thereby enabling those skilled in the art to better understand and utilize this disclosure. This disclosure is limited only by the claims and their full scope and equivalents.

Claims

Claims 1. A data processing method, comprising: Receive a molecular processing task, where the molecular processing task includes molecular data to be processed and molecular domain information corresponding to the molecular data to be processed; input the molecular data to be processed and the molecular domain information into a molecular processing model, and obtain a molecular processing result output by the molecular processing model. The molecular processing model determines at least one sub-data to be processed corresponding to the molecular data to be processed and data type information of each sub-data to be processed according to the molecular domain information, determines feature information of each sub-data to be processed according to each sub-data to be processed and the data type information of each sub-data to be processed, and generates a molecular processing result according to the feature information of each sub-data to be processed.

2. The method according to claim 1, wherein the molecular processing model comprises a classification layer, an embedding layer, an encoding layer, and a downstream task layer, and the molecular data to be processed comprises at least one sub-data to be processed; inputting the molecular data to be processed and the molecular domain information into the molecular processing model to obtain a molecular processing result output by the molecular processing model, including: Input the molecular data to be processed and the molecular domain information into the classification layer, and obtain at least one sub-data to be processed determined by the classification layer according to the molecular domain information and data type information corresponding to each sub-data to be processed; Input each sub-data to be processed and the data type information corresponding to each sub-data to be processed into the embedding layer, and obtain embedding feature information of each sub-data to be processed corresponding to each sub-data to be processed output by the embedding layer according to each sub-data to be processed and the data type information corresponding to each sub-data to be processed; Input the embedding feature information of each sub-data to be processed into the encoding layer, and obtain encoded feature information of the molecule to be processed output by the encoding layer, where the encoded feature information of the molecule to be processed is obtained according to each data type information; input the encoded feature information of the molecule to be processed into the downstream task layer, and obtain a molecular processing result output by the downstream task layer.

3. The method according to claim 2, wherein the encoding layer comprises a shared self-attention module and a type mixing module; embedding the feature information of each sub-data to be processed into the encoding layer to obtain the molecular encoding feature information to be processed output by the encoding layer, including: Input each sub-data to be processed into the shared self-attention module, and obtain initial molecular feature information corresponding to each sub-data to be processed; Input the initial molecular feature information into the type mixing module, and obtain type mixing feature information corresponding to the initial molecular feature information; generate encoded feature information of the molecule to be processed according to the type mixing feature information.

4. The method according to claim 3, wherein the type mixing module comprises at least one type feature sub-module; inputting the initial molecular feature information into the type mixing module to obtain type mixing feature information corresponding to the initial molecular feature information, including: Determine the data type information corresponding to each initial molecular feature information and the type feature sub-module corresponding to each data type information; Input the initial molecular feature information into the corresponding type feature sub-module, and obtain type mixing feature information corresponding to the initial molecular feature information.

5. The method according to claim 4, wherein the type feature sub-module includes a small molecule sub-module and a protein sub-module; the sub-data to be processed includes small molecule data to be processed, protein data to be processed or complex data to be processed.

6. The method according to any one of claims 1 to 5, wherein the molecular processing model is obtained by training through the following steps: obtaining a first training sample, wherein, The first training sample includes first sample molecular data and sample molecular domain information corresponding to the first sample molecular data; Train a molecular processing pre-training model according to the first sample molecular data and the sample molecular domain information corresponding to the first sample molecular data. The molecular processing pre-training model includes a classification layer, an embedding layer, and an encoding layer. Obtain a second training sample, where the second training sample includes second sample molecular data and sample label data corresponding to the second sample molecular data. Train a molecular processing model according to the second sample molecular data and the sample label data corresponding to the second sample molecular data. The molecular processing model includes a molecular processing pre-training model and a downstream task layer.

7. The method according to claim 6, training a molecular processing pre-trained model based on the first sample molecular data and the sample molecular domain information corresponding to the first sample molecular data, including: Input the first sample molecular data and the sample molecular domain information into the classification layer, and obtain at least one sample sub-data determined by the classification layer according to the sample molecular domain information and sample data type information corresponding to each sample sub-data. Input each sample sub-data and the sample data type information corresponding to each sample sub-data into the embedding layer, and obtain sample sub-data embedding feature information corresponding to each sample sub-data. Input the sample sub-data embedding feature information into the encoding layer, and obtain sample molecular encoding feature information output by the encoding layer, where the sample molecular encoding feature information is obtained according to each sample data type information. Input the sample molecular encoding feature information into the output layer, and obtain a molecular preprocessing result output by the output layer. Perform multi-task learning of continuous atomic coordinates and classification atomic types based on the molecular preprocessing result, and obtain a first model loss value. Adjust the model parameters of the molecular processing pre-training model based on the first model loss value until the model training stop condition is reached.

8. The method according to claim 7, wherein inputting each sample sub-data and the sample data type information corresponding to each sample sub-data into the embedding layer to obtain the sample sub-data embedding feature information corresponding to each sample sub-data, includes: Perform embedding processing on the sample sub-data and the sample data type information to obtain sample sub-data feature information and sample data type embedding information. Determine a structural position embedding strategy for the sample sub-data according to a preset structural control strategy, and determine structural position embedding information corresponding to the sample sub-data according to the structural embedding strategy. The structural position embedding strategy includes a 2D strategy or a 3D strategy. Generate sample sub-data embedding feature information corresponding to the sample sub-data according to the sample sub-data feature information, sample data type embedding information, and the structural position embedding information.

9. The method according to any one of claims 6 to 8, training a molecular processing model based on the second sample molecular data and the sample label data corresponding to the second sample molecular data, comprising: Input the second sample molecular data into the molecular processing pre-training model, and obtain molecular encoding feature information output by the molecular processing pre-training model. Input the molecular encoding feature information into the downstream task layer, and obtain predicted label data output by the downstream task layer. Calculate a second model loss value according to the predicted label data and the sample label data. Adjust the model parameters of the molecular processing pre-training model and the model parameters of the downstream task layer according to the second model loss value, and continue to train the molecular processing model until the model training stop condition is reached.

10. A data processing method, comprising: Receive a small molecule property prediction task, where the small molecule property prediction task includes small molecule data to be processed and molecular domain information corresponding to the small molecule data to be processed; input the small molecule data to be processed and the molecular domain information into a molecular property prediction model to obtain a molecular property prediction result output by the molecular property prediction model, where the molecular property prediction model determines data type information corresponding to the small molecule data to be processed according to the molecular domain information, determines sub-data feature information to be processed according to the data type information and the small molecule data to be processed, and generates a molecular property prediction result according to the sub-data features to be processed. Receive a virtual screening task, where the virtual screening task includes complex data to be processed and molecular domain information corresponding to the complex data to be processed; input the complex data to be processed and the molecular domain information into a molecular virtual screening model to obtain a molecular virtual screening result output by the molecular virtual screening model, where the molecular virtual screening model determines small molecule data to be processed and protein data to be processed corresponding to the complex data to be processed according to the molecular domain information, determines sub-data feature information to be processed according to the small molecule data to be processed and the protein data to be processed, and generates a molecular virtual screening result according to the sub-feature data to be processed.

11. A data processing method, comprising: Obtain a first training sample, where the first training sample includes first sample molecule data and sample molecular domain information corresponding to the first sample molecule data; train a molecular processing pre-training model according to the first sample molecule data and the sample molecular domain information corresponding to the first sample molecule data, where the molecular processing pre-training model includes a classification layer, an embedding layer, and an encoding layer; obtain a second training sample, where the second training sample includes second sample molecule data and sample label data corresponding to the second sample molecule data; train a molecular processing model according to the second sample molecule data and the sample label data corresponding to the second sample molecule data to obtain model parameters of the molecular processing model, where the molecular processing model includes a molecular processing pre-training model and a downstream task layer; send the model parameters of the molecular processing model to an edge device.

12. A training method for a molecular processing model, applied to a cloud-side device, includes: A memory and a processor; 13. A computing device, comprising: The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented. A computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented. A computer program, when executed on a computer, causes the computer to execute the steps of the method according to any one of claims 1 to 12. ​ 19

Citation Information

Patent Citations

  • Method for predicting drug target associativity based on combined cross-domain attention model

    CN116646001A