Methods, apparatus, devices, and readable storage media for molecular characterization

CN122822136APending Publication Date: 2026-09-25BEIJING DP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610931704.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25

Smart Images

  • Figure CN122822136A_ABST
    Figure CN122822136A_ABST
Patent Text Reader

Abstract

According to an implementation of the present disclosure, a method for molecular characterization is provided. According to the method, first, atomic composition of a target molecule is determined. Second, based on the atomic composition, a structural feature representation corresponding to the atomic composition is determined, the structural feature representation indicating element abundance and atomic characterization distribution corresponding to the atomic composition. Third, based on the atomic composition and the structural feature representation, a chemical feature representation corresponding to the atomic composition is determined. Finally, based on the structural feature representation and the chemical feature representation, a characterization result of the target molecule is determined. The above method can make the molecular characterization result not only reflect the structural information of the molecule, but also reflect the chemical function information of the molecule, thereby enriching the information dimension contained in the molecular characterization and improving the description ability of the molecular characterization result to the characteristics of the molecule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and more particularly to methods, apparatus, devices, and readable storage media for molecular characterization. Background Technology

[0002] With the popularization of artificial intelligence, performing molecular analysis based on AI has gradually become a trend. Typically, molecular analysis involves constructing molecular representations based on relevant molecular information and then using these representations to perform analytical tasks such as prediction, classification, or retrieval. As the dimensions and information sources of molecular representations increase, how to effectively integrate different types of feature information to enhance molecular representation capabilities has become a pressing technical problem to be solved. Summary of the Invention

[0003] In a first aspect, according to an implementation of this disclosure, a method for molecular characterization is proposed, which may include the following steps: First, determining the atomic composition of a target molecule. Second, based on the atomic composition, determining a structural feature representation corresponding to the atomic composition, wherein the structural feature representation indicates the elemental abundance and atomic characterization distribution corresponding to the atomic composition. Third, based on the atomic composition and the structural feature representation, determining a chemical feature representation corresponding to the atomic composition. Finally, based on the structural feature representation and the chemical feature representation, determining the characterization result of the target molecule.

[0004] In this way, structural and chemical feature representations of the target molecule can be constructed separately, and the characterization results of the target molecule can be determined based on the structural and chemical feature representations. This allows the obtained molecular characterization results to reflect not only the structural information of the molecule but also its chemical functional information, thereby enriching the information dimensions contained in the molecular characterization, improving the ability of the molecular characterization results to describe molecular properties, and providing a more comprehensive feature basis for subsequent molecular prediction, analysis, and retrieval applications.

[0005] In a second aspect, according to an implementation of this disclosure, a molecular characterization apparatus is proposed. This apparatus may include: a target molecule analysis module for determining the atomic composition of a target molecule; a structural feature representation determination module for determining a structural feature representation corresponding to the atomic composition, wherein the structural feature representation indicates the elemental abundance and atomic characterization distribution corresponding to the atomic composition; a chemical feature representation determination module for determining a chemical feature representation corresponding to the atomic composition based on the atomic composition and the structural feature representation; and a molecular characterization module for determining the characterization result of the target molecule based on the structural feature representation and the chemical feature representation.

[0006] In a third aspect, according to an implementation of this disclosure, an electronic device is proposed. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.

[0007] In a fourth aspect, according to an implementation of this disclosure, a computer-readable storage medium is proposed. A computer program is stored on the medium, which, when executed by a processor, implements the method of the first aspect.

[0008] In a fifth aspect, according to an implementation of this disclosure, a computer program product is proposed. This computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.

[0009] It should be understood that the description in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] Figure 1 A block diagram of an example environment in which molecular characterization of the present disclosure can be carried out is shown;

[0011] Figure 2 A schematic diagram of some implementations of the molecular characterization process according to this disclosure is shown;

[0012] Figure 3 The diagram illustrates some application scenarios of molecular characterization implemented according to this disclosure;

[0013] Figure 4 The following diagram illustrates the application of molecular characterization according to some implementations of this disclosure, as well as a flowchart of post-training of chemical codebooks;

[0014] Figure 5 A schematic structural block diagram of a molecular characterization apparatus according to some implementations of this disclosure is shown; and

[0015] Figure 6 A block diagram of an electronic device that can implement one or more embodiments of the present disclosure is shown. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0018] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0019] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0020] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization from relevant users should be obtained. Among them, relevant users can include any type of rights holder, such as individuals, enterprises, and groups.

[0021] As an example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information. This allows the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of the technical solution disclosed herein, based on the prompt message.

[0022] As an optional but non-restrictive implementation, a method of sending a prompt message to a relevant user in response to a proactive request can be exemplified by a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.

[0023] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0024] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0025] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, environment 100 may include electronic device 110. In this example environment 100, electronic device 110 can acquire relevant information about target molecule 102. As an example, the relevant information about target molecule 102 may include a three-dimensional molecular diagram of target molecule 102, or descriptive information indicating the three-dimensional structure of target molecule 102. Taking the three-dimensional molecular diagram of target molecule 102 as an example, the three-dimensional molecular diagram may include atomic nodes, chemical bond edges, atomic features, bond features, and atomic three-dimensional coordinates. Based on the relevant information about target molecule 102, electronic device 110 can generate a characterization result 104 of target molecule 102 using target model 115. Figure 1 The diagram only shows one target model 115 as an example; in practice, multiple different target models 115 may be used in collaboration to complete the characterization of the target molecule 102.

[0026] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 may also support any type of user-facing interface (such as "wearable" circuitry). Server-side equipment (not shown) can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc. As an example, server-side equipment can provide backend services for the applications of electronic device 110.

[0027] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0028] Utilizing artificial intelligence to perform molecular analysis requires defining molecular characterization. The goal of molecular characterization is to transform information such as atomic types, chemical bonds, three-dimensional coordinates, and conformational relationships into feature representations that can be processed by machine learning models. Conventional approaches typically fall into three categories: pre-training methods based on one-dimensional linear molecular representations, graph neural network methods based on two-dimensional molecular diagrams, and structural pre-training methods based on three-dimensional molecular conformations. Among these, three-dimensional structural pre-training methods, which utilize spatial information such as atomic coordinates, bond lengths, bond angles, and torsion angles, have achieved good results in molecular property prediction tasks in recent years.

[0029] On the other hand, the advantage of three-dimensional structure pre-training methods is that they can capture rich geometric structure information, but their training objectives mainly revolve around three-dimensional structures, and they pay insufficient attention to chemical functional information such as pharmacophores, functional groups, hydrogen bond donors and acceptors, acid and base groups, aromatic rings and metal binding sites.

[0030] Furthermore, when introducing new modalities, common practices include full fine-tuning of the entire pre-trained large model, or directly concatenating, adding, or fusing continuous features extracted from different modalities. Full fine-tuning consumes a large amount of GPU memory and computational resources, and downstream biomedical tasks often only have hundreds to thousands of samples, making it difficult to stably train large-scale models. Simple continuous feature fusion is prone to the problem of one modality signal being too strong while another modality is overwhelmed, resulting in structural information and chemical functional information not being effectively complementary. Some methods use vector quantization to map continuous vectors to several codewords in a discrete codebook, thereby obtaining compressed discrete representations. This type of method can reduce storage and computational overhead, but it is prone to codebook collapse during training, i.e., a large number of inputs use only a few codewords, leading to a decrease in representational ability. At the same time, existing vector quantization methods usually do not have a specific design for the correspondence between molecular elemental abundance, atomic-level structural representation, and pharmacophore functional information.

[0031] In embodiments of this disclosure, a method for molecular characterization is proposed. In this method, an electronic device determines the atomic composition of a target molecule. Based on the atomic composition, a structural feature representation corresponding to the atomic composition is determined, the structural feature representation indicating the elemental abundance and atomic characterization distribution corresponding to the atomic composition. Based on the atomic composition and the structural feature representation, a chemical feature representation corresponding to the atomic composition is determined. Based on the structural feature representation and the chemical feature representation, the characterization result of the target molecule is determined.

[0032] Through the above process, structural and chemical feature representations of the target molecule can be constructed separately, and the characterization results of the target molecule can be determined based on the structural and chemical feature representations. This allows the obtained molecular characterization results to reflect not only the structural information of the molecule but also its chemical functional information, thereby enriching the information dimensions contained in the molecular characterization, improving the ability of the molecular characterization results to describe molecular properties, and providing a more comprehensive feature basis for subsequent molecular prediction, analysis, and retrieval applications.

[0033] Figure 2 An example flow diagram of a method 200 for molecular characterization according to some embodiments of the present disclosure is shown. For ease of discussion, reference will be made to... Figure 1 The process 200 is described in the context of environment 100. In environment 100, molecular characterization can be performed by electronic device 110, but some of these operations can be performed by requesting a server device (not shown) (e.g., determining the continuous structural feature representation of atoms, calling the structure codebook, calling the chemical codebook, or the training process of the chemical codebook can be implemented at the server device).

[0034] In box 201, electronic device 110 can first determine the atomic composition of target molecule 102.

[0035] The electronic device 110 can use the acquisition of relevant information about the target molecule 102 as a trigger condition. The relevant information about the target molecule 102 may include a three-dimensional molecular diagram of the target molecule 102, or descriptive information indicating the three-dimensional structure of the target molecule 102. Taking a three-dimensional molecular diagram of the target molecule 102 as an example, the three-dimensional molecular diagram may include information such as atomic nodes, chemical bond edges, atomic features, bond features, and atomic three-dimensional coordinates.

[0036] After acquiring relevant information about the target molecule 102, the electronic device 110 can analyze the relevant information to determine the atomic composition of the target molecule 102. The atomic composition can be used to characterize each atom contained in the target molecule 102 and its corresponding element type. In some embodiments, the electronic device 110 can identify the types of atoms such as carbon, nitrogen, oxygen, sulfur, and halogen atoms contained in the target molecule 102 based on the atomic node information in the three-dimensional molecular diagram, and determine the connection relationships and spatial coordinates between the atoms, thereby obtaining the atomic composition of the target molecule 102.

[0037] It should be understood that this application does not limit the specific method for obtaining the atomic composition. In addition to extracting directly from a three-dimensional molecular diagram, the atomic composition of the target molecule 102 can also be determined based on molecular structure files, molecular descriptors, chemical structural expressions, or other information that can characterize the molecular structure.

[0038] In box 202, electronic device 110 can determine a structural feature representation corresponding to the atomic composition based on the atomic composition, the structural feature representation indicating the elemental abundance and atomic characterization distribution corresponding to the atomic composition.

[0039] In some embodiments, the electronic device 110 can encode the structural information of the target molecule 102 using an encoder based on the atomic nodes, chemical bond edges, atomic features, bond features, and three-dimensional coordinates of each atom contained in the target molecule 102, in order to generate a corresponding structural feature representation. This structural feature representation can be used to characterize the structural properties of the target molecule 102 and reflect the composition of different elements and the distribution characteristics of various types of atoms in the target molecule 102.

[0040] As an example, structural feature representation can indicate the relative distribution of different elements such as carbon, nitrogen, oxygen, sulfur, and halogen atoms in target molecule 102, thus reflecting the elemental abundance of target molecule 102. Simultaneously, structural feature representation can also characterize the distribution patterns of different atoms in the structural space and the relationships between atoms, thereby reflecting the atomic characterization distribution of target molecule 102.

[0041] In box 203, electronic device 110 can determine the chemical feature representation corresponding to the atomic composition based on the atomic composition and structural feature representation.

[0042] In some embodiments, the electronic device 110 can combine the atomic composition of the target molecule 102 with the structural properties reflected in the structural feature representation to extract and characterize the chemical functional information of the target molecule 102, thereby generating a corresponding chemical feature representation. As an example, the chemical feature representation can be used to characterize the chemical properties, functional attributes, or potential action features corresponding to the pharmacophore in the target molecule 102.

[0043] As an example, chemical characterization can reflect whether the atoms in target molecule 102 possess pharmacophore-related properties such as hydrogen bond donor / acceptor characteristics, acid-base properties, aromaticity, ring structure characteristics, or metal-binding properties. Furthermore, chemical characterization can also reflect the chemical action characteristics corresponding to the functional groups formed by multiple atoms, thereby characterizing the potential functions of target molecule 102 in areas such as molecular recognition, molecular binding, or biological activity.

[0044] In box 204, electronic device 110 can determine the characterization results of target molecule 102 based on structural feature representation and chemical feature representation.

[0045] In some embodiments, the electronic device 110 can fuse structural feature representations and chemical feature representations to generate corresponding molecular characterization results. The structural feature representation is primarily used to characterize the structural properties of the target molecule 102, while the chemical feature representation is primarily used to characterize the chemical functional properties of the target molecule 102. By jointly characterizing both, the final molecular characterization result can simultaneously include information at both the structural and chemical functional levels.

[0046] As an example, the electronic device 110 can fuse structural feature representations and chemical feature representations through feature representation mapping, feature representation combination, feature representation aggregation, or other methods to obtain the characterization results of the target molecule 102. This disclosure does not limit the specific fusion method of structural feature representations and chemical feature representations, as long as it can generate characterization results for the target molecule 102 based on the structural feature representations and chemical feature representations.

[0047] Through the above process, structural and chemical feature representations of the target molecule 102 can be constructed separately, and the characterization results of the target molecule 102 can be determined based on these representations. This allows the obtained molecular characterization results to reflect not only the structural information of the molecule but also its chemical functional information, thereby enriching the information dimensions contained in the molecular characterization, improving the ability of the molecular characterization results to describe molecular properties, and providing a more comprehensive feature foundation for subsequent molecular prediction, analysis, and retrieval applications. As an example, after obtaining the characterization results of the target molecule 102, the electronic device 110 can perform corresponding molecular analysis tasks based on the characterization results. For example, the characterization results can be used to perform molecular property prediction, molecular function analysis, molecular similarity retrieval, drug screening, biomedical-related prediction tasks, or other processing tasks related to molecular characterization.

[0048] In some embodiments of this disclosure, the electronic device 110 can determine the structural feature representation corresponding to the atomic composition in the following manner. As an example, the electronic device 110 can acquire a pre-constructed structure codebook, which includes structure content items, each indicating at least a discrete structural feature representation of an atom. For each atom in the atomic composition, three-dimensional composition information corresponding to each atom is determined. A discrete structural feature representation matching the three-dimensional composition information of each atom is determined from the structure codebook as the structural feature representation of each atom, wherein the discrete structural feature representation of each atom is part of the structural feature representation corresponding to the atomic composition.

[0049] The structural content items in the structural codebook contain multiple discrete structural feature representations. Each discrete structural feature representation can be obtained by performing clustering processing on the continuous structural feature representations of each molecular sample in the molecular sample set.

[0050] The structure codebook can be pre-built; the construction process of the structure codebook is described below. As an example, electronic device 110 can acquire a molecular sample set and use a molecular coding model to encode the molecular samples in the set, obtaining a continuous structural feature representation corresponding to the molecular samples. In some embodiments, the molecular coding model can be a three-dimensional molecular encoder, used to generate the corresponding continuous structural feature representation based on the atomic information, chemical bond information, and spatial structure information in the molecular samples.

[0051] After obtaining the continuous structural feature representations corresponding to the molecular sample set, the electronic device 110 can perform clustering processing on the continuous structural feature representations to obtain multiple cluster centers, and use these multiple cluster centers as multiple discrete structural feature representations in the structural codebook. In other words, the discrete structural feature representations in the structural codebook can be composed of cluster centers formed by clustering continuous structural feature representations, thereby establishing a mapping relationship between the continuous structural space and the discrete structural space.

[0052] In some embodiments, the molecular sample set may include a large number of molecular samples. For example, millions of molecular samples can be used to generate corresponding continuous structural feature representations, and a structural codebook can be constructed based on these continuous structural feature representations. It is readily understood that this disclosure does not limit the specific size of the molecular sample set or the specific number of discrete structural feature representations in the structural codebook.

[0053] By constructing a structure codebook using clustering, the discrete structural feature representations in the codebook can be naturally adapted to the elemental abundance and atomic representation distribution in the molecular sample set. Since different elements appear at different frequencies in the molecular sample set, the structural patterns corresponding to different elements can form different numbers of discrete structural feature representations based on the actual data distribution, thereby improving the structure codebook's ability to represent structural information.

[0054] In some embodiments, the structure codebook can be constructed by the electronic device 110 based on a molecular sample set. It should be understood that the structure codebook can also be pre-constructed by a third-party device, server, or other training platform and stored in a local or remote storage medium of the electronic device 110. The electronic device 110 can directly obtain the pre-constructed structure codebook when executing embodiments of this disclosure and use the structure codebook to determine the discrete structural feature representation of the target molecule 102.

[0055] In some embodiments, the electronic device 110 can acquire a pre-built structure codebook. The structure codebook includes structure content items, which indicate at least a plurality of discrete structure feature representations. Different discrete structure feature representations can be used to characterize different atomic structure patterns.

[0056] For each atom in the target molecule 102, the electronic device 110 can determine the corresponding three-dimensional composition information of the atom. The three-dimensional composition information may include the spatial coordinates of the atom, the spatial relationship between the atoms, bond lengths, bond angles, torsion angles, or other information that can characterize the spatial configuration of the atom.

[0057] For each atom in the atomic composition, after obtaining the three-dimensional composition information of the atom, the electronic device 110 can generate a corresponding continuous structural feature representation based on the three-dimensional composition information. The continuous structural feature representation can be used to characterize the spatial structural features of the atom and reflect the position of the atom in the continuous structural space.

[0058] Electronic device 110 can perform matching between a continuous structural feature representation and multiple discrete structural feature representations in a structural codebook to determine a target discrete structural feature representation corresponding to the continuous structural feature representation. In some embodiments, electronic device 110 can calculate the degree of matching between the continuous structural feature representation and each discrete structural feature representation separately, and select the discrete structural feature representation with the highest degree of matching as the target discrete structural feature representation. As an example, the degree of matching can be similarity, distance value, or other matching metrics.

[0059] The matching process can be represented as: i = argmin ||v - C_struct[i]||, where v can represent the continuous structural feature representation of atoms in the atomic composition, and C_struct[i] can represent the content item at the i-th position in the structure codebook. argmin ||v - C_struct[i]|| can be represented as traversing the content items in the structure codebook, selecting the content item with the highest matching degree to the continuous structural feature representation of atoms in the atomic composition, and recording the target index item i of that content item. Through the above process, the atomic representation in the continuous structure space can be mapped to the discrete structure space.

[0060] The electronic device 110 can use the determined discrete structural feature representation of the target as the structural feature representation of the corresponding atom. For multiple atoms in the target molecule 102, the corresponding discrete structural feature representation can be determined separately, and the multiple discrete structural feature representations can be used together as the structural feature representation of the target molecule 102.

[0061] By mapping the continuous structural feature representation of atoms to the discrete structural feature representation in the structural codebook, a large number of atomic structural patterns can be characterized using a finite number of discrete structural feature representations. This allows for the discretization of structural information while preserving structural information, and provides a foundation for determining chemical feature representations based on structural feature representations.

[0062] In some embodiments, the process by which the electronic device 110 determines the chemical feature representation corresponding to the atomic composition may include acquiring a pre-constructed chemical codebook, which includes chemical content items, each indicating at least the chemical feature representation of an atom; the chemical codebook and the structure codebook are constructed to have the same index space, and the chemical codebook and the structure codebook share index items. Target index items corresponding to the discrete structural feature representations that match the three-dimensional composition information of each atom are acquired. Based on the target index items, chemical content items corresponding to each atom are determined in the chemical codebook as the chemical feature representation of each atom.

[0063] In some embodiments, the chemical codebook can be cascaded after the structure codebook. The chemical codebook and the structure codebook can be constructed to have the same index space, and they share index entries. In other words, a correspondence can be established between the structural content items (discrete feature representations) in the structure codebook and the chemical content items (chemical feature representations) in the chemical codebook through the same index. Furthermore, the chemical content items in the chemical codebook can have the same data structure as the structural content items in the structure codebook.

[0064] In some embodiments, the electronic device 110 may first determine the corresponding continuous structural feature representation based on the three-dimensional composition information of the target atom, and then perform a matching process between the continuous structural feature representation and multiple discrete structural feature representations in the structural codebook to determine the corresponding target discrete structural feature representation (target structural content item), and record the target index item i of the target discrete structural feature representation. That is, the target index item i of the target discrete structural feature representation can be used to characterize the position of the target atom in the discrete structural space.

[0065] After determining the target index term, the electronic device 110 can use the target index term i to access the chemical codebook and read the chemical content term corresponding to the target index term in the chemical codebook as a chemical feature representation of the target atom. In other words, the electronic device 110 can directly access the corresponding chemical content term in the chemical codebook based on the structure matching result determined by the structure codebook, without having to perform independent matching processing on the chemical features again.

[0066] In some embodiments, chemical feature representation may primarily employ pharmacophore-related properties. For example, chemical feature representation may be used to characterize whether a target atom possesses hydrogen bond acceptor properties, hydrogen bond donor properties, acidic properties, basic properties, aromatic properties, ring structure properties, or metal-binding properties. For different atoms, multidimensional labels, vector representations, or other methods may be used to characterize the corresponding pharmacophore-related properties.

[0067] By cascading and reusing the target index entries of the structure codebook and chemical codebook, the corresponding chemical feature representation can be directly determined using the structure matching results in the structure codebook. This reduces the computational overhead of establishing mapping relationships for different modes and improves the correlation between information from different modes and the ability of molecular characterization results to jointly characterize structural and chemical functional information.

[0068] Since the chemical codebook is accessed through the target index entries corresponding to the structural codebook, different chemical content entries in the chemical codebook can progressively learn the chemical functional information corresponding to different structural modes. By sharing an index space between the structural and chemical codebooks, a stable correspondence can be established between structural and chemical modes in a unified discrete space, thereby achieving alignment between structural information and pharmacophore information.

[0069] The construction process of the chemical codebook is described below. Electronic device 110 can construct an initialized chemical codebook with the same index space as the structure codebook. The initialized chemical codebook contains multiple initialized chemical content items, the number of which is the same as the number of structure content items in the structure codebook. Based on the atomic composition sample, the structure content item corresponding to the atomic composition sample is determined. Based on the atomic composition sample and the corresponding structure content item, the predicted chemical content item corresponding to the atomic composition sample is determined. Based on the difference between the predicted chemical content item and the standard chemical content item of the atomic composition sample, the multiple initialized chemical content items are adjusted until predetermined conditions are met to obtain the chemical feature representation of the atom. The difference between the predicted chemical content item and the standard chemical content item of the atomic composition sample includes at least one of the following: a first difference indicating the pharmacophore prediction result, a second difference indicating the degree of masked atom reconstruction, and a third difference indicating the degree of denoising of the atom's three-dimensional coordinates.

[0070] In some embodiments, the electronic device 110 can construct an initialized chemical codebook with the same index space as the structural codebook. The initialized chemical codebook may contain multiple initialized chemical content items, meaning that multiple initialized chemical content items can correspond to multiple initialized pharmacophore feature representations. The number of initialized chemical content items can be the same as the number of structural content items in the structural codebook. In other words, the chemical codebook and the structural codebook can have the same number of content items and the same index space, thereby enabling an index correspondence between the chemical content items in the chemical codebook and the structural content items in the structural codebook.

[0071] In some embodiments, multiple initialized chemical content items in the initialized chemical codebook can be generated by random initialization, predefined initialization, or other initialization methods, and are mainly used for subsequent learning of chemical functional information corresponding to the structural pattern.

[0072] After generating the initial chemical codebook, the electronic device 110 can acquire atomic composition samples (molecular samples) and determine the structural feature representation corresponding to each atomic sample in the atomic composition samples based on the atomic composition samples. In some embodiments, the electronic device 110 can use the structural codebook to perform discretization processing on the continuous structural feature representation corresponding to the atomic composition samples to obtain the corresponding discrete structural feature representation and the corresponding target index item.

[0073] After obtaining the structural feature representation corresponding to the atomic composition sample, the electronic device 110 can access the chemical codebook to be iteratively updated based on the target index term and obtain the (initialized) chemical content term corresponding to the target index term. Since the chemical codebook and the structural codebook share the same index space, the target index term corresponding to the discrete structural feature representation (structural content term) determined in the structural codebook can be directly used to access the chemical content term in the chemical codebook as a predicted chemical content term. The predicted chemical content term can indicate pharmacophore features, for example, whether the atomic sample belongs to a pharmacophore category such as hydrogen bond acceptor, hydrogen bond donor, acidic group, basic group, aromatic ring, ring structure, functional group, or metal binding site. For different atoms, multidimensional labels, vector representations, or other methods can be used to characterize the corresponding pharmacophore-related attributes.

[0074] Subsequently, the electronic device 110 can perform parameter adjustment processing on multiple initial chemical content items in the initial chemical codebook based on the difference between the predicted chemical content items and the standard chemical content items corresponding to the atomic composition sample, in order to update the chemical content items in the chemical codebook. The aforementioned parameter adjustment can stop after a predetermined condition is met. As an example, the predetermined condition could be the number of adjustments, the duration of the entire adjustment process, or the change between multiple consecutive predicted chemical content items being less than a threshold, etc.

[0075] In some embodiments, the difference between the predicted chemical content item and the standard chemical content item of the atomic composition sample may include at least one of a first difference indicating the pharmacophore prediction result, a second difference indicating the degree of masked atom recovery, and a third difference indicating the degree of denoising of the atom's three-dimensional coordinates. These three differences can be determined through a pharmacophore prediction task, a masked atom recovery task, and a coordinate denoising task. The pharmacophore prediction task can be used to predict the pharmacophore category of the target atom. The masked atom recovery task can be used to predict the atom type of the masked atom. The coordinate denoising task can be used to predict the true three-dimensional coordinates or coordinate offset of the atom.

[0076] Electronic device 110 can construct a joint training loss based on some or all of the differences in the first difference, second difference, and third difference, and perform parameter updates on multiple initial pharmacophore feature representations in the initial chemical codebook based on the joint training loss.

[0077] Taking all differences as an example, the joint training loss can be expressed as:

[0078] L_total = L_phar + L_atom + L_coord

[0079] Here, L_phar can represent the loss corresponding to pharmacophore prediction, L_atom can represent the loss corresponding to mask atom recovery, and L_coord can represent the loss corresponding to coordinate denoising.

[0080] Electronic device 110 can perform backpropagation processing based on joint training loss to adjust relevant parameters in the initialized chemical codebook, thereby reducing the joint training loss. As the training process continues, the chemical content items in the chemical codebook can gradually learn chemical functional information corresponding to different structural patterns. The above training process uses three types of tasks—pharmacophore prediction, masked atom recovery, and coordinate denoising—as examples. In addition, tasks such as bond type prediction, interatomic distance prediction, fragment reconstruction, conformational energy prediction, molecular property-assisted prediction, or contrastive learning can also be used as differences.

[0081] In some embodiments, the structure codebook can be kept in a frozen or weakly updated state during the post-multi-task training phase, thereby allowing the structure codebook to continue to retain the structural modal knowledge learned during the pre-training phase. The chemical codebook, on the other hand, is continuously updated based on joint training loss to learn supplementary pharmacophore functional information corresponding to different structural modalities.

[0082] In some embodiments, the chemical codebook can be constructed by the electronic device 110 based on a set of atomic composition samples (molecular samples). It should be understood that the chemical codebook can also be pre-constructed by a third-party device, server, or other training platform and stored in a local or remote storage medium of the electronic device 110. The electronic device 110 can directly obtain the pre-constructed chemical codebook when executing embodiments of this disclosure and use the chemical codebook to determine the chemical characteristic representation of the target molecule 102.

[0083] By adopting a shared index space for structural and chemical codebooks, and combining a joint training approach involving pharmacophore prediction, masked atom recovery, and coordinate denoising tasks, the model's ability to represent chemical functional information can be enhanced while preserving structural knowledge. This improves the joint representation capability between structural and chemical functional information, as well as the transferability of molecular representation results to downstream tasks.

[0084] In some embodiments, the electronic device 110 determines the characterization result of the target molecule 102 by performing a first mapping process on the structural feature representation and a second mapping process on the chemical feature representation, so as to map the structural feature representation after the first mapping process and the chemical feature representation after the second mapping process to the same feature space. The mapped structural feature representation and the mapped chemical feature representation are then fused to determine the characterization result of the target molecule 102.

[0085] In some embodiments, the first mapping process can be used to perform feature representation projection processing on structural feature representations, and the second mapping process can be used to perform feature representation projection processing on chemical feature representations. The ascent process can map different feature representations to a feature space of a unified dimension, thereby improving the compatibility between different modal information.

[0086] As an example, the first mapping layer can correspond to the structure projection layer, and the second mapping layer can correspond to the chemical projection layer. The parameters in the first and second mapping layers can be updated in the post-training stage through joint training loss. The electronic device 110 can generate corresponding enhanced atom representations based on the structural feature representations (structural content items) in the structure codebook and the chemical feature representations (chemical content items) in the chemical codebook. As an example, if the target structural content item in the structure codebook can be represented as C_struct[i], and the target chemical content item in the chemical codebook can be represented as C_chem[i], then the electronic device 110 can determine the enhanced atom representations in the following way:

[0087] v_r = θ1(C_struct[i]) + θ2(C_chem[i]).

[0088] Where θ1 can represent the first mapping process for structural feature representation, and θ2 can represent the second mapping process for chemical feature representation.

[0089] In some embodiments, the structural feature representation in the structural codebook is primarily used to retain the three-dimensional structural knowledge learned during the pre-training phase, while the chemical feature representation in the chemical codebook can be primarily used to learn the chemical functional information associated with the corresponding structural patterns. By employing a first mapping process and a second mapping process, the structural feature representation and the chemical feature representation can be fused in a unified feature space, thereby reducing conflicts between different modal information. Compared to directly splicing continuous features, the target molecule 102 characterization result generated based on a unified index space and mapping fusion method in this embodiment can improve the correlation and stability between structural information and chemical functional information.

[0090] After determining the characterization results of the target molecule 102, the electronic device 110 can also perform at least one downstream task of molecular property prediction and molecular function retrieval based on the characterization results of the target molecule 102. Molecular property prediction includes at least one of biological interaction prediction, drug perturbation prediction, mass spectrometry feature simulation, and molecular physicochemical property prediction; molecular function retrieval includes at least one of drug screening and molecular functional similarity retrieval.

[0091] In some embodiments, the electronic device 110 can perform optimization processing on the characterization results of the target molecule 102. As an example, the optimization processing may include methods such as average pooling, summative pooling, virtual node aggregation, or attention aggregation to aggregate the characterization results of multiple atoms in the target molecule 102, obtaining a molecular characterization corresponding to the entire target molecule 102. Virtual node aggregation may refer to introducing virtual nodes connected to multiple atomic nodes in a graph neural network and aggregating information between multiple atomic nodes through these virtual nodes; attention aggregation may refer to assigning different weights to different atomic characterizations based on their importance before performing the aggregation process.

[0092] Combination Figure 3 As shown, after optimization, the electronic device 110 can perform at least one downstream task of molecular characteristic prediction task 302 and molecular function retrieval task 304 based on the optimized molecular characterization. Molecular characteristic prediction task 302 may include at least one of biological interaction prediction, drug perturbation prediction, mass spectrometry feature simulation, and molecular physicochemical property prediction; molecular function retrieval task 304 may include at least one of drug screening and molecular functional similarity retrieval.

[0093] As an example, in a biological interaction prediction task, electronic device 110 can splice, fuse, or jointly encode molecular characterization with enzyme protein characterization, target protein characterization, or receptor characterization, and input the fused features into a classification model or prediction model to predict whether there is an interaction relationship between the target molecule 102 and the corresponding biological object. In one embodiment, the biological interaction prediction task may include an enzyme-substrate interaction prediction task.

[0094] As an example, in a drug perturbation prediction task, electronic device 110 can combine molecular characterization with gene characterization, cell characterization, or biological state characterization, and predict changes in gene expression, cell state, or biological response after the action of target molecule 102 based on the combined features. In one embodiment, the drug perturbation prediction task may include a single-cell drug perturbation prediction task.

[0095] As an example, in a mass spectrometry feature simulation task, electronic device 110 can fuse molecular characterization with mass spectrometry experimental condition information, and predict the mass-to-charge ratio distribution, peak intensity distribution, or other mass spectrometry feature information corresponding to the target molecule 102 based on the fused features. In one embodiment, the mass spectrometry feature simulation task may include a molecular mass spectrometry simulation task.

[0096] As an example, in a molecular physicochemical property prediction task, electronic device 110 can predict the solubility, stability, polarity, toxicity, lipophilicity, or other physicochemical property parameters of target molecule 102 based on molecular characterization.

[0097] As an example, in a drug screening task, the electronic device 110 can screen out candidate drug molecules that meet preset conditions from the candidate molecule set based on the characterization similarity, biological activity prediction results, or interaction relationship prediction results between the target molecule 102 and the candidate molecules.

[0098] As an example, in a molecular functional similarity retrieval task, electronic device 110 can identify candidate molecules that are similar to target molecule 102 in terms of functional properties, biological activity, or chemical mechanism of action based on the characterization similarity between target molecule 102 and candidate molecules.

[0099] The characterization results of target molecule 102 in this embodiment simultaneously include structural modal information and chemical functional modal information. Therefore, compared to molecular characterizations generated solely based on three-dimensional structural information, it can improve the chemical function perception capability in downstream tasks. Furthermore, compared to performing full fine-tuning on a complete pre-trained model, the characterization results of target molecule 102 in this embodiment can be directly transferred to corresponding downstream tasks, thereby reducing the scale of training parameters and computational resource consumption. While the above examples provide some tasks, actual tasks are not limited to these. Examples can also include and be extended to other downstream tasks such as molecular toxicity prediction, drug-target interaction prediction, molecule generation, virtual screening, metabolic stability prediction, molecular similarity retrieval, and reaction product prediction.

[0100] Furthermore, compared to molecular graphical models or graph neural network models trained from scratch, the characterization results of target molecule 102 in this embodiment can simultaneously utilize the structural knowledge obtained during the large-scale pre-training phase and the pharmacophore functional knowledge obtained during the post-training phase. Therefore, this embodiment is particularly suitable for biomedical prediction scenarios with a small number of samples but high sensitivity to chemical functional information.

[0101] Figure 4A schematic diagram illustrating molecular characterization applications according to some embodiments of this disclosure is shown. The entire process can be implemented in an electronic device 110. The electronic device 110 acquires relevant information about a target molecule 102 and, using a molecular encoder 402, obtains a continuous structural characterization of each atom in the target molecule 102. The electronic device 110 discretizes the structure using a pre-constructed structure codebook. The structure codebook is obtained by clustering the atomic characterizations of a large number of molecular samples, with each codeword corresponding to a representative atomic structural pattern. For each atomic characterization, the nearest codeword (structural content item, i.e., discrete structural feature representation) is found in the structure codebook, and the index of that codeword is recorded. This index is used to represent the category of the atom in the discrete structural space and also serves as a bridge to access subsequent chemical codebooks. The chemical codebook has the same number of codewords and index space as the structure codebook, but the codewords (chemical content items) in the chemical codebook are trainable parameters. For an index selected from the structure codebook, the electronic device 110 can read the chemical content item at the same index position in the chemical codebook. The structural and chemical content items are then linearly mapped and summed to obtain an enhanced atomic representation that combines structural and pharmacophore functional information for downstream tasks. Furthermore, for the chemical content items in the chemical codebook, a multi-task post-training approach can be used to simultaneously maintain structure-awareness and acquire chemical function-awareness. The first type of task is the pharmacophore prediction task, which predicts whether each atom belongs to a single-atom hydrogen bond acceptor, single-atom hydrogen bond donor, acidic group, basic group, aromatic ring, ring structure, functional group, or metal binding site based on the enhanced atomic representation. This task uses a multi-label binary classification loss to enable the chemical codebook to learn atomic-level chemical function information.

[0102] To avoid focusing solely on pharmacophores and neglecting structural information during training, structure-related post-training tasks can be retained simultaneously. The second type of task is masked atom type recovery, which randomly masks some atom types and requires recovery of the masked atom types based on context. The third type of task is coordinate denoising, which introduces perturbation noise into the three-dimensional coordinates of atoms and then requires prediction or recovery of the true coordinates. This design constrains the model to continue focusing on the spatial conformational patterns of molecules. For training the chemical codebook, the total training loss can be expressed as the sum of pharmacophore prediction loss, masked atom recovery loss, and coordinate denoising loss. In one embodiment, post-training can be performed on approximately 13 million molecules with three-dimensional conformations, and pharmacophore labels can be generated by a chemical feature recognition tool. After training, the chemical content items can simultaneously include molecular representations of both structural and chemical functional modalities.

[0103] Figure 5A schematic structural block diagram of a molecular characterization apparatus 500 according to some embodiments of the present disclosure is shown. Apparatus 500, by way of example, may be implemented in or included in electronic device 110. Various modules / components in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0104] like Figure 5 As shown, the device 500 may include a target molecule analysis module 501 for determining the atomic composition of the target molecule; a structural feature representation determination module 502 for determining the structural feature representation corresponding to the atomic composition, wherein the structural feature representation indicates the elemental abundance and atomic characterization distribution corresponding to the atomic composition; a chemical feature representation determination module 503 for determining the chemical feature representation corresponding to the atomic composition based on the atomic composition and the structural feature representation; and a molecular characterization module 504 for determining the characterization result of the target molecule based on the structural feature representation and the chemical feature representation.

[0105] In some embodiments of this disclosure, the structural feature representation determination module 502 may be configured to acquire a pre-built structural codebook, which includes structural content items that at least indicate discrete structural feature representations of atoms. For each atom in the atomic composition, three-dimensional composition information corresponding to each atom is determined. A discrete structural feature representation matching the three-dimensional composition information of each atom is determined from the structural codebook as the structural feature representation of each atom, wherein the discrete structural feature representation of each atom is part of the structural feature representation corresponding to the atomic composition.

[0106] In some embodiments of this disclosure, the structural feature representation determination module 502 may also be configured to determine the continuous structural feature representation corresponding to the three-dimensional composition information of each atom. The continuous structural feature representation is compared with each structural content item in the structural codebook to obtain the structural content item that matches the continuous structural feature representation based on the degree of matching, and this structural content item is used as the structural feature representation of each atom.

[0107] In some embodiments of this disclosure, the structural content items in the structural codebook indicate multiple discrete structural feature representations, each of which is obtained by performing clustering on the continuous structural feature representations of each molecular sample in the molecular sample set.

[0108] In some embodiments of this disclosure, the chemical feature representation determination module 503 can be configured to acquire a pre-constructed chemical codebook, which includes chemical content items, each indicating at least the chemical feature representation of an atom. The chemical codebook and the structure codebook are constructed to have the same index space, and the chemical codebook and the structure codebook share index items. Target index items corresponding to the structure content items that match the three-dimensional composition information of each atom are acquired. Based on the target index items, the chemical content items corresponding to each atom are determined in the chemical codebook as the chemical feature representation of each atom.

[0109] In some embodiments of this disclosure, a chemical codebook construction module is also included. This module can be configured to construct an initial chemical codebook with the same index space as the structure codebook. The initial chemical codebook contains multiple initialized chemical content items, the number of which is the same as the number of structural content items in the structure codebook. Based on the atomic composition sample, structural content items corresponding to the atomic composition sample are determined. Based on the atomic composition sample and its corresponding structural content items, predicted chemical content items corresponding to the atomic composition sample are determined. Based on the differences between the predicted chemical content items and the standard chemical content items of the atomic composition sample, the multiple initialized chemical content items are adjusted until predetermined conditions are met to obtain a chemical feature representation of the atom.

[0110] In some embodiments of this disclosure, the difference between the predicted chemical content item and the standard chemical content item of the atomic composition sample includes at least one of a first difference indicating the pharmacophore prediction result, a second difference indicating the degree of mask atom restoration, and a third difference indicating the degree of denoising of the atom three-dimensional coordinates.

[0111] In some embodiments of this disclosure, the molecular characterization module 504 can be configured to perform a first mapping process on the structural feature representation and a second mapping process on the chemical feature representation, so as to map the structural feature representation after the first mapping process and the chemical feature representation after the second mapping process to the same feature space. The mapped structural feature representation and the mapped chemical feature representation are then fused to determine the characterization result of the target molecule.

[0112] In some embodiments of this disclosure, a target molecule analysis module is also included. This module can be configured to perform at least one downstream task, namely molecular property prediction and molecular function retrieval, based on the characterization results of the target molecule. Molecular property prediction includes at least one of biological interaction prediction, drug perturbation prediction, mass spectrometry feature simulation, and molecular physicochemical property prediction; molecular function retrieval includes at least one of drug screening and molecular functional similarity retrieval.

[0113] Figure 6A block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 6 The electronic device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 6 The illustrated electronic device 600 may include or be implemented as Figure 1 Electronic devices 110 or Figure 5 The device 500.

[0114] like Figure 6 As shown, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.

[0115] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.

[0116] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0117] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0118] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (such as a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0119] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0120] According to an exemplary implementation of this disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform... Figure 2 The methods provided are among the various optional methods available in the code, so they will not be elaborated upon here.

[0121] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0122] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0123] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. As an example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0125] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for molecular characterization, comprising: Determine the atomic composition of the target molecule; Based on the atomic composition, a structural feature representation corresponding to the atomic composition is determined, wherein the structural feature representation indicates the elemental abundance and atomic characterization distribution corresponding to the atomic composition; Based on the atomic composition and the structural feature representation, the chemical feature representation corresponding to the atomic composition is determined; as well as Based on the structural and chemical feature representations, the characterization results of the target molecule are determined.

2. The method according to claim 1, wherein determining the structural feature representation corresponding to the atomic composition comprises: Obtain a pre-constructed structure codebook, the structure codebook including structure content items, the structure content items including at least discrete structural feature representations of atoms; For each atom in the atomic composition, determine the three-dimensional composition information corresponding to each atom; A discrete structural feature representation that matches the three-dimensional composition information of each atom is determined from the structural codebook, and is used as the structural feature representation of each atom, wherein the discrete structural feature representation of each atom is part of the structural feature representation corresponding to the composition of the atom.

3. The method of claim 2, wherein determining a discrete structural feature representation from the structure codebook that matches the three-dimensional composition information of each atom comprises: Determine the continuous structural feature representation corresponding to the three-dimensional composition information of each atom; The continuous structural feature representation is compared with the discrete structural feature representation corresponding to each structural content item in the structural codebook, so as to obtain a discrete structural feature representation that matches the continuous structural feature representation based on the degree of matching, which is then used as the structural feature representation of each atom.

4. The method of claim 2, wherein each discrete structural feature representation in the discrete structural feature representation is obtained by performing clustering processing on the continuous structural feature representations of each molecular sample in the molecular sample set.

5. The method according to claim 2, wherein determining the chemical characteristic representation corresponding to the atomic composition includes: Obtain a pre-constructed chemical codebook, the chemical codebook including chemical content items, the chemical content items indicating at least the chemical characteristic representation of atoms; The chemical codebook and the structural codebook are constructed to have the same index space, and the chemical codebook and the structural codebook share index entries; Obtain the target index item corresponding to the structural content item that matches the three-dimensional composition information of each atom; Based on the target index entry, a chemical content entry corresponding to each atom is determined in the chemical codebook, serving as a chemical feature representation of each atom.

6. The method of claim 5, wherein the chemical codebook is pre-constructed in the following manner: Construct an initial chemical codebook with the same index space as the structure codebook. The initial chemical codebook contains multiple initial chemical content items, and the number of the initial chemical content items is the same as the number of structure content items in the structure codebook. Based on the atomic composition sample, determine the structural content items corresponding to the atomic composition sample; Based on the atomic composition sample and the structural content item corresponding to the atomic composition sample, the predicted chemical content item corresponding to the atomic composition sample is determined; Based on the difference between the predicted chemical content items and the standard chemical content items of the atomic composition sample, the plurality of initialized chemical content items are adjusted until predetermined conditions are met to obtain the chemical characteristic representation of the atoms.

7. The method of claim 6, wherein the difference between the predicted chemical content item and the standard chemical content item of the atomic composition sample includes at least one of a first difference indicating the pharmacophore prediction result, a second difference indicating the degree of mask atom restoration, and a third difference indicating the degree of denoising of the atom three-dimensional coordinates.

8. The method according to claim 1, wherein determining the characterization result of the target molecule comprises: A first mapping process is performed on the structural feature representation, and a second mapping process is performed on the chemical feature representation, so as to map the structural feature representation after the first mapping process and the chemical feature representation after the second mapping process to the same feature space; The structural feature representations and chemical feature representations after fusion mapping are used to determine the characterization results of the target molecule.

9. The method according to claim 1, further comprising: Based on the characterization results of the target molecule, perform at least one downstream task in molecular property prediction and molecular function retrieval.

10. The method according to claim 9, wherein the molecular property prediction includes at least one of biological interaction prediction, drug perturbation prediction, mass spectrometry feature simulation, and molecular physicochemical property prediction; and the molecular function retrieval includes at least one of drug screening and molecular functional similarity retrieval.

11. An apparatus for molecular characterization, comprising: The target molecule analysis module is used to determine the atomic composition of the target molecule; The structural feature representation determination module is used to determine the structural feature representation corresponding to the atomic composition based on the atomic composition, wherein the structural feature representation indicates the elemental abundance and atomic characterization distribution corresponding to the atomic composition; A chemical feature representation determination module is used to determine the chemical feature representation corresponding to the atomic composition based on the atomic composition and the structural feature representation; as well as A molecular characterization module is used to determine the characterization results of the target molecule based on the structural feature representation and the chemical feature representation.

12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.

13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method of any one of claims 1 to 10.