Molecular image processing method and device based on encoder and attention mechanism enhancement

By introducing a single-layer Transformer encoder and a GNN module with enhanced multi-head attention mechanism into the Transformer encoder architecture, the problem of high computational complexity in large-scale molecular data processing is solved, and efficient molecular graph representation and local feature extraction are achieved, which is suitable for large-scale drug screening tasks.

CN120220144AActive Publication Date: 2025-06-27SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510229802.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-27
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The current Transformer encoder architecture has high computational complexity when processing large-scale molecular data, making it difficult to adapt to large-scale molecular-level drug screening application scenarios.

Method used

The single-layer Transformer encoder module and the GNN module enhanced by the multi-head attention mechanism are used to realize the global information dissemination of the molecular graph through the self-attention mechanism, and local neighborhood information is extracted through the multi-head diagram attention mechanism to generate graph-level representation of the molecular graph.

Benefits of technology

It significantly reduces the complexity and computational cost of the model, improves the local information expression ability of the molecular map under complex structures, enhances the expression ability of the global characteristics of the molecule, and is suitable for drug screening tasks for large-scale molecular data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220144A_ABST
    Figure CN120220144A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a molecular image processing method and device based on encoder and attention mechanism enhancement, and the method comprises the steps: mapping a node feature matrix # imgabs0 # of a molecular image to a hidden feature space to generate an initial node feature # imgabs1 #; the initial node feature # imgabs2 # is transmitted to a single-layer encoder module; the attention weight is calculated through the dot product of a query # imgabs3 # space and a key # imgabs4 # space in the single-layer encoder module; on the basis of the attention weight weighted value vector # imgabs5 # obtained through calculation, node features are updated; extracting local neighborhood information in an initial node feature # imgabs6 # through a multi-head attention mechanism enhanced GNN module; and based on the updated node features and the extracted local neighborhood information, generating a graph-level representation of the molecular graph. In this way, the local information expression ability of the molecular graph under a complex structure is improved, and richer and more accurate data support is provided for overall expression of molecules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, for example, to a molecular image processing method and device enhanced based on an encoder and an attention mechanism. Background Art

[0002] The core task of molecular representation learning is to transform the chemical structure and physical properties of molecules into numerical representations that can be processed by a computer, thereby supporting the reasoning and prediction of various downstream tasks.

[0003] In recent years, with the rapid development of deep learning technology, deep learning-based molecular representation learning methods have become a research hotspot. These methods directly extract features from data through an end-to-end learning mode, breaking through the dependence on artificial rules in traditional methods. Compared with traditional methods, deep learning methods exhibit stronger expressive power and generalization performance, especially when dealing with large-scale molecular data and complex chemical structures.

[0004] The current Transformer encoder architecture usually captures global information by stacking multiple attention modules, which further exacerbates its computational complexity.

[0005] Therefore, in the artificial intelligence-assisted drug discovery process, there are problems of low computational efficiency and difficulty in adapting to large-scale molecular-level drug screening application scenarios.

[0006] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application. Summary of the Invention

[0007] To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments, but rather serves as a preface to the subsequent detailed description.

[0008] The embodiments of the present disclosure provide a molecular image processing method enhanced based on an encoder and an attention mechanism. This method is applied to a neural network model including a single-layer Transformer encoder module, a GNN module enhanced with multi-head attention, and an information fusion module. The method includes: Mapping the node feature matrix of the molecular image to a hidden feature space to generate initial node features ; Transmitting the initial node features to the single-layer Transformer encoder module; Through the query in the single-layer Transformer encoder module and the key Calculate the attention weights using the dot product in space; Based on the calculated weighted value vector of attention weights , update the node features; The GNN module enhanced by the multi-head attention mechanism extracts the initial node features in the local neighborhood information; Based on the updated node features and the extracted local neighborhood information, generate the graph-level representation of the molecular graph.

[0009] In some embodiments, map the node feature matrix of the molecular image to the hidden feature space to generate the initial node features , satisfying the formula: , where, is the weight matrix, F represents the feature dimension of each node, is the dimension of the hidden layer, is the bias vector, is the non-linear activation function.

[0010] In some embodiments, pass the initial node features to the single-layer Transformer encoder module, including: Map the initial node features to the query , key and value spaces, satisfying the formula: , , , where, , , , is the dimension of the hidden layer.

[0011] In some embodiments, calculate the attention weights using the dot product in the query and key spaces of the single-layer Transformer encoder module, satisfying the formula: , where, represents the transposed matrix of the key space, is the normalization function.

[0012] In some embodiments, based on the calculated weighted value vector of attention weights , update the node features, satisfying the formula: , Among them, is a layer normalization operation, which is used to enhance the stability and convergence of training.

[0013] In some embodiments, the GNN module enhanced by the multi-head attention mechanism extracts the initial node features in the local neighborhood information, including: For any node and its neighbor nodes , the unnormalized attention score of the th attention head satisfies the formula: , Among them, is a weight matrix with learning ability, is the weight vector of the th attention head, represents the feature concatenation operation, is a non-linear activation function, and ⊤ represents the transpose operation of the matrix; Normalize the calculated attention scores through the function to obtain the attention weight of node to node , which satisfies the formula: , , The normalized attention weight is used to weighted aggregate the neighbor node features, which satisfies the formula: , Among them, the neighbor set of node represents all the directly connected nodes of node ; The final node features are obtained by weighted average of all attention heads, which satisfies the formula: , Among them, represents a non-linear activation function.

[0014] In some embodiments, based on the updated node features and the extracted local neighborhood information, a graph-level representation of the molecular graph is generated, including: Based on the updated node features and the extracted local neighborhood information, weighted summation is performed for fusion to obtain the fused node features, which satisfies the formula: , Among them, Denote the molecular graph feature matrix extracted by the GNN module enhanced with multi-head attention mechanism, where each row corresponds to the final feature of a node. , which is used to dynamically adjust the ratio of global and local features; is the updated node feature.

[0015] Based on the fused node features, generate the graph-level representation of the molecular graph through global average pooling operation.

[0016] The embodiments of the present disclosure provide a data analysis device based on an encoder and an enhanced attention mechanism. The device includes: A processing module for mapping the node feature matrix to the hidden feature space to generate the initial node features ; The processing module is also used to transfer the initial node features to the single-layer Transformer encoder module; The single-layer Transformer encoder module is used to calculate the attention weights based on the dot product of the query and the key spaces; The single-layer Transformer encoder module is also used to update the node features based on the calculated attention weight weighted value vector ; The GNN module enhanced with multi-head attention mechanism is used to extract the local neighborhood information in the initial node features ; The information fusion module is used to generate the corresponding image analysis representation result based on the updated node features and the extracted local neighborhood information.

[0017] The embodiments of the present disclosure provide an electronic device, which includes at least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can execute the above-mentioned molecular image processing method based on an encoder and an enhanced attention mechanism.

[0018] The embodiments of the present disclosure provide a storage medium storing program instructions, which when running, execute the above-mentioned molecular image processing method based on an encoder and an enhanced attention mechanism.

[0019] The molecular image processing method, device, device and storage medium provided by the embodiments of the present disclosure can achieve the following technical effects: The present disclosure realizes the global information propagation of molecular graphs through the self-attention mechanism of a single-layer Transformer encoder, avoiding the computational complexity problems caused by multi-layer stacking in traditional Transformer architectures. At the same time, this mechanism can efficiently model the dependencies between long-distance nodes in molecular graphs, thus significantly enhancing the expression ability of molecular global features. Moreover, a multi-head graph attention network is introduced to enhance the model's expression ability for local features of molecular graphs. The multi-head graph attention mechanism captures diverse interaction characteristics between nodes through parallel computing, thus ensuring the efficient extraction and aggregation of local chemical features. It improves the local information expression ability of molecular graphs under complex structures, providing richer and more accurate data support for the overall representation of molecules.

[0020] The above general description and the following description are only exemplary and explanatory, and are not used to limit this application. Brief Description of the Drawings

[0021] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and among them: Figure 1 is a schematic flowchart of a method for enhancing molecular image processing based on an encoder and an attention mechanism provided by an embodiment of the present disclosure; Figure 2 is a schematic diagram of the overall architecture of a model provided by an embodiment of the present disclosure; Figure 3 is a molecular graph of the model input provided by an embodiment of the present disclosure; Figure 4 is a schematic diagram of a single-layer Transformer encoder module provided by an embodiment of the present disclosure; Figure 5 is a schematic diagram of a GNN module enhanced by multi-head attention provided by an embodiment of the present disclosure; Figure 6 is a schematic diagram of the structure of a device for enhancing molecular image processing based on an encoder and an attention mechanism provided by an embodiment of the present disclosure; Figure 7 is a schematic diagram of the structure of a device for enhancing molecular image processing based on an encoder and an attention mechanism provided by an embodiment of the present disclosure. Detailed Description of the Embodiments

[0022] In order to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and explanation, and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to provide a thorough understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner to simplify the drawings.

[0023] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.

[0024] Unless otherwise specified, the term "plurality" means two or more.

[0025] In the embodiments of the present disclosure, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.

[0026] The term "and / or" is a description of the associated relationship of an object, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B these three relationships.

[0027] The term "corresponding" may refer to an associated relationship or a binding relationship. A corresponding to B means that there is an associated relationship or a binding relationship between A and B.

[0028] The core task of molecular representation learning is to transform the chemical structure and physical properties of molecules into numerical representations that can be processed by computers, thereby supporting the reasoning and prediction of various downstream tasks. The key challenge of this task lies in how to effectively capture the local and global characteristics of molecules while meeting the requirements of computational efficiency and generalization ability. In the early stage of molecular representation learning, traditional methods such as Simplified Molecular Input Line Entry System (SMILES) and Extended-Connectivity Fingerprints (ECFP) were widely used. SMILES represents the molecular structure in the form of a string. Its advantage is its simplicity, but due to its linear structure, it cannot reflect the complex topological relationships and spatial characteristics between atoms in the molecule. ECFP encodes the molecular structure into a fixed-length vector through predefined rules. Although it can capture local chemical features, its regular design is difficult to flexibly adapt to diverse molecular structures and heavily relies on prior knowledge. In addition, these methods show limited generalization ability when dealing with new molecules unknown to the model. When facing complex molecules and modern large-scale datasets, their expressive ability and applicability gradually show deficiencies, which have promoted the exploration of more advanced data-driven representation methods.

[0029] In recent years, with the rapid development of deep learning technology, deep learning-based molecular representation learning methods have become a research hotspot. These methods directly extract features from data through an end-to-end learning mode, breaking through the dependence on artificial rules of traditional methods. Compared with traditional methods, deep learning methods show stronger expressive ability and generalization performance, especially when dealing with large-scale molecular data and complex chemical structures. An important progress is the proposal of molecular graph representation. This method represents molecules as graph structures, where atoms are represented as nodes in the graph and chemical bonds are represented as edges. In this way, molecular graphs can more naturally describe the topological characteristics and connection relationships in molecules and provide a more suitable input format for deep learning models.

[0030] Graph Neural Networks (GNNs) are a type of deep learning models for processing graph data that have emerged in recent years and have achieved remarkable success in molecular representation learning. GNNs learn features of the relationships between nodes and their neighbors through a message-passing mechanism. The core idea is to update the representation of the central node by aggregating the features of neighboring nodes. The multi-layer stacked structure of GNNs enables it to gradually integrate the local information of nodes and finally generate an embedded representation of the entire molecule. This method is not only flexibly adaptable to graph structure inputs of different scales but also performs excellently in capturing the topological properties in molecular graphs. For example, GNNs can effectively extract the crucial local chemical features in molecular property prediction and at the same time show good scalability and are applicable to a variety of molecular task scenarios.

[0031] However, despite the excellent performance of GNNs in molecular representation learning, there are also inherent limitations in their design. The current mainstream GNN models mainly rely on local aggregation mechanisms, that is, the representation of nodes is updated by aggregating the features of neighboring nodes in each layer of the network. Although this design performs well in capturing local information, it is insufficient in dealing with long-range dependencies in complex molecular graphs. Especially in molecules containing long-chain chemical structures or complex cyclic structures, the signals between distant nodes are easily over-compressed during multi-layer propagation, resulting in information loss. Another issue worthy of attention is the expressive power of GNNs. When dealing with molecules with similar structures but different properties, traditional GNNs show certain deficiencies in discrimination ability. For example, their performance is not precise enough in the graph isomorphism problem. This deficiency mainly stems from the limitation of its message-passing mechanism in extracting global graph information.

[0032] To overcome the limitations of GNNs, in recent years, researchers have begun to explore other deep learning models for processing graph-structured data. Models based on the Transformer architecture can effectively capture the global dependencies in the graph through the self-attention mechanism, making up for the deficiency of GNNs in long-range information modeling. In molecular representation learning, the Transformer encoder generates a more expressive molecular embedded representation by modeling the relationships between any node pairs. This method can not only capture the long-range dependency information of molecules but also effectively handle the heterogeneity and complexity of molecular graphs through a flexible attention distribution mechanism.

[0033] However, the Transformer model still faces the challenge of high computational costs when dealing with large-scale data, and the consumption of its full-pair computational complexity on resources has become a bottleneck. The current Transformer encoder architecture usually captures global information through the stacking of multiple attention modules, which further exacerbates its computational complexity. Therefore, in the process of artificial intelligence-assisted drug discovery, it is necessary to optimize the Transformer architecture to improve its computational efficiency and adapt to the actual application scenarios of large-scale drug screening.

[0034] To solve the above problems, the present disclosure provides a molecular image processing method, apparatus, device, and storage medium enhanced based on an encoder and an attention mechanism.

[0035] The following describes the molecular image processing method, apparatus, device, and storage medium enhanced based on an encoder and an attention mechanism provided by the embodiments of the present disclosure with reference to the accompanying drawings.

[0036] Figure 1 It is a schematic flowchart of a molecular image processing method enhanced based on an encoder and an attention mechanism provided by an embodiment of the present disclosure. This method is applied to a neural network model including a single-layer Transformer encoder module, a graph neural network (GNN) module enhanced with multi-head attention, and an information fusion module.

[0037] As Figure 1 shown, the molecular image processing method enhanced based on an encoder and an attention mechanism may include: S101, mapping the node feature matrix of the molecular image to a hidden feature space to generate initial node features ; S102, passing the initial node features to the single-layer Transformer encoder module; S103, calculating the attention weights through the dot product in the query and key spaces of the single-layer Transformer encoder module; S104, updating the node features based on the weighted value vector calculated from the attention weights; S105, extracting the local neighborhood information in the initial node features through the GNN module enhanced with multi-head attention mechanism; S106, generating a graph-level representation of the molecular graph based on the updated node features and the extracted local neighborhood information.

[0038] It should be noted that the molecular graph is modeled as a graph data structure , where the set of nodes represents the atoms in the molecule, and the set of edges represents the chemical bonds between the atoms. The input of the model of the present invention is the node feature matrix and the adjacency matrix , where the node feature matrix each row represents the attributes of a node (atom), including the physical and chemical properties of each atom in the molecule, such as atom type, number of valence electrons, etc., represents the number of nodes in the molecular graph, represents the feature dimension of each node. The adjacency matrix is used to describe the connection relationship between the atoms in the molecule.

[0039] In some embodiments, the node feature matrix of the molecular image is mapped to a hidden feature space to generate the initial node features , satisfying the formula: , wherein, is the weight matrix, F represents the feature dimension of each node, is the dimension of the hidden layer, is the bias vector, is the non-linear activation function.

[0040] In some embodiments, the initial node features are passed to a single-layer Transformer encoder module, including: The initial node features are mapped to the query , key and value spaces, satisfying the formula: , , , wherein, , , , is the dimension of the hidden layer.

[0041] In some embodiments, the attention weights are calculated by the dot product of the query and key spaces in the single-layer Transformer encoder module, satisfying the formula: , wherein, represents the key The transposed matrix of space is a normalization function.

[0042] In some embodiments, based on the calculated attention weight weighted value vector , update the node features to satisfy the formula: , where is the layer normalization operation, which is used to enhance the stability and convergence of training.

[0043] In some embodiments, the GNN module enhanced by the multi-head attention mechanism extracts the initial node features in the local neighborhood information, including: For any node and its neighbor node , the unnormalized attention score of the th attention head satisfies the formula: , where is a weight matrix with learning ability, is the weight vector of the th attention head, represents the feature concatenation operation, is a non-linear activation function, and ⊤ represents the matrix transpose operation; Normalize the calculated attention scores through the function to obtain the attention weight of node to node , satisfying the formula: , The normalized attention weight is used to weight and aggregate the neighbor node features, satisfying the formula: , where the neighbor set of node represents all the directly connected nodes of node ; The final node features are obtained by weighted averaging through all attention heads, satisfying the formula: , where represents a non-linear activation function.

[0044] In some embodiments, based on the updated node features and the extracted local neighborhood information, generate the graph-level representation of the molecular graph, including: Based on the updated node features and the extracted local neighborhood information, weighted summation is performed for fusion to obtain the fused node features, which satisfy the formula: , where, represents the molecular graph feature matrix extracted by the GNN module enhanced by the multi-head attention mechanism, and each row corresponds to the final feature of a node. , which is used to dynamically adjust the ratio of global and local features; is the updated node feature.

[0045] Based on the fused node features, a graph-level representation of the molecular graph is generated through global average pooling operation.

[0046] Specifically, as Figure 2 shown, the overall architecture of the model consists of a lightweight Transformer encoder module 201, a multi-head attention enhanced GNN module 201, and an information fusion module 203. As Figure 3 shown, the input molecular graph of the model is represented by the node feature matrix and the adjacency matrix , where the node feature matrix describes the physicochemical properties of atoms (such as atom type, bond type, aromaticity, etc.), and the adjacency matrix defines the chemical bond connection relationship between atoms. First, the node features are mapped to the hidden space through a linear transformation to generate the initial feature representation, and these features are then passed to the global feature modeling module and the local feature extraction module.

[0047] The Transformer model was initially proposed in the field of natural language processing. By stacking multiple layers of attention mechanisms, it can effectively capture complex semantic relationships and significantly improve the expressive ability of the model. As the Transformer architecture is gradually applied to graph representation learning tasks, many designs directly follow the concept of stacking multiple layers of attention in order to enhance the modeling ability of complex relationships in the graph structure. However, in molecular graph tasks, our experimental results show that the stacking of multiple layers of attention does not significantly improve the prediction performance, and even shows a performance decline in some cases. This may be due to the fact that the molecular graph has a fixed topological structure and limited dependencies between nodes. The stacking of multiple layers of attention is prone to information redundancy and exacerbates the problem of gradient disappearance or degradation, thus having a negative impact on the model performance. Based on the experimental findings, the present invention proposes a lightweight single-layer Transformer encoder module. As Figure 4As shown, this module adopts a single-layer self-attention mechanism, which significantly reduces the complexity and computational cost of the model while retaining the global information modeling ability. Experimental results show that this module has good applicability and significant performance advantages in molecular graph tasks. Specifically, this module generates queries, keys, and values through node features, and calculates attention weights using these mappings to effectively capture the correlations between global nodes. Subsequently, by performing a weighted sum operation on the node features, the node representations are updated and the efficient propagation of global information is achieved. This design not only simplifies the model architecture but also improves the efficiency and expressive power of information integration, especially when dealing with large-scale molecular datasets.

[0048] As Figure 5 shown, the GNN module enhanced by multi-head attention extracts local features of the molecular graph through the multi-head attention mechanism. This module uses multiple attention heads to separately capture diverse interaction characteristics between nodes to ensure the efficient extraction and expression of local chemical features. Among them, each attention head independently calculates the importance weights of neighbor nodes and generates local feature representations of nodes through weighted aggregation. Subsequently, the results of multiple attention heads are combined through feature concatenation and fusion operations to generate more expressive local feature embeddings. The multi-head mechanism significantly enhances the adaptability of the model to complex molecular structures and shows good robustness especially when dealing with cyclic structures and long-chain chemical structures.

[0049] The information fusion module integrates global features and local features through a dynamically learned weight mechanism to generate the final molecular graph embedding representation. Specifically, this module adaptively adjusts the contribution ratios of global features and local features in the embedding representation according to task requirements, thus ensuring the expression quality and task applicability of the molecular graph embedding. The fused molecular embedding representation is not only applicable to molecular property prediction and toxicity assessment but also can be widely applied to various downstream tasks such as virtual drug screening.

[0050] Through the collaborative optimization design of the above modules, combined with Figure 1 the corresponding molecular image processing method enhanced by the encoder and attention mechanism, the present invention achieves efficient global information modeling and local feature extraction of molecular graph representations while maintaining a low computational complexity. Experimental results show that this model demonstrates a significant performance improvement when dealing with complex chemical structures and large-scale molecular datasets, and at the same time has extremely high computational efficiency, meeting the requirements of actual application scenarios in drug discovery.

[0051] The molecular image processing method enhanced by an encoder and an attention mechanism proposed in the embodiments of the present disclosure uses a molecular graph representation learning model that can effectively capture global and local features in the molecular structure, demonstrating high accuracy and efficiency in related tasks of artificial intelligence-assisted drug discovery, and can provide reliable technical support for related research.

[0052] In the embodiments of the present disclosure, the self-attention mechanism of a single-layer Transformer encoder is used to achieve the global information propagation of the molecular graph, avoiding the computational complexity problem caused by the multi-layer stacking in the traditional Transformer architecture. At the same time, this mechanism can efficiently model the dependency relationships between long-distance nodes in the molecular graph, thereby significantly enhancing the expression ability of molecular global features.

[0053] Moreover, a multi-head graph attention network is introduced to enhance the model's expression ability for local features of the molecular graph. The multi-head graph attention mechanism captures diverse interaction characteristics between nodes through parallel computing, thereby ensuring the efficient extraction and aggregation of local chemical features. This design significantly improves the local information expression ability of the molecular graph under complex structures, providing richer and more accurate data support for the overall representation of molecules.

[0054] An information fusion module is designed to integrate local and global feature representations. This module adaptively balances the contributions of local and global features through a weight learning mechanism, thereby generating a more accurate molecular graph embedding representation to meet the requirements of downstream tasks.

[0055] Corresponding to Figure 1 the molecular image processing method enhanced by an encoder and an attention mechanism in Figure 6 this disclosure also provides a molecular image processing device enhanced by an encoder and an attention mechanism. As shown, this device may specifically include: A processing module 601 for mapping the node feature matrix to a hidden feature space to generate initial node features The processing module 601 is also used to transfer the initial node features to a single-layer Transformer encoder module; A single-layer Transformer encoder module 602 for calculating attention weights based on the dot product of the query and key spaces; The single-layer Transformer encoder module 602 is also used to update the node features based on the calculated attention weight weighted value vector ; A GNN module 603 enhanced by a multi-head attention mechanism for extracting initial node features Local neighborhood information in The information fusion module 604 is configured to generate a corresponding image analysis representation result based on the updated node features and the extracted local neighborhood information.

[0056] In some embodiments, the node feature matrix of the molecular image is mapped to a hidden feature space to generate initial node features , satisfying the formula: , where is the weight matrix, F represents the feature dimension of each node, is the dimension of the hidden layer, is the bias vector, is the non-linear activation function.

[0057] In some embodiments, the initial node features are passed to a single-layer Transformer encoder module, including: The initial node features are mapped to the query , key and value spaces, satisfying the formula: , , , where , , , is the dimension of the hidden layer.

[0058] In some embodiments, the attention weights are calculated through the dot product of the query and key spaces in the single-layer Transformer encoder module, satisfying the formula: , where represents the transposed matrix of the key space, is the normalization function.

[0059] In some embodiments, based on the calculated attention weight weighted value vector , the node features are updated, satisfying the formula: , where is the layer normalization operation, which is used to enhance the stability and convergence of training.

[0060] In some embodiments, a GNN module enhanced by a multi-head attention mechanism extracts initial node features of the local neighborhood information in, including: For any node and its neighbor nodes , the unnormalized attention scores of the -th attention head satisfy the formula: , where, is a weight matrix with learning ability, is the weight vector of the -th attention head, represents the feature concatenation operation, is a non-linear activation function, and ⊤ represents the transpose operation of the matrix; Normalize the calculated attention scores through the function to obtain the attention weight of node for node , satisfying the formula: , , The normalized attention weight is used for weighted aggregation of neighborhood node features, satisfying the formula: , where, the neighbor set of node represents all directly connected nodes of node ; The final node features are obtained by weighted averaging through all attention heads, satisfying the formula: , where, represents a non-linear activation function.

[0061] In some embodiments, based on the updated node features and the extracted local neighborhood information, a graph-level representation of the molecular graph is generated, including: Based on the updated node features and the extracted local neighborhood information, weighted summation is performed for fusion to obtain the fused node features, satisfying the formula: , where, represents the molecular graph feature matrix extracted by the GNN module enhanced by the multi-head attention mechanism, and each row corresponds to the final feature of a node, , which is used to dynamically adjust the ratio of global and local features; is the updated node feature.

[0062] Based on the fused node features, a graph-level representation of the molecular graph is generated through global average pooling operation.

[0063] Combine Figure 7 As shown, an embodiment of the present disclosure further provides a molecular image processing device 700 enhanced based on an encoder and an attention mechanism, including a processor 704 and a memory 701. Optionally, the system may further include a communication interface 702 and a bus 703. Among them, the processor 704, the communication interface 702, and the memory 701 can communicate with each other through the bus 703. The communication interface 702 can be used for information transmission. The processor 704 can call the logical instructions in the memory 701 to execute the molecular image processing method based on the encoder and the attention mechanism in the above embodiment.

[0064] In addition, when the logical instructions in the above-mentioned memory 701 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0065] The memory 701, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor 704 executes functional applications and data processing by running the program instructions / modules stored in the memory 701, that is, implements the molecular image processing method based on the encoder and the attention mechanism in the above embodiment.

[0066] The memory 701 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device. In addition, the memory 701 may include a high-speed random access memory and may also include a non-volatile memory.

[0067] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are set as a molecular image processing method based on an encoder and an attention mechanism.

[0068] The above-mentioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.

[0069] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or it may also be a transitory storage medium.

[0070] The above description and the drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. Embodiments merely represent possible variations. Unless explicitly required, separate components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or substituted for parts and features of other embodiments. As used in the description of the embodiments, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groupings of these. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device including the element. In this article, each embodiment may focus on the differences from other embodiments, and the same or similar parts between the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.

[0071] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0072] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components can be or can not be physically separated. The components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. In addition, the functional units in the embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur out of the order disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0074] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0075] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0076] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0077] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0078] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or in a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0079] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0080] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitations are imposed herein.

[0081] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A molecular image processing method based on encoder and attention mechanism enhancement, characterized in that: The method is applied to a neural network model including a single-layer Transformer encoder module, a multi-head attention enhanced GNN module, and an information fusion module; the method includes: The node feature matrix of the molecular image Mapping to the latent feature space to generate initial node features ; The initial node features Passed to a single-layer Transformer encoder module; Through the query in a single-layer Transformer encoder module and key The dot product of the space calculates the attention weight; Based on the calculated attention weight vector , update node features; The GNN module enhanced by the multi-head attention mechanism extracts the initial node features Local neighborhood information in ; Based on the updated node features and the extracted local neighborhood information, a graph-level representation of the molecular graph is generated.

2. The method according to claim 1, characterized in that The node feature matrix of the molecular image Mapping to the latent feature space to generate initial node features , satisfying the formula: , in, is the weight matrix, F represents the feature dimension of each node, is the dimension of the hidden layer, is the bias vector, is a non-linear activation function.

3. The method according to claim 1, characterized in that The initial node features Passed to a single-layer Transformer encoder module, including: The initial node features Mapping to Query ,key Sum Space, satisfying the formula: , , , in, , , , is the dimension of the hidden layer.

4. The method according to claim 3, characterized in that The query in the single-layer Transformer encoder module and key The dot product of the space calculates the attention weight, satisfying the formula: , in, Display key The transposed matrix of the space, is a normalization function.

5. The method according to claim 4, characterized in that The attention weighted value vector calculated based on , update the node features to satisfy the formula: , in, It is a layer normalization operation used to enhance the stability and convergence of training.

6. The method according to claim 1, characterized in that The GNN module enhanced by the multi-head attention mechanism extracts the initial node features The local neighborhood information in includes: For any node and its neighbor nodes , No. The non-normalized attention score of the attention head satisfies the formula: , in, is a weight matrix with learning ability, It is The weight vector of the attention head, represents the feature concatenation operation, is a nonlinear activation function, and ⊤ represents the transpose operation of the matrix; The calculated attention score is passed Normalize the function and get the node For Node The attention weight , satisfying the formula: , Normalized attention weights It is used to weight the aggregation of neighborhood node features, satisfying the formula: , Among them, the node The neighbor set Representation Node All directly connected nodes; The final node features are obtained by weighted averaging of all attention heads, satisfying the formula: , in, represents a non-linear activation function.

7. The method according to claim 1, characterized in that The step of generating a graph-level representation of a molecular graph based on the updated node features and the extracted local neighborhood information includes: Based on the updated node features and the extracted local neighborhood information, weighted summation is performed to obtain the fused node features, which satisfies the formula: , in, Represents the molecular graph feature matrix extracted by the GNN module enhanced by the multi-head attention mechanism, where each row corresponds to the final feature of a node. , used to dynamically adjust the ratio of global and local features; is the updated node feature. Based on the fused node features, a graph-level representation of the molecular graph is generated through a global average pooling operation.

8. A data analysis device based on an encoder and an enhanced attention mechanism, characterized in that: The device comprises: Processing module, used to transform node feature matrix Mapping to the latent feature space to generate initial node features ; The processing module is also used to convert the initial node features Passed to a single-layer Transformer encoder module; A single-layer Transformer encoder module for query-based and key The dot product of the space calculates the attention weight; A single-layer Transformer encoder module is also used to weight the value vector based on the calculated attention weights. , update node features; Multi-head attention mechanism enhanced GNN module for extracting initial node features Local neighborhood information in ; The information fusion module is used to generate corresponding image analysis representation results based on the updated node features and the extracted local neighborhood information.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; It is characterized in that the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-granularity video retrieval method and device

    CN117194710A

  • Artificial intelligence-based method and apparatus for determining drug feature information

    WO2023134061A1