Information acquisition method, storage medium and electronic device
By aggregating the features of multi-level structural objects in different dimensions and obtaining their interactive attribute information, the problem of insufficient accuracy of information acquisition in the prior art is solved, and the accuracy of information acquisition and prediction capabilities in complex scenarios are improved.
Patent Information
- Application Number
- CN202211007668.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-08-22
AI Technical Summary
In the information acquisition scenario of multi-level structural objects, the prior art adopts structural analysis in a single dimension, resulting in low accuracy in information acquisition and the inability to fully capture structural features in multiple dimensions.
By obtaining the features of the target composition structure of the multi-level structural object in the first dimension and aggregating these features, the features in the second dimension are obtained, and the interactive attribute information of the multi-level structural object is obtained, and the accuracy of information acquisition is improved using multi-dimensional features.
It achieves the accuracy of information acquisition in complex scenarios and ensures that the performance prediction of multi-level structural objects during interaction is more accurate.
Smart Images

Figure CN115359328B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and more specifically, to an information acquisition method, device, storage medium, and electronic device. Background Art
[0002] When acquiring information from single-level objects, a single-dimensional structure can be used to obtain information. However, for multi-level objects, the complexity is much greater than for single-level objects. Therefore, if a single-dimensional structure is still used to acquire relevant information, analysis of the structure in one or more dimensions will often be missed, thus affecting the accuracy of the acquired information. Therefore, the accuracy of information acquisition is low.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide an information acquisition method, device, storage medium, and electronic device to at least solve the technical problem of low accuracy in information acquisition.
[0005] According to one aspect of an embodiment of the present application, an information acquisition method is provided, including: acquiring first-level features corresponding to a target component structure of a multi-level structure object, wherein the multi-level structure object is composed of multiple target component structures connected, and the first-level features are features of the target component structure under the first dimension; aggregating the first-level features corresponding to each target component structure in the multiple target component structures to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure under the second dimension, the target component structure is composed of multiple target component structures under the second dimension, and the target component structure under the second dimension is composed of multiple target component structures under the first dimension; based on the second-level features, acquiring interaction attribute information corresponding to the multi-level structure object, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects.
[0006] According to another aspect of an embodiment of the present application, an information acquisition device is also provided, including: a first acquisition unit, used to acquire first-level features corresponding to a target component structure of a multi-level structure object, wherein the multi-level structure object is composed of multiple target component structures connected, and the first-level features are features of the target component structure under the first dimension; an aggregation unit, used to aggregate the first-level features corresponding to each target component structure in the multiple target component structures to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure under the second dimension, the target component structure is composed of multiple target component structures under the second dimension, and the target component structure under the second dimension is composed of multiple target component structures under the first dimension; a second acquisition unit, used to acquire interaction attribute information corresponding to the multi-level structure object based on the second-level features, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects.
[0007] As an optional solution, the above-mentioned first acquisition unit includes: an extraction module for extracting image features in the target image corresponding to the above-mentioned multi-level structure object; a first acquisition module for obtaining initial first-level features corresponding to each first component structure in the multiple target component structures under the above-mentioned first dimension based on the above-mentioned image features; and a second acquisition module for obtaining first-level features corresponding to each of the above-mentioned first component structures based on the above-mentioned initial first-level features.
[0008] As an optional solution, the above-mentioned second acquisition module includes: a first acquisition sub-module, used to acquire multiple first component structure pairs composed of the above-mentioned first component structures; a calculation sub-module, used to calculate the first structure pair features of each first component structure pair in the above-mentioned multiple first component structure pairs, wherein the above-mentioned first structure pair features include the structural features corresponding to each first component structure in the above-mentioned first component structure pairs, and the relative distance features between each of the above-mentioned first component structures; the second acquisition sub-module, used to acquire the first-level features corresponding to the above-mentioned each first component structure based on the above-mentioned first structure pair features.
[0009] As an optional solution, the above-mentioned second acquisition submodule includes: an execution subunit, which is used to execute the following steps until the convergence condition is reached: aggregating the first structure pair features of the above-mentioned each first component structure pair to obtain the target structure pair feature; merging the above-mentioned target structure pair feature with the structural features corresponding to the above-mentioned each first component structure to obtain the new structural feature corresponding to the above-mentioned each first component structure; when the above-mentioned convergence condition is met, determining the new structural feature corresponding to the above-mentioned each first component structure as the first-level feature corresponding to the above-mentioned each first component structure; when the above-mentioned convergence condition is not met, recalculating the above-mentioned first structure pair feature using the new structural feature corresponding to the above-mentioned each first component structure to obtain the new first structure pair feature of the above-mentioned each first component structure pair, and determining the new first structure pair feature of the above-mentioned each first component structure pair as the first structure pair feature of the above-mentioned each first component structure pair, and determining the new atomic feature corresponding to the above-mentioned each first component structure as the structural feature corresponding to the above-mentioned each first component structure.
[0010] As an optional solution, the above-mentioned first acquisition module includes: a third acquisition sub-module, which is used to obtain the first structural features corresponding to the above-mentioned each first component structure and the target structural features corresponding to the target component structure in the above-mentioned second dimension where the above-mentioned each first component structure is located based on the above-mentioned image features; a splicing sub-module, which is used to splice the above-mentioned first structural features and the above-mentioned target structural features to obtain the second structural features corresponding to the above-mentioned each first component structure, wherein the above-mentioned initial first-level features include the above-mentioned second structural features.
[0011] As an optional solution, the above-mentioned second acquisition unit includes at least one of the following: a third acquisition module, used to acquire a plurality of second component structure pairs consisting of each second component structure in the above-mentioned plurality of target component structures under the above-mentioned second dimension; based on the above-mentioned second-level features, calculating the first target structure feature of each second component structure pair in the above-mentioned plurality of second component structure pairs, wherein the first target structure feature is used to represent the feature inner product similarity between each second component structure in the above-mentioned second component structure pair; and obtaining the above-mentioned interaction attribute information based on the above-mentioned first target structure feature; a fourth acquisition module, used to acquire the above-mentioned plurality of second component structure pairs; based on the above-mentioned second-level features, calculating the second target structure feature of each second component structure pair, wherein the second target structure feature is used to represent the relative distance between each second component structure when linearly mapped on the sequence; and obtaining the above-mentioned interaction attribute information based on the above-mentioned second target structure feature; a fifth acquisition module, used to acquire the above-mentioned plurality of second component structure pairs; based on the above-mentioned second-level features, calculating the third target structure feature of each second component structure pair, wherein the third target structure feature is used to represent the carbon atom distance between each second component structure; and obtaining the above-mentioned interaction attribute information based on the above-mentioned third target structure feature.
[0012] As an optional solution, the above-mentioned second acquisition unit includes: a first aggregation module, which is used to aggregate the above-mentioned first target structural features, the above-mentioned second target structural features, and the above-mentioned third target structural features to obtain aggregated structural features when the above-mentioned first target structural features, the above-mentioned second target structural features, and the above-mentioned third target structural features are obtained; a merging module, which is used to merge the above-mentioned aggregated structural features with the above-mentioned second-level features through vector splicing to obtain output structural features; and a sixth acquisition module, which is used to obtain the above-mentioned interactive attribute information based on the above-mentioned output structural features.
[0013] As an optional solution, the above-mentioned second acquisition module includes: a first input submodule, which is used to input the above-mentioned initial first-level features into the first target model, and transfer information through the first substructure in the above-mentioned first target model to obtain the first-level features corresponding to the above-mentioned each first component structure, wherein the above-mentioned first target model is a neural network model obtained by training with the first sample and used for feature processing; the above-mentioned aggregation unit includes: a second aggregation module, which is used to aggregate the first-level features corresponding to the above-mentioned each first component structure through the second substructure in the above-mentioned first target model to obtain the second-level features corresponding to the above-mentioned target component structure.
[0014] As an optional solution, the above-mentioned second acquisition unit includes: an input module, which is used to input the above-mentioned second-level features into the second target model, and transmit information through the third substructure of the above-mentioned second target model to obtain the input structural features corresponding to the above-mentioned each second component structure, wherein the above-mentioned second target model is a neural network model obtained by training with the second sample and used for feature processing; an output module, which is used to predict and output the input structural features corresponding to the above-mentioned each second component structure through the fourth substructure of the above-mentioned second target model to obtain the above-mentioned interactive attribute information.
[0015] As an optional solution, the above-mentioned device includes: a third acquisition unit, used to obtain the number of target structures of the above-mentioned multi-level structure object before obtaining the first-level features corresponding to the target component structures of the multi-level structure object, wherein the above-mentioned number of target structures includes at least one of the following: the first number corresponding to the target component structures in the above-mentioned multiple target component structures, the second number corresponding to the second component structures in the above-mentioned multiple target component structures under the above-mentioned second dimension, and the third number corresponding to the first component structure in the above-mentioned multiple target component structures under the above-mentioned first dimension; a calculation unit, used to calculate the acquisition complexity of the above-mentioned interactive attribute information based on the above-mentioned number of target structures, wherein the above-mentioned acquisition complexity is used to represent the complexity of the operations to be performed when obtaining the above-mentioned interactive attribute information.
[0016] As an optional solution, the above-mentioned second acquisition unit includes: a seventh acquisition module, used to obtain the above-mentioned interactive attribute information based on the above-mentioned second-level characteristics and the physical and chemical characteristics corresponding to each first component structure in the above-mentioned multiple target component structures under the above-mentioned first dimensions, wherein the above-mentioned physical and chemical characteristics include at least one of the following: a first feature for indicating the hydrophilicity of the target substructure in the above-mentioned first component structure that does not participate in dehydration bonding, a second feature for indicating the hydrophobicity of the above-mentioned target substructure, and a third feature for indicating the charge carried by the above-mentioned target substructure.
[0017] As an optional solution, the above-mentioned first acquisition unit includes: an eighth acquisition module, which is used to obtain the molecular features corresponding to the amino acids in the protein to be detected; the above-mentioned aggregation unit includes: a third aggregation module, which is used to aggregate the above-mentioned molecular features to obtain the atomic features corresponding to the above-mentioned amino acids; the above-mentioned second acquisition unit includes: a ninth acquisition module, which is used to obtain the interactive attribute information corresponding to the above-mentioned protein to be detected based on the above-mentioned atomic features.
[0018] According to another aspect of the embodiments of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above information acquisition method.
[0019] According to another aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the information acquisition method through the computer program.
[0020] In an embodiment of the present application, first-level features corresponding to a target component structure of a multi-level structure object are obtained, wherein the multi-level structure object is composed of multiple target component structures connected together, and the first-level features are features of the target component structure in a first dimension; the first-level features corresponding to each target component structure in the multiple target component structures are aggregated to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure in a second dimension, the target component structure is composed of multiple target component structures in the second dimension, and the target component structure in the second dimension is composed of multiple target component structures in the first dimension; based on the second-level features, interaction attribute information corresponding to the multi-level structure object is obtained, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects. By first obtaining features in one dimension and then using the features in the one dimension to obtain features in the next dimension, multi-dimensional reference features are provided for obtaining interaction attribute information, thereby achieving the purpose of using multi-dimensional reference features to obtain interaction attribute information in complex scenarios, thereby achieving the technical effect of improving the accuracy of information acquisition, and thus solving the technical problem of low accuracy of information acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0022] Figure 1 is a schematic diagram of an application environment of an optional information acquisition method according to an embodiment of the present application;
[0023] Figure 2 is a schematic diagram of a process of an optional information acquisition method according to an embodiment of the present application;
[0024] Figure 3is a schematic diagram of an optional information acquisition method according to an embodiment of the present application;
[0025] Figure 4 is a schematic diagram of another optional information acquisition method according to an embodiment of the present application;
[0026] Figure 5 is a schematic diagram of another optional information acquisition method according to an embodiment of the present application;
[0027] Figure 6 is a schematic diagram of another optional information acquisition method according to an embodiment of the present application;
[0028] Figure 7 is a schematic diagram of another optional information acquisition method according to an embodiment of the present application;
[0029] Figure 8 is a schematic diagram of another optional information acquisition method according to an embodiment of the present application;
[0030] Figure 9 is a schematic diagram of another optional information acquisition method according to an embodiment of the present application;
[0031] Figure 10 is a schematic diagram of an optional information acquisition device according to an embodiment of the present application;
[0032] Figure 11 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0036] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0037] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision, where cameras and computers replace the human eye in identifying and measuring objects, performing further image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0038] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0039] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0040] The solutions provided in the embodiments of this application involve artificial intelligence computer vision technology, machine learning and other technologies, which are specifically illustrated by the following embodiments:
[0041] According to one aspect of the embodiments of the present application, a method for obtaining information is provided. Optionally, as an optional implementation, the above information obtaining method can be applied to, but is not limited to, Figure 1 In the environment shown, the environment may include, but is not limited to, a user device 102 and a server 112 . The user device 102 may include, but is not limited to, a display 108 , a processor 106 , and a memory 1004 . The server 112 includes a database 114 and a processing engine 116 .
[0042] The specific process can be as follows:
[0043] Step S102 , the user device 102 obtains first-level features corresponding to a target component structure of a multi-level structure object;
[0044] Steps S104-S106, sending the first level features to the server 112 via the network 110;
[0045] In step S108, the server 112 aggregates the first-level features through the processing engine 116 to obtain second-level features corresponding to the target component structure, and further obtains interactive attribute information corresponding to the multi-level structure object based on the second-level features;
[0046] In steps S110 - S112 , the interaction attribute information is sent to the user device 102 via the network 110 . The user device 102 displays the interaction attribute information on the display 108 via the processor 106 and stores the interaction attribute information in the memory 104 .
[0047] remove Figure 1 In addition to the examples shown, the above steps can be completed with the assistance of a server, that is, the server performs steps such as aggregating the first-level features and obtaining interactive attribute information corresponding to the multi-level structure objects, thereby reducing the processing pressure on the server. The user device 102 includes but is not limited to a handheld device (such as a mobile phone), a laptop computer, a desktop computer, an in-vehicle device, etc., and this application does not limit the specific implementation of the user device 102.
[0048] Alternatively, as an optional implementation, Figure 2 As shown, the information acquisition method includes:
[0049] S202, obtaining first-level features corresponding to a target component structure of a multi-level structure object, wherein the multi-level structure object is formed by connecting multiple target component structures, and the first-level features are features of the target component structure in a first dimension;
[0050] S204: Aggregate the first-level features corresponding to each target component structure in the multiple target component structures to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure in a second dimension, the target component structure is composed of multiple target component structures in the second dimension, and the target component structure in the second dimension is composed of multiple target component structures in the first dimension;
[0051] S206 , based on the second-level features, obtaining interaction attribute information corresponding to the multi-level structure object, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects.
[0052] Optionally, in this embodiment, the above-mentioned information acquisition method can be applied to, but is not limited to, various prediction tasks of protein interactions, such as first representing the protein as a point cloud or k-nearest neighbor graph in 3D space, and then using a geometric neural network or a graph neural network to represent the point cloud or graph as a feature vector, and further obtaining the interaction property information of the protein based on the feature vector; the obtained interaction property information of the protein can also be applied to, but is not limited to, specific landing scenarios such as relatively fast and low-cost drug analysis and screening involving proteins.
[0053] Optionally, in this embodiment, the multi-level structural object can be understood as, but not limited to, a structural object composed of multiple levels of structures, such as a protein composed of amino acids, a media video composed of video frames, a neural network built by multiple structural layers, etc., and the multi-level structural object can also be understood as, but not limited to, an object in an application scenario that interacts with other structural objects. Predicting the performance of the object during interaction is a challenge to the accuracy of the information, which in turn creates a need for information to ensure a certain degree of accuracy.
[0054] Optionally, in this embodiment, the relationship between the multi-level structure object, the target component structure, the target component structure under the second dimension, and the target component structure under the first dimension may include, but is not limited to: the multi-level structure object is composed of multiple target component structures connected together, the target component structure is composed of multiple target component structures under the second dimension, and the target component structure under the second dimension is composed of multiple target component structures under the first dimension;
[0055] For example, proteins are composed of multiple amino acids, amino acids are composed of multiple molecules, and molecules are composed of atoms. Specifically, in proteins, amino acid molecules are chemically bonded and exist in the form of residues, so amino acid molecules can also be understood as amino acid residues. Further assuming that proteins are multi-level structural objects, amino acids can be understood, but not limited to, as target component structures, amino acid molecules can be understood, but not limited to, as target component structures in the first dimension (molecular dimension), and amino acid atoms can be understood, but not limited to, as target component structures in the second dimension (atomic dimension).
[0056] For example, a video is composed of multiple frames, a video frame is composed of multiple images, and an image is composed of multiple pixels. Assuming that a video is regarded as a multi-level structure object, the video frame can be understood as, but not limited to, the target component structure, the image can be understood as, but not limited to, the target component structure in the first dimension (image dimension), and the pixel can be understood as, but not limited to, the target component structure in the second dimension (pixel dimension).
[0057] To further illustrate, this embodiment can be optionally applied in a video processing scenario, such as obtaining pixel-level features (first-level features) corresponding to video frames (target component structures) of a video A (multi-level structural object); aggregating the frame-level features to obtain frame-level features (second-level features) corresponding to the video frames (target component structures); and obtaining interaction attribute information corresponding to a video (multi-level structural object) based on the frame-level features (second-level features), wherein the interaction attribute information is used to predict the performance of a video A (multi-level structural object) when interacting with other videos (such as video fusion, AI face-changing, etc.) (such as the degree of matching during video fusion, the degree of authenticity during AI face-changing, etc.).
[0058] Optionally, in this embodiment, since the target component structure in the second dimension is composed of multiple target component structures in the first dimension, the characteristics of the target component structure in the first dimension can be first obtained, and then the characteristics of the target component structure in the second dimension can be further inferred from the characteristics of the target component structure in the first dimension, thereby obtaining different characteristics in multiple dimensions, and further using the different characteristics in the multiple dimensions to obtain more accurate interactive attribute information;
[0059] Further in this embodiment, the first-level features corresponding to each target component structure in the multiple target component structures are aggregated to obtain the second-level features corresponding to the target component structures, which can be understood, but not limited to, as inferring the features of each target component structure in the second dimension based on the feature processing of each target component structure in the first dimension through aggregation processing; wherein, aggregation processing can be understood, but not limited to, as an operation of aggregating scattered objects together, and can also be understood, but not limited to, aggregating multiple objects according to preset conditions to obtain target objects in a collective form, such as chain aggregation, step-by-step aggregation, body aggregation, suspended aggregation, etc.
[0060] Optionally, in this embodiment, the first dimension and the second dimension can be understood as, but are not limited to, unit levels. For example, proteins appear as amino acid molecules at the molecular level, or proteins in the molecular dimension are amino acid molecules; similarly, proteins appear as amino acid atoms at the atomic level, or proteins in the atomic dimension are amino acid atoms.
[0061] Optionally, in this embodiment, interaction attribute information is used to predict the performance of a multi-level structural object when interacting with other structural objects, or it can be understood as limiting multi-level structural objects to structural objects that need to interact with other structural objects. In this embodiment, the purpose is to solve the technical problem of low accuracy in obtaining information corresponding to structural objects. However, not all structural objects have this technical problem. Therefore, this embodiment can be understood, but is not limited to, as the technical problem of low accuracy in obtaining interaction attribute information corresponding to multi-level structural objects. That is, multi-level structural objects have characteristic factors with more dimensions. If some dimensions are ignored during the information acquisition process, the acquired information will further lack analysis under the corresponding dimensions, reducing the accuracy of the information acquisition. As for interaction attribute information, compared with the information representation when the structural object exists statically, the information generated during interaction and dynamic representation is more complex, making it more difficult to ensure the accuracy of the information.
[0062] It should be noted that the method of first obtaining features in one dimension and then using the features in the one dimension to obtain features in the next dimension provides multi-dimensional reference features for the acquisition of interactive attribute information, and by adopting multi-dimensional reference features to obtain interactive attribute information in complex scenarios, the accuracy of obtaining interactive attribute information is improved.
[0063] To further illustrate, the optional Figure 3 As shown, the second component structure 304 (the target component structure under the second dimension) is composed of multiple first component structures 302 (the target component structure under the first dimension), the target component structure 306 is composed of multiple second component structures 304 (the target component structure under the second dimension), and the multi-level structure object 308 is composed of multiple target component structures 306 connected together;
[0064] Specifically, in this embodiment, the first-level features corresponding to the target component structure 306 of the multi-level structure object 308 are obtained, wherein the first-level features can also be understood as, but not limited to, the features of the first component structure 302 of the multi-level structure object 308; the first-level features corresponding to each target component structure 306 are aggregated to obtain the second-level features corresponding to the target component structure 306, wherein the second-level features can also be understood as, but not limited to, the features of the second component structure 304 of the multi-level structure object 308; based on the second-level features, the interaction attribute information corresponding to the multi-level structure object 308 is obtained, wherein the interaction attribute information is used to predict the performance of the multi-level structure object 308 when interacting with other structure objects.
[0065] Through the embodiments provided by the present application, first-level features corresponding to the target component structure of a multi-level structure object are obtained, wherein the multi-level structure object is composed of multiple target component structures connected together, and the first-level features are features of the target component structure under the first dimension; the first-level features corresponding to each target component structure in the multiple target component structures are aggregated to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure under the second dimension, the target component structure is composed of multiple target component structures under the second dimension, and the target component structure under the second dimension is composed of multiple target component structures under the first dimension; based on the second-level features, interaction attribute information corresponding to the multi-level structure object is obtained, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects, and by first obtaining features under one dimension and then using the features under the one dimension to obtain features under the next dimension, multi-dimensional reference features are provided for the acquisition of interaction attribute information, thereby achieving the purpose of using multi-dimensional reference features to obtain interaction attribute information in complex scenarios, thereby achieving the technical effect of improving the accuracy of information acquisition.
[0066] As an optional solution, obtaining the first-level features corresponding to the target component structure of the multi-level structure object includes:
[0067] S1, extracting image features from the target image corresponding to the multi-level structure object;
[0068] S2, obtaining initial first-level features corresponding to each first component structure in the target component structures under multiple first dimensions based on the image features;
[0069] S3, based on the initial first-level features, obtaining first-level features corresponding to each first component structure.
[0070] Optionally, in this embodiment, the initial first-level features may be, but are not limited to, independent features corresponding to each first component structure that can be directly obtained from the image features. The independent features may be, but are not limited to, being used only to represent the structural features of the first component structure itself, without involving relative features between different first component structures.
[0071] Optionally, in this embodiment, a graph-based representation learning method can be used, but is not limited to it, to use a graph modeling method to deeply understand the structural characteristics of multi-level structure objects, such as representing the multi-level structure objects as point clouds or k-nearest neighbor graphs in 3D space, and then using geometric neural networks or graph neural networks to represent the point clouds or graphs into vectors;
[0072] To further illustrate, optionally, for example, a target component structure is represented as a node of a graph, and the distance between target component structures is represented as an edge; or, the target component structure in the second dimension is represented as a node of a graph, and the distance between target component structures in the second dimension is represented as an edge; or, the target component structure in the first dimension is represented as a node of a graph, and the distance between target component structures in the first dimension is represented as an edge.
[0073] It should be noted that, whether it is a modeling method based on the target composition structure as the core, a modeling method based on the target composition structure under the second dimension as the core, or a modeling method based on the target composition structure under the first dimension as the core, it is often unable to well represent the rich content of the multi-level structure object in the multi-level dimensions. Therefore, in this embodiment, the target composition structure under the first dimension and the target composition structure under the second dimension are hierarchically represented from the two dimensions of the target composition structure under the first dimension and the target composition structure under the second dimension in the multi-level structure object.
[0074] Through the embodiments provided in the present application, image features in a target image corresponding to a multi-level structural object are extracted; initial first-level features corresponding to each first component structure in a target component structure under multiple first dimensions are obtained based on the image features; based on the initial first-level features, first-level features corresponding to each first component structure are obtained, thereby achieving the purpose of acquiring the first-level features in image form, thereby realizing the technical effect of improving the accuracy of acquiring the first-level features.
[0075] As an optional solution, based on the initial first-level features, obtaining the first-level features corresponding to each first component structure includes:
[0076] S1, obtaining a plurality of first component structure pairs consisting of respective first component structures;
[0077] S2, calculating first structure pair features of each first component structure pair in the plurality of first component structure pairs, wherein the first structure pair features include structural features corresponding to each first component structure in the first component structure pair and relative distance features between each first component structure;
[0078] S3, based on the first structure pair features, obtaining first-level features corresponding to each first component structure.
[0079] Optionally, in this embodiment, multiple first component structure pairs can be understood, but are not limited to, as non-repetitive combinations that can be formed by each first component structure. For example, assuming that the first component structure A, the first component structure B, and the first component structure C, the combinations that can be formed (first component structure pairs) include at least the first component structure A and the first component structure B, the first component structure B and the first component structure C, and the first component structure A and the first component structure C.
[0080] Optionally, in this embodiment, the structural features corresponding to each first component structure can be used, but are not limited to, to represent the independent structural properties of each first component structure itself, and the relative distance features between each first component structure can be used, but are not limited to, to represent the relative distance properties between each two first component structures.
[0081] It should be noted that, multiple first component structure pairs consisting of various first component structures are obtained; the first structure pair features of each first component structure pair in the multiple first component structure pairs are calculated, wherein the first structure pair features include the structural features corresponding to each first component structure in the first component structure pair, and the relative distance features between each first component structure; based on the first structure pair features, the first-level features corresponding to each first component structure are obtained.
[0082] To illustrate further, it is optional to assume that given the current M structural features, and as shown in the following formula (1), m ij(First structure pair feature) constructs the edge feature of the first component structure as a pass to the first component structure a i The news, m ij Contains the structural features of each first component structure pair, such as (a i ,a j ), and the relative distance characteristics of each first component structure pair, such as the distance square |c i -c j | 2 .
[0083]
[0084] As an optional solution, based on the first structure pair feature, obtaining the first-level features corresponding to each first component structure includes:
[0085] Perform the following steps until convergence is achieved:
[0086] S1, performing aggregation processing on the first structure pair features of each first component structure pair to obtain the target structure pair features;
[0087] S2, merging the target structure pair feature with the structural features corresponding to each first component structure to obtain a new structural feature corresponding to each first component structure;
[0088] S3-1, when a convergence condition is met, determining the new structural features corresponding to each first component structure as the first-level features corresponding to each first component structure;
[0089] S3-2. When the convergence condition is not met, the first structure pair features are recalculated using the new structural features corresponding to each first component structure to obtain new first structure pair features of each first component structure pair, and the new first structure pair features of each first component structure pair are determined as the first structure pair features of each first component structure pair, and the new atomic features corresponding to each first component structure are determined as the structural features corresponding to each first component structure.
[0090] Optionally, in this embodiment, the convergence condition may be, but is not limited to, related to the new structural feature. For example, if the new structural feature satisfies the feature condition, it means that the convergence condition is satisfied. The convergence condition may also be, but is not limited to, related to factors such as time and number of times. For example, if the number of convergences reaches 3 times, it means that the convergence condition is satisfied.
[0091] It should be noted that the following steps are performed until the convergence condition is met: aggregating the first structure pair features of each first component structure pair to obtain a target structure pair feature; merging the target structure pair feature with the structural features corresponding to each first component structure to obtain a new structural feature corresponding to each first component structure; when the convergence condition is met, determining the new structural feature corresponding to each first component structure as the first-level feature corresponding to each first component structure; when the convergence condition is not met, recalculating the first structure pair feature using the new structural feature corresponding to each first component structure to obtain a new first structure pair feature of each first component structure pair, and determining the new first structure pair feature of each first component structure pair as the first structure pair feature of each first component structure pair, and determining the new atomic feature corresponding to each first component structure as the structural feature corresponding to each first component structure;
[0092] Optionally, in this embodiment, if Figure 4 As shown, the first structure pair features of each first component structure pair are aggregated to obtain a target structure pair feature 404; the target structure pair feature 404 is merged with the structural features corresponding to each first component structure to obtain a new structural feature corresponding to each first component structure. Taking the first component structure pair 402 as an example, the target structure pair feature 404 is merged with the structural features corresponding to the first component structure A1 to obtain a new structural feature A2, and the target structure pair feature 404 is merged with the structural features corresponding to the first component structure B1 to obtain a new structural feature B2;
[0093] When the convergence condition is met, the new structural feature A2 is determined as the first-level feature corresponding to the first component structure A1, and the new structural feature B2 is determined as the first-level feature corresponding to the first component structure B1; when the convergence condition is not met, the new structural feature A2 is used as the structural feature corresponding to the first component structure A1, and the new structural feature B2 is used as the structural feature corresponding to the first component structure A1, and the new target structure pair feature is recalculated and obtained, and the new target structure pair feature and the new structural feature A2 are merged to obtain a new structural feature A3, and the new target structure pair feature and the new structural feature B2 are merged to obtain a new structural feature B3, and it is judged again whether the convergence condition is met.
[0094] Further examples are given, optionally based on the above formula (1), as shown in formula (2) and formula (3), specifically by aggregating the first structure pair features of each first component structure pair, the structural features of the first layer are obtained; when the convergence condition is not met, the first structure pair features of each first component structure pair are aggregated based on the structural features of the first layer, and the structural features of the previous layer are merged (i.e. the structural features of the first layer) to generate the structural features of a new layer If the convergence condition ends at this time, the atomic features of a new layer will be output As the first-level features corresponding to each first component structure.
[0095]
[0096]
[0097] As an optional solution, initial first-level features corresponding to each first component structure in the target component structures under multiple first dimensions are obtained based on the image features, including:
[0098] S1, obtaining first structural features corresponding to each first component structure and target structural features corresponding to the target component structure in the second dimension where each first component structure is located based on the image features;
[0099] S2, performing splicing processing on the first structural features and the target structural features to obtain second structural features corresponding to each first component structure, wherein the initial first-level features include the second structural features.
[0100] Optionally, in this embodiment, as shown in the following formula (4), each initial first-level feature It is determined by its structural characteristics and the target composition structure under the second dimension.
[0101]
[0102] Among them, I G (a i ) and I A (a i ) are the first component structure a under one-hot representation i The first component structure type and the second component structure type, the second component structure can be but is not limited to the target component structure under the second dimension, It is a vector concatenation operation.
[0103] As an optional solution, based on the second-level features, the interactive attribute information corresponding to the multi-level structure objects is obtained, including:
[0104] The second-level features corresponding to each target component structure in the multiple target component structures are summed up, and optionally based on the above formulas (1), (2), (3) and (4), the latest second-level features are obtained as shown in the following formula (5).
[0105]
[0106] Among them, the generation method of the initial representation can be optimized but not limited to according to different prediction tasks, such as only aggregating the representation of the skeleton first component structure (target component structure under the first dimension) in the second component structure (target component structure under the second dimension), using the self-attention mechanism to learn the aggregation weight of each first component structure (target component structure under the first dimension), or using the self-attention mechanism to select the top-k important first component structures (target component structure under the first dimension) for aggregation.
[0107] As an optional solution, based on the second-level features, obtaining the interaction attribute information corresponding to the multi-level structure object includes at least one of the following:
[0108] S1, obtaining a plurality of second component structure pairs consisting of respective second component structures in target component structures under a plurality of second dimensions; calculating a first target structure feature of each second component structure pair in the plurality of second component structure pairs based on the second-level features, wherein the first target structure feature is used to represent the feature inner product similarity between each second component structure in the second component structure pair; obtaining interaction attribute information based on the first target structure feature;
[0109] S2, obtaining a plurality of second component structure pairs; calculating a second target structure feature for each second component structure pair based on the second-level features, wherein the second target structure feature is used to represent the relative distance between each second component structure when linearly mapped on the sequence; and obtaining interaction attribute information based on the second target structure feature;
[0110] S3, obtaining multiple second component structure pairs; calculating the third target structure feature of each second component structure pair based on the second-level feature, wherein the third target structure feature is used to represent the carbon atom distance between each second component structure; and obtaining interaction attribute information based on the third target structure feature.
[0111] It should be noted that a plurality of second component structure pairs consisting of respective second component structures in target component structures under a plurality of second dimensions are obtained; based on the second-level features, a first target structure feature of each second component structure pair in the plurality of second component structure pairs is calculated, wherein the first target structure feature is used to represent the feature inner product similarity between each second component structure in the second component structure pair; and interaction attribute information is obtained based on the first target structure feature;
[0112] To further illustrate, it is optional to assume that the second component structure is G i ,G j , and then G i ,G j The feature similarity bias between is calculated by the following formula (6):
[0113]
[0114] Among them, β ij By calculating G i ,G j The feature inner product similarity is obtained.
[0115] It should be noted that a plurality of second component structure pairs are obtained; based on the second-level features, a second target structure feature of each second component structure pair is calculated, wherein the second target structure feature is used to represent the relative distance between each second component structure when linearly mapped on the sequence; and interaction attribute information is obtained based on the second target structure feature;
[0116] To further illustrate, it is optional to assume that the second component structure is G i ,G j , and then G i ,G j The relative distance between the sequences after linear mapping is calculated by the following formula (7):
[0117]
[0118] Among them, z ij G i ,G j The relative distance on the sequence after linear mapping.
[0119] It should be noted that a plurality of second component structure pairs are obtained; based on the second-level features, a third target structure feature of each second component structure pair is calculated, wherein the third target structure feature is used to represent the carbon atom distance between each second component structure; and interaction attribute information is obtained based on the third target structure feature;
[0120] To further illustrate, it is optional to assume that the second component structure is G i ,G j , and then G i ,G j The distance between the beta carbon atoms is calculated by the following formula (8):
[0121]
[0122] in, G i ,G j The distance between the beta carbon atoms.
[0123] As an optional solution, based on the second-level features, the interactive attribute information corresponding to the multi-level structure objects is obtained, including:
[0124] S1, when a first target structural feature, a second target structural feature, and a third target structural feature are obtained, performing aggregation processing on the first target structural feature, the second target structural feature, and the third target structural feature to obtain an aggregated structural feature;
[0125] S2, through vector splicing, merges the aggregated structural features and the second-level features to obtain the output structural features;
[0126] S3, obtains interaction attribute information based on the output structure features.
[0127] Optionally, in order to consider multiple characteristics, in this embodiment, based on the above formulas (6), (7) and (8), as shown in the following formula (9), the three characteristics of feature similarity bias, spatial distance similarity bias, and sequence relative distance bias are integrated, and the final propagation attention weight α is calculated through the Softmax function. ij , and then use the propagation attention weight α ij Output interactive attribute information:
[0128]
[0129] Optionally, in this embodiment, to improve the output accuracy of the interactive attribute information, a graph attention mechanism may be used, but is not limited to, to add learned weights to the acquisition of the interactive attribute information. Specifically, an Attention Bias mechanism may be used to add multiple similarities as biases to the attention module.
[0130] To further illustrate, optionally based on the above formula (9), continue as shown in the following formulas (10), (11), (12), and (13), the K-layer message passing transmits the first target structure feature, the second target structure feature, and the third target structure feature to propagate the attention weight α ij Aggregate the weights separately to generate corresponding aggregate structural features, such as u i ,v i ,w i ; Then through vector splicing and structural features of the current layer Merge and generate output features through network φ3
[0131]
[0132]
[0133]
[0134]
[0135] Among them, when calculating the aggregation of spatial features in the above formula (12), G is used i The orientation matrix O i , can be calculated by, but not limited to, the following formula (13) and formula (14), where × is the vector outer product:
[0136]
[0137]
[0138] As an optional solution, based on the initial first-level features, obtaining first-level features corresponding to each first component structure includes: inputting the initial first-level features into a first target model, and performing information transfer through a first substructure in the first target model to obtain first-level features corresponding to each first component structure, wherein the first target model is a neural network model trained using the first sample and used for feature processing;
[0139] As an optional solution, the first-level features corresponding to each target component structure in multiple target component structures are aggregated to obtain the second-level features corresponding to the target component structure, including: aggregating the first-level features corresponding to each first component structure through the second substructure in the first target model to obtain the second-level features corresponding to the target component structure.
[0140] As an optional solution, based on the second-level features, the interactive attribute information corresponding to the multi-level structure objects is obtained, including:
[0141] S1, inputting the second-level features into the second target model, transferring information through the third substructure of the second target model, and obtaining input structure features corresponding to each second component structure, wherein the second target model is a neural network model trained using the second sample and used for feature processing;
[0142] S2, through the fourth substructure of the second target model, predict and output the input structure features corresponding to each second component structure to obtain interactive attribute information.
[0143] Optionally, in this embodiment, to improve the efficiency of information acquisition, a neural network model may be used to participate in the process, but is not limited to;
[0144] To further illustrate, alternatively, for example Figure 5As shown, the first target model 502 is an Equivalent Graph Neural Network (EGNN), and the second target model 504 is a Geometric Attention Network (GAN). EGNN accepts two inputs, and the features of all the first component structures are and the corresponding coordinates c, such as the first-level features and coordinates; the first-level features are obtained through K-layer message passing Such as the updated first-level features; secondly, the summation Readout layer aggregates each updated first-level feature to generate the corresponding second-level features Such as second-level feature 1, second-level feature 2, third-level feature, etc. Finally, GAN receives the aggregated second-level feature The relative distance z of the second structural pair in the sequence ij , the coordinates of the β carbon atom of the second component structure are taken as input, and the updated second-level features are obtained through K-layer message passing Among them, the updated second-level features It can be used for prediction of specific tasks through the downstream Readout layer to obtain prediction results;
[0145] Specifically, in this embodiment, based on Figure 5 The scenario shown, continuing with e.g. Figure 6 As shown, the first target model 502 is a three-layer EGNN neural network architecture, which extracts the input coordinates and feature 1 to obtain relative distance, Src (sparse expression), and Dst (combined expression), and inputs the feature extraction result into the first MLP to obtain multiple side information; the above multiple side information is aggregated to obtain aggregated information, which is then input into the second MLP to obtain feature 2; further through aggregation and other processing methods, feature 3 is obtained, and when the iteration end condition is not met at present, the iteration is started from the beginning, wherein the two MLPs can be understood as, but not limited to, φ in the above formula (1). e and φ in the above formula (3) h .
[0146] Optionally, in this embodiment, based on Figure 5 The scenario shown, continuing with e.g. Figure 7As shown, the second target model 504 is a three-layer GAN neural network architecture, which first obtains input information such as features, distances, and sequences, integrates the similarities of the above-mentioned multiple input information, and calculates the corresponding propagation attention weights through the Softmax function; then uses the propagation attention weights to perform weight distribution processing on each input information to obtain new input information, such as features*, distances*, sequences*, etc.; further, the new input information such as features*, distances*, sequences*, etc. is merged to obtain target features, and based on the target features, it is determined whether the convergence conditions are met. If not, the above steps are continued to be iteratively executed. If so, the target features are output.
[0147] As an optional solution, before obtaining the first-level features corresponding to the target component structure of the multi-level structure object, the method includes:
[0148] S1, obtaining the number of target structures of a multi-level structure object, wherein the number of target structures includes at least one of the following: a first number corresponding to a target component structure in a plurality of target component structures, a second number corresponding to a second component structure in a plurality of target component structures under a second dimension, and a third number corresponding to a first component structure in a plurality of target component structures under a first dimension;
[0149] S2, based on the number of target structures, calculates the acquisition complexity of the interaction attribute information, where the acquisition complexity is used to represent the complexity of the operations to be performed when acquiring the interaction attribute information.
[0150] Optionally, in this embodiment, it is assumed that the multi-level structure object contains L target component structures in the second dimension and M target component structures in the first dimension, and it is assumed that a neural network model is used to participate in the process, and the complexity of vector addition, multiplication, concatenation and inner product calculation of fixed dimensions in the neural network model is constant. Further, for the feature representation learning in the first dimension, the computational complexity of the above formula (1), formula (2) and formula (3) is O(M 2 ), the complexity of K1 layer iteration is O(K1M 2 ), the aggregation complexity of the above formula (5) is O(M). Similarly, in the aggregation calculation under the second dimension in the above formulas (6), (7), (8), (9), (10), (11), and (12), the complexity of the K2 layer iteration is O(K2L 2 ), and finally, the total computational complexity is O(K1M 2 +K2L 2 );
[0151] Optionally, in this embodiment, during the participation of the EGNN model, the computational interaction complexity is relatively high (O(K1M 2), we can consider only calculating the message passing of the k nearest neighbor atoms, and balance the modeling effect of the time and space complexity of the model to improve the efficiency of information acquisition.
[0152] As an optional solution, based on the second-level features, the interactive attribute information corresponding to the multi-level structure objects is obtained, including:
[0153] Based on the second-level features and the physical and chemical features corresponding to each first component structure in the target component structures under multiple first dimensions, interactive attribute information is obtained, wherein the physical and chemical features include at least one of the following: a first feature for representing the hydrophilicity of the target substructure in the first component structure that does not participate in dehydration bonding, a second feature for representing the hydrophobicity of the target substructure, and a third feature for representing the charge carried by the target substructure.
[0154] Optionally, in this embodiment, taking the target composition structure including amino acids as an example, in the modeling of atomic particle size and amino acid particle size, the physicochemical characteristics of amino acids, such as the hydrophilicity and hydrophobicity of the residues and the charge of the residues, can be additionally considered.
[0155] As an optional solution, obtaining first-level features corresponding to target constituent structures of a multi-level structural object includes: obtaining molecular features corresponding to amino acids in a protein to be detected;
[0156] As an optional solution, aggregating the first-level features corresponding to each target component structure in the multiple target component structures to obtain the second-level features corresponding to the target component structure, including: aggregating the molecular features to obtain the atomic features corresponding to the amino acids;
[0157] As an optional solution, based on the second-level features, the interactive attribute information corresponding to the multi-level structure object is obtained, including: based on the atomic features, obtaining the interactive attribute information corresponding to the protein to be detected.
[0158] Optionally, in this embodiment, the above-mentioned information acquisition method is applied to various prediction tasks of protein interactions, such as first representing the protein as a point cloud or k-nearest neighbor graph in 3D space, and then using a geometric neural network or a graph neural network to represent the point cloud or graph as a feature vector, and further obtaining the interaction property information of the protein based on the feature vector; the obtained interaction property information of the protein can also be applied to, but is not limited to, specific landing scenarios involving relatively fast and low-cost drug analysis and screening of proteins.
[0159] Optionally, in this embodiment, a multi-level structure object can be understood as, but not limited to, a structure object composed of multiple levels of structures, such as a protein composed of amino acids, a media video composed of video frames, a neural network constructed by multiple structural layers, etc., and a multi-level structure object can also be understood as, but not limited to, an object in an application scenario that interacts with other structure objects. Predicting the performance of the object during interaction is a challenge to the accuracy of the information, which in turn creates a need for information to ensure a certain degree of accuracy.
[0160] To further illustrate, proteins are optionally composed of multiple amino acids, amino acids are composed of multiple molecules, and molecules are composed of atoms. Assuming that proteins are taken as multi-level structural objects, amino acids can be understood as, but not limited to, target composition structures, amino acid molecules can be understood as, but not limited to, target composition structures in the first dimension (molecular dimension), and amino acid atoms can be understood as, but not limited to, target composition structures in the second dimension (atomic dimension).
[0161] Optionally, in this embodiment, amino acids are the basic building blocks of proteins. Assume that an amino acid molecule G i Represented as a 3D point cloud, namely G i =(A i ,C i ). Among them, A i ={a1,a2,…a m} is the atom in amino acid, C i ={c1,c2,…c m} is the three-dimensional coordinate of the atom;
[0162] Alternatively, in this embodiment, a protein is composed of one or more amino acid sequences. It refers to a protein composed of an amino acid sequence, or a protein composed of multiple amino acid sequences;
[0163] Further Figure 8 As shown in Figure 1, a protein can be represented as a two-level 3D point cloud: the first level is an atomic-level point cloud consisting of all atoms in all amino acids, and the second level is a point cloud consisting of amino acids, where the spatial positions of amino acids are oriented by the coordinates of backbone carbon atoms.
[0164] It should be noted that this embodiment proposes to use graph neural networks to hierarchically represent the interactions between atoms and amino acids in protein structures at two granularities: atoms and amino acids. For details, please refer to Figure 9 As shown in Figure 2. The modeling process mainly consists of the following steps:
[0165] Step S1: Modeling each atom in the protein, using an equivariant graph neural network to model all interactions between atoms, ultimately generating a unique representation for each atom. Step S1 aims to encode all interactions and spatial relationships between atoms, such as possible forces and bonds, and distances in space, including atoms that make up the same amino acid and atoms in different amino acids.
[0166] Step S2, aggregating the representations after atomic interactions to generate an initial representation of the amino acid molecule;
[0167] In step S3, the geometric attention network is used to further represent the interactions and spatial relationships between amino acid molecules, combining the relative distances between amino acids on the chain and their distances in three-dimensional space. This ultimately generates an (L, d) matrix representing the protein, where L is the number of amino acids and d is the dimension of the representation vector.
[0168] Specifically, in this embodiment, the two main components of the neural network architecture are EGNN and GAN. First, EGNN accepts two inputs, all atomic features and the corresponding atomic coordinates c, and the updated atomic features are obtained through K-layer message passing Secondly, the summation readout layer aggregates the atomic features in each amino acid to generate the corresponding amino acid features Finally, GAN receives the aggregated amino acid features The relative distance z between amino acid pairs in the sequence ij , the coordinates of the amino acid's β carbon atom are used as input, and the updated amino acid features are obtained through K-layer message passing All amino acid features can be used for prediction of specific tasks through the downstream Readout layer;
[0169] Furthermore, based on the atomic granularity representation learning of EGNN, in this embodiment, given the current M atomic features m ij The edge feature is constructed (refer to the above formula 1) as the edge feature passed to atom a i Specifically, m ij Contains any atom pair in the protein (a i ,a j ) characteristics and the square of the distance between atomic pairs |c i -c j | 2 By aggregating all the atoms in all amino acids (refer to the above formula 2) and merging the atomic features of the previous layer (Refer to formula 3 above) Generate a new layer of atomic features
[0170] Initial characteristics of each atom It is determined by its atomic characteristics and the characteristics of the amino acid in which it is located (refer to the above formula 4). G (a i ) and I A (a i ) are atoms a in one-hot representation i The amino acid type and atom type, It is a vector concatenation operation.
[0171] In generating the final atomic features Then, the atomic features in each amino acid are summed (refer to the above formula 5) to generate the corresponding amino acid features.
[0172] Optionally, the amino acid molecule granularity representation learning based on GAN is used. In this embodiment, after obtaining the representation of L amino acids, The interaction between amino acids needs to be further represented. Since amino acids are the basic building blocks of proteins, it is necessary to more finely encode the feature similarity, spatial distance and relative distance between amino acids in the protein sequence. This embodiment uses the graph attention mechanism to add learned weights to the message passing-aggregation mechanism of amino acids. Since it is necessary to consider multiple aspects of information (feature similarity, spatial distance and relative distance between amino acids in the protein sequence), the Attention bias mechanism is used to add multiple similarities as bias to the attention module. Specifically, the amino acid G i ,G j The feature similarity bias, spatial distance similarity bias, and sequence relative distance bias are determined by β ij , γ ij ,δ ij Calculated separately. ij By calculating the similarity of the inner product of the characteristics of two amino acids, z ij is the relative distance on the sequence after linear mapping. is the distance between the beta carbon atoms of the two amino acids. Further integrating the similarities of these three aspects (refer to the above formula 9), the final propagation attention weight α is calculated through the Softmax function ij .
[0173] K-layer message passing combines amino acid features, spatial distance features, and sequence distance features with α ij Aggregate the weights separately to generate the corresponding features u i ,v i ,w i , by vector concatenation and current layer amino acid features Merge and generate output features through network φ3 (Please refer to the above formula 13).
[0174] Among them, when calculating the aggregation of spatial features in the above formula 12, the amino acid G is used i The orientation matrix O i , can be calculated by the following formula, where × is the vector outer product. For details, please refer to the above formula 14 and formula 15.
[0175] Alternatively, assume that there are L amino acids and M atoms in a protein. For simplicity, assume that the computational complexity of fixed-dimensional vector addition, multiplication, concatenation, and inner product in a neural network is constant. For atomic-scale representation learning, the computational complexity of formulas (1) to (3) is O(M 2 ), the complexity of K1 layer iteration is O(K1M 2 ), the aggregation complexity of the above formula (5) is O(M). Similarly, in the aggregation calculation of amino acids in the above formulas (6) to (12), the complexity of the K2 layer iteration is O(K2L 2 ), finally, the total computational complexity is O(K1M 2 +K2L 2 ).
[0176] Optionally, in this embodiment, the graph-based representation learning method models at the atomic granularity or at the amino acid granularity. The atomic granularity method aims to extract the features of the interactions and spatial relationships between atoms on the protein interaction surface. Similarly, the amino acid granularity method aims to extract the features of the interactions and spatial relationships between amino acids on the protein interaction surface. However, modeling at a single granularity ignores the multi-level structural representation of the protein itself and cannot simultaneously capture the complex features of the multi-level structure at multiple granularities.
[0177] In addition, graph-based modeling methods use traditional graph neural networks based on message passing. Generally speaking, the initial features of atoms or amino acids, including atom or amino acid types, physicochemical characteristics, and 3D coordinates, are transformed through neural networks and message passing to achieve feature extraction and fusion. Traditional graph neural networks have been widely used in general graph data. However, when directly applied to 3D graph data, an important drawback is that the feature transformation of 3D coordinates is variable. That is, after the atomic coordinates are translated, rotated, rearranged, and other transformations, the vector representation of the atoms or amino acids generated by the model in Euclidean space cannot be guaranteed to be unchanged. This is inconsistent with the inductive prior required for modeling protein interactions, which greatly affects the modeling effect;
[0178] In this example, considering the shortcomings of existing graph-based representation learning methods, this example aims to hierarchically model the interactions between atoms and amino acids in protein structures at both atomic and amino acid granularity levels, while ensuring that the resulting protein representation is invariant to translation, rotation, and rearrangement in 3D space. This can serve as a basic component in deep neural networks to achieve attribute prediction tasks for various protein interactions.
[0179] To illustrate further, the optional hypothetical neural network function The space The input x in the space is mapped to The output y in the transformation of x If φ(T g (x))=φ(x), it is said that φ is related to the transformation T g Invariance. For m inputs in n-dimensional Euclidean space This embodiment aims to propose a model that maintains the following three invariants:
[0180] 1. Translation invariance: T g is specified by an arbitrary n-dimensional translation vector, T g (X)={x1+g,…,x m +g}, abbreviated as X+g, the translation invariance is φ(X+g)=φ(X);
[0181] 2. Rotational invariance: T g is an arbitrary orthogonal matrix Specify, T g (X) = {Qx1, ..., Qx m}, abbreviated as QX, translation invariance is φ(QX)=φ(X);
[0182] 3. Permutation invariance: T g is specified by an arbitrary permutation of length m, i.e. T g (X)=P(X), and the permutation invariance is φ(P(X))=φ(X).
[0183] Optionally, in this embodiment, the invariance to atomic coordinate translation and rotation and the invariance to atomic arrangement are achieved because:
[0184] |c i +g-(c j +g)| 2 =|c i -c j | 2
[0185] |Qc i -Qc j |2 =|c i -c j | T Q T Q|c i -c j |=|c i -c j | 2
[0186] Among them, m in the above formula (1) ij The summation operation in formula (4) is invariant to the arrangement of atoms, so the initial representation of the amino acid is It is invariant to the translation and rotation of atomic coordinates and invariant to the arrangement of atoms.
[0187] Secondly, we prove that GAN is invariant to atomic coordinate translation and rotation. Due to the invariance of the basis vector (v1, v2) to atomic coordinate translation, the orientation matrix O remains unchanged to translation. Since:
[0188]
[0189] It can be seen that w i It remains unchanged for translation transformation. For any orthogonal transformation Q, it can be proved that the orientation matrix O is equivariant to rotation:
[0190] Q N -Qc=Q(c N -c)=Qv1
[0191] Q C -Qc=Q(c C -c)=Qv2
[0192]
[0193] Therefore, we can get w i Invariant to rotations:
[0194]
[0195] From this we can get the final output of HGIN, It is invariant to the translation and rotation of atomic coordinates and invariant to the arrangement of atoms.
[0196] Through the embodiments provided in the present application, molecular features corresponding to amino acids in the protein to be detected are obtained; the molecular features are aggregated to obtain atomic features corresponding to the amino acids; based on the atomic features, interaction property information corresponding to the protein to be detected is obtained, and the protein is modeled at both the atomic and amino acid levels by representing the protein as a hierarchical 3D graph. The generated protein representation vector is invariant to the translation, rotation, and atomic arrangement of the input atomic coordinate system, thereby achieving the purpose of embedding the above-mentioned information acquisition method as a basic representation module in the learning task of predicting protein interaction properties (such as affinity) and changes in interaction properties into a more complex neural network, thereby achieving the technical effect of improving the accuracy of predicting protein interaction properties.
[0197] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0198] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0199] According to another aspect of the embodiments of the present application, an information acquisition device for implementing the above-mentioned information acquisition method is also provided. Figure 10 As shown, the device includes:
[0200] A first acquisition unit 1002 is configured to acquire first-level features corresponding to a target component structure of a multi-level structure object, wherein the multi-level structure object is composed of multiple target component structures connected together, and the first-level features are features of the target component structure in a first dimension;
[0201] an aggregation unit 1004 for aggregating first-level features corresponding to each target component structure in the plurality of target component structures to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure in a second dimension, the target component structure is composed of the plurality of target component structures in the second dimension, and the target component structure in the second dimension is composed of the plurality of target component structures in the first dimension;
[0202] The second acquiring unit 1006 is configured to acquire interaction attribute information corresponding to the multi-level structure object based on the second-level features, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects.
[0203] For specific embodiments, reference may be made to the examples shown in the above-mentioned information acquisition device, which will not be described in detail in this example.
[0204] As an optional solution, the first acquiring unit 1002 includes:
[0205] An extraction module, used for extracting image features in a target image corresponding to a multi-level structure object;
[0206] A first acquisition module is configured to obtain, based on the image features, initial first-level features corresponding to each first component structure in the target component structures under a plurality of first dimensions;
[0207] The second acquisition module is used to acquire the first-level features corresponding to each first component structure based on the initial first-level features.
[0208] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0209] As an optional solution, the second acquisition module includes:
[0210] A first acquisition submodule, configured to acquire a plurality of first component structure pairs consisting of respective first component structures;
[0211] a calculation submodule, configured to calculate first structure pair features of each first component structure pair in the plurality of first component structure pairs, wherein the first structure pair features include structural features corresponding to each first component structure in the first component structure pair and relative distance features between each first component structure;
[0212] The second acquisition submodule is used to acquire the first-level features corresponding to each first component structure based on the first structure pair features.
[0213] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0214] As an optional solution, the second acquisition submodule includes:
[0215] The execution subunit is used to perform the following steps until the convergence condition is reached:
[0216] Aggregating the first structure pair features of each first component structure pair to obtain target structure pair features;
[0217] Merging the target structure pair feature with the structural features corresponding to each first component structure to obtain a new structural feature corresponding to each first component structure;
[0218] When the convergence condition is met, the new structural features corresponding to the respective first component structures are determined as the first-level features corresponding to the respective first component structures;
[0219] When the convergence conditions are not met, the first structure pair features are recalculated using the new structural features corresponding to each first component structure to obtain new first structure pair features of each first component structure pair, and the new first structure pair features of each first component structure pair are determined as the first structure pair features of each first component structure pair, and the new atomic features corresponding to each first component structure are determined as the structural features corresponding to each first component structure.
[0220] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0221] As an optional solution, the first acquisition module includes:
[0222] A third acquisition submodule is configured to obtain, based on the image features, first structural features corresponding to each first component structure and target structural features corresponding to the target component structure in the second dimension where each first component structure is located;
[0223] The splicing submodule is used to splice the first structural features and the target structural features to obtain the second structural features corresponding to each first component structure, wherein the initial first-level features include the second structural features.
[0224] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0225] As an optional solution, the second obtaining unit 1006 includes at least one of the following:
[0226] a third acquisition module, configured to acquire a plurality of second component structure pairs consisting of respective second component structures in the target component structures under the plurality of second dimensions; calculate, based on the second-level features, a first target structure feature of each second component structure pair in the plurality of second component structure pairs, wherein the first target structure feature is used to represent the feature inner product similarity between each second component structure in the second component structure pair; and acquire interaction attribute information based on the first target structure feature;
[0227] a fourth acquisition module, configured to acquire a plurality of second component structure pairs; calculate a second target structure feature for each second component structure pair based on the second-level features, wherein the second target structure feature is used to represent the relative distance between each second component structure when linearly mapped on the sequence; and acquire interaction attribute information based on the second target structure feature;
[0228] The fifth acquisition module is used to obtain multiple second component structure pairs; based on the second-level features, calculate the third target structure features of each second component structure pair, wherein the third target structure features are used to represent the carbon atom distance between each second component structure; and obtain interactive attribute information based on the third target structure features.
[0229] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0230] As an optional solution, the second obtaining unit 1006 includes:
[0231] a first aggregation module configured to aggregate the first target structural feature, the second target structural feature, and the third target structural feature to obtain an aggregated structural feature when the first target structural feature, the second target structural feature, and the third target structural feature are acquired;
[0232] The merging module is used to merge the aggregated structural features and the second-level features through vector splicing to obtain the output structural features;
[0233] The sixth acquisition module is used to acquire interaction attribute information based on the output structure features.
[0234] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0235] As an optional solution, the second acquisition module includes: a first input submodule, configured to input the initial first-level features into the first target model, and to obtain the first-level features corresponding to each first component structure through information transmission through the first substructure in the first target model, wherein the first target model is a neural network model trained using the first sample and used for feature processing;
[0236] The aggregation unit 1004 includes: a second aggregation module, configured to aggregate the first-level features corresponding to each first component structure through the second substructure in the first target model to obtain the second-level features corresponding to the target component structure.
[0237] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0238] As an optional solution, the second obtaining unit 1006 includes:
[0239] an input module, configured to input the second-level features into a second target model, and to transmit information through the third substructure of the second target model to obtain input structural features corresponding to each second component structure, wherein the second target model is a neural network model trained using the second sample and used for feature processing;
[0240] The output module is used to predict and output the input structure features corresponding to each second component structure through the fourth substructure of the second target model to obtain interactive attribute information.
[0241] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0242] As an optional solution, the device includes:
[0243] a third acquiring unit, configured to acquire a target structure quantity of the multi-level structure object before acquiring the first-level features corresponding to the target component structure of the multi-level structure object, wherein the target structure quantity includes at least one of the following: a first quantity corresponding to the target component structure in the plurality of target component structures, a second quantity corresponding to the second component structure in the plurality of target component structures under the second dimension, and a third quantity corresponding to the first component structure in the plurality of target component structures under the first dimension;
[0244] The calculation unit is used to calculate the acquisition complexity of the interaction attribute information based on the number of target structures, wherein the acquisition complexity is used to represent the complexity of the operation to be performed when acquiring the interaction attribute information.
[0245] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0246] As an optional solution, the second obtaining unit 1006 includes:
[0247] The seventh acquisition module is used to obtain interactive attribute information based on the second-level features and the physical and chemical features corresponding to each first component structure in the target component structures under multiple first dimensions, wherein the physical and chemical features include at least one of the following: a first feature for indicating the hydrophilicity of the target substructure in the first component structure that does not participate in dehydration bonding, a second feature for indicating the hydrophobicity of the target substructure, and a third feature for indicating the charge carried by the target substructure.
[0248] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0249] As an optional solution, the first acquisition unit 1002 includes: an eighth acquisition module for acquiring molecular features corresponding to amino acids in the protein to be detected;
[0250] The aggregation unit 1004 includes: a third aggregation module for performing aggregation processing on the molecular features to obtain atomic features corresponding to the amino acids;
[0251] The second acquisition unit 1006 includes: a ninth acquisition module, configured to acquire interaction attribute information corresponding to the protein to be detected based on the atomic features.
[0252] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0253] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned information acquisition method is also provided. Figure 11 As shown, the electronic device includes a memory 1102 and a processor 1104. The memory 1102 stores a computer program, and the processor 1104 is configured to execute the steps in any of the above method embodiments through the computer program.
[0254] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0255] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0256] S1, obtaining the first-level features corresponding to the target component structure of the multi-level structure object, wherein the multi-level structure object is composed of multiple target component structures connected, and the first-level features are the features of the target component structure in the first dimension;
[0257] S2, aggregating the first-level features corresponding to each target component structure in the multiple target component structures to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure in the second dimension, the target component structure is composed of multiple target component structures in the second dimension, and the target component structure in the second dimension is composed of multiple target component structures in the first dimension;
[0258] S3, based on the second-level features, obtaining interaction attribute information corresponding to the multi-level structure object, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects.
[0259] Alternatively, those skilled in the art will appreciate that Figure 11The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 11 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 11 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 11 Different configurations shown.
[0260] Among them, the memory 1102 can be used to store software programs and modules, such as program instructions / modules corresponding to the information acquisition method and device in the embodiments of the present application. The processor 1104 executes various functional applications and data processing by running the software programs and modules stored in the memory 1102, that is, realizing the above-mentioned information acquisition method. The memory 1102 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1102 may further include a memory remotely located relative to the processor 1104, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1102 can be used to store, but is not limited to, information such as multi-level structural objects, target component structures, first-level features, second-level features, and interactive attribute information. As an example, such as Figure 11 As shown, the memory 1102 may include, but is not limited to, the first acquisition unit 1002, aggregation unit 1004, and second acquisition unit 1006 in the information acquisition device. In addition, it may also include, but is not limited to, other module units in the information acquisition device, which will not be repeated in this example.
[0261] Optionally, the transmission device 1106 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1106 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1106 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0262] In addition, the above-mentioned electronic device also includes: a display 1108, used to display information such as the above-mentioned multi-level structural objects, target component structures, first-level features, second-level features, and interactive attribute information; and a connection bus 1110, used to connect the various module components in the above-mentioned electronic device.
[0263] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.
[0264] According to one aspect of the present application, a computer program product is provided, comprising a computer program / instructions containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions provided in the embodiments of the present application are performed.
[0265] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0266] It should be noted that the computer system of the electronic device is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0267] A computer system includes a central processing unit (CPU), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from the storage unit into random access memory (RAM). The RAM also stores various programs and data required for system operation. The CPU, the read-only memory, and the RAM are connected to each other via a bus. Input / output interfaces (I / O interfaces) are also connected to the bus.
[0268] The following components are connected to the input / output interface: an input section including a keyboard, mouse, etc.; an output section including a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section including a hard disk; and a communication section including a network interface card such as a local area network card and a modem. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the input / output interface as needed. Removable media such as magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc. are installed in the drive as needed so that computer programs read from them can be installed into the storage section as needed.
[0269] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication portion, and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions defined in the system of the present application are performed.
[0270] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementations described above.
[0271] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0272] S1, obtaining the first-level features corresponding to the target component structure of the multi-level structure object, wherein the multi-level structure object is composed of multiple target component structures connected, and the first-level features are the features of the target component structure in the first dimension;
[0273] S2, aggregating the first-level features corresponding to each target component structure in the multiple target component structures to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure in the second dimension, the target component structure is composed of multiple target component structures in the second dimension, and the target component structure in the second dimension is composed of multiple target component structures in the first dimension;
[0274] S3, based on the second-level features, obtaining interaction attribute information corresponding to the multi-level structure object, wherein the interaction attribute information is used to predict the performance of the multi-level structure object when interacting with other structure objects.
[0275] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0276] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0277] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.
[0278] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0279] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0280] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0281] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0282] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. An information acquisition method, characterized in that: include: Obtaining a first-level feature corresponding to a target component structure of a multi-level structure object, wherein the multi-level structure object is composed of multiple target component structures connected together, the first-level feature is a feature of the target component structure in a first dimension, and the multi-level structure object is a protein composed of amino acids; aggregating first-level features corresponding to each target component structure in the multiple target component structures to obtain second-level features corresponding to the target component structure, wherein the second-level features are features of the target component structure in a second dimension, the target component structure is composed of multiple target component structures in the second dimension, and the target component structure in the second dimension is composed of multiple target component structures in the first dimension; Based on the second-level features, interaction attribute information corresponding to the multi-level structure object is obtained, wherein the interaction attribute information is used to predict performance of the multi-level structure object when interacting with other structure objects.
2. The method according to claim 1, characterized in that The obtaining of the first-level features corresponding to the target component structure of the multi-level structure object includes: Extracting image features in a target image corresponding to the multi-level structure object; Obtaining, based on the image features, initial first-level features corresponding to each first component structure in a plurality of target component structures under the first dimension; Based on the initial first-level features, first-level features corresponding to the respective first component structures are obtained.
3. The method according to claim 2, characterized in that The acquiring, based on the initial first-level features, the first-level features corresponding to the respective first component structures includes: Acquire a plurality of first component structure pairs consisting of the first component structures; Calculating a first structure pair feature of each first component structure pair in the plurality of first component structure pairs, wherein the first structure pair feature includes a structure feature corresponding to each first component structure in the first component structure pair and a relative distance feature between each first component structure; Based on the first structure pair features, first-level features corresponding to the respective first component structures are obtained.
4. The method according to claim 3, characterized in that The acquiring, based on the first structure pair features, the first-level features corresponding to the respective first component structures includes: Perform the following steps until convergence is achieved: Aggregating the first structure pair features of each of the first component structure pairs to obtain target structure pair features; Merging the target structure pair feature with the structural features corresponding to each of the first component structures to obtain new structural features corresponding to each of the first component structures; When the convergence condition is met, determining the new structural features corresponding to the respective first component structures as the first-level features corresponding to the respective first component structures; If the convergence condition is not met, the first structure pair features are recalculated using the new structural features corresponding to the respective first component structures to obtain new first structure pair features of the respective first component structure pairs, and the new first structure pair features of the respective first component structure pairs are determined as the first structure pair features of the respective first component structure pairs, and the new atomic features corresponding to the respective first component structures are determined as the structural features corresponding to the respective first component structures.
5. The method according to claim 2, characterized in that The obtaining, based on the image features, an initial first-level feature corresponding to each first component structure in the plurality of target component structures in the first dimension includes: Obtaining, based on the image features, first structural features corresponding to the respective first component structures and target structural features corresponding to target component structures in the second dimension where the respective first component structures are located; The first structural features and the target structural features are spliced to obtain second structural features corresponding to the respective first component structures, wherein the initial first-level features include the second structural features.
6. The method according to claim 1, characterized in that The acquiring, based on the second-level features, the interaction attribute information corresponding to the multi-level structure object includes at least one of the following: Acquire a plurality of second component structure pairs consisting of each second component structure in the plurality of target component structures under the second dimension; calculate a first target structure feature of each second component structure pair in the plurality of second component structure pairs based on the second-level features, wherein the first target structure feature is used to represent the feature inner product similarity between each second component structure in the second component structure pair; and acquire the interaction attribute information based on the first target structure feature; Obtaining the plurality of second component structure pairs; calculating, based on the second-level features, a second target structure feature of each second component structure pair, wherein the second target structure feature is used to represent a relative distance between each second component structure when linearly mapped on a sequence; and obtaining the interaction attribute information based on the second target structure feature; Obtain the multiple second component structure pairs; calculate the third target structure feature of each second component structure pair based on the second-level feature, wherein the third target structure feature is used to represent the carbon atom distance between each second component structure; and obtain the interaction attribute information based on the third target structure feature.
7. The method according to claim 6, characterized in that The acquiring, based on the second-level features, the interaction attribute information corresponding to the multi-level structure object includes: When the first target structural feature, the second target structural feature, and the third target structural feature are obtained, performing aggregation processing on the first target structural feature, the second target structural feature, and the third target structural feature to obtain an aggregated structural feature; Merging the aggregated structural features with the second-level features through vector concatenation to obtain output structural features; The interaction attribute information is obtained based on the output structural features.
8. The method according to claim 2, characterized in that The obtaining, based on the initial first-level features, the first-level features corresponding to the respective first component structures includes: inputting the initial first-level features into a first target model, and performing information transfer through a first substructure in the first target model to obtain the first-level features corresponding to the respective first component structures, wherein the first target model is a neural network model trained using the first sample and used for feature processing; The aggregating the first-level features corresponding to each target component structure in the multiple target component structures to obtain the second-level features corresponding to the target component structure includes: aggregating the first-level features corresponding to each first component structure through the second substructure in the first target model to obtain the second-level features corresponding to the target component structure.
9. The method according to claim 6, characterized in that The acquiring, based on the second-level features, the interaction attribute information corresponding to the multi-level structure object includes: Inputting the second-level features into a second target model, and performing information transfer through the third substructure of the second target model to obtain input structure features corresponding to each of the second component structures, wherein the second target model is a neural network model trained using the second sample and used for feature processing; The input structural features corresponding to each of the second component structures are predicted and outputted through the fourth substructure of the second target model to obtain the interactive attribute information.
10. The method according to any one of claims 1 to 9, characterized in that Before obtaining the first-level features corresponding to the target component structure of the multi-level structure object, the method includes: Obtaining the number of target structures of the multi-level structure object, wherein the number of target structures includes at least one of the following: a first number corresponding to target component structures in the plurality of target component structures, a second number corresponding to the second component structure in the plurality of target component structures under the second dimension, and a third number corresponding to the first component structure in the plurality of target component structures under the first dimension; Based on the number of target structures, the acquisition complexity of the interaction attribute information is calculated, wherein the acquisition complexity is used to represent the complexity of operations to be performed when acquiring the interaction attribute information.
11. The method according to any one of claims 1 to 9, characterized in that The acquiring, based on the second-level features, the interaction attribute information corresponding to the multi-level structure object includes: The interactive attribute information is obtained based on the physical and chemical characteristics corresponding to the second-level characteristics and each first component structure in the multiple target component structures under the first dimension, wherein the physical and chemical characteristics include at least one of the following: a first feature for indicating the hydrophilicity of the target substructure in the first component structure that does not participate in dehydration bonding, a second feature for indicating the hydrophobicity of the target substructure, and a third feature for indicating the charge carried by the target substructure.
12. The method according to any one of claims 1 to 9, characterized in that The obtaining of the first-level features corresponding to the target constituent structure of the multi-level structure object includes: obtaining atomic features corresponding to amino acids in the protein; The aggregating the first-level features corresponding to each target component structure in the multiple target component structures to obtain the second-level features corresponding to the target component structure includes: aggregating the atomic features to obtain the molecular features corresponding to the amino acids; The acquiring of the interaction attribute information corresponding to the multi-level structure object based on the second-level feature includes: acquiring the interaction attribute information corresponding to the protein to be detected based on the molecular feature.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein the program can be executed by a terminal device or a computer to execute the method described in any one of claims 1 to 12.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 12 through the computer program.