Compound structure generation method and device
By constructing a compound database and performing fragment replacement based on basic structural information, the problems of generating invalid structures and synthetic feasibility in compound structure generation were solved, thereby improving the diversity and novelty of compound structures.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies have problems in compound structure generation, such as generating invalid structures, failing to guarantee synthetic feasibility, and being unable to flexibly control the diversity, novelty, synthetic complexity, and chemical type of compound structures.
By constructing a compound database, determining the actual environment of the fragment to be replaced, screening reference environments and fragments in the compound database, and replacing fragments based on basic structural information, the target compound structure is generated.
It improves the speed and diversity of compound structure generation, ensures the synthetic feasibility and chemical rationality of newly generated compounds, and enhances the novelty of compound structures.
Smart Images

Figure CN121862233A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of biomedical technology, and in particular to a method and apparatus for generating compound structures. Background Technology
[0002] In the field of pharmaceutical research and development, structure generation methods are widely used in de novo drug design, and the performance of the generated compound structures has a significant impact on the overall research outcome. Techniques based on deep learning models and traditional atomic-level methods may produce invalid structures and fail to address the issue of synthetic feasibility. Therefore, providing a more efficient method for compound structure generation, ensuring that the generated compounds are chemically valid while allowing flexible control over the diversity, novelty, synthetic complexity, and chemical type of the generated compounds, is a pressing technical problem that needs to be solved. Summary of the Invention
[0003] In view of this, the present disclosure proposes a method and apparatus for generating compound structures.
[0004] According to one aspect of this disclosure, a method for generating a compound structure is provided, the method comprising:
[0005] The replacement fragment to be replaced in the first compound and its corresponding actual environment are identified. The actual environment refers to at least one actual connection fragment in the first compound that is directly connected to the replacement fragment.
[0006] Based on the structural character encoding corresponding to each actual connection segment, the reference environment in the compound database is filtered to determine the first environment that is the same as the actual environment. The compound database stores the reference environment information of the reference environment, which includes the first identifier of the reference environment and the structural character encoding corresponding to the reference connection segment contained in the reference environment.
[0007] Based on the first identifier of the first environment, a first association table corresponding to the first environment is determined from the association table of the compound database. The association table is used to record the first identifier of the reference environment and the second identifier of a reference fragment in the reference environment.
[0008] Based on the basic structural information of the fragment to be replaced and the reference fragment information of the reference fragment corresponding to the second identifier in each of the first association tables, a plurality of first fragments are selected from the reference fragments corresponding to the first association tables. The compound database stores the reference fragment information of the reference fragments, which includes the second identifier of the reference fragment, the basic structural information of the reference fragment, and the structural character encoding corresponding to the reference fragment.
[0009] Based on the basic structural information of the first fragment and the basic structural information of the fragment to be replaced, the fragment to be replaced in the first compound is replaced with each of the first fragments to obtain multiple target compound structures.
[0010] In one possible implementation, the replacement fragment to be replaced in the first compound and the corresponding actual environment are determined, including:
[0011] Based on the replacement requirements, the replacement fragment to be replaced in the first compound is determined, and the basic structural information of the replacement fragment is determined;
[0012] The actual connecting segments in the first compound that are directly connected to the segment to be replaced are identified, and the structural character codes corresponding to each actual connecting segment are identified. The structural character codes corresponding to the actual connecting segments include the structural character codes corresponding to the feature structures in the actual connecting segments whose radius distance from the segment to be replaced is less than or equal to a preset radius.
[0013] In one possible implementation, based on the basic structural information of the segment to be replaced and the reference segment information of the reference segment corresponding to the second identifier in each of the first association tables, a plurality of first segments are selected from the reference segments corresponding to the first association table, including:
[0014] The reference segments that meet the filtering conditions in the reference segments corresponding to the second identifier in the first association table are sequentially identified as the first segments, until the number of the plurality of first segments reaches the replacement quantity.
[0015] In one possible implementation, the basic structural information includes the number of heavy atoms and / or the shortest cutting distance, wherein the number of heavy atoms is the number of heavy atoms in the fragment; and the shortest cutting distance is the number of bonds on the shortest path between the cutting positions in the fragment.
[0016] The screening criteria include at least one of the following: the difference between the number of heavy atoms in the reference fragment and the number of heavy atoms in the fragment to be replaced is less than or equal to a first difference; the difference between the shortest cutting distance of the reference fragment and the shortest cutting distance of the fragment to be replaced is less than or equal to a second difference.
[0017] In one possible implementation, the replacement fragment to be replaced in the first compound is determined based on the replacement requirement, and the basic structural information of the replacement fragment is determined, including:
[0018] The first compound is segmented at least once based on the prohibited and / or replaceable objects specified in the first compound to determine the replacement fragments to be replaced;
[0019] The prohibited substitution objects and the replaceable objects are determined based on at least one of the following factors: the active group of the first compound, the compound skeleton, and the atoms that affect the properties of the first compound.
[0020] In one possible implementation, the method further includes:
[0021] Upon receiving a first update request based on a second compound, the second compound is segmented multiple times, and the segment to be recorded and the recording environment of the segment to be recorded in the second compound are determined based on the result of each segmentation. The recording environment refers to each connecting segment in the second compound that is directly connected to the segment to be recorded.
[0022] The compound database is updated based on the fragment to be recorded and the environment to be recorded.
[0023] In one possible implementation, updating the compound database based on the fragment to be recorded and the recording environment includes at least one of the following operations:
[0024] If the reference environment recorded in the compound database does not include the environment to be recorded and the segment to be recorded, the environment to be recorded is determined as the reference environment and the segment to be recorded is determined as the reference segment. A corresponding association table, reference environment information and reference segment information are added to the compound database.
[0025] If the reference environment recorded in the compound database includes the environment to be recorded but does not include the segment to be recorded, the segment to be recorded is determined as the reference segment, and a corresponding association table and reference segment information are added to the compound database.
[0026] If the reference environment recorded in the compound database does not include the environment to be recorded but includes the fragment to be recorded, the environment to be recorded is determined as the reference environment, and a corresponding association table and reference environment information are added to the compound database.
[0027] If the reference environment recorded in the compound database includes both the environment to be recorded and the fragment to be recorded, and if the compound database does not record an association table corresponding to the environment to be recorded and the fragment to be recorded, then a corresponding association table is added to the compound database.
[0028] In one possible implementation, the method further includes:
[0029] Upon receiving a second update request based on an existing database, the compound database is updated based on each data record in the existing database;
[0030] The data record includes: a first structural character code corresponding to the environment to be merged, a second structural character code corresponding to the segment to be merged in the environment to be merged, a third structural character code indicating the connection status between the segment to be merged and the connecting segment in the environment to be merged, and basic structural information of the segment to be merged.
[0031] Updating the compound database based on each data record in the existing database includes at least one of the following operations:
[0032] If the reference environment recorded in the compound database does not include the environment to be merged and the fragment to be merged, the environment to be merged is determined as the reference environment and the fragment to be merged is determined as the reference fragment. Based on the corresponding data records, a corresponding association table, reference environment information and reference fragment information are added to the compound database.
[0033] If the reference environment recorded in the compound database includes the environment to be merged but does not include the fragment to be merged, the fragment to be merged is determined as the reference fragment, and a corresponding association table and corresponding reference fragment information are added to the compound database based on the corresponding data record;
[0034] If the reference environment recorded in the compound database does not include the environment to be merged but includes the fragment to be merged, the environment to be merged is determined as the reference environment, and a corresponding association table and reference environment information are added to the compound database based on the corresponding data record;
[0035] If the reference environment recorded in the compound database includes both the environment to be merged and the fragment to be merged, and if the compound database does not record an association table corresponding to the environment to be merged and the fragment to be merged, then a corresponding association table is added to the compound database.
[0036] According to another aspect of this disclosure, a compound structure generation apparatus is provided, comprising:
[0037] The replacement determination module is used to determine the replacement fragment to be replaced in the first compound and the corresponding actual environment. The actual environment refers to at least one actual connection fragment in the first compound that is directly connected to the replacement fragment.
[0038] An environment filtering module is used to filter reference environments in a compound database based on the structural character codes corresponding to each actual connection fragment, and determine a first environment that is the same as the actual environment. The compound database stores reference environment information of the reference environment, and the reference environment information includes a first identifier of the reference environment and the structural character codes corresponding to the reference connection fragments contained in the reference environment.
[0039] The association table determination module is used to determine the first association table corresponding to the first environment from the association table of the compound database based on the first identifier of the first environment. The association table is used to record the first identifier of the reference environment and the second identifier of a reference fragment in the reference environment.
[0040] The first fragment filtering module is used to filter out multiple first fragments from the reference fragments corresponding to the first association table based on the basic structural information of the fragment to be replaced and the reference fragment information of the reference fragment corresponding to the second identifier in each of the first association tables. The compound database stores the reference fragment information of the reference fragments, and the reference fragment information includes the second identifier of the reference fragment, the basic structural information of the reference fragment, and the structural character encoding corresponding to the reference fragment.
[0041] The replacement module is used to replace the fragment to be replaced in the first compound with each of the first fragments based on the basic structural information of the first fragment and the basic structural information of the fragment to be replaced, so as to obtain multiple target compound structures.
[0042] In one possible implementation, the replacement determining module includes:
[0043] The first determining submodule is used to determine the replacement fragment to be replaced in the first compound according to the replacement requirements, and to determine the basic structural information of the replacement fragment;
[0044] The second determining submodule is used to determine the actual connecting segments in the first compound that are directly connected to the segment to be replaced, and to determine the structural character codes corresponding to each actual connecting segment. The structural character codes corresponding to the actual connecting segments include the structural character codes corresponding to the feature structures in the actual connecting segments whose radius distance from the segment to be replaced is less than or equal to a preset radius.
[0045] In one possible implementation, the first segment filtering module includes:
[0046] The sequential filtering submodule is used to sequentially identify the reference segments that meet the filtering conditions in the reference segments corresponding to the second identifier in the first association table as the first segment, until the number of the plurality of first segments reaches the replacement quantity.
[0047] In one possible implementation, the basic structural information includes the number of heavy atoms and / or the shortest cutting distance, wherein the number of heavy atoms is the number of heavy atoms in the fragment; and the shortest cutting distance is the number of bonds on the shortest path between the cutting positions in the fragment.
[0048] The screening criteria include at least one of the following: the difference between the number of heavy atoms in the reference fragment and the number of heavy atoms in the fragment to be replaced is less than or equal to a first difference; the difference between the shortest cutting distance of the reference fragment and the shortest cutting distance of the fragment to be replaced is less than or equal to a second difference.
[0049] In one possible implementation, the replacement fragment to be replaced in the first compound is determined based on the replacement requirement, and the basic structural information of the replacement fragment is determined, including:
[0050] The first compound is segmented at least once based on the prohibited and / or replaceable objects specified in the first compound to determine the replacement fragments to be replaced;
[0051] The prohibited substitution objects and the replaceable objects are determined based on at least one of the following factors: the active group of the first compound, the compound skeleton, and the atoms that affect the properties of the first compound.
[0052] In one possible implementation, the device further includes:
[0053] The first receiving module is configured to, upon receiving a first update request based on the second compound, perform multiple segments on the second compound, and determine the segment to be recorded and the recording environment of the segment to be recorded in the second compound based on the result of each segmentation. The recording environment refers to each connecting segment in the second compound that is directly connected to the segment to be recorded.
[0054] The first update module is used to update the compound database based on the fragment to be recorded and the environment to be recorded.
[0055] In one possible implementation, the first update module includes at least one of the following sub-modules:
[0056] The first submodule is used to determine the environment to be recorded as a reference environment and the fragment to be recorded as a reference fragment when the reference environment recorded in the compound database does not include the environment to be recorded and the fragment to be recorded, and to add corresponding association tables, reference environment information and reference fragment information in the compound database.
[0057] The second submodule is used to determine the segment to be recorded as a reference segment when the reference environment recorded in the compound database includes the environment to be recorded but does not include the segment to be recorded, and to add a corresponding association table and reference segment information in the compound database.
[0058] The third submodule is used to determine the environment to be recorded as the reference environment when the reference environment recorded in the compound database does not include the environment to be recorded but includes the fragment to be recorded, and to add a corresponding association table and reference environment information in the compound database.
[0059] The fourth submodule is used to add a corresponding association table in the compound database if the reference environment recorded in the compound database includes both the environment to be recorded and the fragment to be recorded.
[0060] In one possible implementation, the device further includes:
[0061] The second update module is used to update the compound database based on each data record in the existing database upon receiving a second update request based on the existing database.
[0062] The data record includes: a first structural character code corresponding to the environment to be merged, a second structural character code corresponding to the segment to be merged in the environment to be merged, a third structural character code indicating the connection status between the segment to be merged and the connecting segment in the environment to be merged, and basic structural information of the segment to be merged.
[0063] Updating the compound database based on each data record in the existing database includes at least one of the following operations:
[0064] If the reference environment recorded in the compound database does not include the environment to be merged and the fragment to be merged, the environment to be merged is determined as the reference environment and the fragment to be merged is determined as the reference fragment. Based on the corresponding data records, a corresponding association table, reference environment information and reference fragment information are added to the compound database.
[0065] If the reference environment recorded in the compound database includes the environment to be merged but does not include the fragment to be merged, the fragment to be merged is determined as the reference fragment, and a corresponding association table and corresponding reference fragment information are added to the compound database based on the corresponding data record;
[0066] If the reference environment recorded in the compound database does not include the environment to be merged but includes the fragment to be merged, the environment to be merged is determined as the reference environment, and a corresponding association table and reference environment information are added to the compound database based on the corresponding data record;
[0067] If the reference environment recorded in the compound database includes both the environment to be merged and the fragment to be merged, and if the compound database does not record an association table corresponding to the environment to be merged and the fragment to be merged, then a corresponding association table is added to the compound database.
[0068] According to another aspect of this disclosure, a compound structure generation apparatus is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described method.
[0069] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.
[0070] According to another aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0071] The compound structure generation method and apparatus provided in this disclosure pre-constructs a compound database, and then generates compound structures based on the compound database. Due to fragment replacement based on the compound database, the optimization of the compound database reduces the data size compared to existing technologies without loss of content. This ensures fragment diversity while increasing the speed of generating new compound structures and guaranteeing the feasibility of actual synthesis of the newly generated compound structures, thus enhancing the novelty and diversity of the newly generated compound structures. Furthermore, during the fragment replacement process, the first newly introduced fragment exists in existing compounds, ensuring that the newly generated compound structure is based on a chemically reasonable mutation (CReM).
[0072] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0073] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0074] Figure 1A flowchart is shown for a method of generating a compound structure according to an embodiment of the present disclosure.
[0075] Figure 2 This diagram illustrates compound segmentation in a compound structure generation method according to an embodiment of the present disclosure.
[0076] Figure 3 A schematic diagram of a compound database in a compound structure generation method according to an embodiment of the present disclosure is shown.
[0077] Figure 4 This diagram illustrates a method for generating a compound structure according to an embodiment of the present disclosure, in which a second compound is slicing to update a compound database.
[0078] Figure 5 This is a block diagram illustrating an apparatus 1900 for generating compound structures according to an exemplary embodiment. Detailed Implementation
[0079] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0080] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.
[0081] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.
[0082] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0083] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0084] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0085] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions.
[0086] The chemical space of drugs is extremely large, estimated to be greater than 10. 33 There are a vast number of compounds. Given this sheer number, enumerating the entire space or conducting any form of exhaustive search would be impossible. Therefore, developing effective methods and strategies for exploring this space is a pressing technical problem. Among related technologies, model-based generation of new compound structures with predictive properties is employed. There are two main strategies for structure generation:
[0087] The first method involves iteratively generating a structure that matches the model's predictions.
[0088] The process includes: generating or selecting an initial structure; evaluating the generated structure using a model; selecting the most promising candidate; generating a new structure based on the selected structure and evaluating the model again, followed by subsequent steps. This process is repeated until a compound with desirable properties is generated. In this case, the structure generation and property estimation steps are separate. In related technologies, there are three main approaches to generating compound structures: atom-based, fragment-based, and reaction-based structure generation methods.
[0089] Atom-based structure generation methods use simple rules such as "adding / removing / replacing atoms / bonds" to modify the input structure and generate new structures. While theoretically this approach can generate all possible structures, resulting in high novelty and diversity, it requires numerous generation steps and involves enormous computational demands, limiting its applicability (it is more suitable for systematically exploring local chemical spaces). Furthermore, chemical validity must be controlled during structure generation to avoid erroneous structural changes, and the feasibility of synthesizing compound structures using atom-based methods remains a concern.
[0090] In reaction-based compound structure generation, a pre-constructed library of rules-based reaction compounds can produce higher novelty and diversity within just a few generations, with lower computational complexity compared to atom-based methods. Reaction-based methods significantly alter the structure during compound generation. A comprehensive reaction library can also enumerate approximate analogs of reference compounds, enabling local exploration of the chemical space. Because reaction-based compound generation provides feasible and readily available synthetic routes, the number of rules for constructing the reaction library and the size of the library itself significantly limit the novelty and diversity of the generated compounds.
[0091] Fragment-based compound generation falls between atom-based and reaction-based approaches. It is achieved by replacing or adding entire atomic groups at once, so the initial set of fragments directly determines the novelty and diversity of the generated compounds. However, in related technologies, fragment libraries are extremely large to achieve fragment diversity, resulting in lengthy processes for generating compound structures based on these libraries, and the sheer size of the fragment libraries also presents challenges in their management.
[0092] The second method involves directly generating structures with ideal properties using machine learning models.
[0093] The development of deep learning and generative models has reignited interest in the second strategy of structure enumeration. Compound structures can be generated in an unsupervised or supervised manner. However, model-based generation suffers from the following problems: to make the generation more focused, models trained on a large number of different compounds can be retrained on a small subset of compounds targeting specific target activities. This can lead to model bias, causing the generated compounds to resemble the active compounds more closely. The percentage of effective structures in the generated compounds can vary greatly depending on the model architecture, ranging from almost 100% to 4%. Furthermore, the synthetic feasibility of the generated compound structures also presents challenges.
[0094] To address the aforementioned technical problems, this disclosure provides a method and apparatus for generating compound structures. A compound database is pre-constructed. Then, during the compound structure generation process, a replacement fragment to be replaced in a first compound and its corresponding actual environment are determined. The actual environment refers to at least one actual connection fragment in the first compound that is directly connected to the replacement fragment. Based on the structural character codes corresponding to each actual connection fragment, reference environments in the compound database are filtered to determine a first environment identical to the actual environment. The compound database stores reference environment information for the reference environment, including a first identifier of the reference environment and the structural character codes corresponding to the reference connection fragments contained in the reference environment. Based on the first identifier of the first environment, the method and apparatus are derived from the compound database. The first association table corresponding to the first environment is determined in the association table of the library. The association table is used to record the first identifier of the reference environment and the second identifier of a reference fragment in the reference environment. According to the basic structural information of the fragment to be replaced and the reference fragment information of the reference fragment corresponding to the second identifier in each of the first association tables, multiple first fragments are selected from the reference fragments corresponding to the first association tables. The compound database stores the reference fragment information of the reference fragments. The reference fragment information includes the second identifier of the reference fragment, the basic structural information of the reference fragment, and the structural character code corresponding to the reference fragment. Based on the basic structural information of the first fragments and the basic structural information of the fragments to be replaced, the fragments to be replaced in the first compound are replaced with each of the first fragments to obtain multiple target compound structures. Due to the fragment replacement performed by the compound database, the optimization of the compound database has reduced the data volume compared with the prior art without loss of content. This ensures the diversity of fragments, improves the speed of generating new compound structures, and also ensures the feasibility of actual synthesis of the newly generated compound structures, thus making the new compound structures novel and diverse. Furthermore, during the replacement process, the first newly replaced fragment exists in the existing compound, ensuring that the structure of the newly generated compound is based on chemically reasonable mutations (CReM).
[0095] like Figure 1 As shown, the compound structure generation method provided in this disclosure includes steps S101-S105. This method can be applied to a terminal or a server, and this disclosure does not limit it.
[0096] In step S101, the replacement fragment to be replaced in the first compound and the corresponding actual environment are determined. The actual environment refers to at least one actual connection fragment in the first compound that is directly connected to the replacement fragment.
[0097] In this embodiment, the first compound can be a compound targeted for research and development, such as drug development. This first compound can be either actually existing or designed and generated, possessing certain properties required for research and development. Therefore, the first compound may contain specific functional groups, specific heavy atom framework structures, etc., all of which represent structural characteristics that meet the research and development requirements. Fragment replacement is performed to generate a new compound whose performance better meets the research and development requirements than the first compound. Therefore, in step S101, the replacement requirements can be determined based on the research and development needs. These replacement requirements may include: atoms, fragments (such as functional groups), and / or framework structures in the first compound that cannot be replaced; atoms and fragments (such as functional groups) in the first compound that need to be replaced; and the preset radii, etc., which are not limited in this disclosure.
[0098] In this embodiment, step S101 may include: determining the replacement fragment to be replaced in the first compound according to the replacement requirements, and determining the basic structural information of the replacement fragment.
[0099] The basic structural information of the replacement fragment can describe its structural features, including the number of heavy atoms, the shortest cutting distance, and the number of cutting positions. The number of heavy atoms refers to the number of heavy atoms (heavy atoms can be atoms with relatively large atomic masses located on the fragment's framework, such as atoms other than hydrogen atoms on the fragment's framework) in the replacement fragment. The shortest cutting distance is the number of bonds in the shortest path between cutting positions in the fragment. The number of cutting positions can refer to the number of cuts required to cleave the replacement fragment from the first compound; in some embodiments, the number of cutting positions can be set to 1, 2, or 3. Figure 2 In the example shown, the basic structural information of the fragment to be replaced includes 4 heavy atoms, 1 shortest cutting distance, and 2 cutting positions. The first compound can be cut multiple times, each cutting yielding a fragment to be replaced and its corresponding actual environment.
[0100] In some embodiments, determining the replacement fragment in the first compound based on replacement requirements and determining the basic structural information of the replacement fragment may include: slicing the first compound at least once based on specified prohibited and / or replaceable objects to determine the replacement fragment. The prohibited and replaceable objects are determined based on at least one of the following factors: the active group of the first compound, the compound skeleton, and atoms affecting the properties of the first compound.
[0101] In this embodiment, step S101 may further include: determining the actual connecting segments in the first compound that are directly connected to the segment to be replaced, and determining the structural character codes corresponding to each actual connecting segment. The structural character codes corresponding to the actual connecting segments include the structural character codes corresponding to the feature structures in the actual connecting segments whose radius distance from the segment to be replaced is less than or equal to a preset radius.
[0102] Wherein, the actual environment of the fragment to be replaced in the first compound is the actual connecting fragment in the first compound that is directly connected to the first compound, such as for Figure 2 In the example shown, the actual environment of the fragment to be replaced includes actual linker fragment 1 and actual linker fragment 2. However, since the actual linker fragments may be relatively large, in order to simplify the generation steps and promote the diversity of target compound structures, it is possible to represent each actual linker fragment only with its characteristic structure. Therefore, after determining the actual linker fragments, based on the preset radius indicated in the replacement requirements, the portion of the structure of each actual linker fragment from the cutting position with the fragment to be replaced, with a radius distance less than or equal to the preset radius, can be taken as the characteristic structure, and the structural character code corresponding to the characteristic structure can be used as the structural character code corresponding to the actual linker fragment. The structural character code can be SMILES (Simplified Molecular-Input Line-Entry System), t-SMILES, DeepSMILES, InChI (International Chemical Identifier), etc., and this disclosure does not limit it. The preset radius and radius distance can be represented by the number of keys. A smaller preset radius allows for the retrieval of more reference fragments corresponding to the first environment. This can be set according to actual research and development needs and the radius distance of the structures corresponding to the structural character codes of the reference link fragments in each reference environment of the compound database. For example, it can be set to 1, 2, 3, etc., and this disclosure does not impose any limitations on this. For example, Figure 2 In the example shown, if the preset radius is 3, then the structural character codes corresponding to actual connection fragment 1 and actual connection fragment 2 are the structural character codes corresponding to feature structure 1 and feature structure 2, respectively.
[0103] In step S102, reference environments in the compound database are filtered according to the structural character encoding corresponding to each actual connection fragment to determine a first environment that is the same as the actual environment. Then, a first identifier corresponding to the first environment is determined.
[0104] In this embodiment, as Figure 3As shown, the compound database stores reference environment information for each reference environment, reference fragment information for each reference fragment, and association tables corresponding to each reference environment. The reference environment information includes a first identifier for the reference environment and the structural character encoding corresponding to the reference connection fragment contained within the reference environment, wherein the structural character encoding corresponding to the reference connection fragment can be the structural character encoding of a feature structure within the reference connection fragment. The reference fragment information includes a second identifier for the reference fragment, basic structural information of the reference fragment, and the structural character encoding corresponding to the reference fragment. Different reference environments have different first identifiers, different reference fragments have different second identifiers, and the first and second identifiers are also different, so that all reference environments and reference fragments can be distinguished by using the first and second identifiers. The compound database records multiple association tables, each used to record the first identifier of a reference environment and the second identifier of a reference fragment within that reference environment.
[0105] In one possible implementation, the method may further include an update step for the compound database. Depending on the basis for the update, the update step may have the following two possible implementations.
[0106] Method 1: Update the compound database based on the new compounds.
[0107] The update steps include: upon receiving a first update request based on a second compound, performing multiple segments on the second compound, determining the segment to be recorded and the recording environment of the segment in the second compound based on the result of each segmentation, wherein the recording environment refers to each connection segment in the second compound that is directly connected to the segment to be recorded; and updating the compound database based on the segment to be recorded and the recording environment.
[0108] Furthermore, based on the radius distance (e.g., 3) of the feature structure expressed by the structural character encoding of the connecting fragments in the reference environment in the compound database, the structural character encoding corresponding to the feature structure (also with a radius distance of 3) of each connecting fragment in the environment to be recorded is determined. Figure 3 In the example shown, after the second compound A is segmented, fragment 1 to be recorded and connecting fragments 1-1 and 1-2 in the recording environment 1 are obtained. Given that "the radius distance of the feature structure expressed by the structural character encoding of the connecting fragment in the reference environment of the compound database is 3", the feature structures in connecting fragment 1-1 and connecting fragment 1-2 are determined to be feature structures 1 and 2, respectively. Then, the structural character encodings corresponding to feature structures 1 and 2 are obtained. Furthermore, it is also necessary to determine the structural character encoding corresponding to the fragment to be recorded.
[0109] After determining the structural character encoding corresponding to the fragment to be recorded and the structural character encoding corresponding to the feature structures of each connected fragment in the environment to be recorded, the system further determines, based on the structural character encoding, whether the reference environment recorded in the compound database already includes the environment to be recorded, and whether the reference fragment already includes the fragment to be recorded. Then, the compound database is updated based on the fragment to be recorded and the environment to be recorded, which may include at least one of the following operations:
[0110] Operation 1: If the reference environment recorded in the compound database does not include the environment to be recorded and the fragment to be recorded, then the environment to be recorded is identified as the reference environment, and the fragment to be recorded is identified as the reference fragment. Corresponding association tables, reference environment information, and reference fragment information are added to the compound database. That is, a first identifier corresponding to the environment to be recorded is determined, a second identifier corresponding to the fragment to be recorded is determined, and then the basic structural information of the fragment to be recorded (including the number of heavy atoms and the shortest cutting distance) is determined. Then, an association table corresponding to the environment to be recorded (which includes the first identifier corresponding to the environment to be recorded and the second identifier corresponding to the fragment to be recorded) and reference environment information, as well as reference fragment information corresponding to the fragment to be recorded, are added to the compound database.
[0111] Operation 2: If the reference environment recorded in the compound database includes the environment to be recorded but does not include the fragment to be recorded, the fragment to be recorded is determined as the reference fragment, and corresponding reference fragment information (adding reference fragment information corresponding to the fragment to be recorded) and the association table corresponding to the environment to be recorded are added to the compound database.
[0112] Operation 3: If the reference environment recorded in the compound database does not include the environment to be recorded but includes the fragment to be recorded, the environment to be recorded is determined as the reference environment, and an association table corresponding to the environment to be recorded and reference environment information are added to the compound database (as in Operation 1).
[0113] Operation 4: If the reference environment recorded in the compound database includes both the environment to be recorded and the fragment to be recorded, and if the compound database does not record an association table corresponding to the environment to be recorded and the fragment to be recorded (i.e., the compound database does not have an association table including the first identifier of the environment to be recorded and the second identifier of the fragment to be recorded), then a corresponding association table is added to the compound database. If the compound database records an association table corresponding to the environment to be recorded and the fragment to be recorded, then no information is updated for the compound database.
[0114] For example, such as Figure 4The example shown illustrates the process of updating compound data for second compounds A, B, and C. If both the recording fragment 1 and the recording environment 1 are not recorded in the compound database, operation 1 can be executed. Then, since the characteristic structures of connection fragments 2-1 and 2-2 in the recording environment 2 are identical to those in the recording environment 1, and the recording fragment 2 is not recorded in the compound database, operation 2 can be executed. Subsequently, since the characteristic structures of connection fragments 3-1 and 3-2 in the recording environment 3 are identical to those in the recording environment 1, and the recording fragment 3 is not recorded in the compound database, operation 2 can also be executed. In this way, the reference fragment, reference environment, and association table are managed separately in the compound database. When a new second compound appears and the compound database needs to be updated, only incremental updates are performed (as in operations 1-4 above). This significantly reduces the data volume of the compound database compared to existing databases that require "full recording for each new compound." This inevitably increases the speed of retrieving reference fragments from compound databases to determine the replacement fragments, improving the overall speed and efficiency of compound structure generation. Furthermore, given the same amount of data in the database, the compound database of this application clearly possesses more reference fragments and reference environments, enhancing the diversity and novelty of compound structure generation. Moreover, since each data point in the compound database is determined based on existing compounds through segmentation, it ensures the feasibility of synthesizing newly generated compound structures after the reference fragments are replaced, and also guarantees that the newly generated compound structures are based on chemically reasonable mutations.
[0115] Method 2: Update the compound database based on the existing database.
[0116] The update steps include: upon receiving a second update request based on an existing database, updating the compound database based on each data record in the existing database;
[0117] The data record includes: a first structural character code corresponding to the environment to be merged, a second structural character code corresponding to the segment to be merged within the environment, a third structural character code indicating the connection state between the segment to be merged and the connecting segments in the environment, and basic structural information of the segment to be merged. In related technologies, the first structural character code corresponding to the environment to be merged includes the structural character codes corresponding to each connecting segment constituting the environment to be merged. However, since the first structural character code in the prior art is the structural character code corresponding to all connecting segments, the update step further includes: further determining the structural character code corresponding to the "feature structure of each connecting segment" in the environment to be merged based on the first structural character code, using the radius distance of the feature structure expressed by the structural character codes corresponding to the connecting segments in the reference environment in the compound database.
[0118] After determining the structural character codes corresponding to the "feature structures of each connecting segment" in the environment to be merged, the system further determines whether the reference segment recorded in the compound database already includes the segment to be merged based on the second structural character code corresponding to the segment to be merged, and whether the reference environment recorded in the compound database already includes the environment to be merged based on the structural character codes corresponding to the "feature structures of each connecting segment" in the environment to be merged. Then, the compound database is updated based on each data record in the existing database, including at least one of the following operations:
[0119] Operation 5: If the reference environment recorded in the compound database does not include the environment to be merged and the fragment to be merged, the environment to be merged is determined as the reference environment and the fragment to be merged is determined as the reference fragment. Based on the corresponding data records, corresponding association tables, reference environment information and reference fragment information are added to the compound database.
[0120] Operation 6: If the reference environment recorded in the compound database includes the environment to be merged but does not include the fragment to be merged, the fragment to be merged is determined as the reference fragment, and a corresponding association table and reference fragment information are added to the compound database based on the corresponding data record.
[0121] Operation 7: If the reference environment recorded in the compound database does not include the environment to be merged but includes the fragment to be merged, the environment to be merged is determined as the reference environment, and a corresponding association table and reference environment information are added to the compound database based on the corresponding data record.
[0122] Operation 8: If the reference environment recorded in the compound database includes both the environment to be merged and the fragment to be merged, and if the compound database does not record an association table corresponding to the environment to be merged and the fragment to be merged (i.e., the compound database does not have an association table including the first identifier of the merging environment and the second identifier of the fragment to be merged), then a corresponding association table is added to the compound database. If the compound database records an association table corresponding to the environment to be merged and the fragment to be merged, then no information is updated for the compound database.
[0123] The update content of the compound database update in operations 5-8 is the same as that in operations 1-4 above, and can be referred to above; this disclosure does not impose any limitations on it. Through operations 5-8, during the process of merging data records in the existing database, only a portion of the information from each data record in the existing database is merged into the compound database (that is, compared to the full recording method for each compound in the existing database, the compound database only performs incremental recording). This achieves the update of the compound database based on the existing database without loss of effective content, resulting in a significant reduction in the number of compound databases compared to existing databases in the prior art, while ensuring no loss of effective content. Furthermore, since each data record in the existing database is determined based on existing compounds through segmentation, it ensures the feasibility of synthesizing newly generated compound structures after the reference fragments in the updated compound database are replaced, and also ensures that the newly generated compound structures are based on chemically reasonable mutations.
[0124] In step S103, based on the first identifier of the first environment, the first association table corresponding to the first environment is determined from the association tables of the compound database. After determining the first identifier of the first environment, each association table in the compound database can be traversed to determine the association table with the first identifier as the first association table. And as... Figure 3 As shown, the first association table records the second identifier corresponding to a certain reference fragment in the first environment. However, there can be many first association tables that record the first identifier of the first environment, so step S104 needs to be executed to further filter the reference fragments.
[0125] In step S104, based on the basic structural information of the segment to be replaced and the reference segment information of the reference segment corresponding to the second identifier in each of the first association tables, a plurality of first segments are selected from the reference segments corresponding to the first association table.
[0126] In this embodiment, since the number of selected first association tables may be excessive, that is, the number of reference segments corresponding to the second identifier recorded in the selected first association tables may be excessive, the reference segments that meet the filtering conditions in the reference segments corresponding to the second identifier in the first association table can be sequentially determined as the first segments when a replacement quantity is set, until the number of the plurality of first segments reaches the replacement quantity.
[0127] In one possible implementation, the screening criteria may include at least one of the following: the difference between the number of heavy atoms in the reference fragment and the number of heavy atoms in the fragment to be replaced is less than or equal to a first difference; and the difference between the shortest cutting distance of the reference fragment and the shortest cutting distance corresponding to the fragment to be replaced is less than or equal to a second difference.
[0128] The purpose of setting the first difference is to ensure that the selected first fragment and the fragment to be replaced are more similar in the number of heavy atoms, so that the performance of the generated target compound structure can be further improved compared to the first compound. The first difference can be ±2, ±3, etc., and this disclosure does not limit this. The purpose of setting the second difference is to ensure that after the selected first fragment replaces the fragment to be replaced, the connection state between the first fragment and the actual connecting fragment is closer to the connection state between the fragment to be replaced and the actual connecting fragment in the original first compound, again so that the performance of the generated target compound structure can be further improved compared to the first compound.
[0129] Specifically, if the number of first segments determined by traversing the entire first association table for each second identifier is less than the replacement quantity, then reference segments where "the difference between the number of heavy atoms and the number of heavy atoms in the segment to be replaced is less than or equal to a first difference, and the difference between the shortest cutting distance and the shortest cutting distance corresponding to the segment to be replaced is less than or equal to a second difference" are used, until the number of first segments reaches the replacement quantity. If the number of first segments still cannot reach the replacement quantity, the value of the first difference or the second difference can be adjusted to obtain a sufficient number of first segments.
[0130] In step S105, based on the basic structural information of the first fragment and the basic structural information of the fragment to be replaced, the fragment to be replaced in the first compound is replaced with each of the first fragments to obtain multiple target compound structures.
[0131] In this embodiment, since only the fragment to be replaced is replaced with the first fragment, it is necessary to ensure accurate connection between the first fragment and the actual connecting fragment during the replacement process. Therefore, the accurate connection is ensured by environmental matching during the replacement process: this includes matching the reference connecting fragments connected to each cutting position of the first fragment (that is, the position where it can be connected to other structures) with the actual connecting fragments to determine the correct connection method, and finally obtaining multiple target compound structures.
[0132] For example, regarding Figure 2 The first compound shown is the fragment to be replaced. Figure 2 As shown, in the first segment is Figure 4 In the case of the fragment 3 to be recorded shown (which becomes the reference fragment after being recorded in the compound database), since position 1 and position 2 are connected to the reference connecting fragment with characteristic structure 1 and the reference connecting fragment with characteristic structure 2 respectively, and the actual connecting fragment 1 in the first compound has characteristic structure 1 and the actual connecting structure 2 has characteristic structure 2, the position 1 and position 2 of the fragment 3 to be recorded after replacement are connected to the actual connecting fragment 1 and the actual connecting structure 2 respectively.
[0133] This disclosure also provides a compound structure generation apparatus, comprising:
[0134] The replacement determination module is used to determine the replacement fragment to be replaced in the first compound and the corresponding actual environment. The actual environment refers to at least one actual connection fragment in the first compound that is directly connected to the replacement fragment.
[0135] An environment filtering module is used to filter reference environments in a compound database based on the structural character codes corresponding to each actual connection fragment, and determine a first environment that is the same as the actual environment. The compound database stores reference environment information of the reference environment, and the reference environment information includes a first identifier of the reference environment and the structural character codes corresponding to the reference connection fragments contained in the reference environment.
[0136] The association table determination module is used to determine the first association table corresponding to the first environment from the association table of the compound database based on the first identifier of the first environment. The association table is used to record the first identifier of the reference environment and the second identifier of a reference fragment in the reference environment.
[0137] The first fragment filtering module is used to filter out multiple first fragments from the reference fragments corresponding to the first association table based on the basic structural information of the fragment to be replaced and the reference fragment information of the reference fragment corresponding to the second identifier in each of the first association tables. The compound database stores the reference fragment information of the reference fragments, and the reference fragment information includes the second identifier of the reference fragment, the basic structural information of the reference fragment, and the structural character encoding corresponding to the reference fragment.
[0138] The replacement module is used to replace the fragment to be replaced in the first compound with each of the first fragments based on the basic structural information of the first fragment and the basic structural information of the fragment to be replaced, so as to obtain multiple target compound structures.
[0139] In one possible implementation, the replacement determining module includes:
[0140] The first determining submodule is used to determine the replacement fragment to be replaced in the first compound according to the replacement requirements, and to determine the basic structural information of the replacement fragment;
[0141] The second determining submodule is used to determine the actual connecting segments in the first compound that are directly connected to the segment to be replaced, and to determine the structural character codes corresponding to each actual connecting segment. The structural character codes corresponding to the actual connecting segments include the structural character codes corresponding to the feature structures in the actual connecting segments whose radius distance from the segment to be replaced is less than or equal to a preset radius.
[0142] In one possible implementation, the first segment filtering module includes:
[0143] The sequential filtering submodule is used to sequentially identify the reference segments that meet the filtering conditions in the reference segments corresponding to the second identifier in the first association table as the first segment, until the number of the plurality of first segments reaches the replacement quantity.
[0144] In one possible implementation, the basic structural information includes the number of heavy atoms and / or the shortest cutting distance, wherein the number of heavy atoms is the number of heavy atoms in the fragment; and the shortest cutting distance is the number of bonds on the shortest path between the cutting positions in the fragment.
[0145] The screening criteria include at least one of the following: the difference between the number of heavy atoms in the reference fragment and the number of heavy atoms in the fragment to be replaced is less than or equal to a first difference; the difference between the shortest cutting distance of the reference fragment and the shortest cutting distance of the fragment to be replaced is less than or equal to a second difference.
[0146] In one possible implementation, the replacement fragment to be replaced in the first compound is determined based on the replacement requirement, and the basic structural information of the replacement fragment is determined, including:
[0147] The first compound is segmented at least once based on the prohibited and / or replaceable objects specified in the first compound to determine the replacement fragments to be replaced;
[0148] The prohibited substitution objects and the replaceable objects are determined based on at least one of the following factors: the active group of the first compound, the compound skeleton, and the atoms that affect the properties of the first compound.
[0149] In one possible implementation, the device further includes:
[0150] The first receiving module is configured to, upon receiving a first update request based on the second compound, perform multiple segments on the second compound, and determine the segment to be recorded and the recording environment of the segment to be recorded in the second compound based on the result of each segmentation. The recording environment refers to each connecting segment in the second compound that is directly connected to the segment to be recorded.
[0151] The first update module is used to update the compound database based on the fragment to be recorded and the environment to be recorded.
[0152] In one possible implementation, the first update module includes at least one of the following sub-modules:
[0153] The first submodule is used to determine the environment to be recorded as a reference environment and the fragment to be recorded as a reference fragment when the reference environment recorded in the compound database does not include the environment to be recorded and the fragment to be recorded, and to add corresponding association tables, reference environment information and reference fragment information in the compound database.
[0154] The second submodule is used to determine the segment to be recorded as a reference segment when the reference environment recorded in the compound database includes the environment to be recorded but does not include the segment to be recorded, and to add a corresponding association table and reference segment information in the compound database.
[0155] The third submodule is used to determine the environment to be recorded as the reference environment when the reference environment recorded in the compound database does not include the environment to be recorded but includes the fragment to be recorded, and to add a corresponding association table and reference environment information in the compound database.
[0156] The fourth submodule is used to add a corresponding association table in the compound database if the reference environment recorded in the compound database includes both the environment to be recorded and the fragment to be recorded.
[0157] In one possible implementation, the device further includes:
[0158] The second update module is used to update the compound database based on each data record in the existing database upon receiving a second update request based on the existing database.
[0159] The data record includes: a first structural character code corresponding to the environment to be merged, a second structural character code corresponding to the segment to be merged in the environment to be merged, a third structural character code indicating the connection status between the segment to be merged and the connecting segment in the environment to be merged, and basic structural information of the segment to be merged.
[0160] Updating the compound database based on each data record in the existing database includes at least one of the following operations:
[0161] If the reference environment recorded in the compound database does not include the environment to be merged and the fragment to be merged, the environment to be merged is determined as the reference environment and the fragment to be merged is determined as the reference fragment. Based on the corresponding data records, a corresponding association table, reference environment information and reference fragment information are added to the compound database.
[0162] If the reference environment recorded in the compound database includes the environment to be merged but does not include the fragment to be merged, the fragment to be merged is determined as the reference fragment, and a corresponding association table and corresponding reference fragment information are added to the compound database based on the corresponding data record;
[0163] If the reference environment recorded in the compound database does not include the environment to be merged but includes the fragment to be merged, the environment to be merged is determined as the reference environment, and a corresponding association table and reference environment information are added to the compound database based on the corresponding data record;
[0164] If the reference environment recorded in the compound database includes both the environment to be merged and the fragment to be merged, and if the compound database does not record an association table corresponding to the environment to be merged and the fragment to be merged, then a corresponding association table is added to the compound database.
[0165] It should be noted that although the above embodiments have been used as examples to illustrate the method and apparatus for generating compound structures, those skilled in the art will understand that this disclosure is not limited thereto. In fact, users can flexibly set each step and module according to their personal preferences and / or actual application scenarios, as long as the technical solution of this disclosure is applicable.
[0166] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0167] This disclosure also provides a compound structure generation apparatus, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0168] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0169] This disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above method.
[0170] Figure 5 This is a block diagram illustrating an apparatus 1900 for generating compound structures according to an exemplary embodiment. For example, apparatus 1900 may be provided as a server or terminal device. (Refer to...) Figure 5The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0171] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0172] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.
[0173] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0174] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.
[0175] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.
[0176] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0177] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0178] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0179] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0180] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for generating a compound structure, characterized in that, The method includes: The replacement fragment to be replaced in the first compound and its corresponding actual environment are identified. The actual environment refers to at least one actual connection fragment in the first compound that is directly connected to the replacement fragment. Based on the structural character encoding corresponding to each actual connection segment, the reference environment in the compound database is filtered to determine the first environment that is the same as the actual environment. The compound database stores the reference environment information of the reference environment, which includes the first identifier of the reference environment and the structural character encoding corresponding to the reference connection segment contained in the reference environment. Based on the first identifier of the first environment, a first association table corresponding to the first environment is determined from the association table of the compound database. The association table is used to record the first identifier of the reference environment and the second identifier of a reference fragment in the reference environment. Based on the basic structural information of the fragment to be replaced and the reference fragment information of the reference fragment corresponding to the second identifier in each of the first association tables, a plurality of first fragments are selected from the reference fragments corresponding to the first association tables. The compound database stores the reference fragment information of the reference fragments, which includes the second identifier of the reference fragment, the basic structural information of the reference fragment, and the structural character encoding corresponding to the reference fragment. Based on the basic structural information of the first fragment and the basic structural information of the fragment to be replaced, the fragment to be replaced in the first compound is replaced with each of the first fragments to obtain multiple target compound structures.
2. The method according to claim 1, characterized in that, The replacement fragment to be replaced in the first compound and its corresponding actual environment were identified, including: Based on the replacement requirements, the replacement fragment to be replaced in the first compound is determined, and the basic structural information of the replacement fragment is determined; The actual connecting segments in the first compound that are directly connected to the segment to be replaced are identified, and the structural character codes corresponding to each actual connecting segment are identified. The structural character codes corresponding to the actual connecting segments include the structural character codes corresponding to the feature structures in the actual connecting segments whose radius distance from the segment to be replaced is less than or equal to a preset radius.
3. The method according to claim 1, characterized in that, Based on the basic structural information of the segment to be replaced and the reference segment information of the reference segment corresponding to the second identifier in each of the first association tables, multiple first segments are selected from the reference segments corresponding to the first association tables, including: The reference segments that meet the filtering conditions in the reference segments corresponding to the second identifier in the first association table are sequentially identified as the first segments, until the number of the plurality of first segments reaches the replacement quantity.
4. The method according to claim 3, characterized in that, The basic structural information includes the number of heavy atoms and / or the shortest cutting distance, wherein the number of heavy atoms is the number of heavy atoms in the fragment; and the shortest cutting distance is the number of bonds on the shortest path between the cutting positions in the fragment. The screening criteria include at least one of the following: the difference between the number of heavy atoms in the reference fragment and the number of heavy atoms in the fragment to be replaced is less than or equal to a first difference; the difference between the shortest cutting distance of the reference fragment and the shortest cutting distance of the fragment to be replaced is less than or equal to a second difference.
5. The method according to claim 2, characterized in that, Based on the replacement requirements, the replacement fragment to be replaced in the first compound is determined, and the basic structural information of the replacement fragment is determined, including: The first compound is segmented at least once based on the prohibited and / or replaceable objects specified in the first compound to determine the replacement fragments to be replaced; The prohibited substitution objects and the replaceable objects are determined based on at least one of the following factors: the active group of the first compound, the compound skeleton, and the atoms that affect the properties of the first compound.
6. The method according to claim 1, characterized in that, The method further includes: Upon receiving a first update request based on a second compound, the second compound is segmented multiple times, and the segment to be recorded and the recording environment of the segment to be recorded in the second compound are determined based on the result of each segmentation. The recording environment refers to each connecting segment in the second compound that is directly connected to the segment to be recorded. The compound database is updated based on the fragment to be recorded and the environment to be recorded.
7. The method according to claim 6, characterized in that, Updating the compound database based on the fragment to be recorded and the recording environment includes at least one of the following operations: If the reference environment recorded in the compound database does not include the environment to be recorded and the segment to be recorded, the environment to be recorded is determined as the reference environment and the segment to be recorded is determined as the reference segment. A corresponding association table, reference environment information and reference segment information are added to the compound database. If the reference environment recorded in the compound database includes the environment to be recorded but does not include the segment to be recorded, the segment to be recorded is determined as the reference segment, and a corresponding association table and reference segment information are added to the compound database. If the reference environment recorded in the compound database does not include the environment to be recorded but includes the fragment to be recorded, the environment to be recorded is determined as the reference environment, and a corresponding association table and reference environment information are added to the compound database. If the reference environment recorded in the compound database includes both the environment to be recorded and the fragment to be recorded, and if the compound database does not record an association table corresponding to the environment to be recorded and the fragment to be recorded, then a corresponding association table is added to the compound database.
8. The method according to claim 1, characterized in that, The method further includes: Upon receiving a second update request based on an existing database, the compound database is updated based on each data record in the existing database; The data record includes: a first structural character code corresponding to the environment to be merged, a second structural character code corresponding to the segment to be merged in the environment to be merged, a third structural character code indicating the connection status between the segment to be merged and the connecting segment in the environment to be merged, and basic structural information of the segment to be merged. Updating the compound database based on each data record in the existing database includes at least one of the following operations: If the reference environment recorded in the compound database does not include the environment to be merged and the fragment to be merged, the environment to be merged is determined as the reference environment and the fragment to be merged is determined as the reference fragment. Based on the corresponding data records, a corresponding association table, reference environment information and reference fragment information are added to the compound database. If the reference environment recorded in the compound database includes the environment to be merged but does not include the fragment to be merged, the fragment to be merged is determined as the reference fragment, and a corresponding association table and corresponding reference fragment information are added to the compound database based on the corresponding data record; If the reference environment recorded in the compound database does not include the environment to be merged but includes the fragment to be merged, the environment to be merged is determined as the reference environment, and a corresponding association table and reference environment information are added to the compound database based on the corresponding data record; If the reference environment recorded in the compound database includes both the environment to be merged and the fragment to be merged, and if the compound database does not record an association table corresponding to the environment to be merged and the fragment to be merged, then a corresponding association table is added to the compound database.
9. A compound structure generation apparatus, characterized in that, The device includes: The replacement determination module is used to determine the replacement fragment to be replaced in the first compound and the corresponding actual environment. The actual environment refers to at least one actual connection fragment in the first compound that is directly connected to the replacement fragment. An environment filtering module is used to filter reference environments in a compound database based on the structural character codes corresponding to each actual connection fragment, and determine a first environment that is the same as the actual environment. The compound database stores reference environment information of the reference environment, and the reference environment information includes a first identifier of the reference environment and the structural character codes corresponding to the reference connection fragments contained in the reference environment. The association table determination module is used to determine the first association table corresponding to the first environment from the association table of the compound database based on the first identifier of the first environment. The association table is used to record the first identifier of the reference environment and the second identifier of a reference fragment in the reference environment. The first fragment filtering module is used to filter out multiple first fragments from the reference fragments corresponding to the first association table based on the basic structural information of the fragment to be replaced and the reference fragment information of the reference fragment corresponding to the second identifier in each of the first association tables. The compound database stores the reference fragment information of the reference fragments, and the reference fragment information includes the second identifier of the reference fragment, the basic structural information of the reference fragment, and the structural character encoding corresponding to the reference fragment. The replacement module is used to replace the fragment to be replaced in the first compound with each of the first fragments based on the basic structural information of the first fragment and the basic structural information of the fragment to be replaced, so as to obtain multiple target compound structures.
10. A compound structure generation apparatus, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
11. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
12. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.