A method, apparatus, system, and storage medium for clustering compounds

By segmenting compound samples into subsets and generating sample legends, and combining the initial and legend label recognition results, the inefficiency of traditional compound clustering methods is solved, achieving more efficient and accurate clustering of small molecule compounds.

CN115049866BActive Publication Date: 2025-11-25HUIYI KEJI (SHANGHAI) LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210537018.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-11-25
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

Traditional cheminformatics-based compound clustering methods are inefficient and slow, with excessive computational and storage requirements. This prevents the clustering of small molecule compounds from being scaled up to the level of hundreds of thousands, resulting in vague and inaccurate identification results and an inability to effectively utilize compound information.

Method used

By acquiring samples of compounds to be identified and segmenting them into sample subsets containing initial identification labels, generating sample legends using the sample subsets, and combining the initial identification labels and legend labels, the target identification result is identified. A coarse-grained clustering method is used to reduce the amount of features used and improve computational efficiency and accuracy.

Benefits of technology

While reducing the amount of features used, it significantly improves the accuracy and processing speed of small molecule compound clustering, achieving more efficient and larger-scale compound clustering and breaking through the limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115049866B_ABST
    Figure CN115049866B_ABST
Patent Text Reader

Abstract

The application provides a compound clustering method, device, system and storage medium, by obtaining a to-be-identified compound sample, and dividing the to-be-identified compound sample into a sample subset containing an initial identification label; obtaining a sample legend according to the sample subset; obtaining a target identification result corresponding to the to-be-identified compound sample according to the sample legend and the identification label; wherein the identification label includes the initial identification label. The application provides a high-efficiency, fast and accurate small-molecule compound clustering method based on statistical compound clustering, improves the accuracy of small-molecule compound clustering, reduces the processing space of clustering, breaks through the limitations of small-molecule clustering, and makes the processing of small-molecule compound clustering more efficient and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, specifically to a method, apparatus, system, and storage medium for compound clustering. Background Technology

[0002] We usually call molecules composed of a few or dozens of atoms small molecules, which are substances that can exist in solid, gaseous, or liquid states at room temperature. Common small organic molecule compounds include ethanol, glucose, and methane.

[0003] Clustering is used to subdivide large datasets of compounds into smaller groups of similar compounds. It is commonly used to analyze high-throughput screening results, virtual screening, or docking studies. Traditional cheminformatics-based clustering methods are inefficient and slow. Even using compound fingerprint similarity for identification requires excessive computational and storage space, limiting the number of compounds that can be identified.

[0004] Therefore, a new solution is needed. Summary of the Invention

[0005] In view of this, embodiments of this specification provide a method, apparatus, system, and storage medium for clustering compounds, used in the clustering process of small molecule compounds.

[0006] The embodiments in this specification provide the following technical solutions:

[0007] This specification provides an embodiment of a method for clustering compounds, including:

[0008] Obtain a sample of the compound to be identified, and segment the sample into a subset containing an initial identification label;

[0009] Based on the aforementioned sample subset, a sample legend is obtained.

[0010] Based on the sample illustration and the identification label, the target identification result corresponding to the compound sample to be identified is obtained; wherein, the identification label includes the initial identification label.

[0011] This specification also provides an apparatus for compound clustering, comprising:

[0012] An acquisition module is used to acquire a sample of a compound to be identified and to segment the sample into a subset containing an initial identification label;

[0013] The module is used to obtain a sample legend based on the sample subset.

[0014] The output module is used to obtain the target identification result corresponding to the compound sample to be identified based on the sample legend and the identification label; wherein the identification label includes the initial identification label.

[0015] This specification also provides a system for compound clustering, including: a memory, a processor, and a computer program. The computer program is stored in the memory, and the processor runs the computer program to perform the following steps: acquiring a sample of a compound to be identified, and segmenting the sample of the compound to be identified into a sample subset containing an initial identification label; obtaining a sample legend based on the sample subset; and obtaining a target identification result corresponding to the sample of the compound to be identified based on the sample legend and the identification label; wherein the identification label includes the initial identification label.

[0016] This specification also provides a readable storage medium storing a computer program that, when executed by a processor, performs the following steps: acquiring a sample of a compound to be identified and segmenting the sample into a subset containing an initial identification label; obtaining a sample legend based on the sample subset; and obtaining a target identification result corresponding to the sample of the compound to be identified based on the sample legend and the identification label; wherein the identification label includes the initial identification label.

[0017] Compared with existing technologies, the beneficial effects achieved by at least one of the above-mentioned technical solutions adopted in the embodiments of this specification include at least the following: acquiring a sample of a compound to be identified and segmenting the sample into a subset containing an initial identification label; obtaining a sample legend based on the sample subset; and obtaining the target identification result corresponding to the sample of the compound to be identified based on the sample legend and the identification label. Adding compound image recognition and detection to the coarse-grained identification of small molecule compound clustering can improve the accuracy of small molecule compound clustering, reduce the processing space of clustering, and overcome the limitations of small molecule clustering, thereby making the processing of small molecule compound clustering more efficient and accurate. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram illustrating an application of compound clustering provided in the embodiments of this specification;

[0020] Figure 2This specification provides a method flow for compound clustering in its embodiments. Figure 1 ;

[0021] Figure 3 This specification provides a method flow for compound clustering in its embodiments. Figure 2 ;

[0022] Figure 4 This is a schematic diagram of a compound clustering apparatus provided in the embodiments of this specification;

[0023] Figure 5 This is a schematic diagram of a system structure for compound clustering provided in the embodiments of this specification. Detailed Implementation

[0024] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0025] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0027] Additionally, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that practice can be carried out without these specific details.

[0028] Small molecule compound clustering is commonly used for analyzing high-throughput screening results, virtual screening, or docking studies. Traditional cheminformatics-based clustering methods suffer from low identification efficiency and slow speed. Even when using compound fingerprint features for identification, the computational and storage requirements are excessive, limiting the number of compounds that can be identified.

[0029] In view of this, the inventors found that the results of machine learning clustering of small molecule compounds in existing technologies are often vague and inaccurate, making it impossible to identify fixed types in the compound information, resulting in processing results that are of no use. Even when using the similarity of compound fingerprint features, the high processing and storage space consumption prevents small molecule compound clustering from being expanded to a compound library of more than 100,000 compounds, thus causing limitations in small molecule compound clustering.

[0030] Based on this, the embodiments of this specification propose a compound clustering processing scheme: Figure 1 This is a schematic diagram illustrating an application of compound clustering provided in the embodiments of this specification. For example... Figure 1 As shown, the sample includes a compound sample 11 to be identified, such as a large compound. The compound sample 11 to be identified is segmented into a sample subset containing an initial identification label; based on the sample subset, a sample legend is obtained; based on the sample legend and the identification label, the target identification result corresponding to the compound sample to be identified is obtained (e.g., including category 1, category 2... category n-1 and category n).

[0031] In practice, it can be executed by a single entity, such as server 10. The server includes terminal devices capable of running software, including but not limited to computers, tablets, and mobile phones.

[0032] This specification proposes a compound clustering method. The method involves acquiring samples of compounds to be identified and segmenting them into subsets containing initial identification labels. Based on these subsets, a sample legend is generated. Finally, based on the sample legend and the identification labels, the target identification result corresponding to the compound sample is obtained. This coarse-grained clustering method, combined with compound image processing, significantly reduces the molecular feature extraction process with minimal feature usage, resulting in a substantial improvement in computational efficiency. Furthermore, by utilizing compound image processing based on the preliminary clustering results to complete the clustering of small molecule compounds, a more efficient, data-rich, and accurate small molecule compound clustering process is achieved.

[0033] The application scenarios described above are merely illustrative for the purpose of understanding this application, and the embodiments described herein are not limited in any way. Rather, the embodiments described herein can be applied to any applicable scenario.

[0034] The technical solutions provided by the various embodiments of this application are described below with reference to the accompanying drawings.

[0035] Figure 2 This specification provides a method flow for compound clustering in its embodiments. Figure 1 .like Figure 2As shown, the method may include steps S210 to S230. Step S210 involves obtaining a sample of the compound to be identified and segmenting the sample into a subset containing an initial identification label.

[0036] In this embodiment, the compound samples to be identified include large compounds, and the aim is to subdivide them into small groups of similar small molecule compounds using a special clustering method. In some embodiments, the compound samples to be identified are represented in the SMILES sequence text format, which can display the chemical properties of the compounds (e.g., molecular and atomic characteristics). In the case of large amounts of data, the compound samples to be identified are initially divided into multiple sample subsets, thereby initially segmenting the complex sample and facilitating subsequent processing. Small molecule compounds, from a chemical perspective, generally refer to biologically functional molecules with a molecular weight of less than 1000 Daltons; from a biological perspective, they generally refer to bioactive small peptides, oligopeptides, oligosaccharides, oligonucleotides, vitamins, minerals, small water clusters, etc.; and from a nutritional perspective, small molecules can also be classified as proteins, fats, sugars, etc.

[0037] In some embodiments, existing stock of compounds to be identified can be clustered according to their compound attribute characteristics using general statistical clustering methods. This not only segments the samples into multiple subsets but also generates initial identification labels (i.e., representations of the types and numbers of similar small molecule compounds identified in the clustering). To achieve more accurate compound clustering, simple dimensionality reduction is performed to facilitate faster identification of the final target results. Statistical clustering methods include K-means (k-means clustering algorithm) and OPTICS (Ordering points to identify the clustering structure). In other embodiments, with large volumes of compounds to be identified, no preliminary clustering process is performed. Therefore, statistical clustering methods are used to initially divide all the samples into multiple subsets, generating initial identification labels during the segmentation process. The implementation process is similar to that of K-means or OPTICS and will not be elaborated here.

[0038] Step S220: Obtain the sample legend based on the sample subset.

[0039] In conjunction with the above embodiments, after segmenting the sample of compounds to be identified into sample subsets, it is necessary to obtain sample legends corresponding to all the sample of compounds to be identified in each sample subset. Specifically, the sample legends are converted into corresponding discrete mathematical graphs according to the representation format of the sample of compounds to be identified. This transforms the processable compounds to be identified into more easily identifiable discrete mathematical graphs based on their attribute characteristics. Furthermore, the fingerprint features of each compound to be identified are more prominent in the discrete data graph, thus facilitating the identification of the most similar feature parts among multiple compounds and improving the accuracy of small molecule compound clustering.

[0040] Step S230: Based on the sample illustration and the identification label, obtain the target identification result corresponding to the compound sample to be identified.

[0041] Specifically, identification tags are used to identify the category of compounds based on their characteristics, including at least one or more. Sample legends use these identification tags to identify the category of a single group of similar small molecule compounds, thereby obtaining the target identification results for all sample legends corresponding to the compounds to be identified in the sample subset. The identification tags include the type and number of small molecule compounds, etc.

[0042] To improve the accuracy of small molecule compound clustering, it is necessary to accurately obtain identification labels for the initially segmented sample subsets, thereby obtaining more accurate target identification results for all the compound samples to be identified.

[0043] In some embodiments, the identification label includes a legend label, which can better highlight the fingerprint features corresponding to each class of similar small molecule compounds. Therefore, after initially segmenting the compound samples to be identified into multiple coarse-grained sample subsets, and converting each compound sample to be identified in the sample subset, represented in sequence text format, into a sample legend, when obtaining the target identification result corresponding to the compound sample to be identified based on the sample legend, it is not limited to the initial identification label judging the category of the compound to be identified based on the molecular features or atomic properties of the compound. This not only improves the accuracy of small molecule compound clustering, but also, based on the coarse-grained initial identification segmentation of the compound samples to be identified, divides a large amount of data of compound samples to be identified into a relatively small range of data processing, realizing data dimensionality reduction, and making room for subsequent sample legend identification process. That is, data dimensionality reduction improves the utilization rate of data processing space, thereby accelerating the processing speed of small molecule compound clustering.

[0044] In some embodiments, feature extraction is performed on sample legends, combined with the fingerprint features corresponding to the compounds, and legend labels are obtained through training with large datasets. These legend labels highlight specific ranges of fingerprint features within the connected space corresponding to each class of small molecule compounds, and are used to cluster compounds with the most similar substructures into the same class. The most similar substructure represents the connected components of the small molecule compound graph, which may include, for example, one or more identical atoms and chemical bonds.

[0045] Specifically, image feature points are extracted from the sample images. For each sample image, a topological comparison is performed on each molecule in the compound to be identified. Combined with the fingerprint features of the corresponding compound, the most similar subgraph feature is obtained from the image features. This means that the most similar subgraph feature among compounds can be found by identifying prominent and unique discriminant features in the image features. Simultaneously, a similarity score is calculated, and a threshold is used to determine whether they belong to the same category. After training with a large amount of data, if the similarity score of the most similar subgraph feature in any pair of sample images meets the threshold, the two sample images are determined to belong to the same category. Therefore, compounds with a common most similar subgraph in all sample images within the sample subset are assigned to the same category, and a label corresponding to that category is obtained. The label includes the most similar subgraph and the threshold corresponding to the similarity score.

[0046] In some embodiments, obtaining the target identification result corresponding to the compound sample to be identified based on the sample legend and the identification label includes: clustering the compounds to be identified in the sample legend that satisfy the initial identification label and the legend label into the same category based on the sample legend, the initial identification label and the legend label; and obtaining the identification category corresponding to all the compound samples to be identified based on the initial identification label and the legend label corresponding to different categories of compounds.

[0047] In this embodiment, the identification labels include initial identification labels and legend identification labels. Each category of compounds corresponds to a set of initial identification labels and legend identification labels. Therefore, during the small molecule compound clustering process of the sample legends, by combining the initial identification labels (e.g., Ti1) and legend identification labels (e.g., Pt1), the compounds in the sample legends within the sample subset that simultaneously satisfy both the initial identification labels and legend labels are clustered into the same category (e.g., compound 1). Then, based on the initial identification labels and legend labels corresponding to different categories of compounds (e.g., [Ti1, Pt1], [Ti2, Pt2], ..., [Ti5, Pt5]...), the identification categories (e.g., compound 1, compound 2...) corresponding to all the compounds to be identified are obtained respectively.

[0048] In some embodiments, clustering the compounds to be identified in the sample legends that satisfy the initial identification label and the legend label into the same category includes: obtaining the sample standard image corresponding to the initial identification label and the legend label reaching a preset threshold; performing similarity calculation on all sample legends and the sample standard image; if the sample legend matches the sample standard image, then clustering the compounds to be identified corresponding to the sample legend into the same category.

[0049] In conjunction with the above embodiments, when the compounds to be identified corresponding to both the initial identification label and the legend label in the sample legend simultaneously satisfy the condition that they are clustered into the same category, it is necessary to obtain the sample standard image corresponding to the same set of initial identification labels and legend labels reaching a preset threshold. Reaching the preset threshold for the same set of initial identification labels and legend labels may include the sum of the weights corresponding to the initial identification labels and legend labels reaching the preset threshold. Furthermore, clustering small molecule compounds can be achieved more accurately based on the most similar substructure contained in the legend labels. Therefore, after obtaining the sample standard image corresponding to the initial identification label and the legend label reaching the preset threshold, a similarity calculation is performed between all sample legends and the sample standard image corresponding to each category. If a sample legend matches a sample standard image, the compounds to be identified corresponding to the sample legend are clustered into the same category as the compounds corresponding to the sample standard image. The calculation of matching between sample legends and sample standard images can be performed using the following formula:

[0050] S = G1∩G2 (Formula 1)

[0051] Here, G1 and G2 are the input sample legend and sample standard image, respectively. The similarity algorithm in Formula 2 can be used to determine whether they match.

[0052]

[0053] Here, A and B represent the number of nodes in the sample legend and the standard sample image, respectively, and |A∩B| represents the number of common nodes between images A and B. That is, when J(A,B) equals a preset threshold, the sample legend matches the standard sample image. Finally, the compounds to be identified corresponding to the sample legend are clustered into the same category as the compounds corresponding to the standard sample image.

[0054] In some embodiments, obtaining a sample legend based on the sample subset includes:

[0055] Based on the attribute characteristics of the compounds in the sample subset, each compound sample to be identified in the sample subset is converted into a corresponding sample legend.

[0056] Specifically, based on the attribute characteristics of the compounds to be identified in the sample subset, such as the molecular characteristics, logP (oil-water partition coefficient), ring number, and atomic characteristics of the compounds, the data is converted into a data graph with atoms as nodes and chemical bonds as edges, based on the data representation of the compounds to be identified (e.g., compounds represented by SMILES sequence text). In some examples, nodes include atom number attributes, and edges include single, double, and triple bond attributes.

[0057] In some embodiments, after obtaining the target identification result corresponding to the compound sample to be identified, the method further includes: outputting the target identification result and storing the compound sample to be identified corresponding to the target identification result. This enables downstream applications or data storage of small molecule compound clustering. Therefore, the compound clustering method of the present invention can not only quickly and accurately achieve small molecule compound clustering, but also improve the efficiency and accuracy of clustering applications of associated small molecule compounds. Further illustrative examples are provided below.

[0058] Figure 3 This is a method flow for compound clustering provided in an embodiment of the present invention. Figure 2 .like Figure 3 As shown, the preferred steps of this embodiment of the invention are as follows: Input the compound sample to be identified in SMILES sequence text format; segment the compound sample to be identified into a sample subset containing initial identification labels according to statistical clustering; convert the compound sample to be identified in SMILES sequence text format in the sample subset into a discrete mathematical graph; calculate the most similar structure and obtain the legend label by topological comparison of the discrete mathematical graph in the sample subset; obtain the target identification result corresponding to the compound sample to be identified based on the sample legend (i.e., discrete mathematical graph), the initial identification label and the legend label; output the clustering result of the SMILES sequence compound sample to be identified.

[0059] Example 1:

[0060] Step 1: Obtain the sample of the compound to be identified, represented by SMILES sequence text. Step 2: For a sample subset corresponding to a small molecule compound cluster obtained through statistical clustering, for example, the initial identification label is Ti2. Statistical clustering methods include K-means and OPTICS. Step 3: Convert the compound to be identified in the sample subset, represented by SMILES sequence text, into a corresponding sample legend. Step 4: Obtain the target identification result corresponding to the sample of the compound to be identified based on the sample legend and the identification label. For example, based on the initial identification label (Ti2), compounds that do not belong to the Ti2 class are removed from the sample subset. This also includes obtaining the standard sample image corresponding to the initial identification label and the standard sample image reaching a preset threshold based on the sample legend, the initial identification label, and the legend label; performing similarity calculations on all sample legends and the standard sample images; if the sample legend matches the standard sample image, the compound to be identified corresponding to the sample legend is clustered into the same category, for example, identifying compounds belonging to the Ti2 class in the sample subset; ultimately, compounds belonging to the Ti2 class are removed more accurately. This yields the target identification results for all compounds to be identified. Step 5: Output the clustered sample subset for subsequent downstream analysis applications or data storage.

[0061] Example 2:

[0062] Step 1: Obtain the compound samples to be identified, represented by SMILES sequence text. Step 2: Calculate chemical properties using functions within the Python RDKit package (a programming language), and obtain sample subsets corresponding to at least two small molecule compound clusters through statistical clustering. For example, the initial identification labels may be at least two of Ti1, Ti2, Ti3, and Ti4. Statistical clustering methods include K-means and OPTICS. Step 3: Convert the compounds to be identified in the sample subsets, represented by SMILES sequence text, into corresponding sample legends. Step 4: Based on the sample legends and identification labels, obtain the target identification results corresponding to the compound samples. For example, based on the initial identification labels (e.g., Ti1, Ti2, and Ti3), the compounds in the sample subset may be accurately identified as belonging to the three categories Ti1, Ti2, and Ti4. This includes clustering compounds in the sample legend that satisfy both the initial identification label and the legend label into the same category; and obtaining the identification category corresponding to all compounds in different categories based on the initial identification label and the legend label. Step 5: Output the clustered sample subset for subsequent downstream analysis applications or data storage.

[0063] Example 3:

[0064] Step 1: Obtain the compound samples to be identified, represented by SMILES sequence text. Step 2: Obtain the features of the compounds to be identified, such as Morgan Fingerprints (molecular fingerprints) or deep learning-trained embedding variables; and use the PCA algorithm to reduce the data dimensionality to 10-100 dimensions. Then, use statistical clustering to obtain sample subsets corresponding to at least two small molecule compound clusters, for example, initial identification labels of at least two of Ti1, Ti2, Ti3, and Ti4. Statistical clustering methods include, for example, K-means. Step 3: Convert the compounds to be identified in the sample subsets, represented by SMILES sequence text, into corresponding sample legends. Step 4: Based on the sample legends and identification labels, obtain the target identification results corresponding to the compound samples. For example, based on the initial identification labels (Ti1, Ti2, and Ti3), the compounds in the sample subset are ultimately accurately identified as belonging to the four categories Ti1, Ti2, Ti4, and Ti5. This includes clustering compounds in the sample legend that satisfy both the initial identification label and the legend label into the same category; and obtaining the identification category corresponding to all compounds in different categories based on the initial identification label and the legend label. Step 5: Output the clustered sample subset for subsequent downstream analysis applications or data storage.

[0065] Figure 4 This is a schematic diagram of a compound clustering apparatus provided in the embodiments of this specification, as shown below. Figure 4 As shown, the device 30 includes:

[0066] The acquisition module 31 is used to acquire a sample of the compound to be identified and to segment the sample of the compound to be identified into a sample subset containing an initial identification label;

[0067] Module 32 is used to obtain a sample legend based on the sample subset.

[0068] The output module 33 is used to obtain the target identification result corresponding to the compound sample to be identified based on the sample legend and the identification label; wherein the identification label includes the initial identification label.

[0069] Figure 4 The apparatus of the illustrated embodiment can be used to perform corresponding actions. Figure 2 The steps in the method embodiments shown are implemented in a similar manner and have similar technical effects, and will not be repeated here.

[0070] Figure 5This is a schematic diagram of a system structure for compound clustering provided in the embodiments of this specification, such as... Figure 5 As shown, the system 40 includes: a processor 41, a memory 42, and a computer program; wherein

[0071] The memory 42 is used to store the computer program, and the memory may also be flash memory. The computer program is, for example, an application program or functional module that implements the above method.

[0072] The processor 41 is configured to execute the computer program stored in the memory to implement the various steps performed by the device in the above method. For details, please refer to the relevant descriptions in the preceding method embodiments.

[0073] Alternatively, the memory 42 can be either standalone or integrated with the processor 41.

[0074] When the memory 42 is a device independent of the processor 41, the device may further include:

[0075] Bus 43 is used to connect the memory 42 and the processor 41.

[0076] The present invention also provides a readable storage medium storing a computer program, which, when executed by a processor, is used to implement the methods provided in the various embodiments described above.

[0077] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0078] The present invention also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the device to implement the methods provided in the various embodiments described above.

[0079] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.

[0080] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the product embodiments described later are relatively simple since they correspond to the methods; relevant parts can be referred to the descriptions in the system embodiments.

[0081] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for clustering compounds, characterized in that, The method includes: Obtain a sample of the compound to be identified, and segment the sample into a subset containing an initial identification label; Based on the sample subset, a sample legend is obtained; wherein the sample legend is obtained by converting the compounds in the sample subset into discrete mathematical graphs; Based on the sample legend and the identification label, the target identification result corresponding to the compound sample to be identified is obtained; wherein, the identification label includes the initial identification label and the legend label; clustering the compounds to be identified in the sample legend that satisfy the initial identification label and the legend label into the same category includes: obtaining the sample standard image corresponding to the initial identification label and the legend label reaching a preset threshold; performing similarity calculation on all sample legends and the sample standard image, and if the sample legend matches the sample standard image, then clustering the compounds to be identified corresponding to the sample legend into the same category; In this process, feature extraction is performed on the sample legends, and fingerprint features corresponding to the compounds are combined to train legend labels, which are used to highlight the fingerprint features corresponding to each type of small molecule compound.

2. The method according to claim 1, characterized in that, Based on the sample legend and identification label, the target identification result corresponding to the compound sample to be identified is obtained, including: Based on the sample legend, the initial identification label, and the legend label, the compounds to be identified in the sample legend that satisfy the initial identification label and the legend label are clustered into the same category; Based on the initial identification labels and legend labels corresponding to different categories of compounds, the identification categories corresponding to all compounds to be identified are obtained.

3. The method according to claim 1, characterized in that, Based on the sample subset; The sample legend is obtained, including: Based on the attribute characteristics of the compounds in the sample subset, each compound sample to be identified in the sample subset is converted into a corresponding sample legend.

4. The method according to claim 1, characterized in that, After obtaining the target identification result corresponding to the compound sample to be identified, the method further includes: The target identification result is output, and the corresponding compound sample to be identified is stored.

5. A device for clustering compounds, characterized in that, The device includes: An acquisition module is used to acquire a sample of a compound to be identified and to segment the sample into a subset containing an initial identification label; A module is configured to obtain a sample legend based on the sample subset; wherein the sample legend is obtained by converting compounds in the sample subset into discrete mathematical graphs; The output module is used to obtain the target identification result corresponding to the compound sample to be identified based on the sample legend and the identification label; wherein, the identification label includes the initial identification label and the legend label; clustering the compounds to be identified in the sample legend that satisfy the initial identification label and the legend label into the same category includes: obtaining the sample standard image corresponding to the initial identification label and the legend label reaching a preset threshold; performing similarity calculation on all sample legends and the sample standard image, and if the sample legend matches the sample standard image, then clustering the compounds to be identified corresponding to the sample legend into the same category; In this process, feature extraction is performed on the sample legends, and fingerprint features corresponding to the compounds are combined to train legend labels, which are used to highlight the fingerprint features corresponding to each type of small molecule compound.

6. A system for clustering compounds, characterized in that, include: The device includes a memory, a processor, and a computer program, the computer program being stored in the memory, and the processor running the computer program to perform the method for clustering compounds according to any one of claims 1 to 4.

7. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, is used to implement the method for clustering compounds according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Machine learning systems for automated pharmaceutical molecule screening and scoring

    US11127488B1