Data management program, data management system, and data management method
By generating and storing a list of compounds, the problem of low efficiency in compound information management and retrieval in existing technologies is solved, and automated management and convenient retrieval of compound information are realized.
Patent Information
- Application Number
- CN202380096846.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-11-21
AI Technical Summary
现有技术难以自动生成与关注事项对应的化合物列表,无法高效管理和检索化合物相关信息。
Through data management programs and systems, information on compounds of interest is obtained, interest IDs and compound IDs are generated, and the associations are established. A list of compounds is generated and stored in the database, thereby achieving automated management and retrieval of compound information.
It enables the automatic generation of compound lists corresponding to matters of interest, improving the efficiency of compound information management and the convenience of retrieval.
Smart Images

Figure CN121002579A_ABST
Abstract
Description
Technical Field
[0001] One aspect of the present invention relates to a data management program, a data management system, and a data management method. Background Technology
[0002] Various information processing methods related to compounds are known. For example, Patent Document 1 describes a material property prediction device that effectively generates effective compound characteristic quantities reflecting expert insights and predicts the physical properties of unknown compounds with high accuracy. This device uses a material database containing multiple case databases to predict material properties. The records included in the case databases correlate structural information related to the material's structure with material properties related to the material's characteristics.
[0003] Previous technical documents
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent Application Publication No. 2021-39534 Summary of the Invention
[0006] The technical problem to be solved by the invention
[0007] The desired structure is to automatically generate a list of compounds corresponding to the areas of interest.
[0008] means for solving technical problems
[0009] The data management program according to one aspect of the present invention enables a computer to perform the following steps: acquiring concern information for collecting one or more compounds; accessing a prescribed information source and collecting compound information representing one or more compounds corresponding to the concern information from the information source; generating a concern ID as an identifier for uniquely identifying the concern information; generating a compound ID as an identifier for uniquely identifying each of the one or more compounds represented by the collected compound information; associating each of the one or more compound IDs with the concern ID to generate a compound list; and storing the compound list in a database.
[0010] The data management system according to one aspect of the present invention includes at least one processor, which performs the following processes: acquiring concern information for collecting one or more compounds; accessing a predetermined information source and collecting compound information representing one or more compounds corresponding to the concern information from the information source; generating a concern ID as an identifier for uniquely identifying the concern information; generating a compound ID as an identifier for each of the one or more compounds represented by the collected compound information; associating each of the one or more compound IDs with the concern ID to generate a compound list; and storing the compound list in a database.
[0011] One aspect of the present invention relates to a data management method executed by a data management system having at least one processor. The data management method includes the following steps: acquiring concern information for collecting one or more compounds; accessing a predetermined information source and collecting compound information representing one or more compounds corresponding to the concern information from the information source; generating a concern ID as an identifier to uniquely determine the concern information; generating a compound ID as an identifier to uniquely determine each of the one or more compounds represented by the collected compound information; associating each of the one or more compound IDs with a concern ID to generate a compound list; and storing the compound list in a database.
[0012] In this approach, one or more compounds corresponding to a concern are collected, and the concern information and each compound are linked together by identifiers. The set of compounds corresponding to the concern information is generated and saved as a compound list. With this structure, a list of compounds corresponding to a specified concern can be automatically generated.
[0013] Invention Effects
[0014] According to one aspect of the present invention, a list of compounds corresponding to matters of interest can be automatically generated. Attached Figure Description
[0015] Figure 1 This is a diagram illustrating an example of a data management system.
[0016] Figure 2 This is a diagram representing an example of a data structure.
[0017] Figure 3 This is a diagram showing an example of the structure of compound ID.
[0018] Figure 4 This is a flowchart illustrating an example of a process that generates data related to compounds.
[0019] Figure 5 This is a diagram illustrating an example of a method for generating compound IDs.
[0020] Figure 6 This is a flowchart illustrating an example of the process for retrieving compounds.
[0021] Figure 7 This is a diagram showing examples of search results.
[0022] Figure 8 This is a diagram showing examples of search results.
[0023] Figure 9 This is a diagram showing examples of search results. Detailed Implementation
[0024] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are labeled with the same symbols and repeated descriptions are omitted.
[0025] [System Structure]
[0026] The data management system involved in this invention is a computer system for managing information related to compounds such as organic compounds. In one example, the data management system has at least one of a data generation function that generates and stores data related to compounds, and a data retrieval function that retrieves compounds corresponding to search criteria. The data retrieval function is also referred to as filtering.
[0027] A data management system consists of one or more computers. When multiple computers are used, they are connected via communication networks such as the Internet and intranets, thus logically forming a unified data management system.
[0028] A computer constituting a data management system typically comprises a processor, memory, and communication interfaces as hardware devices. The processor is, for example, a CPU. Memory consists of storage devices such as flash memory and hard disks. The various functions of the data management system are implemented by the processor executing programs stored in memory.
[0029] A data management program used to enable a computer to function as a data management system includes program code for implementing the various functional modules of the data management system. This data management program may be provided after being non-transitoryly recorded on a tangible recording medium such as a CD-ROM, DVD-ROM, or semiconductor memory. Alternatively, the data management program may be provided via a communication network as a data signal superimposed on a carrier wave. The provided data management program may be, for example, recorded in memory.
[0030] Figure 1This is a diagram illustrating a data management system 10 as described in an example. In this example, the data management system 10 is connected to an information source 20, a database 30, and one or more user terminals 40 via a communication network. The communication network may include at least one of the Internet and an intranet.
[0031] Information source 20 is a collection of at least one computer system accessed through the functional modules of data management system 10 for obtaining information related to the compound. In one example, information source 20 may be at least one collection of publicly available databases, websites, and search engines. For example, information source 20 may include databases such as SciFinder, PubChem, NIST ChemistryWebBook, and CAS, and search engines such as Google Patents.
[0032] Database 30 is a device for non-temporarily storing various types of data used by the data management system 10. Database 30 can be a component of the data management system 10, or it can be located in a computer system different from the data management system 10. In short, database 30 is accessed through the functional modules of the data management system 10.
[0033] User terminal 40 is a computer operated by a user of data management system 10. A user is a person who wants to use data management system 10 to prepare or obtain information related to the compound. User terminal 40 can be various computers such as personal computers, workstations, tablet terminals, smartphones, and wearable terminals.
[0034] The data management system 10 includes a processor 101. In one example, the processor 101 functions as a collection unit 11, a generation unit 12, and a retrieval unit 13. The collection unit 11 is a functional module that collects information related to compounds from the information source 20. The generation unit 12 is a functional module that generates new data records related to compounds based on the collected information and stores these data records in the database 30. The retrieval unit 13 is a functional module that extracts more than one data record related to one or more compounds corresponding to the search conditions received from the user terminal 40 from the database 30 and provides the search results based on the extracted data records to the user terminal 40.
[0035] [Data Structures]
[0036] (Relationships between data)
[0037] refer to Figure 2 The structure of the various data used in the data management system 10 is explained. Figure 2 This is a diagram illustrating an example of this data structure. In one example, database 30 stores concern data 31, compound data 32, and resource data 33.
[0038] Data 31 on matters of concern is electronic data representing various matters of concern. A matter of concern refers to a criterion used to collect and group one or more compounds. A matter of concern can also be described as the category to which one or more compounds belong. Alternatively, a matter of concern can be described as a condition satisfied by the compounds to be collected. Matters of concern can be related to the characteristics or properties of a compound. For example, matters of concern can represent elements, functional groups, or other components; molecular weight, polarity, or other physical properties; the person or organization that discovered or synthesized the compound; or various other matters such as year, country, academic society, paper, media, etc. In one example, each data record in Data 31 on matters of concern includes a matter of concern ID and matter of concern information as data items. The matter of concern ID is a unique identifier that identifies each matter of concern. The matter of concern information is information representing the matter of concern. The matter of concern information can be described using one or more keywords (meta-information) such as "epoxy resin, 2010-2012, Japan".
[0039] Compound data 32 is electronic data representing a list of one or more compounds related to a matter of interest. In this invention, this list is also referred to as a "compound list." Compound data 32 can also be described as a collection of one or more compound lists. Each data record in compound data 32 contains a matter of interest ID, a compound ID, and compound attributes as data items. The compound ID is a unique identifier that identifies each compound. Compound attributes are information representing the characteristics or properties of the compound. In a compound list, one matter of interest ID corresponds to one or more combinations of compound IDs and compound attributes. Therefore, in a compound list, compound IDs and their combinations have a one-to-N relationship.
[0040] Compound attributes are represented as a collection of one or more sub-data items. For example, compound attributes may include sub-data items such as compound name, SMILES (Simplified Molecular Input Line Entry System), molecular weight, chemical formula, composition formula, InChI (International Chemical Identifier), InChIKey, polarity, refractive index, dipole moment, content, and precise mass. SMILES and InChI are both single-line strings that describe the structure of the compound according to prescribed rules. InChIKey is a fixed-length string obtained by hashing InChI. SMILES, InChI, and InChIKey are all methods for representing compounds.
[0041] Resource data 33 is electronic data representing a list of one or more resources related to a matter of concern. In this invention, this list is also referred to as a "resource list." A resource refers to a source of information related to the compound. Examples of resources include databases, literature, and websites. Resource data 33 can also be described as a collection of one or more resource lists. Each data record in resource data 33 contains a matter of concern ID, a resource ID, and resource attributes as data items. The resource ID is a unique identifier that identifies each resource. Resource attributes are information representing the characteristics or properties of a resource. In a resource list, one matter of concern ID corresponds to one or more combinations of resource IDs and resource attributes. Therefore, in a resource list, resource IDs and their combinations have a one-to-N relationship.
[0042] Resource attributes are represented as a collection of one or more sub-data items. For example, a resource attribute can include sub-data items such as resource name, person's name, organization name, and access information. A person's name can be the name of the author, discoverer, developer, or administrator. An organization name can be the name of the organization to which such a person belongs. Access information is the information required by others to access the resource; for example, it can be represented by book information, paper information, URL, etc.
[0043] The data for matters of interest 31 is associated with both the compound data (compound list) 32 and the resource data (resource list) 33 based on the matter of interest ID. Because information about compounds is obtained from resources, each compound can be directly associated with a resource. However, in database 30, compound IDs and resource IDs are not directly associated; instead, the association is established indirectly through the matter of interest ID. That is, the correspondence between compounds and resources is determined by the matter of interest.
[0044] (Compound ID)
[0045] Compound IDs can be represented using SMILES, InChI, InChIKey, or serial numbers. Alternatively, compound IDs can be represented using a coding system different from these well-known techniques.
[0046] In this coding system, a compound ID includes: a first property string (using one or more alphanumeric characters to represent the first property of the compound); a component string (using one or more English letters to represent one or more constituents of the compound); and a second property string (using one or more English letters to represent the second property of the compound). Both the first and second properties are numerical values representing the property of the compound. Examples of properties expressed by the first and second properties include molecular weight, polarity, refractive index, dipole moment, and content. The represented property differs between the first and second properties. The component string indicates whether each of the one or more constituents of interest is included in the compound. Each of the one or more constituents represented by the component string can be an element or a functional group.
[0047] If we consider a compound database storing information on approximately 2 trillion compounds, then by arranging the 26 classes of the English alphabet into 9-digit sequences without distinguishing between uppercase and lowercase letters, it would be possible to distinguish approximately 5 trillion (=26 trillion) compounds. 9 Compounds. By adding at least one digit to this 9-digit number to describe a particular property, it is possible to prepare a compound ID that uniquely identifies approximately 5 trillion compounds and provides information about that property. Therefore, the compound ID is represented as a string of 10 or more characters. The reason for setting the number of characters in the compound ID to 30 or less is to take into account the results of the InChIKey, which has a character limit of 27. The upper limit for the number of characters in the compound ID can be the same as that of the InChIKey, which is 27.
[0048] Figure 3 This diagram illustrates an example of a compound ID structure represented using code system 190. In code system 190, the compound ID consists of 15 characters. The first digit represents the first property string, the second to eighth digits represent the component strings, and the ninth to fifteenth digits represent the second property string.
[0049] The first property string represents the molecular weight of the compound as the first property. The characters used to assign the first property string are 62 alphanumeric characters: 0-9, a-z, and A-Z. A first property string with a length of 1 represents the hundreds and thousands digits of the molecular weight. Therefore, the first property string can represent molecular weights greater than 0 and less than 6200. Figure 3 Based on the example above, if the first property string of a compound ID is "b", then the molecular weight of the compound represented by that compound ID is greater than 1300 and less than 1400. In one example, the molecular weight is associated with each alphanumeric character in ascending order.
[0050] The component string indicates whether a component is present in the compound for each of the seven components of interest (components 1 through 7). In one example, components 1 through 7 are nitrogen (N), fluorine (F), silicon (Si), phosphorus (P), sulfur (S), chlorine (Cl), and carbonyl (O=). The characters assigned to each position in the component string are English letters, specified using either uppercase or lowercase. If the compound contains a component, an uppercase letter is assigned to the position corresponding to that component. If the compound does not contain that component, a lowercase letter is assigned to that position. The uppercase and lowercase letter assignments can be reversed in this example.
[0051] The second property string represents the polarity of the compound as expressed by the octanol / water partition coefficient (LogP). The second property string is represented by a binary number from 0 to 127 (=2). 7 The integer is a binary number. The characters assigned to each position in the second property string are English letters, designated as either uppercase or lowercase. If the binary value of a position is "1", an uppercase letter is assigned to that position; if the value is "0", a lowercase letter is assigned. For example, if LogP = 7, the binary number is labeled "0000111", therefore lowercase letters are assigned to positions 9 through 12, and uppercase letters are assigned to positions 13 through 15. The assignment of uppercase and lowercase letters can also be reversed in this example.
[0052] In this invention, the labeling method of compound IDs represented by code system 190 is called "analytical compound representation method" or "ARKey".
[0053] The first property, composition, and second property are examples of compound attributes. Therefore, a compound ID, as described by code system 190, serves not only as an identifier to uniquely identify each compound but also as representing at least one compound attribute. In contrast, the InChIKey, as a fixed-length string, functions as an identifier but does not represent a compound attribute. While the number of characters in a compound ID representing the first property, composition, and second property is roughly the same as inInChIKey, even if it is longer, this compound ID is more advantageous than InChIKey in that it contains more information.
[0054] [System Actions]
[0055] The following is an example of the operation of the data management system 10, which serves as an example of a data management method.
[0056] (Data generation function)
[0057] refer to Figure 4The data generation function of the data management system 10 is explained. Figure 4 This is a flowchart illustrating an example of the process of generating data related to compounds, shown as process flow S1.
[0058] In step S11, the collection unit 11 acquires information on matters of interest based on instructions from the user terminal 40. The user sends instructions from the user terminal 40 to prepare information on compounds corresponding to the matters of interest. The collection unit 11 can directly acquire the information on matters of interest received from the user terminal 40. Alternatively, the collection unit 11 can generate information on matters of interest based on collection conditions received from the user terminal 40. In one example, the collection unit 11 can acquire more than one keyword (meta-information) as information on matters of interest.
[0059] In step S12, the collection unit 11 collects compound information and resource information corresponding to the information of interest from the information source 20. Compound information represents information about one or more compounds, including the individual compound attributes of those compounds. Resource information represents information about one or more resources, including the individual resource attributes of those resources. The information of interest can be considered the retrieval criteria for compound information and resource information. Therefore, "compound information corresponding to the information of interest" can be said to represent compound information about one or more compounds that satisfy its retrieval criteria, and "resource information corresponding to the information of interest" can be said to represent resource information about one or more resources that satisfy its retrieval criteria. The collection unit 11 accesses the information source 20 and collects compound information and resource information corresponding to the information of interest through methods such as database retrieval and web scraping. In database retrieval, the collection unit 11 can collect compound information or resource information through processing including set operations.
[0060] In step S13, the generation unit 12 generates a data record of the concerns data. The generation unit 12 generates a concerns ID used to identify the acquired concerns information. For example, the generation unit 12 can encode one or more keywords (meta-information) constituting the concerns information using methods such as Base64, and set the string obtained through this encoding as the concerns ID. Alternatively, the generation unit 12 can set a sequence number as the concerns ID. The generation unit 12 associates the concerns information with the concerns ID to generate a new data record.
[0061] In step S14, the generation unit 12 generates a list of compounds. The generation unit 12 determines one or more compounds by referring to the compound information, and sets a compound ID and compound attributes for each of the one or more compounds.
[0062] The generation unit 12 can set data items recorded in the compound information as compound attributes. Alternatively, the generation unit 12 can generate compound attributes by performing operations or transformations based on strings or numerical values recorded in the compound information. For example, the generation unit 12 can generate SMILES using Named Entity Recognition (NER). The generation unit 12 can use libraries such as opsin and ChemDataExtractor, or machine learning models such as a fine-tuned GPT model, to perform NER. As another example, the generation unit 12 can use cheminformatics toolkits such as Rdkit and Openbabel to generate accurate mass or composition formulas.
[0063] The generation unit 12 can set SMILES, InChI, InChIKey, or a serial number as the compound ID. Alternatively, the generation unit 12 can set a compound ID marked by the aforementioned ARKey. In one example, the generation unit 12 generates an ARKey-based compound ID based on the InChIKey determined or generated from the compound information. (See reference...) Figure 5 The generation order of these components will be explained. Figure 5 This is a diagram illustrating an example of a method for generating compound IDs based on InChIKey. Figure 5 In the example, the following are the premises.
[0064] Compound ID follows Figure 3 The code structure shown is 190.
[0065] The first physical property is molecular weight.
[0066] • The component string indicates that N, F, Si, P, S, Cl and O are components 1 to 7.
[0067] The second property is polarity, represented by LogP.
[0068] Figure 5 Compound 200 shown is 3-phenyl-1-benzothiophene-2-carboxylic acid. The InChI Key 210 for compound 200 is “XEYPANZFAQJADI-UHFFFAOYSA-N”. InChI Key 210 includes a first block 211 represented by the first 14 digits, a second block 212 represented by the 16th to 25th digits, and a final block 213 represented by the 27th digit. The first block 211 represents the structure of compound 200 and functions as an identifier. The second block 212 represents information related to isotopes, optical isomers, etc. The final block 213 represents charge information.
[0069] The generation unit 12 extracts the first block 211 from InChIKey 210. Then, the generation unit 12 changes the string arrangement of this first block 211. Figure 5 In the example, generation unit 12 reverses the arrangement of the string in block 211. That is, generation unit 12 moves the first digit of block 211 to the fourteenth digit, the second digit to the thirteenth digit, the third digit to the twelfth digit, ... the thirteenth digit to the second digit, and the fourteenth digit to the first digit. Generation unit 12 can also change the arrangement of the string in block 211 using other methods. Then, generation unit 12 uses the string 220 obtained through this change to generate compound ID 230 representing information related to molecular weight, composition, and polarity.
[0070] The generation unit 12 appends one digit to the converted string 220 as the first property string 231. Then, the generation unit 12 determines the molecular weight of compound 200 with reference to the acquired compound information and sets the alphanumeric value representing that molecular weight as the first property string 231. In one example, the generation unit 12 determines an alphanumeric value corresponding to the molecular weight from a set of alphanumeric values sorted in ascending order and sets the determined alphanumeric value as the first property string 231. Figure 5 In the example, the generation unit 12 sets "2" in the first property string 231 in response to the case where the molecular weight is greater than 200 and less than 300.
[0071] The generation unit 12 embeds information related to the seven components N, F, Si, P, S, Cl, and O into the first to seventh digits of the converted string 220 to set the component string 232. Referring to the acquired compound information, the generation unit 12 determines that compound 200 contains sulfur and a carbonyl group. Then, the generation unit 12 sets the fifth digit corresponding to sulfur and the seventh digit corresponding to the carbonyl group to uppercase letters, and sets the first to fourth digits and the sixth digit to lowercase letters. Through this series of processes, the first to seventh digits of the converted string 220 are converted from “IDAJQAF” to “idajQaF”. These converted seven digits become the component string 232, located at the second to eighth digits of compound ID230.
[0072] The generation unit 12 embeds polarity (LogP) related information into the 8th to 14th digits of the converted string 220, thereby setting the second property string 233. The generation unit 12 determines the polarity (LogP) of the compound 200 by referring to the acquired compound information, and sets the 8th to 14th digits in a manner that displays this polarity. Figure 5In the example, in response to LogP being 27 (=4.27 / 20×128), the generation unit 12 sets the 10th, 11th, 13th, and 14th bits to uppercase letters and the 8th, 9th, and 12th bits to lowercase letters to display the binary number marker "0011011". Figure 5 In the example, the truth value of LogP is 4.27. To standardize and make efficient use of the 7 digits, its truth value is divided by 20 and multiplied by 128. Through this series of processes, the 8th to 14th digits of the converted string 220 are changed from "ZNAPYEX" to "znAPyEX". This converted 7 digits become the second property string 233, located in the 9th to 15th digits of compound ID 230.
[0073] exist Figure 5 In the example, generation unit 12 eventually generates compound ID230 of compound 200, which is called "2idajQaFznAPyEX".
[0074] The generation unit 12 generates one or more combinations of compound IDs and compound attributes. The generation unit 12 then associates these combinations with interest IDs to generate a compound list. This compound list represents information related to one or more compounds corresponding to the acquired interest information.
[0075] In step S15, the generation unit 12 generates a resource list. The generation unit 12 determines one or more resources based on resource information, and sets a resource ID and resource attributes for each of these resources. The generation unit 12 may set a sequence number as the resource ID. The generation unit 12 generates one or more combinations of resource IDs and resource attributes. The generation unit 12 associates these combinations with a concern item ID to generate the resource list. This resource list represents information related to one or more resources corresponding to the acquired concern item information.
[0076] In step S16, the generation unit 12 stores the generated data records and lists in the database 30. The generation unit 12 stores a collection of data records of the matters of interest, a compound list, and a resource list in the database 30. These data records and lists are linked together by the matters of interest ID. Through this process, information related to compounds associated with the acquired matters of interest is stored in the database 30, and the compound can be retrieved later.
[0077] In one example, the data management system 10 executes process flow S1 for each of multiple concerns. As a result, data on the compounds of various concerns are accumulated in database 30.
[0078] (Data retrieval function)
[0079] refer to Figure 6The data retrieval function of the data management system 10 is explained. Figure 6 This is a flowchart illustrating an example of compound retrieval processing as processing flow S2. The data management system 10 executes processing flow S2 in response to receiving search criteria from a user terminal 40. This user terminal 40 may be different from or the same as the user terminal 40 that executes processing flow S1, i.e., the one instructing data generation.
[0080] In step S21, the retrieval unit 13 receives retrieval criteria from the user terminal 40. The user sends the retrieval criteria from the user terminal 40 in order to obtain information about compounds corresponding to matters of interest. The retrieval criteria can be defined by at least one of matters of interest, compound attributes, and resource attributes.
[0081] In step S22, the retrieval unit 13 extracts data related to compounds corresponding to the search criteria from the database 30. The extracted data may be at least one set of the following: data of interest 31, compound data (compound list) 32, and resource data (resource list) 33. When the compound ID is specified by ARKey, the retrieval unit 13 may perform a search that includes a comparison between the search criteria and the compound ID to extract the data.
[0082] In step S23, the retrieval unit 13 generates retrieval results based on the extracted data and sends the retrieval results to the user terminal 40. In one example, the retrieval unit 13 generates and sends retrieval results that represent the extracted data in the form of tables, charts, distribution maps, scatter plots, graphs, etc. The user terminal 40 receives and processes the retrieval results. For example, the user terminal 40 can display the retrieval results on a display screen or store the retrieval results in a designated storage device.
[0083] Figure 7 Here is an example of a search result. In this example, the search unit 13 generates a search result 300 that visualizes one or more compounds 302 and one or more resources 303 corresponding to the matter of interest 301 specified as the search criteria using a graphical structure. In the search result 300, the matter of interest 301 and each compound 302 are connected by a straight line, and the matter of interest 301 and each resource 303 are also connected by a straight line. Compounds 302 and resources 303 are not directly connected. That is, the search unit 13 does not directly describe the relationship between compounds 302 and resources 303, but rather describes it through the connection between the matter of interest 301.
[0084] Figure 8Another example of the search results is shown. In this example, the search unit 13 generates search results 310 that visualize the distribution of multiple compounds corresponding to the search criteria specified by the axis of interest, using the distribution of the component of interest and molecular weight as two axes. In the graph showing the distribution, the horizontal axis represents molecular weight, and the vertical axis represents the number of compounds for each component of interest. The component of interest is, for example, an element or a functional group. Each row of the graph represents the distribution of each molecular weight for one or more compounds having one or more components of interest. The search unit 13 represents the distribution by using one or more blocks that have undergone gradient processing. The search unit 13 sets the concentration of each block according to the number of compounds. Figure 8 In the example, the darker the color of the block, the more compounds it corresponds to.
[0085] Figure 9 Another example of the search results is shown. In this example, the search unit 13 generates search results 320 that visualize multiple compounds corresponding to the search criteria specified by the search terms, based on the distribution of the component of interest and LogP along two axes. In the graph representing the distribution, the horizontal axis represents LogP, and the vertical axis represents the number of compounds for each component of interest. Each row of the graph represents the distribution of each LogP for one or more compounds with one or more components of interest. The search unit 13 represents the distribution by using one or more blocks with gradient processing. The search unit 13 sets the concentration of each block according to the number of compounds. Figure 9 In the example, the darker the color of the block, the more compounds it corresponds to.
[0086] [Variation Example]
[0087] The technology involved in this invention has been described in detail above with reference to various examples. However, this invention is not limited to the above examples. Various modifications can be made to the technology involved in this invention without departing from its spirit.
[0088] In step S16 of processing flow S1, generation unit 12 can display at least one of the compound list and resource list on the display. For example, generation unit 12 can display one or more compounds and one or more resources corresponding to the interest using a chart structure such as search result 300. Alternatively, generation unit 12 can display the distribution of one or more compounds represented by the compound list, such as search results 310 and 320, based on the two axes of the component of interest and physical properties. Data management system 10 can display at least one of the compound list and resource list on the display in both data generation and data retrieval functions.
[0089] In the above example, the data management system 10 manages concern data 31, compound data (compound list) 32, and resource data (resource list) 33. However, the data management system may also choose not to manage at least one of the concern data and resource data.
[0090] The data management system 10 described above is equivalent to the server in a client-server system, but it can also be installed on a computer, which is equivalent to a user terminal. Alternatively, the data management system can be installed on a separate computer.
[0091] The processing order of a method executed by at least one processor is not limited to the example above. For example, some of the steps described above may be omitted, or the steps may be executed in a different order. Furthermore, any two or more of the steps described above may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be executed in addition to those described above.
[0092] In comparing the magnitude of two values in this invention, either "above" or "greater" can be used, or either "below" or "less than" can be used.
[0093] In this invention, the statement "at least one processor executes the first process, executes the second process, ... executes the nth process," or its corresponding statement, represents the concept of a situation where the processor, the executing entity of the n processes from the first process to the nth process, changes midway. That is, this statement represents the concept of two scenarios: one where all n processes are executed by the same processor, and another where the processor changes arbitrarily among the n processes.
[0094] [Postscript]
[0095] As can be understood from the above examples, the present invention includes the following methods.
[0096] (Postscript 1)
[0097] A data management program that causes a computer to perform the following steps:
[0098] Obtain information on concerns for collecting data on more than one compound;
[0099] Access the designated information source and collect compound information from the information source representing one or more compounds corresponding to the information of concern;
[0100] Generate a concern ID as an identifier to uniquely determine the concern information;
[0101] For each of the more than one compounds represented by the collected compound information, a compound ID is generated as an identifier to uniquely identify the compound;
[0102] A compound list is generated by associating each of the more than one compound ID with the interest ID; and
[0103] The list of compounds is stored in a database.
[0104] (Postscript 2)
[0105] According to the data management procedure described in Appendix 1, the computer further performs the following steps:
[0106] Data records that generate attention data by associating the attention information with the attention ID; and
[0107] In addition to the compound list, the data records are also stored in the database.
[0108] (Note 3)
[0109] According to the data management procedure described in Appendix 1 or 2, wherein,
[0110] The step of generating the concern item ID includes the following steps: setting the string obtained by encoding one or more keywords constituting the concern item information as the concern item ID.
[0111] (Note 4)
[0112] The data management procedure according to any one of Appendices 1 to 3, wherein,
[0113] The step of generating the compound ID includes the following steps: for each of the more than one compound, generating a fixed-length string containing a first property string representing the first property of the compound with one or more alphanumeric characters, a component string representing the one or more components constituting the compound with one or more English letters, and a second property string representing the second property of the compound with one or more English letters, and having a length of more than 10 characters and less than 30 characters as the compound ID.
[0114] (Note 5)
[0115] According to the data management procedure described in Appendix 4, wherein,
[0116] The first physical property is molecular weight.
[0117] The length of the first property string is 1.
[0118] The step of generating the fixed-length string as the compound ID includes the following steps: in the first property string, the hundreds and thousands digits of the molecular weight are represented by numbers, uppercase English letters, or lowercase English letters.
[0119] (Note 6)
[0120] According to the data management procedure described in Appendix 4 or 5, wherein,
[0121] The step of generating the fixed-length string as the compound ID includes the following steps: when the ingredient corresponding to the English letter is included in the compound, each of the more than one English letter in the ingredient string is represented by one of an uppercase letter and a lowercase letter; when the ingredient corresponding to the English letter is not included in the compound, each of the more than one English letter in the ingredient string is represented by the other of the uppercase letter and the lowercase letter.
[0122] (Note 7)
[0123] The data management procedure according to any one of Appendices 4 to 6, wherein,
[0124] The second property is polarity.
[0125] The second property string represents the value of the polarity as a binary number.
[0126] The step of generating the fixed-length string as the compound ID includes the following steps: when the value corresponding to the English letter is 1, each of the more than one English letter in the second property string is represented by one of an uppercase letter and a lowercase letter; when the value corresponding to the English letter is 0, each of the more than one English letter in the second property string is represented by the other of the uppercase letter and the lowercase letter.
[0127] (Postscript 8)
[0128] The data management procedure according to any one of Appendices 4 to 7, wherein,
[0129] The step of generating the fixed-length string as the compound ID includes the following steps: generating a string obtained by changing the arrangement of the 14 characters of the first block of the InChIKey constituting the compound, which will be used as the ingredient string and the second property string.
[0130] (Note 9)
[0131] According to the data management procedure described in Appendix 8, wherein,
[0132] The change in the arrangement of the 14 characters is the reversal of the arrangement of the 14 characters.
[0133] (Postscript 10)
[0134] The data management procedure according to any one of Appendices 1 to 9 causes the computer to further perform the following steps:
[0135] For each of the more than one compounds represented by the collected compound information, a compound property including the precise mass and composition of that compound is set;
[0136] For each of the more than one compounds represented by the collected compound information, the compound attribute is associated with the compound ID; and
[0137] Generate a list of compounds that further includes one or more of the compound attributes associated with the compound ID.
[0138] (Postscript 11)
[0139] According to any one of Appendices 1 to 10, the data management procedure causes the computer to further perform the following steps:
[0140] The list of compounds is displayed on the monitor.
[0141] (Postscript 12)
[0142] The data management procedure according to any one of Appendices 1 to 11 causes the computer to further perform the following steps:
[0143] Collect resource information from the information source representing one or more resources corresponding to the information of concern; for each of the one or more resources represented by the collected resource information, generate a resource ID as an identifier to uniquely identify the resource;
[0144] A resource list is generated by associating each of the more than one resource ID with the concern item ID; and
[0145] In addition to the compound list, the resource list is also stored in a database.
[0146] (Postscript 13)
[0147] According to the data management procedure described in Appendix 12, the computer further performs the following steps:
[0148] For each of the more than one resources represented by the collected resource information, a resource attribute is set that includes at least one of the person's name and the organization's name recorded in the resource;
[0149] For each of the more than one resources represented by the collected resource information, the resource attribute is associated with the resource ID; and
[0150] Generate a resource list that further includes one or more resource attributes each associated with the resource ID.
[0151] (Postscript 14)
[0152] According to the data management procedure described in Appendix 12 or 13, the computer further performs the following steps: displaying on the screen the relationship between the one or more compounds represented by the compound list and the one or more resources represented by the resource list in the form of a chart structure that connects the one or more compounds and the information of concern, but does not connect the one or more compounds and the one or more resources, and connects the one or more resources and the information of concern.
[0153] (Postscript 15)
[0154] A data management system includes at least one processor, which performs the following processing:
[0155] Obtain information on concerns for collecting data on more than one compound;
[0156] Access the designated information source and collect compound information from the information source representing one or more compounds corresponding to the information of concern;
[0157] Generate a concern ID as an identifier to uniquely determine the concern information;
[0158] For each of the more than one compounds represented by the collected compound information, a compound ID is generated as an identifier to uniquely identify the compound;
[0159] A compound list is generated by associating each of the more than one compound ID with the interest ID; and
[0160] The list of compounds is stored in a database.
[0161] (Postscript 16)
[0162] A data management method, executed by a data management system having at least one processor, the data management method comprising the following steps:
[0163] Obtain information on concerns for collecting data on more than one compound;
[0164] Access the designated information source and collect compound information from the information source representing one or more compounds corresponding to the information of concern;
[0165] Generate a concern ID as an identifier to uniquely determine the concern information;
[0166] For each of the more than one compounds represented by the collected compound information, a compound ID is generated as an identifier to uniquely identify the compound;
[0167] A compound list is generated by associating each of the more than one compound ID with the interest ID; and
[0168] The list of compounds is stored in a database.
[0169] According to notes 1, 15, and 16, one or more compounds corresponding to the information of concern are collected. The information of concern and each compound are linked together by identifiers, and a set of compounds corresponding to the information of concern is generated and saved as a compound list. This structure allows for the automatic generation of a list of compounds corresponding to a specified concern. Furthermore, the list of compounds corresponding to a concern can be retrieved later.
[0170] According to Note 2, the acquired information of concern is also stored in the database, thus enabling the retrieval of the relationship between the information of concern and the compound ID from the database.
[0171] According to Note 3, by encoding the attention information to set attention IDs, attention IDs can be easily generated and managed.
[0172] According to Appendix 4, since a compound ID that uniquely identifies a compound represents two properties and one or more components, it can be used as a reference when searching for compounds based on these three elements. That is, in addition to serving as an identifier for uniquely identifying a large number of compounds, the compound ID also serves as information for comparison with search criteria. The first property, one or more components, and the second property are represented by letters or numbers; therefore, the compound ID, which performs these two functions, is limited to a string of a fixed length of 10 to 30 characters. Furthermore, since the compound ID is represented by a fixed-length string rather than a variable-length string, it is easy to process on a computer. Through these structures associated with the compound ID, the amount of data regarding a large number of compounds can be reduced, and compound-related information processing can be performed efficiently. For example, the efficiency of retrieving compounds from a database can be improved.
[0173] According to Note 5, the molecular weight, which is the first property, is represented by a single digit. Therefore, it is possible to suppress the number of characters in the first property string while embedding sufficient information related to the molecular weight into the compound ID.
[0174] According to Appendix 6, in the ingredient string, whether a component corresponding to a given letter is included in the compound is not indicated by the letter itself or other strings, but by the distinction between uppercase and lowercase letters. Therefore, the ingredient string serves both as an identifier and as information indicating the presence or absence of a component. By using this type of ingredient string, the number of characters in the string can be reduced. This conserves space required for storing compound IDs in the database and improves the efficiency of computer processing of compound IDs.
[0175] According to Appendix 7, the second property string represents the polarity value expressed in binary numbers. The values of each digit of this binary number, i.e., 0 or 1, are distinguished by uppercase and lowercase letters in the English letters of the second property string. The second property string serves both as part of an identifier and as information representing the value of the second property. By using this second property string, the number of characters in the second property string can be reduced. Therefore, the area required for storing compound IDs in the database can be saved, improving the processing efficiency of compound IDs by computers.
[0176] According to Appendix 8, by using the first block of the InChIKey, which functions as an identifier, a compound ID can be easily prepared as an identifier for the compound. Furthermore, by changing the character arrangement within this first block, compound IDs that are difficult to immediately deduce from the InChIKey can be prepared.
[0177] According to Note 9, by reversing the first block of InChIKey, the ingredient string and the second property string can be obtained, thus making it easier and more reliable to prepare compound IDs that function as identifiers of compounds.
[0178] According to Appendix 10, compound properties, including precise quality and compositional information important from an analytical perspective, are stored as part of a compound list in the database. This structure allows for the automatic generation of a list representing detailed information about compounds corresponding to matters of interest. Furthermore, detailed information about compounds corresponding to matters of interest can be retrieved later.
[0179] According to Note 11, the list of compounds is displayed on the screen, thus providing users with information about compounds corresponding to their concerns.
[0180] According to Appendix 12, more than one resource corresponding to the concern information is further collected. The concern information and each resource are associated with each other through identifiers, and a collection of resources corresponding to the concern information is generated and saved as a resource list. Through this structure, a list of resources corresponding to a concern can be automatically generated simply by specifying the concern. Furthermore, the list of resources corresponding to a concern can be retrieved later.
[0181] According to Appendix 13, resource attributes, including names of persons and organizations for convenience from the perspective of reference resources, are stored in the database as part of the resource list. This structure allows for the automatic generation of a list representing detailed information about resources corresponding to matters of interest. Furthermore, it enables the retrieval of detailed information about resources corresponding to matters of interest later.
[0182] According to Appendix 14, the relationship between compounds and resources is not based on a direct link, but rather expressed through a connection mediated by concerns. This graphical structure allows the relationship to be visually and easily understood by the user.
[0183] Symbol Explanation
[0184] 10-Data Management System, 11-Collection Department, 12-Generation Department, 13-Retrieval Department, 20-Information Source, 30-Database, 31-Data of Concerns, 32-Compound Data, 33-Resource Data, 40-User Terminal, 230-Compound ID, 231-First Property String, 232-Component String, 233-Second Property String.
Claims
1. A data management program that causes a computer to perform the following steps: Obtain information on concerns for collecting data on more than one compound; Access the designated information source and collect compound information from the information source representing one or more compounds corresponding to the information of concern; Generate a concern ID as an identifier to uniquely determine the concern information; For each of the more than one compounds represented by the collected compound information, a compound ID is generated as an identifier to uniquely identify the compound; A compound list is generated by associating each of the more than one compound ID with the interest ID; and The list of compounds is stored in a database.
2. The data management program according to claim 1, which causes the computer to further perform the following steps: Data records that generate attention data by associating the attention information with the attention ID; and In addition to the compound list, the data records are also stored in the database.
3. The data management program according to claim 1 or 2, wherein, The step of generating the concern item ID includes the following steps: setting the string obtained by encoding one or more keywords constituting the concern item information as the concern item ID.
4. The data management program according to claim 1 or 2, wherein, The step of generating the compound ID includes the following steps: for each of the more than one compound, generating a fixed-length string containing a first property string representing the first property of the compound with one or more alphanumeric characters, a component string representing the one or more components constituting the compound with one or more English letters, and a second property string representing the second property of the compound with one or more English letters, and having a length of more than 10 characters and less than 30 characters as the compound ID.
5. The data management program according to claim 4, wherein, The first physical property is molecular weight. The length of the first property string is 1. The step of generating the fixed-length string as the compound ID includes the following steps: in the first property string, the hundreds and thousands digits of the molecular weight are represented by numbers, uppercase English letters, or lowercase English letters.
6. The data management program according to claim 4, wherein, The step of generating the fixed-length string as the compound ID includes the following steps: when the ingredient corresponding to the English letter is included in the compound, each of the more than one English letter in the ingredient string is represented by one of an uppercase letter and a lowercase letter; when the ingredient corresponding to the English letter is not included in the compound, each of the more than one English letter in the ingredient string is represented by the other of the uppercase letter and the lowercase letter.
7. The data management program according to claim 4, wherein, The second property is polarity. The second property string represents the value of the polarity as a binary number. The step of generating the fixed-length string as the compound ID includes the following steps: when the value corresponding to the English letter is 1, each of the more than one English letter in the second property string is represented by one of an uppercase letter and a lowercase letter; when the value corresponding to the English letter is 0, each of the more than one English letter in the second property string is represented by the other of the uppercase letter and the lowercase letter.
8. The data management program according to claim 4, wherein, The step of generating the fixed-length string as the compound ID includes the following steps: generating a string obtained by changing the arrangement of the 14 characters of the first block of the InChIKey constituting the compound, which will be used as the ingredient string and the second property string.
9. The data management program according to claim 8, wherein, The change in the arrangement of the 14 characters is the reversal of the arrangement of the 14 characters.
10. The data management program according to claim 1 or 2, which causes the computer to further perform the following steps: For each of the more than one compounds represented by the collected compound information, a compound property including the precise mass and composition of that compound is set; For each of the more than one compounds represented by the collected compound information, the compound attribute is associated with the compound ID; and Generate a list of compounds that further includes one or more of the compound attributes associated with the compound ID.
11. The data management program according to claim 1 or 2, which causes the computer to further perform the following steps: The list of compounds is displayed on the monitor.
12. The data management program according to claim 1 or 2, which causes the computer to further perform the following steps: Collect resource information from the information source representing one or more resources corresponding to the information of concern; For each of the more than one resources represented by the collected resource information, a resource ID is generated as an identifier to uniquely identify the resource. A resource list is generated by associating each of the more than one resource ID with the concern item ID; and In addition to the compound list, the resource list is also stored in a database.
13. The data management program according to claim 12, which causes the computer to further perform the following steps: For each of the more than one resources represented by the collected resource information, a resource attribute is set that includes at least one of the person's name and the organization's name recorded in the resource; For each of the more than one resources represented by the collected resource information, the resource attribute is associated with the resource ID; and Generate a resource list that further includes one or more resource attributes each associated with the resource ID.
14. The data management program according to claim 12, which causes the computer to further perform the following steps: The relationship between the one or more compounds represented by the compound list and the one or more resources represented by the resource list is displayed on the screen in the form of a chart structure that connects the one or more compounds and the concern information, but does not connect the one or more compounds and the one or more resources, and connects the one or more resources and the concern information.
15. A data management system comprising at least one processor, said at least one processor performing the following processing: Obtain information on concerns for collecting data on more than one compound; Access the designated information source and collect compound information from the information source representing one or more compounds corresponding to the information of concern; Generate a concern ID as an identifier to uniquely determine the concern information; For each of the more than one compounds represented by the collected compound information, a compound ID is generated as an identifier to uniquely identify the compound; A compound list is generated by associating each of the more than one compound ID with the interest ID; and The list of compounds is stored in a database.
16. A data management method, executed by a data management system having at least one processor, the data management method comprising the following steps: Obtain information on concerns for collecting data on more than one compound; Access the designated information source and collect compound information from the information source representing one or more compounds corresponding to the information of concern; Generate a concern ID as an identifier to uniquely determine the concern information; For each of the more than one compounds represented by the collected compound information, a compound ID is generated as an identifier to uniquely identify the compound; A compound list is generated by associating each of the more than one compound ID with the interest ID; and The list of compounds is stored in a database.
Citation Information
Patent Citations
Material property prediction device and material property prediction method
JP2021039534A