Molecular index code generation method, molecular characteristic data retrieval method and related devices

By generating a unique index code and using the index code in the molecular component characteristic database for rapid positioning, the problem of time-consuming and low efficiency in molecular characteristic data retrieval is solved, and efficient molecular characteristic data retrieval is achieved.

CN120108485APending Publication Date: 2025-06-06PETROCHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311647303.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems such as long and low retrieval time and low efficiency in retrieval of molecular characteristic data, especially in systems with a large number of molecules, which are difficult to achieve efficient retrieval.

Method used

By naming the 24 SOL structure vector elements in the structure-oriented lumped, a unique vector element name is generated, and a one-to-one association is formed with the vector element values ​​to generate an index code. Use index codes to quickly locate and retrieve molecular characteristic data in molecular component characteristic databases.

Benefits of technology

It greatly compresses the search time, improves the search efficiency, is suitable for systems with a large number of molecules, and solves the problems of long and low efficiency in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108485A_ABST
    Figure CN120108485A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of molecular characteristic retrieval, in particular to a molecular index code generation method, a molecular characteristic data retrieval method and a related device.The molecular index code generation method comprises the steps that 24 SOL structure vector elements in a structure-oriented set are named; forming one-to-one association between the vector element name and the SOL structure vector element value of the target molecule; and scanning all association results, and generating an index code of the target molecule through an index code generation rule. On the basis of the characteristic that an SOL structure vector group is a sparse matrix, non-zero vector element values and vector element names in the structure vectors are mixed to generate index codes of molecules, 24-dimensional structure vectors are converted into one-dimensional character strings, the index codes of the molecules are set, then an index code sequence is taken as the outline, a molecular component characteristic database is established, and the molecular component characteristic database is established. And the characteristic data of the target molecule is quickly obtained in the molecular component characteristic database in a mode of positioning by using the row index code and the column variable name, so that the retrieval time is greatly shortened, and the retrieval efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of molecular property retrieval, and is a molecular index code generation method, a molecular property data retrieval method and a related device. Background Art

[0002] The core idea of ​​the structure oriented lumping (SOL) method is that all complex hydrocarbon molecules in oil products can be decomposed into molecular fragments or molecular structural groups, and this molecular structural group is called a structural vector.

[0003] The 24 groups of the structure-guided lumping method are Figure 1 As shown, A6 is a benzene ring; A4 is a four-carbon aromatic ring increment attached to another aromatic ring; A2 is an aromatic ring increment containing two carbons; N6 and N5 are aliphatic rings of 6 carbons and 5 carbons, respectively; N4, N3, N2, and N1 are aliphatic ring increments representing 4 carbons, 3 carbons, 2 carbons, and 1 carbon, respectively, connected to an aromatic ring or a cycloalkane ring; R is the total number of carbons excluding the carbons on the ring; me refers to the number of methyl groups attached to the aromatic ring or aliphatic ring of the molecule; br is the number of alkyl substituents attached to the alkyl, alkene, or alkyl side chain; AA represents a bridge bond between two rings; IH is a hydrogen increment used to specify the degree of unsaturation of the molecule (except the unsaturation on the aromatic ring); NS, NN, and NO are sulfur, nitrogen, and oxygen atoms connecting two carbon atoms; RS, RN, and RO represent sulfur, nitrogen, and oxygen atoms between carbon and hydrogen, respectively; AN represents a nitrogen atom on an aromatic ring; KO represents a carbonyl or aldehyde oxygen atom; Ni and V represent metal nickel and vanadium atoms. In the calculation program writing, the 24-dimensional SOL structure vector is used to represent the molecules. For example, the SOL structure vector of toluene can be expressed in list form as: [1,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0] The database form established for the 24-dimensional structure vector elements of the molecule is shown in Table 1.

[0004] Table 1 The form of the characteristic database established for the 24 SOL structure vectors of the molecule

[0005]

[0006] When searching for the characteristic data of a molecule in a data form, commonly used search methods include: 1. Searching for the molecule data by the name of the molecule, first searching for "toluene", and then retrieving the data of the characteristic column corresponding to the "toluene" row; 2. First obtaining the sequence number or row number of the molecule in the data form, and then searching for the data of the characteristic column corresponding to the row; the above two methods have the following problems:

[0007] (1) Simple and common substances with a small number of carbon atoms (<10) have common names, such as "benzene", "toluene", "isopropylbenzene", "cyclohexane", "isoheptane", etc., while hydrocarbons with a large number of carbon atoms (>12) often have no names. In particular, there are tens of thousands of crude oil molecules. Naming each molecule is not only time-consuming but also unrealistic. Even if all the names are manually remembered, it is difficult to remember them all, which affects programming efficiency. Therefore, for naphtha systems with relatively simple molecular structures and a small number of molecules, molecular naming may be used to retrieve data. However, for crude oil molecular systems with complex structures and a large number of molecules, it is difficult, cumbersome, and inefficient to implement.

[0008] (2) Retrieval by row number. The row number is just a simple numerical serial number. There is no unique correspondence between the row number and the molecular structure. It is impossible to directly point to a specific molecule through the row number. When the molecular order changes, the row number will change accordingly. The molecule is represented by 24 SOL structure vectors. Once stored in the characteristic database list, it is split into 24 independent columns. No column can fully reflect the structural characteristics of the molecule. It is impossible to obtain the row number of the molecule in the data table through a single column value. Therefore, each time, the 24 SOL structure vector element values ​​of the queried molecule must be compared row by row and column by column with the values ​​of the 24 columns in the data list. Only after they match can the complete molecular structure information and its row number be obtained. Generally, only 2 to 5 vector element values ​​are needed to characterize a molecule, and most of the vector element values ​​are 0. In order to obtain the molecular row number, all 24 vector element values ​​must be scanned each time, even if they are 0, so it is very time-consuming and inefficient to retrieve. Summary of the invention

[0009] The present invention provides a molecular index code generation method, a molecular characteristic data retrieval method and related devices, which overcome the deficiencies of the above-mentioned prior art and can effectively solve the problems of the existing molecular characteristic data retrieval, such as high retrieval time consumption, low efficiency and unsuitability for molecular systems with a large number of molecules.

[0010] One of the technical solutions of the present invention is achieved by the following measures: A method for generating a molecular index code, comprising:

[0011] Name the 24 SOL structure vector elements in the structure-oriented lumping, determine the vector element name of each SOL structure vector element, and the 24 vector element names must not be repeated;

[0012] Obtain the SOL structure vector element value of the target molecule and form a one-to-one association between the vector element name and the vector element value;

[0013] Scan all association results and generate the index code of the target molecule through the index code generation rule, wherein the index code generation rule includes determining the corresponding vector element name based on the one-to-one association between the vector element name and the vector element value when the vector element value is non-zero, and converting the associated vector element value into a string suffix before or after the vector element name, and skipping the pair of vector element name and vector element value when the vector element value is zero.

[0014] The following are further optimizations and / or improvements to the above technical solutions:

[0015] The above-mentioned one-to-one association between the vector element name and the vector element value includes: the vector element names of the 24 SOL structure vectors and the SOL structure vector element values ​​of the target molecules form 24 key-value pairs one-to-one, wherein the vector element name serves as the key value in the key-value pair, and the SOL structure vector element value serves as the numerical value in the key-value pair.

[0016] The 24 SOL structure vectors in the structure-oriented lumping are named by arbitrary characters to determine corresponding vector element names, wherein the vector element names of the 24 SOL structure vectors must not be repeated.

[0017] The second technical solution of the present invention is achieved by the following measures: A method for constructing a molecular component characteristic database, comprising:

[0018] Based on the vector element names and vector element values ​​of the molecular component SOL structure vector, obtaining index codes of several molecules, wherein the index code of each molecule is obtained using the molecular index code generation method;

[0019] Determine the relevant characteristic types and characteristic data of each molecule based on the molecular component SOL structure vector;

[0020] Set the row label as the index code and the column name as the variable name to generate a molecular component characteristic data table, where the variable name is the characteristic type.

[0021] The third technical solution of the present invention is achieved by the following measures: a molecular component characteristic data retrieval method, comprising:

[0022] Obtaining 24 SOL structure vectors of the molecule to be searched, and obtaining an index code of the molecule to be searched based on the vector element names and vector element values ​​of the 24 SOL structure vectors, wherein the index code of the molecule to be searched is obtained using the molecular index code generation method;

[0023] The required property data is retrieved from the molecular component property database using the index code and the property type to be retrieved, wherein the molecular component property database is obtained using the molecular component property database construction method.

[0024] The fourth technical solution of the present invention is achieved by the following measures: a structure-oriented lumped molecular index code generation device, comprising:

[0025] A naming unit is used to name the 24 SOL structure vector elements in the structure-oriented lumping, and the vector element name of each SOL structure vector element is determined, wherein the 24 vector element names must not be repeated;

[0026] An association unit, which obtains the SOL structure vector element value of the target molecule and forms a one-to-one association between the vector element name and the vector element value;

[0027] The index code generation unit scans all the association results and generates the index code of the target molecule according to the index code generation rule, wherein the index code generation rule includes determining the corresponding vector element name in combination with the one-to-one association relationship between the vector element name and the vector element value when the vector element value is non-zero, and converting the vector element value associated therewith into a string suffix before or after the vector element name, and skipping the pair of vector element name and vector element value when the vector element value is zero.

[0028] The fifth technical solution of the present invention is achieved by the following measures: a molecular component characteristic database construction device, comprising:

[0029] A first analysis unit obtains index codes of a plurality of molecules based on the vector element names and vector element values ​​of the molecular component SOL structure vectors, wherein the index code of each molecule is obtained using the molecular component characteristic database construction method;

[0030] A second analysis unit determines the relevant characteristic types and characteristic data of each molecule based on the molecular component SOL structure vector;

[0031] The database construction unit sets the row label as the index code and the column name as the variable name to generate a molecular component characteristic data table, in which the variable name is the characteristic type.

[0032] The sixth technical solution of the present invention is achieved by the following measures: a molecular component characteristic data retrieval device, characterized in that it includes:

[0033] A data acquisition unit, which obtains 24 SOL structure vectors of the molecule to be searched, and obtains an index code of the molecule to be searched based on the vector element names and vector element values ​​of the 24 SOL structure vectors, wherein the index code of the molecule to be searched is obtained using the molecular component characteristic database construction method;

[0034] The retrieval unit retrieves the required characteristic data in the molecular component characteristic database by using the index code and the characteristic type to be retrieved, wherein the molecular component characteristic database is obtained by using the molecular component characteristic database construction method.

[0035] The seventh technical solution of the present invention is achieved through the following measures: an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement a molecular component characteristic data retrieval method.

[0036] The present invention is based on the characteristic that the SOL structure vector group of a complex molecular system is a sparse matrix. By mixing non-zero values ​​and key symbols in vector elements, a new index code that uniquely corresponds to the molecular structure and is easy for machine automatic recognition is generated, and the 24-dimensional structure vector is converted into a 1-dimensional string. Then, a molecular component characteristic database is established based on the index code sequence. Furthermore, when querying molecular characteristic data, the 24-dimensional structure vector of the target molecule is first converted into a string index code, and then the characteristic data of the target molecule is quickly obtained in the molecular component characteristic database by positioning using "row index code" and "column variable name". That is, if the queried molecule is in the Nth row of the database, the retrieval process is equivalent to scanning only an N×1-dimensional matrix, and there is no need to retrieve the entire N×24-dimensional matrix, thereby greatly compressing the retrieval time and improving the retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Attached Figure 1 This is a diagram of 24 SOL structure vector elements in the structure-oriented aggregation of the present invention.

[0038] Attached Figure 2 The figure is a flow chart of an index code generation method of the present invention.

[0039] Attached Figure 3 The figure is a flow chart of another index code generation method of the present invention.

[0040] Attached Figure 4 The figure is a schematic diagram of the process of constructing the molecular component characteristic database of the present invention.

[0041] Attached Figure 5 The figure is a schematic flow chart of the molecular component characteristic data retrieval method of the present invention.

[0042] Attached Figure 6 It is a schematic diagram of the structure of the index code generating device of the present invention.

[0043] Attached Figure 7 This is a schematic diagram of the structure of the device for building the molecular component characteristic database of the present invention.

[0044] Attached Figure 8 It is a schematic diagram of the structure of the molecular component characteristic data retrieval device of the present invention. DETAILED DESCRIPTION

[0045] The present invention is not limited by the following embodiments, and specific implementation methods can be determined based on the technical solution of the present invention and actual conditions.

[0046] The following is an explanation of the professional terms that appear in this embodiment:

[0047] The core idea of ​​the structure oriented lumping (SOL) method is that all complex hydrocarbon molecules in oil products can be decomposed into molecular fragments or molecular structural groups, and this molecular structural group is called a structural vector.

[0048] The present invention will be further described below in conjunction with embodiments and drawings:

[0049] Embodiment 1: As attached Figure 2 As shown, an embodiment of the present invention discloses a method for generating a molecular index code, comprising:

[0050] Step S110, naming the 24 SOL structure vector elements in the structure-oriented lumping, determining the vector element name of each SOL structure vector element, wherein the 24 vector element names need not be repeated;

[0051] Step S120, obtaining the SOL structure vector element value of the target molecule, and forming a one-to-one association between the vector element name and the vector element value; the one-to-one association between the vector element name and the vector element value can be formed based on different programming languages ​​and a data association method can be selected;

[0052] Step S130, scan all association results, and generate an index code for the target molecule through an index code generation rule, wherein the index code generation rule includes determining the corresponding vector element name based on a one-to-one association between the vector element name and the vector element value when the vector element value is non-zero, and converting the vector element value associated therewith into a string suffix before or after the vector element name, and skipping the pair of vector element name and vector element value when the vector element value is zero.

[0053] The present invention discloses a method for generating a molecular index code. According to the characteristic of SOL characterizing molecular structure, that is, only about 1 to 9 (5 on average) vector element values ​​are usually needed to characterize a molecule, and most of the other vector element values ​​are 0. By mixing non-zero values ​​in vector elements and SOL structure vectors, a new index code that uniquely corresponds to the molecular structure and is easy for machine automatic recognition is generated, so as to improve retrieval efficiency and reduce time consumption.

[0054] Embodiment 2: As attached Figure 3 As shown, an embodiment of the present invention discloses a method for generating a molecular index code, comprising:

[0055] Step S210, naming the 24 SOL structure vector elements in the structure-oriented lumping, determining the vector element name of each SOL structure vector element, wherein the 24 vector element names need not be repeated;

[0056] The above arbitrary characters may be, but are not limited to, any characters such as English, text, symbols, etc., and may be selected according to actual needs.

[0057] For example:

[0058] (1) Name the 24 SOL structure vector elements in the structure-oriented lumping with English characters, as shown in the following table;

[0059] Table 2 Example 1 of the naming table of 24 SOL structure vectors

[0060] Vector element name A B C D E F G H I J K L M N O P Q R S T U V W X SOL Structure Vector A6 A4 A2 N6 N5 N4 N3 N2 N1 R br me IH AA NS RS AN NN RN NO RO KO Ni V

[0061] (2) Name the 24 SOL structure vector elements in the structure-oriented lumping with arbitrary characters, as shown in the following table;

[0062] Table 3 Example 2 of the naming table of 24 SOL structure vectors

[0063] Vector element name A First C D E F G H I # K L Ω N O P Q R S T U V W X SOL Structure Vector A6 A4 A2 N6 N5 N4 N3 N2 N1 R br me IH AA NS RS AN NN RN NO RO KO Ni V

[0064] Step S220, obtaining the SOL structure vector element value of the target molecule, the vector element names of the 24 SOL structure vectors and the SOL structure vector element values ​​of the target molecule form 24 key-value pairs one-to-one, wherein the vector element names serve as the key values ​​in the key-value pairs, and the SOL structure vector element values ​​serve as the numerical values ​​in the key-value pairs;

[0065] The key-value pair is composed of a key value Key and a value Value, and is used to indicate the corresponding relationship between the value and the key value. In this embodiment, it indicates the corresponding relationship between the vector element name and the vector element value.

[0066] For example, the 24 SOL structure vector elements in the structure-oriented lumping are named with English characters. The target molecule is isoheptane, and its SOL structure vector element values ​​are shown in Table 4:

[0067] Table 4 Comparison table of SOL structure vector and SOL structure vector element values ​​of target molecules Example 1

[0068]

[0069] At this time, combined with the above table, 24 key-value pairs can be formed. The key values ​​are: A, B, C,..., X, and the values ​​are: 0, 0, 0, 0, 0, 0, 0, 0, 7, 1..., 0.

[0070] Step S230, scan all key-value pairs, and generate the index code of the target molecule through the index code generation rule, wherein the index code generation rule includes when the vector element value is non-zero, combining the key-value pair to determine the corresponding vector element name, and converting the associated vector element value into a string suffix before or after the vector element name, and when the vector element value is zero, skipping the pair of vector element name and vector element value.

[0071] In the above, all key-value pairs are scanned, and the index code of the target molecule is generated by the index code generation rule. For example, the SOL structure vector and the SOL structure vector element value comparison table of the target molecule are shown in Table 5. The combination table consists of 24 key-value pairs of isoheptane and 24 key-value pairs of toluene. Then all key-value pairs are scanned, and the index code of the target molecule is generated by the index code generation rule. After isoheptane forms 24 key-value pairs, all key-value pairs are scanned. The index code of isoheptane generated according to the index code generation rule is J7K1M1. After toluene forms 24 key-value pairs, all key-value pairs are scanned. The index code of toluene generated according to the index code generation rule is A1J1. Here, when the vector element value is non-zero, the corresponding vector element name is determined in combination with the key-value pair, and the vector element value associated with it is converted into a string suffix after the vector element name, but the vector element value can also be converted into a string suffix before the vector element name, that is, the index code of isoheptane is 7J1K1M, and the index code of toluene is 1A1J.

[0072] Table 5 Example 2 of comparison table of SOL structure vector and SOL structure vector element values ​​of target molecules

[0073]

[0074] Here, the naming characters of the vector element names of the SOL structure vector are different, and the index codes are also different. For example, the comparison table of the SOL structure vector and the SOL structure vector element values ​​of the target molecule is shown in Table 6. The combination table consists of 24 key-value pairs of isoheptane and 24 key-value pairs of toluene, and then all key-value pairs are scanned. The index code of the target molecule is generated according to the index code generation rule. After isoheptane forms 24 key-value pairs, all key-value pairs are scanned. The index code of isoheptane generated according to the index code generation rule is #7K1Ω1. After toluene forms 24 key-value pairs, all key-value pairs are scanned. The index code of toluene generated according to the index code generation rule is A1#1. Here, you can also choose to convert the vector element value into a string suffix before the vector element name. The specific details are as described in the above example and will not be repeated here.

[0075] Table 6 Comparison table of SOL structure vector and SOL structure vector element values ​​of target molecules Example 3

[0076]

[0077] Embodiment 3: As attached Figure 4 As shown, the embodiment of the present invention discloses a method for constructing a molecular component characteristic database, comprising:

[0078] Step S310, based on the vector element name and vector element value of the molecular component SOL structure vector, obtain the index codes of several molecules, wherein the index code of each molecule is obtained by the method described in the above embodiments 1 and 2;

[0079] Based on the molecular component SOL structure vector, the index codes of several molecules are obtained. The molecules here must meet the requirements of the application system and exhaust all molecules in the system, where the system is the existing naphtha system with a small number of molecules, the crude oil molecule system with a large number of molecules, etc.

[0080] Step S320, determining the relevant characteristic type and characteristic data of each molecule based on the molecular component SOL structure vector;

[0081] The characteristic data of each molecule mentioned above can be calculated and obtained according to the SOL structure vector of the molecular component using existing methods.

[0082] Step S330, set the row label as the index code and the column name as the variable name to generate a molecular component characteristic data table, wherein the variable name is the characteristic type.

[0083] The above-mentioned generation of the molecular component characteristic data table can be implemented using different programming languages. This embodiment can be implemented using, but not limited to, the Python programming language, as described below:

[0084] When programming in Python, you can call the pandas data analysis package to generate a feature data table using the pandas.DataFrame() structure. Here, DataFrame is a tabular data structure with both row labels (index) and column labels (columns). The specific code is as follows:

[0085] property_dataframe=pandas.DataFrame(property,

[0086] columns = ['variable name 1', 'variable name 2', 'variable name 3', ...],

[0087] index = ['index code'])

[0088] Among them: property_dataframe is the name of the predicted characteristic data table variable;

[0089] Pandas is a data analysis package for Python;

[0090] DataFrame is the name of the built-in tabular data structure function of Pandas;

[0091] Property is the name of the predicted characteristic variable, which is a list type;

[0092] Columns is the column variable parameter of the DataFrame function;

[0093] index is the row index parameter of the DataFrame function.

[0094] The molecular component characteristic data table generated by the above steps can be shown in Table 6:

[0095] Table 7 Molecular component characteristic data table example 1

[0096]

[0097] Here, if the characteristic types include molecular weight, boiling point, critical temperature, critical temperature, and eccentricity factor, that is, the variable names include molecular weight, boiling point, critical temperature, critical temperature, and eccentricity factor, the specific writing is as follows:

[0098] predict_property=pandas.DataFrame(property,

[0099] columns = ['molecular weight', 'boiling point', 'critical temperature', 'critical temperature', 'eccentricity factor'],

[0100] index = ['obtained index code'])

[0101] The molecular component characteristic data table generated through the above steps can be shown in Table 7:

[0102] Table 8 Molecular component characteristic data table example 2

[0103]

[0104] Embodiment 4: As attached Figure 5 As shown, the embodiment of the present invention discloses a method for retrieving molecular component characteristic data, comprising:

[0105] Step S410, obtaining 24 SOL structure vectors of the molecule to be searched, and obtaining the index code of the molecule to be searched based on the vector element names and vector element values ​​of the 24 SOL structure vectors, wherein the index code of the molecule to be searched is obtained using the method described in the above embodiments 1 to 2;

[0106] Step S420: Retrieve the required characteristic data from the molecular component characteristic database using the index code and the characteristic type to be retrieved, wherein the molecular component characteristic database is obtained using the method described in Example 3.

[0107] The molecular component characteristic data retrieval method can be implemented using different programming languages. This embodiment can be implemented using, but not limited to, the Python programming language, as described below:

[0108] When programming in Python, you can call the loc() function of the pandas data analysis package to directly locate and retrieve the required feature data based on the row index code and column variable name. The specific code is as follows:

[0109] property_value = property_dataframe.loc('row index code', 'column variable name')

[0110] in:

[0111] property_value is the name of the target variable to be retrieved. The required value is located in the property_dataframe data table through the property_dataframe.loc() function;

[0112] property_dataframe is the name of the property data table variable;

[0113] loc() is a positioning function for DataFrame type.

[0114] For example, the molecule to be searched is toluene. Based on Table 5, the index code of toluene is A1J1, and a molecular component characteristic database as described in Table 8 is established. Then, according to the row index code 'A1J1' and the column variable name 'boiling point', the characteristic data of 'toluene' can be directly located and retrieved in the molecular component characteristic database, that is, property_value = property_dataframe.loc('A1J1', 'boiling point'). The search results are shown in Table 9, and the data in the circle are the search results.

[0115] Table 9 shows the molecular component property data table with search results

[0116]

[0117] In summary, based on the characteristic that the SOL structure vector group of a complex molecular system is a sparse matrix, the present invention generates a new index code that uniquely corresponds to the molecular structure and is easy for the machine to automatically identify by mixing the non-zero vector element values ​​in the structure vector and the corresponding vector element names, and converts the 24-dimensional structure vector into a 1-dimensional string. Then, based on the index code sequence, a molecular component characteristic database is established. Furthermore, when querying the molecular characteristic data, the 24-dimensional structure vector of the target molecule is first converted into a string index code, and then the characteristic data of the target molecule is quickly obtained in the molecular component characteristic database by positioning using the "row index code" and "column variable name". That is, if the queried molecule is in the Nth row of the database, the retrieval process is equivalent to scanning only an N×1-dimensional matrix, and there is no need to retrieve the entire N×24-dimensional matrix, thereby greatly shortening the retrieval time and improving the retrieval efficiency.

[0118] Embodiment 5: As attached Figure 6 As shown, the embodiment of the present invention discloses a structure-oriented lumped molecular index code generation device, comprising:

[0119] Naming unit, name the 24 SOL structure vector elements in the structure-oriented lumps, determine the vector element name of each SOL structure vector element, and the 24 vector element names must not be repeated; here, you can use any characters to name the 24 SOL structure vector elements in the structure-oriented lumps to determine the corresponding vector element names.

[0120] The association unit obtains the SOL structure vector element value of the target molecule and forms a one-to-one association between the vector element name and the vector element value; here, the key-value pairs can be used to form a one-to-one association relationship between the vector element names of the 24 SOL structure vectors and the SOL structure vector element values ​​of the target molecule, thereby forming 24 key-value pairs, in which the vector element name serves as the key value in the key-value pair, and the SOL structure vector element value serves as the numerical value in the key-value pair.

[0121] The index code generation unit scans all the association results and generates the index code of the target molecule through the index code generation rule, wherein the index code generation rule includes when the vector element value is non-zero, combining the one-to-one association between the vector element name and the vector element value to determine the corresponding vector element name, and converting the vector element value associated therewith into a string suffix before or after the vector element name, and skipping the pair of vector element name and vector element value when the vector element value is zero. Here, if the vector element names of the 24 SOL structure vectors and the SOL structure vector element value of the target molecule form 24 key-value pairs, then all the key-value pairs are scanned, and the index code of the target molecule is generated through the index code generation rule, wherein the index code generation rule includes when the vector element value is non-zero, combining the key-value pair to determine the corresponding vector element name, and converting the vector element value associated therewith into a string suffix after the vector element name, and skipping the pair of vector element name and vector element value (i.e. skipping the key-value pair) when the vector element value is zero.

[0122] Embodiment 6: As attached Figure 7 As shown, the embodiment of the present invention discloses a molecular component characteristic database construction device, comprising:

[0123] A first analysis unit obtains index codes of several molecules based on the vector element names and vector element values ​​of the molecular component SOL structure vector, wherein the index code of each molecule is obtained using the molecular index code generation method;

[0124] A second analysis unit determines the relevant characteristic types and characteristic data of each molecule based on the molecular component SOL structure vector;

[0125] The database construction unit sets the row label as the index code and the column name as the variable name to generate a molecular component characteristic data table, in which the variable name is the characteristic type.

[0126] Embodiment 7: As attached Figure 8 As shown, the embodiment of the present invention discloses a molecular component characteristic data retrieval device, comprising:

[0127] A data acquisition unit obtains 24 SOL structure vectors of the molecule to be searched, and obtains an index code of the molecule to be searched based on the vector element names and vector element values ​​of the 24 SOL structure vectors, wherein the index code of the molecule to be searched is obtained using the molecular index code generation method;

[0128] The retrieval unit retrieves the required characteristic data in the molecular component characteristic database by using the index code and the characteristic type to be retrieved, wherein the molecular component characteristic database is obtained by using the molecular component characteristic database construction method.

[0129] Embodiment 8: The embodiment of the present invention discloses an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement a method for retrieving molecular component characteristic data.

[0130] The processor may be a central processing unit (CPU), a general purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. It may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The memory may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory, a mobile hard disk, a magnetic disk or an optical disk.

[0131] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0132] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0133] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.

[0134] The above technical features constitute the best embodiment of the present invention, which has strong adaptability and best implementation effect. Non-essential technical features can be added or reduced according to actual needs to meet the requirements of different situations.

Claims

1. A method for generating a molecular index code, It is characterized in that include: Name the 24 SOL structure vector elements in the structure-oriented lumping, determine the vector element name of each SOL structure vector element, and the 24 vector element names must not be repeated; Obtain the SOL structure vector element value of the target molecule and form a one-to-one association between the vector element name and the vector element value; Scan all association results and generate the index code of the target molecule through the index code generation rule, wherein the index code generation rule includes determining the corresponding vector element name based on the one-to-one association between the vector element name and the vector element value when the vector element value is non-zero, and converting the associated vector element value into a string suffix before or after the vector element name, and skipping the pair of vector element name and vector element value when the vector element value is zero.

2. The method for generating a molecular index code according to claim 1, It is characterized in that The one-to-one association between the vector element name and the vector element value is formed, including: the vector element names of the 24 SOL structure vectors and the SOL structure vector element values ​​of the target molecule form 24 key-value pairs one-to-one, wherein the vector element name serves as the key value in the key-value pair, and the SOL structure vector element value serves as the numerical value in the key-value pair.

3. The method for generating a molecular index code according to claim 1, It is characterized in that The 24 SOL structure vectors in the structure-guided lumping are named by arbitrary characters to determine corresponding vector element names, wherein the vector element names of the 24 SOL structure vectors must not be repeated.

4. A method for constructing a molecular component characteristics database, It is characterized in that include: Based on the vector element names and vector element values ​​of the molecular component SOL structure vector, obtaining index codes of several molecules, wherein the index code of each molecule is obtained using the method described in any one of claims 1 to 3; Determine the relevant characteristic types and characteristic data of each molecule based on the molecular component SOL structure vector; Set the row label as the index code and the column name as the variable name to generate a molecular component characteristic data table, where the variable name is the characteristic type.

5. A method for retrieving molecular component characteristic data, It is characterized in that include: Obtain 24 SOL structure vectors of the molecule to be searched, and obtain an index code of the molecule to be searched based on the vector element names and vector element values ​​of the 24 SOL structure vectors, wherein the index code of the molecule to be searched is obtained by the method described in any one of claims 1 to 3; The required property data is retrieved from the molecular component property database using the index code and the property type to be retrieved, wherein the molecular component property database is obtained using the method described in claim 4.

6. A structure-oriented lumped molecular index code generation device using the method according to any one of claims 1 to 3, It is characterized in that include: A naming unit is used to name the 24 SOL structure vector elements in the structure-oriented lumping, and the vector element name of each SOL structure vector element is determined, wherein the 24 vector element names must not be repeated; An association unit, which obtains the SOL structure vector element value of the target molecule and forms a one-to-one association between the vector element name and the vector element value; The index code generation unit scans all the association results and generates the index code of the target molecule according to the index code generation rule, wherein the index code generation rule includes determining the corresponding vector element name in combination with the one-to-one association relationship between the vector element name and the vector element value when the vector element value is non-zero, and converting the vector element value associated therewith into a string suffix before or after the vector element name, and skipping the pair of vector element name and vector element value when the vector element value is zero.

7. A molecular component characteristic database construction device using the method as claimed in claim 4, It is characterized in that include: A first analysis unit, based on the vector element name and vector element value of the molecular component SOL structure vector, obtains index codes of several molecules, wherein the index code of each molecule is obtained by the method described in any one of claims 1 to 3; A second analysis unit determines the relevant characteristic types and characteristic data of each molecule based on the molecular component SOL structure vector; The database construction unit sets the row label as the index code and the column name as the variable name to generate a molecular component characteristic data table, in which the variable name is the characteristic type.

8. A molecular component characteristic data retrieval device using the method as claimed in claim 5, It is characterized in that include: A data acquisition unit, which obtains 24 SOL structure vectors of the molecule to be searched, and obtains an index code of the molecule to be searched based on the vector element names and vector element values ​​of the 24 SOL structure vectors, wherein the index code of the molecule to be searched is obtained by the method described in any one of claims 1 to 3; A retrieval unit retrieves the required characteristic data from a molecular component characteristic database using an index code and a characteristic type to be retrieved, wherein the molecular component characteristic database is obtained using the method described in claim 4.

9. An electronic device, It is characterized in that It comprises a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the molecular component characteristic data retrieval method as claimed in claim 5.