Structure-oriented lumped data storage and retrieval method and device
By generating structural codes and building a database, the problems of low data storage efficiency and difficult retrieval of SOL data were solved, realizing efficient storage and flexible retrieval of oil molecular data, and improving the efficiency and product quality of the molecular refining process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RICHFIT INFORMATION TECH
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-01
AI Technical Summary
Existing structure-oriented lumped data storage methods suffer from low storage efficiency and difficulty in data retrieval in the field of molecular refining, which limits their application.
By generating structural codes for oil molecules, classifying them according to preset properties, constructing a database, and using sparse matrices for preprocessing, the data storage and retrieval process is optimized.
It improves the storage efficiency of SOL data, simplifies the data retrieval process, enables SOL data with the same chemical properties to be retrieved and analyzed quickly, and enhances the flexibility and accuracy of the data.
Smart Images

Figure CN121963977A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of molecular refining, and in particular to a structure-oriented lumped data storage and retrieval method and apparatus. Background Technology
[0002] In molecular refining, structure-directed lumped arrays (SOLs) are an important representation method used in petroleum molecule characterization techniques to estimate their molecular composition. They effectively integrate chemical structural information of molecules, facilitating molecular analysis and processing. However, existing SOL data storage methods suffer from several problems, such as low storage efficiency and difficulties in data retrieval, which limit the application of SOL data in the field of molecular refining. Summary of the Invention
[0003] To address the problems of the prior art, embodiments of this specification provide a structure-oriented centralized data storage and retrieval method and apparatus.
[0004] This specification provides a structure-oriented aggregate data storage and retrieval method, which includes: combining the codes corresponding to the basic structural units of each oil molecule according to the chemical structure of the oil molecule to obtain a structural code representing the oil molecule; classifying the structural codes corresponding to the oil molecules according to preset properties to obtain the structural codes of each oil molecule under each type of property; and constructing a database for storing oil molecules based on the structural codes of each oil molecule under each type of property.
[0005] According to one aspect of the embodiments of this specification, the code corresponding to each basic structural unit of each oil molecule is determined by: decomposing the oil molecule according to its chemical structure to obtain the basic structural unit of each oil molecule; and uniquely coding the basic structural unit to obtain the code corresponding to each basic structural unit of each oil molecule.
[0006] According to one aspect of the embodiments of this specification, the preset properties include: molecular type, physical properties, and chemical properties. Classifying the structure codes corresponding to the oil molecules according to the preset properties to obtain the structure codes of each oil molecule under each type of property includes: storing the structure codes of oil molecules with the same physical properties in the same data table to obtain the structure codes corresponding to each oil molecule under multiple physical properties; storing the structure codes of oil molecules with the same molecular type in the same data table to obtain the structure codes corresponding to each oil molecule under each molecular type; and storing the structure codes of oil molecules with the same chemical properties in the same data table to obtain the structure codes corresponding to each oil molecule under the chemical properties.
[0007] According to one aspect of the embodiments of this specification, constructing a database for storing oil molecules based on the structural codes of each oil molecule under various properties includes: converting the structural codes of each oil molecule under various properties into vector data to obtain vectors corresponding to multiple oil molecules under various properties; storing the vectors corresponding to multiple oil molecules under various properties in the database respectively, thereby completing the database construction.
[0008] According to one aspect of an embodiment of this specification, the method further includes: querying a structure code corresponding to a search condition provided by a user from the database, wherein the search condition corresponds to a preset property.
[0009] According to one aspect of an embodiment of this specification, querying the structure code corresponding to the search conditions from the database based on the search conditions provided by the user further includes: obtaining the search conditions, the search conditions including: search properties; and querying the vector corresponding to the search properties and the structure code corresponding to the vector from the database based on the search properties.
[0010] According to one aspect of an embodiment of this specification, the search conditions further include a component to be searched and a vector corresponding to the component to be searched, wherein the vector is a query vector; when a structural code corresponding to the search conditions cannot be matched according to the search conditions, the method further includes: calculating the similarity between the query vector and the vector corresponding to an oil molecule recorded in the database to obtain at least one similarity result; sorting the similarity results and feeding back the vector corresponding to the highest similarity result to the user; replacing the component to be searched with the oil molecule corresponding to the vector corresponding to the highest similarity result to generate a new mixture.
[0011] According to one aspect of the embodiments of this specification, sparse matrices are used to preprocess vector data; the preprocessed vector data is then stored in a database according to a preset storage format.
[0012] This specification provides a structure-oriented aggregate data storage and retrieval device, comprising: a structure coding determination unit, used to combine the codes corresponding to each basic structural unit of each oil molecule according to the chemical structure of the oil molecule to obtain a structure code representing the oil molecule; a classification unit, used to classify the structure codes corresponding to the oil molecules according to preset properties to obtain the structure codes of each oil molecule under each type of property; and a database construction unit, used to construct a database for storing oil molecules based on the structure codes of each oil molecule under each type of property.
[0013] This specification also provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the structure-oriented aggregate data storage and retrieval method.
[0014] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the structure-oriented lumped data storage and retrieval method.
[0015] This invention optimizes the processing, storage, and retrieval of data in molecular refining, improving refining efficiency and product quality. Through a structure encoding generation method, the chemical structure information of molecules becomes easier to represent and process; through classification storage rules, SOL data with the same chemical properties can be quickly retrieved and analyzed; and through retrieval and analysis methods, the application of SOL data becomes more flexible and accurate. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 The diagram shown is a flowchart of a structure-oriented aggregate data storage and retrieval method according to an embodiment of this specification;
[0018] Figure 2 The diagram shown is a flowchart of a method for determining the codes corresponding to each basic structural unit of each oil molecule according to an embodiment of this specification.
[0019] Figure 3 The diagram shown is a flowchart of a method for determining the structural codes corresponding to each oil molecule under various properties, according to an embodiment of this specification.
[0020] Figure 4 The diagram shown is a flowchart of a method for constructing a database according to an embodiment of this specification;
[0021] Figure 5 The diagram shown is a flowchart of a method for constructing a database according to an embodiment of this specification;
[0022] Figure 6 The diagram shown is a flowchart of a method for providing matching results to a user according to an embodiment of this specification.
[0023] Figure 7The diagram shown is a structural schematic of a structure-oriented centralized data storage and retrieval device according to an embodiment of this specification.
[0024] Figure 8 The diagram shown is a schematic representation of SOL data according to an embodiment of this specification.
[0025] Figure 9 The diagram shown is a structural schematic of a computer device according to an embodiment of this specification.
[0026] Explanation of symbols in the attached drawings:
[0027] 701. Structural coding determination unit;
[0028] 702. Classification Unit;
[0029] 703. Database construction unit;
[0030] 902. Computer equipment;
[0031] 904, Processor;
[0032] 906. Memory;
[0033] 908. Drive mechanism;
[0034] 910. Input / Output Module;
[0035] 912. Input devices;
[0036] 914. Output devices;
[0037] 916. Presentation equipment;
[0038] 918. Graphical User Interface;
[0039] 920. Network interface;
[0040] 922. Communication link;
[0041] 924. Communication bus. Detailed Implementation
[0042] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0043] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0044] This specification provides the operational steps of the methods described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the methods shown in the embodiments or drawings can be executed sequentially or in parallel.
[0045] It should be noted that the structure-oriented lumped data storage and retrieval method and apparatus described in this specification can be used in the field of molecular refining, and this specification does not limit the application field of a structure-oriented lumped data storage and retrieval method and apparatus.
[0046] Figure 1 The diagram shown is a flowchart of a structure-oriented centralized data storage and retrieval method according to an embodiment of this specification. The system includes:
[0047] Step 101: Based on the chemical structure of the oil molecules, combine the codes corresponding to each basic structural unit of each oil molecule to obtain the structural code representing the oil molecule.
[0048] In the embodiments of this specification, the chemical structure of oil molecules includes, but is not limited to, the types of atoms in the oil molecules, the chemical bond space of the oil molecules, and the spatial configuration of the oil molecules. Among these, the molecule type typically refers to the molecular structure and composition of the petroleum fraction. The chemical bond space in the oil molecules refers to the three-dimensional spatial structure formed by atoms or ions connected by chemical bonds within the oil molecules; this spatial structure determines the shape, size, stability, and chemical properties of the oil molecules. The spatial configuration of the oil molecules refers to the arrangement and spatial structure of the atoms within the oil molecules; the spatial configuration determines the shape and symmetry of the oil molecules.
[0049] In this specification, based on the chemical structure of oil molecules, the structure-directed lumped method is used to decompose oil molecules, thereby obtaining the basic structural units of the oil molecules. An oil molecule may correspond to at least one basic structural unit. Structure-directed lumpedization (SOL) is a method in molecular refining that represents petroleum molecules based on specific structural features or functional groups. These groups are organized into a vector, where the elements of the vector represent different petroleum molecule structures, and the number of elements indicates the quantity of a specific structural group in the molecule. This method is of great significance for predicting the properties of petroleum products and optimizing the refining process.
[0050] Based on the chemical structure of oil molecules, the codes of the basic structural units corresponding to each oil molecule are combined to obtain the structural code corresponding to the oil molecule. In this specification, after decomposing different oil molecules, the type and number of basic structural units of the oil molecule can be obtained. For some oil molecules, multiple identical basic structural units may be obtained, meaning the number of such basic structural units is more than one. Therefore, it is necessary to combine the different numbers of basic structural units and the number of identical basic structural units to determine the code of each basic structural unit corresponding to the oil molecule. These codes can represent the structure of the entire oil molecule. This specification, through the method of generating structural codes, makes the chemical structural information of oil molecules easier to represent and process.
[0051] Step 102: Classify the structure codes corresponding to the oil molecules according to preset properties to obtain the structure codes of each oil molecule under each type of property.
[0052] In the embodiments of this specification, the preset properties include, but are not limited to, physical properties, molecular types, and chemical properties. By differentiating the structural codes of oil molecules according to these preset properties, the structural codes corresponding to different oil molecules with different properties can be obtained.
[0053] Step 103: Construct a database for storing oil molecules based on their structural codes under various properties. For example, store the structural codes of multiple different oil molecules with the same physical property in the same area, and store the structural codes of oil molecules with different physical properties in different areas; store the structural codes of multiple different oil molecules with the same chemical property in the same area, and store the structural codes of oil molecules with different chemical properties in different areas; store the structural codes of multiple different oil molecules with the same molecular type in the same area, and store the structural codes of oil molecules with different molecular types in different areas.
[0054] This manual uses categorized storage rules to enable the rapid retrieval and analysis of SOL data with the same chemical properties, physical properties, or molecular types.
[0055] This manual improves the storage efficiency of SOL data and simplifies the data retrieval process; through the generation method of structure encoding, it makes the chemical structure information of molecules easier to represent and process; through classification storage rules, it enables SOL data with the same chemical properties to be quickly retrieved and analyzed; and through retrieval and analysis methods, it makes the application of SOL data more flexible and accurate.
[0056] Figure 2 The diagram shown is a flowchart of a method for determining the structural codes of different oil molecules under various properties, according to an embodiment of this specification. The method specifically includes the following steps:
[0057] Step 201: Decompose the oil molecules according to their chemical structures to obtain the basic structural units of each oil molecule. In this step, the molecules are decomposed based on their chemical structure information (including but not limited to: atom type, chemical bond type, spatial configuration, etc.) to obtain at least one basic structural unit of the oil molecule. The basic structural unit is used to characterize the molecular composition or molecular structure of the oil molecule. The basic structural unit includes at least: aromatic rings, ketone or aldehyde groups, biphenyl bridging structures, hydrogen-related structural increments or side-chain alkyl groups, the number of branch nodes on straight-chain alkyl groups or olefins, etc.
[0058] Step 202 involves uniquely encoding the basic structural units to obtain the codes corresponding to each basic structural unit of each oil molecule. In this step, the basic structural units described in step 201 are encoded based on a predetermined algorithm, and the codes are unique identifiers for each basic structural unit.
[0059] This manual lists the codes corresponding to various functional groups in common oil molecules, such as... Figure 8 As shown. For example, code A6 represents a six-carbon aromatic ring, which is a single structural unit; code A4 represents a four-carbon aromatic ring; code A2 represents a two-carbon aromatic ring; code N6 represents a six-carbon cycloalkane; code N5 represents a five-carbon cycloalkane, etc.
[0060] Figure 3 The diagram shown is a flowchart of a method for determining the structural codes corresponding to each oil molecule under various properties, according to an embodiment of this specification.
[0061] In the embodiments of this specification, the preset properties include at least: molecular type, physical properties, and chemical properties. Therefore, classifying the structure codes corresponding to the oil molecules according to the preset properties to obtain the structure codes of each oil molecule under each type of property includes the following steps:
[0062] Step 301: Store the structure codes of oil molecules with the same physical properties in the same data table to obtain the structure codes corresponding to each oil molecule under multiple physical properties.
[0063] In the embodiments of this specification, physical properties include, but are not limited to, solubility, boiling point, mass fraction, viscosity, and yield. The structural codes of oil molecules with the same boiling point, solubility, mass fraction, or yield can be stored in the same data table. Specifically, the structural codes of oil molecules with a boiling point of 170 degrees Celsius can be stored in data table A, which records the structural codes of multiple oil molecules; the structural codes of oil molecules with a boiling point of 220 degrees Celsius can be stored in data table B, which records the structural codes of multiple oil molecules; the structural codes of oil molecules with a boiling point of 350 degrees Celsius can be stored in data table C; and the structural codes of oil molecules with a solubility of 10 mg / L can be stored in data table D, which records the structural codes of multiple oil molecules.
[0064] In the embodiments of this specification, in addition to storing the structure codes of molecules with the same physical properties in the same data table, the structure codes of molecules with similar physical properties can also be stored in the same data table. For example, the structure codes of oil molecules with similar solubility or similar boiling points can be stored in the same data table. For example, the structure codes of oil molecules with boiling points between 170 degrees Celsius and 179 degrees Celsius, or the structure codes of oil molecules with solubility between 10 mg / L and 12 mg / L, can be stored in the same data table.
[0065] Step 302: Store the structure codes of oil molecules with the same molecular type in the same data table to obtain the structure codes of each oil molecule under each molecular type.
[0066] In the embodiments of the specification, the molecular types include, but are not limited to, classification of molecules according to the number of atoms, classification of molecules according to their electronic structure, and classification of molecules according to their molecular mass. Specifically, classification according to the number of atoms includes: monatomic molecules, diatomic molecules, organic molecules, or inorganic molecules; classification according to the electronic structure includes: polar molecules or nonpolar molecules; and classification according to molecular mass includes: polymers or small molecules.
[0067] In this step, the structure codes of molecules with the same molecular type are stored in the same data table. For example, the structure codes of oil molecules that are monatomic molecules are stored in the same data table; the structure codes of oil molecules that are diatomic molecules are stored in the same data table; the structure codes of oil molecules that are organic molecules are stored in the same data table; and the structure codes of oil molecules that are inorganic molecules are stored in the same data table.
[0068] Step 303: Store the structure codes of oil molecules with the same chemical properties in the same data table to obtain the structure codes corresponding to each oil molecule under the chemical properties.
[0069] In the embodiments of this specification, chemical properties include: yield, which refers to the proportion of each molecular component in a mixture. For example, a mixture contains five molecular components: methane, cyclopropane, benzene, p-xylene, and propylene. The basic information corresponding to each molecular component is as follows:
[0070] The yield of methane in the mixture was 0.16378381, with a boiling point of -161.5°C and a liquid molar volume of 0.054 at 15.85°C. Its structural code consists of one R and one IH atom. The yield of cyclopropane in the mixture was 0.23546789, with a boiling point of -33°C and a liquid molar volume of 0.067 at 15.85°C. Its structural code consists of three R atoms. The yield of benzene in the mixture was 0.19362839, with a boiling point of 80.1°C and a liquid molar volume of 0.054 at 15.85°C. The volume at 15.85℃ is 0.089, and its structural code is 1 A6; the yield of p-xylene in the mixture is 0.23457589, the boiling point is 138.4 degrees Celsius, and the liquid molar volume at 15.85℃ is 0.093, with a structural code of 1 A6, 2 R, and 1 me; the yield of propylene in the mixture is 0.17254402, the boiling point is -47.7 degrees Celsius, and the liquid molar volume at 15.85℃ is 0.081, with a structural code of 3 R. This specification classifies and stores the structural codes corresponding to oil molecules according to different preset properties, enriching the classification types of oil molecule structural codes, increasing the types of data tables stored in the classification, further enhancing the diversity of data tables in the database, and improving data storage efficiency. At the same time, it allows users to quickly retrieve and analyze data with the same properties when searching the database according to search criteria; simplifying the data retrieval process and enhancing the data query experience.
[0071] Figure 4 The diagram shown is a flowchart of a method for constructing a database according to an embodiment of this specification, which specifically includes the following steps:
[0072] Step 401: Convert the structural codes of each oil molecule under various properties into vector data to obtain vectors corresponding to multiple oil molecules under various properties.
[0073] like Figure 3Taking a mixture as an example, the mixture contains five molecular components: methane, cyclopropane, benzene, p-xylene, and propylene. The structural codes for each molecular component were determined. Based on the structural codes of each molecular component under the "yield" property, the corresponding vector data for these five molecular components were determined. The yield of methane in the mixture is 0.16378381, and its structural code is 1 R and 1 IH; the yield of cyclopropane in the mixture is 0.23546789, and its structural code is 3 Rs; the yield of benzene in the mixture is 0.19362839, and its structural code is 1 A6; the yield of p-xylene in the mixture is 0.23457589, and its structural code is 1 A6, 2 Rs, and 1 me; the yield of propylene in the mixture is 0.17254402, and its structural code is 3 Rs.
[0074] Therefore, a table as shown in Table 1 can be constructed, where each row of data represents the distribution of the structure codes of methane, cyclopropane, benzene, p-xylene, and propylene under the property of yield.
[0075] Table 1
[0076]
[0077] Further converting the data in Table 1 into corresponding vectors, the vector corresponding to methane is: [0,0,0,0,0,0,0,0,0,1,0,0,1,0,0,0,0,0,0,0,0,0,0,0]; the vector corresponding to cyclopropane is: [0,0,0,0,0,0,0,0,0,3,0,0,0,0,0,0,0,0,0,0,0,0,0,0]; and the vector corresponding to benzene is: [1,0,0,0,0] ,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]; the vector corresponding to xylene is: [1,0,0,0,0,0,0,0,0,2,0,1,0,0,0,0,0,0,0,0,0,0,0,0]; the vector corresponding to propylene is: [0,0,0,0,0,0,0,0,0,3,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0].
[0078] Step 402 involves storing the vectors corresponding to multiple oil molecules under various properties in the database, thus completing the database construction. Following the example in step 401, the vector corresponding to at least one oil molecule in a mixture under other different properties can be determined. Further, the vectors corresponding to different oil molecules under different properties are stored in the database, gradually building up the database.
[0079] In this embodiment of the specification, the structure-oriented aggregate data storage and retrieval further includes: querying the database for the structure code corresponding to the search conditions provided by the user, wherein the search conditions correspond to preset properties. In this specification, the search conditions provided by the user include: the components to be searched and the search properties. The search properties correspond to the preset properties mentioned above. The components to be searched by the user are mixtures, and the mixture contains or corresponds to multiple oil molecules. Based on the search conditions provided by the user, relevant data can be quickly matched from the database.
[0080] Figure 5 The diagram shown is a flowchart of a method for constructing a database according to an embodiment of this specification, which specifically includes the following steps:
[0081] Step 501: Obtain search criteria, which include: search properties.
[0082] In this step, the search criteria are those input or provided by the user, including: the components to be searched and the search properties. Specifically, the components to be searched are the mixtures to be searched, and based on the components to be searched, the constituent parts of the mixture can be determined.
[0083] The search properties correspond to the preset properties, including chemical properties, physical properties, and molecular types. The search properties included in the search criteria also have corresponding values. For example, the user inputs the search criteria as: a mixture and the search property is mass fraction. The components of this mixture include methane, cyclopropane, and propylene. The mass fractions of the three oil molecules to be searched, methane, cyclopropane, and propylene, and the corresponding mass fractions of each oil molecule are 0.23412356, 0.41235671, and 0.35351928, respectively.
[0084] Step 502: Based on the search property, query the database for the vector corresponding to the search property and the structure code corresponding to the vector.
[0085] In this embodiment of the specification, based on the search property, a data table corresponding to the search property is queried from the database. Further, based on the value corresponding to the search property, the data table is checked to see if the value exists. If the data table records the value of the search property, the vector data corresponding to the value of the search property recorded in the database is directly obtained. For example, if the search property is mass fraction, and the mass fraction of the component "methane" in the user-input search condition is 0.23412356, then the mass fraction data table is queried from the database, and it is checked whether the value 0.23412356 is recorded in the mass fraction data table. If it is, the vector corresponding to 0.23412356 is obtained, and the structure encoding is determined based on the vector.
[0086] Table 2 is a table of mass molecules in the database of this embodiment. In this embodiment, one or more vector results can be retrieved from the database table (such as Table 2) based on the value of the search property. When multiple vector results are retrieved, all vector results and their corresponding structure codes can be returned to the user.
[0087] Table 2
[0088]
[0089]
[0090] In this step, based on the mass fractions of the three oil molecules to be searched in the search criteria, the database is queried, and the mass fraction table stored in the database is matched with the mass fraction property, and then the vector data corresponding to the mass fraction 0.23412356 is retrieved.
[0091] Further, based on the vector construction rules, the vector is converted into the corresponding structure code, and then the structure code data is returned to the user.
[0092] For example, if the retrieved vector is [0,0,0,0,1,0,0,0,0,1,0,0,0,0,0,0,2,0,0,0,0,0,0], it can be determined that the corresponding SOL data structure encoding for this vector data is: it has one N5, one R, and one NN.
[0093] This manual categorizes and stores the structural codes corresponding to oil molecules based on different preset properties, enriching the classification types of oil molecule structural codes. This increases the number of data table types for categorized storage, further enhancing the diversity of data tables in the database and improving data storage efficiency. Simultaneously, the presence of numerous different types of data tables in the database makes data retrieval more convenient for users based on search criteria, enhancing the data query experience.
[0094] Figure 6 The diagram shown is a flowchart of a method for providing matching results to a user according to an embodiment of this specification. In this embodiment, the search conditions further include the component to be searched and the vector corresponding to the component to be searched, wherein the vector is a query vector. The component to be searched is a mixture to be searched, and the components constituting the mixture to be searched can be determined based on the component to be searched.
[0095] When the structure code corresponding to the search criteria cannot be matched based on the search criteria, the method specifically includes the following steps:
[0096] Step 601: Calculate the similarity between the query vector and the vector corresponding to the oil molecule recorded in the database, and obtain at least one similarity result.
[0097] In the embodiments of this specification, when the database is not exhaustively compiled to include all oil molecules, there may be situations where the structure code cannot be accurately matched according to the search conditions; that is, the corresponding structure code cannot be obtained based on the search conditions. In this step, the vector corresponding to the component to be searched is known in the search conditions and serves as the query vector. When the component to be searched is difficult to explore and develop and is a scarce resource, similar oil molecules can be matched from the database based on the query vector to replace or substitute the component to be searched. In this step, a suitable similarity calculation method (such as cosine similarity, Euclidean distance, etc.) can be used to calculate the similarity between the query vector and the vector corresponding to the oil molecules recorded in the database. In the embodiments of this specification, the similarity between one or more query vectors and the vectors corresponding to the oil molecules recorded in the database can be calculated to obtain the similarity result.
[0098] Step 602: The similarity results are sorted, and the vector corresponding to the highest similarity result is returned to the user. In this embodiment of the specification, the vectors corresponding to oil molecules are sorted according to the similarity calculation results to obtain the vector with the highest similarity, and the oil molecule corresponding to that vector is determined. The sorting results are displayed to the user in a suitable format so that the user can quickly find the most relevant data.
[0099] Step 603: Replace the component to be searched with the oil molecule corresponding to the vector corresponding to the highest similarity result to generate a new mixture. In the embodiments of this specification, the oil molecule corresponding to the vector corresponding to the highest similarity result is similar to the component to be searched in terms of physicochemical properties and molecular type. Using this oil molecule, it can be mixed with other components of the mixture in the search conditions to form a new mixture, which is easy to obtain or develop, can also save costs, and guide users to improve operational efficiency in oil and gas field exploration and development.
[0100] In the embodiments of this specification, the similarity between one or more query vectors and the vectors corresponding to oil molecules recorded in the database can be calculated to obtain similarity results.
[0101] The embodiments in this specification also include:
[0102] Sparse matrices are used to preprocess vector data;
[0103] The preprocessed vector data is stored in the database according to a preset storage format. Specifically, this includes: preprocessing the SOL vector data using efficient data storage structures (including but not limited to sparse matrices and compressed vectors) to reduce data redundancy and improve data storage efficiency. The preprocessed SOL vector data is then stored in the database according to a specific storage format (such as a multidimensional index structure). This allows for quick location of SOL vector data related to the user's search criteria using the multidimensional index results in the database, and further filtering of the located SOL vector data to meet the user's search needs.
[0104] like Figure 7 The diagram shown is a structural schematic of a structure-oriented centralized data storage and retrieval device according to an embodiment of this specification. The basic structure of the device is illustrated in this diagram. The functional units and modules can be implemented in software, or using general-purpose chips or specific chips to implement structure-oriented centralized data storage and retrieval. Specifically, the device includes:
[0105] The structure coding determination unit 701 is used to combine the codes corresponding to each basic structural unit of each oil molecule according to the chemical structure of the oil molecule to obtain the structure code representing the oil molecule.
[0106] The classification unit 702 is used to classify the structure codes corresponding to the oil molecules according to preset properties, so as to obtain the structure codes of each oil molecule under each type of property.
[0107] Database construction unit 703 is used to construct a database for storing oil molecules based on the structural codes of each oil molecule under various properties.
[0108] like Figure 8 The diagram shown illustrates an example of SOL data from an embodiment of this specification. In molecular refining processes, structure-guided lumping effectively integrates the chemical structure information of molecules, facilitating molecular analysis and processing.
[0109] The definitions and codes of various groups in the SOL notation are as follows: A6: a six-carbon aromatic ring, a structural unit that can exist alone; A4, A2: four-carbon and two-carbon aromatic rings, structural increments; N6, N5: six-carbon and five-carbon cycloalkanes; N4, N3, N2, N1: aliphatic ring structural increments of four, three, two, and one carbon atom, respectively; R: the total number of carbon atoms contained in all alkyl structures attached to the ring structure, or the number of carbon atoms in an aliphatic molecule when no ring structure is present; br: the number of branch nodes on side-chain alkyl, straight-chain alkyl, or olefin; me: the number of methyl groups in the alkyl structure that directly connect to carbon atoms in the aromatic or aliphatic ring; IH: structural increments related to hydrogen to describe the saturation of the molecule (except for aromatics). AA: a biphenyl bridge structure between any two non-structural increment rings (A6, N6, or N5). NS, NN, NO: sulfur, nitrogen, and oxygen atoms located in an aliphatic ring or aliphatic chain and attached to two carbon atoms. (Substitution of -CH2-); RS, RN, RO: Insertion of an S atom, -NH- group, or O atom between a carbon atom and a hydrogen atom. AN: Substitution of a carbon atom with a nitrogen group in an aromatic ring, such as pyridine and quinoline; Ko: Substitution of -CH2- or -CH3 to form a ketone or aldehyde group. Ni, V: Appear in porphyrin molecules.
[0110] like Figure 9 The diagram illustrates a computer device according to an embodiment of this specification. The structure-oriented centralized data storage and retrieval method described in this application can be applied to the computer device. The computer device 902 may include one or more processors 904, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 902 may also include any memory 906 for storing any kind of information such as code, settings, data, etc. Non-limitingly, for example, the memory 906 may include any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory can represent a fixed or removable component of the computer device 902. In one case, when the processor 904 executes associated instructions stored in any memory or combination of memories, the computer device 902 can perform any operation of the associated instructions. The computer device 902 also includes one or more drive mechanisms 908 for interacting with any memory, such as hard disk drive mechanisms, optical disk drive mechanisms, etc.
[0111] Computer device 902 may also include an input / output module 910 (I / O) for receiving various inputs (via input device 912) and providing various outputs (via output device 914). A specific output mechanism may include a presentation device 916 and an associated graphical user interface (GUI) 918. In other embodiments, the input / output module 910 (I / O), input device 912, and output device 914 may be omitted, and the device may function solely as a computer device within a network. Computer device 902 may also include one or more network interfaces 920 for exchanging data with other devices via one or more communication links 922. One or more communication buses 924 couple the components described above together.
[0112] Communication link 922 can be implemented in any way, such as via a local area network (LAN), a wide area network (WAN) (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 922 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0113] Corresponding to Figures 1 to 6 In addition to the methods described above, embodiments of this specification also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the methods described above.
[0114] This specification also provides computer-readable instructions, wherein when a processor executes the instructions, the program therein causes the processor to perform the following... Figures 1 to 6 The method shown.
[0115] It should be understood that in the various embodiments of this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.
[0116] It should also be understood that, in the embodiments of this specification, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this specification generally indicates that the preceding and following related objects have an "or" relationship.
[0117] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.
[0118] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0119] In the several embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, or they may be electrical, mechanical, or other forms of connection.
[0120] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments described in this specification, depending on actual needs.
[0121] Furthermore, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0122] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0123] This specification uses specific embodiments to illustrate the principles and implementation methods of this specification. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this specification. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this specification. Therefore, the content of this specification should not be construed as a limitation of this specification.
Claims
1. A structure-oriented lumped data storage and retrieval method, characterized in that, The method includes: Based on the chemical structure of oil molecules, the codes corresponding to each basic structural unit of each oil molecule are combined to obtain the structural code representing the oil molecule. The structure codes corresponding to the oil molecules are classified according to preset properties to obtain the structure codes of each oil molecule under each type of property. A database for storing oil molecules is constructed based on the structural codes of each oil molecule under various properties.
2. The method according to claim 1, characterized in that, The codes corresponding to the basic structural units of each oil molecule are determined as follows: The oil molecules are decomposed according to their chemical structure to obtain the basic structural units of each oil molecule. The basic structural units are uniquely encoded to obtain the codes corresponding to each basic structural unit of each oil molecule.
3. The method according to claim 2, characterized in that, The preset properties include: molecular type, physical properties, and chemical properties; The structure codes corresponding to the oil molecules are classified according to preset properties, resulting in the following structure codes for each type of oil molecule: By storing the structural codes of oil molecules with the same physical properties in the same data table, the structural codes of oil molecules under various physical properties can be obtained. Store the structure codes of oil molecules with the same molecular type in the same data table to obtain the structure codes of each oil molecule under each molecular type. The structural codes of oil molecules with the same chemical properties are stored in the same data table to obtain the structural codes corresponding to each oil molecule under the stated chemical properties.
4. The method according to claim 1, characterized in that, Based on the structural codes of oil molecules under various properties, a database for storing oil molecules is constructed, including: The structural codes of each oil molecule under various properties are converted into vector data to obtain vectors corresponding to multiple oil molecules under various properties. The vectors corresponding to multiple oil molecules under various properties are stored in the database to complete the database construction.
5. The method according to claim 1, characterized in that, The method further includes: querying the database for a structure code corresponding to the search criteria provided by the user, wherein the search criteria correspond to a preset property.
6. The method according to claim 5, characterized in that, The process of retrieving the structure code corresponding to the search criteria provided by the user from the database further includes: The search criteria include: search properties; Based on the search property, query the database for the vector corresponding to the search property and the structure code corresponding to the vector.
7. The method according to claim 6, characterized in that, The search criteria further include the component to be searched and the vector corresponding to the component to be searched, wherein the vector is a query vector; When the structure code corresponding to the search criteria cannot be matched according to the search criteria, the method further includes: Calculate the similarity between the query vector and the vector corresponding to the oil molecules recorded in the database, and obtain at least one similarity result; The similarity results are sorted, and the vector corresponding to the highest similarity result is fed back to the user. The oil molecule corresponding to the vector of the highest similarity result is used to replace the component to be searched, and a new mixture is generated.
8. The method according to claim 1, characterized in that, The method further includes: Sparse matrices are used to preprocess vector data; The preprocessed vector data is stored in the database according to the preset storage format.
9. A structure-oriented lumped data storage and retrieval device, characterized in that, The device includes: The structure coding determination unit is used to combine the codes corresponding to each basic structural unit of each oil molecule according to the chemical structure of the oil molecule to obtain the structure code representing the oil molecule. The classification unit is used to classify the structure codes corresponding to the oil molecules according to preset properties, so as to obtain the structure codes of each oil molecule under each type of property; The database construction unit is used to build a database for storing oil molecules based on the structural codes of each oil molecule under various properties.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 8.