Horizontal Scaling Method, Device and Computer Equipment for Distributed Database

By constructing geological data trees and constraint models, the problem of low scaling efficiency of traditional distributed databases in geological drilling data processing is solved, efficient horizontal scaling and query optimization are achieved, and resource occupancy is reduced.

CN114676194BActive Publication Date: 2025-07-25JIANGMEN POLYTECHNIC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210252936.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-07-25
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

Traditional distributed databases have low horizontal scaling efficiency in geological drilling data processing and have high resource occupancy, making it difficult to meet the needs of large data volume and efficient query.

Method used

By constructing a geological data tree and constraint model, we can obtain the data element characteristics of the geological drilling database, calculate the change parameters of similarity coefficients and attribute characteristics, establish a horizontal scaling model, and optimize query rules to improve scaling efficiency.

Benefits of technology

It improves the horizontal scaling efficiency of the geological drilling database, reduces resource occupancy, and realizes efficient data processing and query optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114676194B_ABST
    Figure CN114676194B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method, device, and computer device for horizontal expansion of a distributed database, which relates to the field of databases. The method includes: obtaining a geological drilling database, where the geological drilling database includes a plurality of geological drilling data elements, and the geological drilling data elements include a plurality of data attributes; obtaining data element features corresponding to the plurality of geological drilling data elements according to the geological drilling data elements; constructing a geological data tree according to the query rules of the geological drilling database and the data element features; obtaining a constraint model according to the geological data tree and the plurality of geological drilling data elements; and obtaining a horizontal expansion model according to the constraint model. By constructing a geological data tree and a constraint model, the efficiency of horizontal expansion of the geological drilling database can be improved, and the resource occupancy rate can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of databases, and particularly to a method, device, and computer device for horizontally expanding a distributed database. Background Art

[0002] Traditional database technologies mostly provide database services in a single-machine mode, where the database service resides on a single computer. The single-machine database model is simple. With the development of the Internet and the advent of the big data era, database technologies have developed rapidly, and distributed databases are used to solve the problems of massive storage and concurrent access of databases. Geological drilling data is an important resource in the national geological industry, with characteristics such as a large amount of data, rich variety, and high value. For geological drilling data, the horizontal expansion efficiency of traditional distributed databases is relatively low. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems of the existing technology to at least a certain extent, and provide a method, device, and computer device for horizontally expanding a distributed database. By constructing a geological data tree and a constraint model, the horizontal expansion efficiency of the geological drilling database can be improved, and the resource occupancy rate can be reduced.

[0004] The technical solutions of the embodiments of the present invention are as follows:

[0005] In a first aspect, the present invention provides a method for horizontally expanding a distributed database, the method comprising:

[0006] Obtain a geological drilling database, the geological drilling database including a plurality of geological drilling data elements, and the geological drilling data elements including a plurality of data attributes;

[0007] According to the geological drilling data elements, obtain data element features corresponding to the plurality of geological drilling data elements;

[0008] According to the query rule of the geological drilling database and the data element features, construct a geological data tree;

[0009] According to the geological data tree and the plurality of geological drilling data elements, obtain a constraint model;

[0010] According to the constraint model, obtain a horizontal expansion model.

[0011] According to some embodiments of the present invention, the data element features corresponding to each geological drilling data element are characterized as one of a geological map, a geological form, a geological report, a geological document, and a geological classification;

[0012] When the geological data tree includes at least two nodes, the constructing a geological data tree according to the query rule of the geological drilling database and the data element features includes:

[0013] Construct the geological data tree with the geological classification as the parent node and at least one of the geological map, the geological form, the geological report, the geological document, and the geological classification as the child node, and / or

[0014] Construct the geological data tree with the geological report as the parent node and at least one of the geological map, the geological form, and the geological document as the child node.

[0015] According to some embodiments of the present invention, obtaining the constraint model based on the geological data tree and a plurality of the geological drilling data elements includes:

[0016] Obtain the attribute characteristics corresponding to each of the data attributes;

[0017] Calculate the similarity between a plurality of the geological drilling data elements according to the geological data tree to obtain a similarity coefficient;

[0018] According to the similarity coefficient, obtain each of the data attributes corresponding to the geological drilling data elements, and calculate the distance between the attribute characteristics corresponding to each of the obtained data attributes to obtain a change parameter of the attribute characteristics;

[0019] Classify each of the geological drilling data elements according to the similarity coefficient and the change parameter of the attribute characteristics to obtain the constraint model.

[0020] According to some embodiments of the present invention, obtaining the horizontal expansion model based on the constraint model includes:

[0021] Calculate the update speed corresponding to the geological drilling data elements according to the constraint model and a preset statistical model;

[0022] Calculate the spatial position parameters corresponding to each of the geological drilling data elements according to the update speed, where the spatial position parameters characterize the spatial position of the geological drilling data elements in the geological drilling database;

[0023] Calculate the expansion efficiency corresponding to the geological drilling database according to each of the geological drilling data elements, the update speed, and the spatial position parameters to obtain the horizontal expansion model.

[0024] According to some embodiments of the present invention, calculating the expansion efficiency corresponding to the geological drilling database according to each of the geological drilling data elements, the update speed, and the spatial position parameters to obtain the horizontal expansion model includes:

[0025] Calculate the expansion efficiency corresponding to the geological drilling database through the expansion efficiency algorithm to obtain the horizontal expansion model; wherein, the expansion efficiency algorithm is expressed as:

[0026]

[0027] μ represents the spatial position parameter, e i represents the i-th geological drilling data element, h j represents the j-th data attribute corresponding to the geological drilling data element, i and j represent positive integers, ω represents the data volume of the geological drilling database, and η represents the update speed.

[0028] According to some embodiments of the present invention, the method of obtaining each of the data attributes corresponding to the geological drilling data element according to the similarity coefficient, and calculating the distances between the attribute characteristics corresponding to each of the obtained data attributes to obtain the change parameter of the attribute characteristics includes:

[0029] Calculate the distances between the attribute characteristics corresponding to each of the data attributes obtained through the distance algorithm to obtain the change parameter of the attribute characteristics;

[0030] Wherein, the distance algorithm is expressed as:

[0031] e(y,z) = ((y - z) 2 B(y - z) 1 / 2 )

[0032] e(y,z) represents the change parameter of the attribute characteristics, both y and z represent the attribute characteristics, and B represents the non-negative definite matrix of the data in the geological drilling database.

[0033] According to some embodiments of the present invention, after obtaining the horizontal expansion model according to the constraint model, the method further includes:

[0034] Obtain the query range of the geological drilling database;

[0035] According to the query range, the horizontal expansion model and a preset query rule, obtain a query optimizer.

[0036] In a second aspect, the present invention provides a horizontal expansion device for a distributed database, and the device includes:

[0037] A data acquisition module, configured to acquire a geological drilling database, where the geological drilling database has a plurality of geological drilling data elements, and the geological drilling data elements include a plurality of data attributes;

[0038] A first processing module, configured to obtain data element features corresponding to a plurality of the geological drilling data elements according to the geological drilling data elements;

[0039] A second processing module, configured to construct a geological data tree according to the query rule of the geological drilling database and the data element characteristics;

[0040] A third processing module, configured to obtain a constraint model according to the geological data tree and a plurality of the geological drilling data elements;

[0041] A fourth processing module, configured to obtain a horizontal expansion model according to the constraint model.

[0042] In a third aspect, the present invention provides a computer device, which includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by one or more of the processors, the one or more processors are caused to execute the steps of any one of the methods described in the first aspect above.

[0043] In a fourth aspect, the present invention further provides a computer-readable storage medium, which can be read and written by a processor. The storage medium stores computer instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of any one of the methods described in the first aspect above.

[0044] The technical solutions provided by the embodiments of the present invention have the following beneficial effects:

[0045] In the embodiments of the present invention, a geological drilling database is obtained. The geological drilling database includes a plurality of geological drilling data elements, and the geological drilling data elements include a plurality of data attributes. Then, according to the geological drilling data elements, data element characteristics corresponding to the plurality of geological drilling data elements are obtained. According to the query rule of the geological drilling database and the data element characteristics, a geological data tree is constructed, and the association relationship between data is obtained by constructing the geological data tree. According to the geological data tree and a plurality of geological drilling data elements, a constraint model is obtained, and the data categories that meet the conditions are judged through the constraint model. According to the constraint model, a horizontal expansion model is obtained, and the geological drilling database is horizontally expanded through the horizontal expansion model, improving the efficiency of horizontal expansion of the distributed database. By constructing the geological data tree and the constraint model, the embodiments of the present invention can improve the efficiency of horizontal expansion of the geological drilling database and reduce the resource occupancy rate.

[0046] Other features and advantages of the present invention will be described in the subsequent description, and, in part, will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the description, the claims, and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation to the technical solution of the present invention.

[0048] Figure 1 It is a schematic structural diagram of a distributed database horizontal expansion device provided by an embodiment of the present invention;

[0049] Figure 2 It is a schematic flow diagram of a distributed database horizontal expansion method provided by an embodiment of the present invention;

[0050] Figure 3 is Figure 2 a schematic sub-step flow diagram of step S130 in

[0051] Figure 4 is Figure 2 a schematic sub-step flow diagram of step S140 in

[0052] Figure 5 is Figure 2 a schematic sub-step flow diagram of step S150 in

[0053] Figure 6 It is a schematic flow diagram of a distributed database horizontal expansion method provided by another embodiment of the present invention;

[0054] Figure 7 It is a schematic diagram of a geological data tree of a distributed database horizontal expansion method provided by an embodiment of the present invention;

[0055] Figure 8 It is a schematic diagram of a geological data tree of a distributed database horizontal expansion method provided by another embodiment of the present invention

[0056] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] It should be noted that although the logical order is shown in the flow chart, in some cases, the steps shown or described may be executed in a different order than that in the flow chart. Terms such as "first" and "second" in the description and claims and the above accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0059] Horizontal expansion, also known as scale-out, uses more nodes to support a large number of requests. Based on this, embodiments of the present invention provide a method, apparatus, and computer device for horizontal expansion of a distributed database. The method for horizontal expansion of the distributed database obtains a geological drilling database, which includes a plurality of geological drilling data elements, and the geological drilling data elements include a plurality of data attributes. Then, based on the geological drilling data elements, data element features corresponding to the plurality of geological drilling data elements are obtained. According to the query rules of the geological drilling database and the data element features, a geological data tree is constructed, and the association relationship between data is obtained by constructing the geological data tree. According to the geological data tree and the plurality of geological drilling data elements, a constraint model is obtained, and the constraint model is used to judge data categories that meet the conditions. According to the constraint model, a horizontal expansion model is obtained, and the geological drilling database is horizontally expanded through the horizontal expansion model, improving the efficiency of horizontal expansion of the distributed database. The above method can improve the efficiency of horizontal expansion of the geological drilling database and reduce the resource occupancy rate by constructing a geological data tree and a constraint model.

[0060] The following further elaborates on the embodiments of the present invention with reference to the accompanying drawings.

[0061] See Figure 1 , Figure 1 which shows a schematic structural diagram of a device for horizontal expansion of a distributed database provided by an embodiment of the present invention. In the Figure 1 example, the device obtains a geological drilling database through a data acquisition module 110. The geological drilling database includes a plurality of geological drilling data elements, and the geological drilling data elements include a plurality of data attributes. Then, a first processing module 120 is used to obtain data element features corresponding to the plurality of geological drilling data elements based on the geological drilling data elements. A second processing module 130 is used to construct a geological data tree according to the query rules of the geological drilling database and the data element features, and the association relationship between data is obtained by constructing the geological data tree. A third processing module 140 is used to obtain a constraint model according to the geological data tree and the plurality of geological drilling data elements, and the constraint model is used to judge data categories that meet the conditions. A fourth processing module 150 is used to obtain a horizontal expansion model according to the constraint model, and the geological drilling database is horizontally expanded through the horizontal expansion model, which can improve the efficiency of horizontal expansion of the geological drilling database and reduce the resource occupancy rate.

[0062] It should be noted that the data acquisition module 110 is respectively connected to the first processing module 120, the second processing module 130, the third processing module 140, and the fourth processing module 150. The first processing module 120 is connected to the second processing module 130, the second processing module 130 is connected to the third processing module 140, and the third processing module 140 is connected to the fourth processing module 150. Among them, the first processing module 120, the second processing module 130, the third processing module 140, and the fourth processing module 150 are all central processing units, and a central processing unit generally consists of a logical operation unit, a control unit, and a storage unit. According to different inputs received by the operation unit, different processes are performed, and calculations are carried out through the arithmetic unit in the computer, which improves the calculation efficiency and saves a large amount of human resources.

[0063] The device and application scenarios described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation to the technical solutions provided by the embodiments of the present invention. Those skilled in the art can know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0064] Those skilled in the art can understand that Figure 1 the distributed database horizontal expansion device shown in does not constitute a limitation to the embodiments of the present invention, and may include more or fewer modules than shown, or combine certain components, or have different component arrangements.

[0065] According to the above distributed database horizontal expansion device, the following describes each embodiment of the distributed database horizontal expansion method of the present invention.

[0066] As Figure 2 shown, Figure 2 shows a schematic flowchart of a distributed database horizontal expansion method provided by an embodiment of the present invention. This distributed database horizontal expansion method is applied to a distributed database horizontal expansion device. This distributed database horizontal expansion method includes but is not limited to steps S110, S120, S130, S140, and S150.

[0067] Step S110, obtain a geological drilling database, where the geological drilling database includes multiple geological drilling data elements, and the geological drilling data elements include multiple data attributes.

[0068] It is understandable that a data model element in a geological database is defined as a geological data element. Geological drilling data elements are also known as geological borehole data. Geological drilling data elements include multiple attributes. Exemplarily, they include: groundwater level, borehole diameter, low-level information, borehole information, etc. Geological drilling data has the characteristics of large data volume, rich variety, and high value. By obtaining the geological drilling database, it is beneficial to perform data processing on geological drilling data and facilitate the subsequent establishment of a geological data tree.

[0069] Step S120: Obtain data element characteristics corresponding to multiple geological drilling data elements according to the geological drilling data elements.

[0070] It is understandable that by analyzing geological and mineral resource data, according to the document classification method, geological data elements can be abstractly represented as five basic data: geological maps, geological forms, geological reports, geological documents, and document classifications. These five basic data are used as the data element characteristics corresponding to the geological data elements. Among them, geological maps, geological documents, and geological forms are physical entities that contain entity data and attribute data, while geological reports and geological classifications are abstract concepts that only contain attribute data and do not contain entity data. These five data elements are the basic data units of geological and mineral data. For the sake of easy description, these five data are defined in the form of a triad. The specific definition is as follows:

[0071] DZdataElement = <Type, MetaData, EntiyData>

[0072] Among them, Type is a type identifier used to determine the type of geological data unit, corresponding to one of the above five basic data units. For a certain geological data element, it can only be one of a geological map, a geological document, a geological state, a geological report, and a geological classification, and has certainty. MetaData is the attribute data of the geological data element, which describes the attribute information of the geological data element. EntiyData is the entity data of the geological information. Geological maps, geological forms, and geological documents can contain entity data. Geological maps, geological documents, geological states, geological reports, and geological classifications have good application value. For the extensive and large-volume geological and mineral data, through the above method, five basic data element characteristics of the data can be extracted, which can be used for data classification and data description. These five representations of geological data are the model elements of the comprehensive geological data model and can be used to describe the static structure and dynamic operation of the model.

[0073] Step S130: Construct a geological data tree according to the query rules of the geological drilling database and the data element characteristics.

[0074] In one embodiment, the data element characteristics corresponding to each geological drilling data element are characterized as one of a geological map, geological morphology, geological report, geological document, and geological classification. Five representation methods corresponding to the data element characteristics are obtained through step S120, and the above five representation methods are referred to as geological data elements. Through the query rules of the geological data elements and the geological drilling database, a geological data tree is constructed. The five basic geological data elements serve as the nodes of the tree data structure and are used to describe the characteristics of geological mineral resource information data and the mutual constraint relationships between them. Different from an ordinary tree, the geological data tree is actually a deformed tree-like structure. Its geological data elements are not only related to the parent node and child nodes, but also in the geological data tree, there are the following three constraint relationships between geological data elements: (1) As the entity data of the smallest unit in the geological data elements, the geological map, geological document, and geological morphology, the geological map cannot be extracted from other types of geological data elements. If a geological data element is one of the geological map, geological morphology, and geological document types, it cannot have a successor; (2) As two conceptual entities, the geological report and geological classification can have successors and predecessors. If they have predecessors, then their predecessors can only be the geological classification, rather than other geological data elements; (3) If the geological report data element has a successor, the successor can only be one of the three types of geological map, geological document, and geological morphology, rather than the geological classification. The successor of the geological classification element can be any of the five geological data elements. By constructing the geological data tree, not only can the association relationships between data elements be obtained, but also the static representation and dynamic operation of each geological data element can be represented.

[0075] It should be noted that using the above three constraint relationships, a formal definition of the geological data tree can be given: DZDataTree = (D, R). Where: D represents the geological drilling database, which contains multiple geological drilling data elements, and R represents the set of relationships on the geological drilling data elements in D, including but not limited to the following three cases: (1) If the geological data tree is empty, indicating that there is no data in D, then the geological data tree is an empty geological data tree; (2) If there is only one geological data element in D but no definition of relationships, then R is empty; (3) If D contains two or more geological data elements, then there is a set of relationships R = {H} on the geological drilling data elements, where H has the following binary relationships: A geological data element in D serves as the root node of the geological data tree. If the geological data element is a geological map, geological morphology, or geological document, then it has no successor elements; if it is a geological report element, then the successor can only be a geological map, geological document, or geological morphology. Exemplarily, there is a partition of D - {root}, D1, D2, …, D m , where m > 0; for any j ≠ k (1 ≤ j, k ≤ m), there is For any i ≠ k (1 ≤ i ≤ m), there exists a geological data element x i ∈ D i such that <root, x i > ∈ H. For the partition of D - {root} from H - {<root, x1>, …, <root, x m >}, there is a unique partition H1, H2, … H m (m > 0). For any j ≠ k (1 ≤ j, k ≤ m), there is For any i (1 ≤ i ≤ m), H i is a binary relation on D i , and (D i , {H i}) is the geological data tree root defined according to this.

[0076] Reference Figure 3 、 Figure 7 and Figure 8 , in the case where the geological data tree includes at least two nodes, according to the query rules and data element characteristics of the geological drilling database, constructing a geological data tree includes, but is not limited to, the following steps:

[0077] Step 131, using the geological classification as the parent node and at least one of the geological map, geological morphology, geological report, geological document, and geological classification as the child node to construct the geological data tree and / or

[0078] Step 132, using the geological report as the parent node and at least one of the geological map, geological morphology, and geological document as the child node to construct the geological data tree.

[0079] It should be noted that in the case where the geological data tree includes at least two nodes, in a geological data tree, the node without a parent node is the root node. If a node has no parent node, it is said that this node has no predecessor. If a node has no child node, it is said that this node has no successor. According to the above three constraint relationships, when using the geological classification as the parent node, the geological map, geological morphology, geological report, geological document, and geological classification can be used as its child nodes to construct the geological data tree; when using the geological report as the parent node, at least one of the geological map, geological morphology, and geological document is used as the child node to construct the geological data tree. It can be obtained that the geological map, geological morphology, and geological document can only be used as the leaf nodes of the geological data tree and have no successors, which conforms to the above constraint conditions. By constructing the geological data tree, not only can the association relationship between data elements be obtained, but also the static representation and dynamic operation of each geological data element can be represented.

[0080] Step S140, obtaining a constraint model according to the geological data tree and multiple geological drilling data elements.

[0081] It is understandable that data attributes related to known geological drilling data types are extracted from the geological data tree to obtain geological drilling data elements corresponding to the data attributes. Based on the geological data tree and multiple geological drilling data elements, a constraint model is obtained. By constructing the constraint model, it is possible to judge data categories that meet the conditions, facilitating subsequent horizontal expansion of the database and query optimization of the expanded database.

[0082] Reference Figure 4 , based on the geological data tree and multiple geological drilling data elements, to obtain a constraint model, including but not limited to the following steps:

[0083] Step S141, obtain the attribute characteristics corresponding to each data attribute.

[0084] It is understandable that according to the above steps S110 and S120, the geological drilling data elements include multiple data attributes. Obtain the attribute characteristics corresponding to the multiple data attributes in the manner of step S120. Similar to the above steps, it will not be elaborated here. By obtaining the attribute characteristics corresponding to each data attribute, it is convenient to construct the subsequent constraint model.

[0085] Step S142, according to the geological data tree, calculate the similarity between multiple geological drilling data elements to obtain a similarity coefficient.

[0086] It is understandable that data attributes related to known geological drilling data types are extracted from the geological data tree. According to the data attributes, multiple geological drilling data elements corresponding to the data attributes are obtained. Calculate the similarity between the multiple geological drilling data elements through a similarity formula to obtain a similarity coefficient. By calculating the similarity between each geological drilling data element and performing data processing, it is convenient to construct the subsequent constraint model.

[0087] Among them, the similarity formula is expressed as:

[0088]

[0089] s(w b ,w c ) represents the similarity coefficient, w b and w c both represent geological drilling data elements, w bc represents the vector between w b and w c , h g represents the data attribute, η represents the update speed, and m represents the quantity.

[0090] Step S143, according to the similarity coefficient, obtain each data attribute corresponding to the geological drilling data element, calculate the distance between the obtained attribute characteristics corresponding to each data attribute, and obtain a change parameter of the attribute characteristics.

[0091] In one embodiment, the change parameter of the attribute characteristics corresponding to the borehole data under different geological conditions is calculated using the following formula:

[0092]

[0093] where e(y, z) represents the change parameter of the attribute characteristics, y and z both represent the attribute characteristics, and m represents the quantity.

[0094] It can be understood that according to the above, if the value of s(w b , w c ) is changed, the difference in the data attribute characteristics is significant. Assuming s(w b , w c ) = 1, in the above formula for solving the change parameter of the attribute characteristics, the obtained attribute characteristic values of the distributed geological borehole data may be negative. Therefore, the following formula needs to be used for calculation: where e(y, z) represents the change parameter of the attribute characteristics, y and z both represent the attribute characteristics, and m represents the quantity.

[0095] In one embodiment, according to the similarity coefficient, each data attribute corresponding to the geological drilling data element is obtained, and the distance between the attribute characteristics corresponding to each calculated data attribute is obtained to obtain the change parameter of the attribute characteristics, including:

[0096] The distance between the attribute characteristics corresponding to each data attribute obtained by calculating through the distance algorithm is used to obtain the change parameter of the attribute characteristics;

[0097] where the distance algorithm is expressed as:

[0098] e(y, z) = ((y - z) 2 B(y - z) 1 / 2 )

[0099] e(y, z) represents the change parameter of the attribute characteristics, y and z both represent the attribute characteristics, and B represents the non - negative definite matrix of the data in the geological drilling database.

[0100] It should be noted that assuming B is the identity matrix, the above formula can be transformed into:

[0101]

[0102] It should be noted that assuming B is a diagonal matrix, the above formula can be transformed into:

[0103]

[0104] Through the above method, the data in the distributed geological drilling database can be processed to obtain data variables, and more than 90% of the information characteristics of the distributed geological drilling database can be obtained.

[0105] Step S144: Classify each geological drilling data element according to the similarity coefficient and the change parameter of the attribute characteristics to obtain a constraint model.

[0106] It can be understood that, according to the similarity coefficient and the change parameter of the attribute characteristics, the differences between geological drilling data elements can be obtained, and the geological drilling data elements in the geological drilling database can be classified. According to the characteristics of the geological drilling data elements, the information in the geological drilling database can be divided into several different categories, providing an accurate data basis for the expansion of the geological drilling database. Establish multiple constraint models to judge the data categories that meet the query conditions, and perform data queries according to different data categories to achieve the expansion and optimization of the geological drilling database.

[0107] In one embodiment, multiple constraint models are established through the following formula:

[0108]

[0109] where represents the difference between two vectors, Y represents the data element feature vector, represents the feature vector for classifying geological drilling data elements, T represents the median of a set of data, K represents the quantity, and Y j represents the feature of a certain data element. By establishing multiple constraint models, it can be used to optimize the horizontal expansion of the distributed geological drilling database.

[0110] Step S150: Obtain a horizontal expansion model according to the constraint model.

[0111] Reference Figure 5 , according to the constraint model, obtaining a horizontal expansion model includes but is not limited to the following steps:

[0112] Step S151: Calculate the update speed corresponding to the geological drilling data element according to the constraint model and the preset statistical model.

[0113] It should be noted that the data growth is affected by the constraint model. The preset statistical model includes a timer and a counter. Among them, the timer is used to count the time used for the actual data volume growth, and the counter is used to count the number of actual data volume growth. Thus, the update speed corresponding to the geological drilling data element, denoted as η, can be obtained by solving according to the number of data volume growth and the time of data volume growth. By calculating the update speed, it is convenient to calculate the horizontal expansion of the geological drilling database later.

[0114] Step S152: Calculate the spatial position parameters corresponding to each geological drilling data element according to the update speed, where the spatial position parameters represent the spatial positions of the geological drilling data elements in the geological drilling database.

[0115] It should be noted that according to the above Step S151, based on the update speed, the updated geological drilling data elements corresponding to the update speed are obtained. In the order of the increase of the geological drilling data elements, they are successively added to the dataset composed of the geological drilling data elements to obtain the subscripts of their positions in the dataset, so as to obtain the spatial position parameters corresponding to each geological drilling data element. Calculating the position parameters in space facilitates the subsequent horizontal expansion of the distributed geological drilling database.

[0116] Step S153: Calculate the expansion efficiency corresponding to the geological drilling database according to each geological drilling data element, the update speed, and the spatial position parameters to obtain a horizontal expansion model.

[0117] It can be understood that based on each geological drilling data element, by calculating the update speed corresponding to the geological drilling data element, the spatial position parameters corresponding to the geological drilling data element are obtained, and thus the spatial position of each geological drilling data element is obtained. According to the above calculation method, a horizontal expansion model is constructed, and the horizontal expansion model can be used for the horizontal expansion of the distributed geological drilling database.

[0118] It should be noted that the data volume of the geological drilling database is set as w1, the number of data attributes is g, and the set of geological drilling data elements composed of all geological drilling databases is {e1, e2, …, e w}, where e i is the i-th piece of data in the geological drilling database. The dataset composed of the data attributes of all geological drilling data elements is {h1, h2, …, h g}, where h j is the j-th data attribute in the geological drilling database, and the update speed of the geological drilling data is η.

[0119] In one embodiment, calculating the expansion efficiency corresponding to the geological drilling database according to each geological drilling data element, the update speed, and the spatial position parameters to obtain a horizontal expansion model includes:

[0120] Calculate the expansion efficiency corresponding to the geological drilling database through an expansion efficiency algorithm to obtain a horizontal expansion model;

[0121] Among them, the expansion efficiency algorithm is expressed as:

[0122]

[0123] μ represents the spatial position parameter, e i represents the i-th geological drilling data element, h j represents the j-th data attribute corresponding to the geological drilling data element. i and j represent positive integers. ω represents the data volume after the expansion of the geological drilling database, w1 represents the initial data volume of the geological drilling database, and η represents the update speed. According to the above method, the horizontal expansion of the geological drilling database can be carried out, so as to provide data support for different industries.

[0124] Reference Figure 6 , after obtaining the horizontal expansion model according to the constraint model, the method further includes but is not limited to the following steps:

[0125] Step S210, obtain the query range of the geological drilling database.

[0126] It should be noted that the query range includes the set of alternative execution plans for query requests that can produce the same result. Each query plan has an execution order, and different implementation methods of various operations may lead to different performances of the execution plan. The execution plan is usually abstracted as an operation tree, where the nodes are operations, and the shape of the tree determines the execution order of the operations. These operation trees are generated according to the transformation rules of the query request. These query operation trees are equivalent because they can generate the same result set. It is necessary to query all these query operation trees to obtain the query range. By obtaining the query range, it is convenient to construct the query optimizer subsequently.

[0127] Step S220, obtain the query optimizer according to the query range, the horizontal expansion model, and the preset query rules.

[0128] It can be understood that the query optimizer of the geological drilling database consists of three parts: the query range of the geological database, the generation of the horizontal expansion model of the geological drilling database, and the query rules of the geological drilling database. The query range is obtained according to step S210, the horizontal expansion model is obtained through the horizontal expansion method of the distributed database according to the query range, and then combined with the preset rules to obtain the query optimizer. The distributed horizontal expansion query of the geological drilling database is realized through the obtained query optimizer.

[0129] It should be noted that the horizontal expansion model includes the cost model of the query optimizer. The cost model of the query optimizer includes cost functions, data statistical information, intermediate result set estimation tools, etc. The main measurement standard of the execution cost is the execution time of the query. The preset query rule is to use the cost model to detect the execution plan, filter out the optimal solutions that do not generate the execution plan through dynamic programming or simulated annealing strategies in the search space, and it is necessary to check the plan according to the cost model to predict the response time, comprehensively consider all execution schemes, and finally obtain the best execution scheme.

[0130] It should also be noted that the distributed geological drilling database query optimizer can be used to construct a distributed geological drilling database query optimizer according to the horizontal expansion model of the drilling database and the element characteristics corresponding to the geological drilling data elements. By analyzing the data element characteristics, the similarity of the cumbersome and repetitive data in the geological borehole database can be distinguished, so as to describe the relevant data with high similarity in the geological drilling database with feature vectors, and realize the horizontal expansion query of the distributed geological drilling database. The expansion optimization process of the distributed geological drilling database refers to the process of generating a query execution plan (QEP), and the generated query execution plan should minimize the objective function, that is, the time required for query execution in a distributed environment. The horizontal expansion method provided by the present invention can also be applied to the horizontal expansion of the distributed geological drilling database, and its query optimizer is also applicable. In the search space stage, there is no need to consider whether the expansion of the distributed geological drilling database is centralized or distributed, because the execution plan is generated according to certain conversion rules, which is applicable in any case and can provide data support for different industries.

[0131] Reference Figure 9 , Figure 9 FIG. shows a computer device 900 provided by an embodiment of the present invention. The computer device 900 may be a server or a terminal, and the internal structure of the computer device 900 includes, but is not limited to:

[0132] A memory 910 for storing programs;

[0133] A processor 920 for executing the programs stored in the memory 910. When the processor 920 executes the programs stored in the memory 910, the processor 920 is used to execute the above-mentioned distributed database horizontal expansion method.

[0134] The processor 920 and the memory 910 may be connected by a bus or other means.

[0135] The memory 910, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the distributed database horizontal expansion method described in any embodiment of the present invention. The processor 920 realizes the above-mentioned distributed database horizontal expansion method by running the non-transitory software programs and instructions stored in the memory 910.

[0136] The memory 910 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store the distributed database horizontal scaling method described above. In addition, the memory 910 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 910 may optionally include a memory remotely located relative to the processor 920, and these remote memories may be connected to the processor 920 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0137] The non-transitory software programs and instructions required to implement the above-described distributed database horizontal scaling method are stored in the memory 910. When executed by one or more processors 920, the distributed database horizontal scaling method provided in any embodiment of the present invention is executed.

[0138] An embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions for executing the above-described distributed database horizontal scaling method.

[0139] In one embodiment, the storage medium stores computer-executable instructions that are executed by one or more control processors 920. For example, when executed by a processor 920 in the above computer device 900, the one or more processors 920 can be caused to execute the distributed database horizontal scaling method provided in any embodiment of the present invention.

[0140] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0141] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as or other transmission mechanisms, and can include any information delivery medium.

[0142] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention. Under the condition of sharing, these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. A method for horizontal scaling of a distributed database, characterized in that The method includes: Obtaining a geological drilling database, where the geological drilling database includes multiple geological drilling data elements, and the geological drilling data elements include multiple data attributes; Obtaining data element features corresponding to the multiple geological drilling data elements according to the geological drilling data elements; Constructing a geological data tree according to the query rule of the geological drilling database and the data element features; Obtaining a constraint model according to the geological data tree and the multiple geological drilling data elements; Obtaining a horizontal expansion model according to the constraint model; Wherein, the data element feature corresponding to each geological drilling data element is characterized as one of a geological map, a geological form, a geological report, a geological document, and a geological classification; When the geological data tree includes at least two nodes, the constructing a geological data tree according to the query rule of the geological drilling database and the data element features includes: Using the geological classification as the parent node and using at least one of the geological map, the geological form, the geological report, the geological document, and the geological classification as the child node to construct the geological data tree and / or Using the geological report as the parent node and using at least one of the geological map, the geological form, and the geological document as the child node to construct the geological data tree; The obtaining a constraint model according to the geological data tree and the multiple geological drilling data elements includes: Obtaining the attribute features corresponding to the respective data attributes; Calculating the similarity between the multiple geological drilling data elements according to the geological data tree to obtain a similarity coefficient; According to the similarity coefficient, obtaining the respective data attributes corresponding to the geological drilling data elements, and calculating the distance between the obtained attribute features corresponding to the respective data attributes to obtain a change parameter of the attribute features; Classifying each geological drilling data element according to the similarity coefficient and the change parameter of the attribute features to obtain the constraint model; The obtaining a horizontal expansion model according to the constraint model includes: Calculating the update speed corresponding to the geological drilling data element according to the constraint model and a preset statistical model; Calculating the spatial position parameter corresponding to each geological drilling data element according to the update speed, where the spatial position parameter characterizes the spatial position of the geological drilling data element in the geological drilling database; Calculating the expansion efficiency corresponding to the geological drilling database according to each geological drilling data element, the update speed, and the spatial position parameter to obtain the horizontal expansion model.

2. The horizontal expansion method of the distributed database according to claim 1, wherein The calculating the expansion efficiency corresponding to the geological drilling database according to each geological drilling data element, the update speed, and the spatial position parameter to obtain the horizontal expansion model includes: Calculating the expansion efficiency corresponding to the geological drilling database through an expansion efficiency algorithm to obtain the horizontal expansion model; Wherein, the expansion efficiency algorithm is expressed as: represents the spatial position parameter, represents the i-th geological drilling data element, represents the j-th data attribute corresponding to the geological drilling data element, where i and j represent positive integers, represents the data volume of the geological drilling database, represents the update speed.

3. The horizontal expansion method of the distributed database according to claim 1, wherein Obtaining each of the data attributes corresponding to the geological drilling data elements according to the similarity coefficient, and obtaining a change parameter of the attribute features by calculating the distances between the attribute features corresponding to each of the calculated data attributes, includes: Obtaining a change parameter of the attribute features by calculating the distances between the attribute features corresponding to each of the data attributes through a distance algorithm; Wherein, the distance algorithm is expressed as: The variation parameter representing the attribute feature, both y and z represent the attribute feature, and B represents the non-negative definite matrix of the data in the geological drilling database.

4. The horizontal expansion method of the distributed database according to claim 1, characterized in that After obtaining the horizontal expansion model according to the constraint model, the method further includes: Obtaining a query range of the geological drilling database; Obtaining a query optimizer according to the query range, the horizontal expansion model, and a preset query rule.

5. A distributed database horizontal expansion device, characterized in that, Including: A data acquisition module, configured to acquire a geological drilling database, where the geological drilling database has a plurality of geological drilling data elements, and the geological drilling data elements include a plurality of data attributes; A first processing module, configured to obtain data element features corresponding to the plurality of geological drilling data elements according to the geological drilling data elements; A second processing module, configured to construct a geological data tree according to the query rule of the geological drilling database and the data element features; A third processing module, configured to obtain a constraint model according to the geological data tree and the plurality of geological drilling data elements; A fourth processing module, obtaining a horizontal expansion model according to the constraint model; Wherein, the data element features corresponding to each of the geological drilling data elements are characterized as one of a geological map, a geological form, a geological report, a geological document, and a geological classification; When the geological data tree includes at least two nodes, constructing the geological data tree according to the query rule of the geological drilling database and the data element features includes: Using the geological classification as a parent node, and using at least one of the geological map, the geological form, the geological report, the geological document, and the geological classification as a child node to construct the geological data tree and / or Using the geological report as a parent node, and using at least one of the geological map, the geological form, and the geological document as a child node to construct the geological data tree; Obtaining the constraint model according to the geological data tree and the plurality of geological drilling data elements includes: Obtaining the attribute features corresponding to each of the data attributes; Calculating the similarity between the plurality of geological drilling data elements according to the geological data tree to obtain a similarity coefficient; Obtaining each of the data attributes corresponding to the geological drilling data elements according to the similarity coefficient, and obtaining a change parameter of the attribute features by calculating the distances between the attribute features corresponding to each of the calculated data attributes; Classifying each of the geological drilling data elements according to the similarity coefficient and the change parameter of the attribute features to obtain the constraint model; Obtaining the horizontal expansion model according to the constraint model includes: Calculating an update speed corresponding to the geological drilling data element according to the constraint model and a preset statistical model; Spatial position parameters corresponding to each of the geological drilling data elements are calculated according to the update speed, wherein the spatial position parameters characterize the spatial positions of the geological drilling data elements in the geological drilling database; According to each of the geological drilling data elements, the update speed, and the spatial position parameters, the expansion efficiency corresponding to the geological drilling database is calculated to obtain the horizontal expansion model.

6. A computer device, characterized in that, The computer device includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by one or more of the processors, one or more of the processors execute the horizontal expansion method of the distributed database according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The storage medium can be read and written by a processor. The storage medium stores computer instructions. When the computer instructions are executed by one or more processors, one or more processors execute the horizontal expansion method of the distributed database according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Continuous full scan data store table and distributed data store featuring predictable answer time for unpredictable workload

    CN102576369A

  • Geologic map database establishment model for describing geologic bodies through single feature

    CN104008167A