Big data-based multi-source geological data standardization processing method

CN115470203BActive Publication Date: 2026-09-18SUQIAN POWER SUPPLY COMPANY OF JIANGSU PROVINCE POWER
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211124230.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2026-09-18
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

地质数据的种类和来源都非常多,目前对地质数据的存储形式和存储媒介都没有严格的标准,各种来源的地质数据形式杂乱,导致在工程实践中采集到的地质数据很难直接调用

Benefits of technology

[0012]Beneficial effects: This invention first uses sample data to determine a standard data framework. The standard data framework includes the sorting of data categories of multiple geological data. Data categories that appear earlier in the standard data framework have larger data volumes and stronger reliability. Data categories with larger data volumes have a wider range of applications and are more likely to be applied. Data categories with stronger reliability are more likely to be used directly without verification, which can effectively improve the application efficiency of geological data. Then, the standard data framework is used to standardize the geological data to be processed to obtain standard data. When calling the standard data, whether manually or automatically, the parts with a wide range of applications and strong reliability can be quickly extracted from the standard data, which can improve the efficiency of geological data calling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470203B_ABST
    Figure CN115470203B_ABST
Patent Text Reader

Abstract

A method for standardizing multi-source geological data based on big data, the method comprising the following steps: S1, collecting a plurality of sample data; S2, determining the confidence of the sample data based on the source of the sample data; S3, clustering the sample data, and determining the richness of the sample data in different data categories according to the clustering result; S4, sorting the data categories of the sample data according to the richness to obtain a basic data framework; S5, correcting the basic data framework according to the confidence to obtain a standard data framework; and S6, standardizing the geological data to be processed by using the standard data framework to obtain standard data. The data amount of the data category in the position of the standard data framework formed by the present application is larger, and the reliability is stronger, the data amount of the data category is larger, the application range is more extensive, the possibility of being applied is higher, and the standardized geological data processed by the standard data framework can be more efficiently called.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geological exploration technology, specifically a method for standardizing multi-source geological data based on big data. Background Technology

[0002] Geological data has wide applications in numerous engineering fields, directly impacting the design and construction of many projects. Geological data comes from a wide variety of sources, and currently there are no strict standards for its storage format and media. The diverse and often chaotic nature of geological data from various sources makes it difficult to directly access data collected in engineering practice. When accessing geological data, manual screening and conversion are often required to extract the necessary information; this process is complex, time-consuming, and has low accuracy. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a method for standardizing multi-source geological data based on big data. In the resulting standard data framework, data categories positioned earlier have larger data volumes and stronger reliability. Data categories with larger data volumes have a wider range of applications and are more likely to be used. Standardized geological data processed through the standard data framework can be accessed more efficiently.

[0004] To achieve the above objectives, the specific solution adopted by the present invention is: a multi-source geological data standardization processing method based on big data, the method comprising the following steps: S1. Collect data from multiple samples; S2. Determine the confidence level of the sample data based on its source; S3. Cluster the sample data and determine the richness of sample data in different data categories based on the clustering results; S4. Sort the data categories of the sample data according to their richness to obtain the basic data framework; S5. Based on the confidence level, the basic data framework is modified to obtain the standard data framework; S6. Standardize the geological data to be processed using the standard data framework to obtain standard data.

[0005] As a further optimization of the above-mentioned multi-source geological data standardization processing method based on big data: In S1, source labels are assigned to sample data when collecting sample data; In S2, the confidence level of the sample data is determined based on the source label corresponding to the sample data.

[0006] As a further optimization of the above-mentioned multi-source geological data standardization processing method based on big data: In S2, the confidence level of the sample data is determined and the confidence level is assigned a first weight.

[0007] As a further optimization of the above-mentioned multi-source geological data standardization processing method based on big data: In S3, while determining the richness of sample data in different data categories, a second weight is assigned to the richness, and the second weight is greater than or less than the first weight.

[0008] As a further optimization of the above-mentioned multi-source geological data standardization processing method based on big data: In S4, all data categories are sorted in order of richness from high to low to obtain the basic data framework.

[0009] As a further optimization of the above-mentioned multi-source geological data standardization processing method based on big data, the specific methods of S5 include: S51. Determine the mode of the first weight of the sample data in each data category as the comparison benchmark; S52. In the basic data framework, each data category is grouped according to the second weight of different data categories to obtain several comparison groups; S53. In each comparison group, for two adjacent data categories, if the benchmark value of the latter data category is greater than the second weight of the former data category, the order of the two data categories shall be adjusted. S54. Repeat S53 to reorder the comparison group to obtain an ordered group; S55. Reassemble all ordered groups into a standard data frame.

[0010] As a further optimization of the above-mentioned multi-source geological data standardization processing method based on big data: In S52, the difference between the richness of the first data category and the richness of the last data category in a comparison group does not exceed a preset threshold.

[0011] As a further optimization of the above-mentioned multi-source geological data standardization processing method based on big data, the specific methods of S6 include: S61. Extract data point values ​​from geological data and determine the data category corresponding to the data point values; S62. Input the data point values ​​into the standard data frame according to the correspondence between data point values ​​and data categories; S63. Assign null values ​​to the parts of the standard data frame where no data points have been entered, and obtain the standard data.

[0012] Beneficial effects: This invention first uses sample data to determine a standard data framework. The standard data framework includes the sorting of data categories of multiple geological data. Data categories that appear earlier in the standard data framework have larger data volumes and stronger reliability. Data categories with larger data volumes have a wider range of applications and are more likely to be applied. Data categories with stronger reliability are more likely to be used directly without verification, which can effectively improve the application efficiency of geological data. Then, the standard data framework is used to standardize the geological data to be processed to obtain standard data. When calling the standard data, whether manually or automatically, the parts with a wide range of applications and strong reliability can be quickly extracted from the standard data, which can improve the efficiency of geological data calling. Attached Figure Description

[0013] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0015] Please see Figure 1 A standardization processing method for multi-source geological data based on big data is proposed, comprising S1 to S6. S1 to S5 can be combined into a standard determination stage, and S6 is the standardization processing stage. In practical applications, the standard determination stage only needs to be executed once, while the standardization processing stage is executed every time geological data is processed.

[0016] S1. Collect multiple sample data. The sample data can come from digital data exported from existing geographic information systems or from paper documents such as geological maps. For easier subsequent processing, the paper data also needs to be converted into digital data, which can be done using existing OCR (Optical Character Recognition) tools. The more sample data there are, the more types and quantities of geological data are included, enabling the formation of more precise standardized processing rules. The process of collecting sample data can be carried out using existing ETL (Extract-Transform-Load) tools, thus enabling more efficient and faster sample data collection.

[0017] S2. Determine the confidence level of the sample data based on its source. Confidence level is used to characterize the reliability of sample data. For sample data that is itself digital data, its confidence level is high. For digital data converted from paper data, because OCR tools have a certain error rate, the converted digital data may contain errors, so its confidence level is low.

[0018] S3. Cluster the sample data and determine the richness of sample data in different data categories based on the clustering results. Because there are many types of geological data, it is necessary to determine the data categories of the sample data in order to standardize the geological data. Using large-scale sample data, more data categories can be identified. The richness can be represented by the total number of sample data corresponding to a data category.

[0019] S4. Sort the data categories of the sample data according to their richness to obtain the basic data framework. Specifically, sort all data categories in descending order of richness to obtain the basic data framework. After sorting the data categories based on richness, the data categories that appear earlier in the list have a larger amount of data, while the data categories that appear later in the list have a smaller amount of data.

[0020] S5. The basic data framework is revised based on the confidence level to obtain the standard data framework. The basic data framework mainly reflects the quantity of sample data corresponding to different data categories and the amount of geological data, but it fails to reflect the reliability of the sample data. Therefore, it is necessary to revise the basic data framework using the confidence level. In the obtained standard data framework, the richness of the data categories that appear earlier in the document and the confidence level of the sample data they contain are both higher.

[0021] S6. Standardize the geological data to be processed using a standard data framework to obtain standard data. The standardization process refers to rewriting the geological data using a standard data framework.

[0022] This invention first uses sample data to determine a standard data framework. The standard data framework includes the sorting of data categories of multiple geological data. Data categories that appear earlier in the standard data framework have larger data volumes and stronger reliability. Data categories with larger data volumes have a wider range of applications and are more likely to be used. Data categories with stronger reliability are more likely to be used directly without verification, which can effectively improve application efficiency. Then, the standard data framework is used to standardize the geological data to be processed to obtain standard data. When calling the standard data, whether manually or automatically, the parts with a wide range of applications and strong reliability can be quickly extracted from the standard data, which can improve the efficiency of geological data retrieval.

[0023] To determine the confidence level of sample data more quickly and efficiently in S2, source labels are assigned to the sample data during collection in S1. These source labels characterize the origin of the sample data and can be determined based on its actual source. For example, if the source is printed data, the source label could be "pap"; if the source is imprecise digital data, such as image data, the source label could be "pic"; and if the source is precise digital data, such as text data, the source label could be "tex". Based on these source labels, S2 determines the confidence level of the sample data according to the corresponding source labels. The correspondence between confidence level and source label can be predetermined; for example, the source label "pap" corresponds to a confidence level of 60%, the source label "pic" corresponds to 80%, and the source label "tex" corresponds to 100%. The number and gradient of confidence levels can be determined based on the number of source labels.

[0024] Because when refining the basic data framework to obtain the standard data framework in S5, both the confidence level and the richness of the sample data need to be considered, in order to better integrate confidence level and richness, in S2, the confidence level of the sample data is determined and a first weight is assigned to it, with a higher confidence level resulting in a larger first weight. In S3, the richness of the sample data in different data categories is determined and a second weight is assigned to it, with the second weight being greater than or less than the first weight; similarly, a higher richness results in a higher second weight. By assigning the first and second weights, the basic data framework can be refined based on these weights. Furthermore, to ensure successful refinement of the basic data framework, the first and second weights should have a relative magnitude; the first weight cannot be completely greater than the second weight, nor can it be completely less than the second weight.

[0025] The specific methods of S5 include S51 to S55.

[0026] S51. Determine the mode of the first weight of the sample data in each data category as the comparison benchmark. Since confidence level corresponds one-to-one with sample data, the first weight also corresponds one-to-one with sample data. The second weight corresponds one-to-one with richness, i.e., with data category. Therefore, the number of first weights far exceeds the number of second weights, and they cannot be directly compared. Thus, the first weights need to be processed first. If the mean of the first weights of all sample data in each data category is directly calculated, some first weights may be too small or too large, causing the mean to fail to reflect the overall distribution of the true first weights of the sample data. Therefore, this invention uses the mode as the comparison benchmark.

[0027] S52. In the basic data framework, data categories are grouped according to their second weights to obtain several comparison groups. The main purpose of modifying the basic data framework is to avoid situations where relying solely on richness to rank data categories might result in excessively low confidence levels for samples from data categories ranked higher in the basic data framework. For two data categories with significantly different positions, since the richness of data categories ranked lower is much lower than that of data categories ranked higher, even if the confidence levels of samples from data categories ranked lower are higher, the number of samples is too small, potentially limiting the applicability of these data categories. Therefore, it is not necessary to adjust the positions of these two data categories. Based on this, grouping data categories according to the second weight allows for adjustments to the positions of data categories only within the comparison groups, significantly improving the speed of adjusting the basic data framework and reducing the difficulty of method execution.

[0028] In S52, the specific division method of the comparison group is as follows: the difference between the richness of the first data category and the richness of the last data category in a comparison group does not exceed a preset threshold, or the difference between the second weight of the first data category and the second weight of the last data category does not exceed a preset threshold value.

[0029] S53. In each comparison group, for two adjacent data categories, if the benchmark value of the latter data category is greater than the second weight of the former data category, the order of the two data categories is adjusted. Because the data categories in the basic data framework are sorted according to richness, the adjustment is mainly based on the confidence level to correct the sorting method. Therefore, the data category with the larger benchmark value needs to be moved to a higher position.

[0030] S54. Repeat S53 to reorder the comparison group to obtain an ordered group. This is done by adjusting the group using bubble sort, which is highly stable.

[0031] S55. Reassemble all ordered groups into a standard data frame. Specifically, this means assembling all ordered groups into a standard data frame according to their positions in the basic data frame.

[0032] The standard data framework takes into account both the richness and confidence of data categories. Therefore, it can better prioritize data categories with wide application and high reliability, and thus quickly extract the widely applied and reliable parts when calling standard data later.

[0033] The specific methods for standardizing geological data according to the standard data framework are as follows: S6 includes S61 to S63.

[0034] S61. Extract data point values ​​from geological data and determine the data category corresponding to the data point values.

[0035] S62. Input the data point values ​​into the standard data frame according to the correspondence between data point values ​​and data categories.

[0036] S63. Assign null values ​​to the parts of the standard data frame where no data point values ​​have been entered, thus obtaining the standard data. It should be noted that assigning null values ​​does not mean assigning 0 values, but rather assigning meaningless values ​​as placeholders; null values ​​can be set to null. When multiple consecutive data categories in the labeled data are all assigned null values, the labeled data can be collapsed to reduce the space occupied by the standard data.

[0037] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for standardizing multi-source geological data based on big data, characterized in that: The method includes the following steps: S1. Collect multiple sample data; In S1, assign source labels to the sample data when collecting sample data. S2. Determine the confidence level of the sample data based on its source; in S2, determine the confidence level of the sample data according to the source label corresponding to the sample data; in S2, assign a first weight to the confidence level while determining the confidence level of the sample data. S3. Cluster the sample data and determine the richness of sample data in different data categories based on the clustering results; In S3, while determining the richness of sample data in different data categories, assign a second weight to the richness, and the second weight is greater than or less than the first weight. S4. Sort the data categories of the sample data according to their richness to obtain the basic data framework; In S4, all data categories are sorted in descending order of richness to obtain the basic data framework; S5. The basic data framework is modified based on the confidence level to obtain the standard data framework; the specific methods of S5 include: S51. Determine the mode of the first weight of the sample data in each data category as the comparison benchmark; S52. In the basic data framework, each data category is grouped according to the second weight of different data categories to obtain several comparison groups; in S52, the difference between the richness of the first data category and the richness of the last data category in a comparison group does not exceed a preset threshold. S53. In each comparison group, for two adjacent data categories, if the benchmark value of the latter data category is greater than the second weight of the former data category, the order of the two data categories shall be adjusted. S54. Repeat S53 to reorder the comparison group to obtain an ordered group; S55. Reassemble all ordered groups into a standard data frame; S6. Standardize the geological data to be processed using the standard data framework to obtain standard data.

2. The method for standardizing multi-source geological data based on big data as described in claim 1, characterized in that, The specific methods of S6 include: S61. Extract data point values ​​from geological data and determine the data category corresponding to the data point values; S62. Input the data point values ​​into the standard data frame according to the correspondence between data point values ​​and data categories; S63. Assign null values ​​to the parts of the standard data frame where no data points have been entered, and obtain the standard data.

Citation Information

Patent Citations

  • Standardization and aggregation system with a plurality of data sources

    CN103413224A

  • Data quality assessment method and device and storage medium

    CN110098961A

  • Method and system for selecting public data sources

    US20160070706A1