A method and system for analyzing unstructured surveying and mapping report data
By pre-analyzing, classifying, and structuring unstructured surveying and mapping report data, and utilizing mapping relationships and data templates, the problem of unstructured data being difficult to utilize was solved, enabling automated data extraction and sharing, and improving the quality of surveying and mapping data organization and sharing services.
Patent Information
- Application Number
- CN202210994247.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-08-18
AI Technical Summary
Unstructured historical surveying reports have diverse data formats and inconsistent standards, making them difficult to standardize. This leads to difficulties in storage, retrieval, and utilization, resulting in resource waste and human error, as well as hindering effective organization and sharing.
By pre-analyzing classification, data parsing, and structured transformation processing, mapping relationships and unstructured data template elements are used to parse unstructured mapping data, identify key information areas, and convert them into structured data.
It has enabled the maximum extraction and automated processing of unstructured surveying and mapping report data, improved the data extraction and organization capabilities of the data sharing resource pool, and enhanced the quality of shared services for surveying and mapping data products.
Smart Images

Figure CN115495544B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of surveying and mapping geographic information technology, specifically relating to a method and system for parsing unstructured surveying and mapping report data. Background Technology
[0002] In recent years, in order to optimize the business environment and accelerate urban development and construction, and with the advent of the big data wave, the technology for processing massive amounts of data has become increasingly mature, data storage costs have decreased, and the application of data analysis has gradually shifted towards unstructured data.
[0003] However, the applicant discovered that under the new circumstances, national and provincial surveying and mapping authorities have successively put forward higher requirements for unified surveying and mapping and the sharing of results for the surveying and mapping industry. With the integration of surveying and mapping operations in various regions and the establishment of resource pools for sharing surveying and mapping results, various units have stored a large amount of unstructured historical surveying and mapping report data containing useful information, but this data cannot be fully and effectively organized and utilized. This is because unstructured historical surveying and mapping report data not only has diverse formats and standards, but also, technically, unstructured information is more difficult to standardize than structured information. Therefore, the storage, retrieval, publication, and utilization of unstructured data require more intelligent IT technologies, such as massive storage, intelligent retrieval, knowledge mining, content protection, and value-added development and utilization of information. Compared with structured data, unstructured data is massive in quantity, generated rapidly, lacks regularity, has low value density, and, coupled with the lack of effective technical means for processing and analysis, is often discarded or ignored. To extract this useful information, various units often need to consume a significant amount of human and material resources, resulting in resource waste and increasing the risk of human error. This also hinders the long-term, stable extraction, storage, and sharing of information. For example, unstructured historical surveying reports contain a large amount of key information such as area and ownership survey data, but their data structure is irregular or incomplete, lacking a predefined data model. To facilitate the extraction of key information, it is necessary to analyze and structure the key information in a large number of unstructured historical surveying reports. However, due to the complex organization, limited labeling, and poor logic of unstructured data, historical data querying and statistical analysis based on document documents are difficult to achieve. Surveying units often face problems such as scattered historical data storage, inconsistent formats, ineffective programmatic content, and significant manual intervention. Therefore, it is urgent to explore and research a method for analyzing unstructured historical surveying report data. Summary of the Invention
[0004] In order to overcome the shortcomings of the prior art, the purpose of this invention is to provide a method for parsing unstructured surveying report data, and a system based on this method for parsing unstructured surveying report data.
[0005] To solve the above problems, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention provides a method for parsing unstructured surveying report data, comprising:
[0007] S1. Pre-analysis and classification processing: Pre-analysis and classification of unstructured mapping data;
[0008] S2. Data parsing and processing: Based on the mapping relationship, the pre-parsed classification data is parsed to obtain intermediate data and binary format original files;
[0009] S3. Structured Transformation Processing: Using the intermediate data obtained from data parsing and the original binary format file as the data source, the corresponding structured table template and mapping relationship are called to transform and output the structured surveying data.
[0010] Furthermore, the method of the present invention also includes establishing a mapping relationship before the pre-analysis classification process. Specifically, the establishment of the mapping relationship involves extracting key information from the shared resource pool of various surveying and mapping services and establishing structural and semantic mappings.
[0011] Furthermore, the establishment of structural and semantic mappings specifically involves: using historical surveying report information mapping technology developed based on Grok syntax to match and reorganize the extracted non-structured, discontinuous, and discrete unit key information to obtain structural and semantic mapping relationships.
[0012] Furthermore, the pre-parsing classification process includes:
[0013] Obtain raw, unstructured mapping data;
[0014] The original unstructured surveying and mapping data was analyzed and pre-classified according to the surveying and mapping report business type.
[0015] Furthermore, the analysis of the original unstructured surveying and mapping data, and the parsing and pre-classification according to the surveying and mapping report business type, specifically involves: selecting the corresponding unstructured data template element according to the surveying and mapping report business type, comparing the original unstructured surveying and mapping data according to the unstructured data template element, locking the key information parsing area, and pre-classifying the data in the non-locked area.
[0016] Furthermore, the step of comparing the original unstructured mapping data with the unstructured data template element to lock the information parsing region specifically involves: using a template matching mechanism based on metadata to compare the differences between the original unstructured mapping data to obtain the key information parsing region of the unstructured mapping data.
[0017] Furthermore, the data parsing specifically includes:
[0018] Select the corresponding mapping relationship from the parsing library according to the classification rules;
[0019] During the parsing process, the data parsing is dynamically triggered from the selected mapping relationships based on the classification data obtained from the pre-parsed classification.
[0020] After parsing, intermediate JSON data and the original binary file are generated.
[0021] Furthermore, after the structured transformation process and the output of the structured mapping data are converted, redundancy analysis is performed on the converted output structured mapping data based on independent template elements to ensure the correctness of the structured mapping data.
[0022] Furthermore, before the pre-parsing and classification process, an unstructured data template element is established. This unstructured data template element does not contain the original structured data from which information is extracted. It is used to identify and process content that deviates from the template during the data parsing process.
[0023] Secondly, the present invention also provides a system based on the above-mentioned unstructured surveying report data parsing method, comprising:
[0024] The pre-analysis classification module is used to pre-analyze and classify unstructured mapping data;
[0025] The data parsing and processing module is used to parse the pre-parsed classification data according to the mapping relationship, and obtain intermediate data and binary format original files;
[0026] Additionally, a structured conversion processing module is used to take the intermediate data obtained from data parsing and the original binary format files as data sources, call the corresponding structured table templates and mapping relationships, and convert and output structured surveying data.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0028] This invention utilizes mapping relationships to analyze key information parsing regions in unstructured historical surveying and mapping data, maximizing the extraction of effective unstructured surveying and mapping data information within these regions. This transforms unstructured surveying and mapping report data into structured data, effectively solving the problem of consuming significant manual extraction due to the fragmented and unusable nature of unstructured surveying and mapping report data. It significantly improves the automation capabilities of the entire data sharing resource pool for data extraction and organization, enhances the internal information processing level of organizations, and ultimately improves the quality of surveying and mapping data product sharing services. Attached Figure Description
[0029] Figure 1This is a flowchart illustrating the unstructured mapping report data parsing method described in this invention;
[0030] Figure 2 This is a schematic diagram of the unstructured surveying report data parsing system described in this invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0032] like Figure 1 As shown, the unstructured mapping report data parsing method of the present invention includes:
[0033] Step S1. Pre-analysis and classification processing: Pre-analysis and classification of unstructured mapping data. Specifically, this includes:
[0034] S1-1. Obtain raw unstructured mapping data, for example: obtain a piece of raw unstructured mapping data, and input it into the parsing system as parameters according to the previous classification rules; where the classification rules are the data type rules pre-defined in the unstructured data resource pool, which classify the data according to the type, scope, field, etc., for more accurate data matching.
[0035] S1-2. Analyze the original unstructured surveying and mapping data, and perform parsing and pre-classification according to the surveying and mapping report business type; specifically: select the corresponding unstructured data template element according to the surveying and mapping report business type, and compare the original unstructured surveying and mapping data with the unstructured data template element to lock the key information parsing area, and perform pre-classification processing on the data in the non-locked area; this pre-classification processing refers to classifying the data type according to the existing conditions and limiting the classification conditions. Here, it means inputting the data in the defined non-locked area and comparing it to determine which category it belongs to, and then processing it separately. The process involves comparing the original unstructured mapping data with the unstructured data template element (i.e., matching and comparing the unstructured data template element with the original unstructured mapping data, and determining whether the information to be extracted is based on comparison principles such as matching rate and similarity) to lock down the information parsing region. Specifically, this involves using a template matching mechanism based on metadata to compare the differences between the original unstructured mapping data to obtain the locked key information parsing region of the unstructured mapping data. The unstructured data template element is the smallest standardized information unit of this type of unstructured data, formulated according to the information composition standards of this type of data, and used to match and extract basic information units from the unstandardized data for convenient subsequent processing.
[0036] Therefore, before pre-parsing and classification, it is necessary to establish unstructured data template elements. These unstructured data template elements do not contain the original structured data from which information is extracted (because the unstructured data template elements are data rules extracted from the surveying and mapping data resource pool, and the data to be parsed cannot be used as input data for establishing the original template elements to avoid contradictions in the data matching process). They are used to identify and process content that deviates from the template during the data parsing process.
[0037] Step S2. Data Parsing and Processing: Based on the mapping relationship, the pre-parsed classification data is parsed to obtain intermediate data and the original binary format file. Specifically, this includes:
[0038] S2-1. Select the corresponding mapping relationship from the parsing library according to the classification rules;
[0039] S2-2. During the parsing process, the mapping relationships are dynamically triggered from the selected mapping relationships based on the classification data obtained from the pre-parsed classification; specifically, the mapping relationships of the pre-parsed classification data are automatically identified, and the index relationship is used to dynamically filter the candidate mapping relationships, select the mapping relationship that meets the requirements, and trigger it to realize data parsing.
[0040] S2-3. After parsing, generate intermediate JSON data and the original binary file.
[0041] If the above process fails to obtain the parsed content, or if the data extraction and parsing results within the key information parsing area are empty, a parsing content location log will be output for the operator to query, and if necessary, the above content will be added to the template library; if the above process successfully obtains the parsed content, the intermediate JSON data will be processed and stored in the database.
[0042] Step S3. Structured Transformation Processing: Using the intermediate data obtained from data parsing and the original binary format file as the data source, the corresponding structured table template and mapping relationship are called to transform and output structured surveying and mapping data. Specifically, the intermediate data and binary format file data do not conform to the final standard data format. After extracting key information, the extracted information is reorganized and transformed through automatic matching and mapping with the structured table template. Finally, according to the definition of the resource pool, it is output as structured surveying and mapping data in a standard format.
[0043] This invention utilizes mapping relationships (including intermediate mapping relationships) to analyze the key information parsing areas of unstructured historical surveying and mapping data. This maximizes the extraction of effective unstructured surveying and mapping data within the locked key information parsing areas, transforming unstructured surveying and mapping report data into structured surveying and mapping report data. This effectively solves the problem of requiring extensive manual extraction due to the fragmented and unusable nature of unstructured surveying and mapping report data. It significantly improves the automation capabilities of the entire data sharing resource pool for data extraction and organization, enhances the internal information processing level of the organization, and ultimately improves the quality of data product sharing services for surveying and mapping units.
[0044] In one possible implementation, the method of the present invention further includes establishing a mapping relationship before the pre-analysis classification process. The establishment of the mapping relationship specifically involves: extracting key information from the shared resource pool of various surveying and mapping business results, establishing structural mapping and semantic mapping, mainly organizing various surveying and mapping business results resources, establishing a resource pool of key information, and defining the logic and method of structural and semantic mapping for the information in the resource pool for subsequent information matching.
[0045] More specifically, by utilizing historical surveying and mapping report information mapping technology developed based on Grok syntax, the extracted unstructured, discontinuous, and discrete key information is matched and recombined to obtain structural mapping relationships and semantic mapping relationships. Among them, historical surveying and mapping reports are a large category of surveying and mapping business results resources. Since this type of data is unstructured, there is a set of rules for extracting and expressing key information for this type of data. After the core information is extracted, it is matched, compared, and recombined with standardized information in the resource pool to form a mapping relationship.
[0046] This invention establishes mapping relationships (structural mapping and semantic mapping) by extracting key information from the shared resource pool of various surveying and mapping services. This not only makes the method applicable to the parsing of unstructured surveying and mapping report data of various surveying and mapping services, thus having strong applicability, but also improves the usability of historical surveying and mapping data.
[0047] In one possible implementation scheme, after the structured transformation process and the output of structured mapping data, redundancy analysis is performed on the output structured mapping data based on independent template elements. Specifically, through redundancy data verification, redundant information is automatically compared with the original unstructured mapping data source, and a location log is output for the operator to verify, further ensuring the correctness of the structured mapping data. The independent template element refers to the smallest standard data template unit defined according to the data definition and data type, which is used as a standard to match unstructured data.
[0048] The following examples further illustrate the unstructured mapping report data parsing method of the present invention.
[0049] Example:
[0050] The machine environment for the unstructured surveying report data parsing method described in this embodiment of the invention is: Windows operating system, .NET Framework, and Oracle database software; its processing flow specifically includes:
[0051] (1) Received a raw, unstructured mapping report data, raw_event;
[0052] (2) Extract key information key(token+appname (if appname does not exist, use hostname), Token refers to key information identification token, which is a kind of authentication token, appname is the application name, hostname is the host server name), find the corresponding parsing rule EventParser from ParserCache (ParserCache is a cache pool that stores parsing rules, which facilitates the quick matching and extraction of rules));
[0053] If the corresponding parsing rule is found and processed successfully, a structured event is returned.
[0054] If the corresponding parsing rule EventParser is not found or the parsing rule EventParser fails to process, it will be processed in ParserContainer. That is, if the specified parsing cache pool cannot be found, it will be put into the original large container of the parser for generalized processing.
[0055] (3) Check if there is a corresponding user Custom configuration in token+appname (if appname does not exist, use hostname);
[0056] If so, the Parser generated according to this configuration will be used for processing. If the processing is successful, an event_parser and structured data structured event will be generated, and the storage cache will be updated. If it is unsuccessful, the DefaultParser will be used (the DefaultParser usually has default processing rules, which are general. If a specific rule cannot be found, the general rule will be used for processing) to retain only the raw_message (unprocessed raw data information) and not extract any fields.
[0057] If not, the Common configuration will be used for processing. Various types of Parser will be used to try processing in turn. If the processing is successful, an event_parser and structured data structured event will be generated, and the storage cache will be updated. If it fails, the DefaultParser will be used to keep only the raw_message and will not extract any fields.
[0058] (4) Returns structured data (structured event).
[0059] Regular expression parsing (data parsing and processing):
[0060] (1) The matching fields are parsed by configuring regular expressions, supporting named grouping, multi-line regular expressions, and Grok syntax;
[0061] (2) KeyValue parsing (i.e., key value parsing method, used to extract key information from relatively regular logs) is suitable for logs containing field names and relatively clear delimiters. Configure KV to extract fields by delimiter and delimiter between KV.
[0062] (3) KeyValue regular expression parsing is suitable for logs with uncertain delimiters and discontinuous KV pairs. It extracts fields by configuring regular expressions for Key, Value, and delimiter.
[0063] (4) Json parsing is applicable to Json log format, and the extracted field structure is consistent with the structure defined in Json;
[0064] (5) XML parsing is applicable to XML log parsing, and the extracted field structure is consistent with the structure defined in the XML;
[0065] (6) CSV parsing is suitable for logs with fixed column order and fixed delimiter. Configure the delimiter and column name to parse the fields;
[0066] (7) Structure parsing is applicable to logs written in fixed byte lengths, and configures the parsing of byte format.
[0067] like Figure 2This invention also provides a system based on the above-mentioned unstructured surveying report data parsing method, including a pre-parsing classification module 100, a data parsing processing module 200, and a structured conversion processing module 300. The pre-parsing classification module 100 is mainly used to pre-parse and classify unstructured surveying data; the data parsing processing module 200 is mainly used to parse the classified data obtained from the pre-parsing classification according to the mapping relationship, obtaining intermediate data and binary format raw files; the structured conversion processing module 300 is mainly used to use the intermediate data and binary format raw files obtained from the data parsing as data sources, call the corresponding structured table template and mapping relationship, and convert and output structured surveying data; before this, a structured table template (i.e., a standard template for structured data) is established, which specifically represents the information of a database structured table, mainly including table name, field name, data type, field length, value constraints, primary and foreign key constraints, etc.
[0068] The system described in this invention is based on the above-mentioned unstructured surveying report data parsing method. Please refer to the above description for the various schemes and expected technical effects, which will not be repeated here.
[0069] Furthermore, the unstructured surveying report data parsing method and system described in this invention also includes establishing a data import / output system and an automatic associated data storage system. The established data import / output system enables memory management and automated allocation of multiple data throughout the entire data parsing process, ensuring memory management during software operation and meeting the requirements for automated multi-data operations. The automatic associated data storage system automatically processes the structured results for storage, reducing manual workload and enabling rapid storage management of extracted data. Moreover, this invention establishes a unified data conversion interface compatible with versions of historical surveying data from various periods. Due to slight version differences in the data to be processed from different periods, only through unified preprocessing and conversion can a unified standard be invoked for processing.
[0070] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.
Claims
1. A method for parsing unstructured surveying report data, characterized in that, include: S1. Pre-analysis and classification processing: Pre-analysis and classification of unstructured mapping data, including: Obtain raw, unstructured mapping data; The original unstructured surveying and mapping data is analyzed, and pre-analyzed and classified according to the surveying and mapping report business type. Specifically, the corresponding unstructured data template element is selected according to the surveying and mapping report business type, and the original unstructured surveying and mapping data is compared with the unstructured data template element to lock the key information parsing area, and the data in the unlocked area is classified and processed. The comparison of the original unstructured surveying and mapping data with the unstructured data template element to lock the key information parsing area is specifically: based on the template matching mechanism of metadata, the differences between the positive and negative comparisons of the original unstructured surveying and mapping data are obtained to lock the key information parsing area of the unstructured surveying and mapping data. S2. Data parsing and processing: Based on the mapping relationship, the pre-parsed classification data is parsed to obtain intermediate data and binary format original files; S3. Structured Transformation Processing: Using the intermediate data obtained from data parsing and the original binary format file as the data source, the corresponding structured table template and mapping relationship are called to transform and output the structured surveying data; The method also includes establishing mapping relationships before pre-analysis and classification processing. Specifically, the establishment of mapping relationships involves: extracting key information from the shared resource pool of various surveying and mapping business results, and establishing structural mapping and semantic mapping. Specifically, the establishment of structural mapping and semantic mapping involves: matching and recombining the extracted unstructured, discontinuous, and discrete key information using historical surveying and mapping report information mapping technology developed based on Grok syntax, to obtain structural mapping relationships and semantic mapping relationships. After the structured transformation process is completed and the output of structured mapping data is converted, redundancy analysis is performed on the converted structured mapping data based on independent template elements to ensure the correctness of the structured mapping data. Specifically, through redundancy data verification, redundant information is automatically compared with the original unstructured mapping data source, and a location log is output for the operator to verify.
2. The method for parsing unstructured surveying report data according to claim 1, characterized in that, The data parsing and processing specifically includes: Select the corresponding mapping relationship from the parsing library according to the classification rules; During the parsing process, the data parsing is dynamically triggered from the selected mapping relationships based on the classification data obtained from the pre-parsed classification. After parsing, intermediate JSON data and the original binary file are generated.
3. The method for parsing unstructured surveying report data according to claim 1 or 2, characterized in that, Before pre-parsing and classification, an unstructured data template element is established. This unstructured data template element does not contain the original structured data from which information is extracted. It is used to identify and process content that deviates from the template during the data parsing process.
4. A system based on the unstructured surveying report data parsing method according to any one of claims 1-3, characterized in that, include: The pre-analysis classification module (100) is used to pre-analyze and classify unstructured mapping data; The data parsing and processing module (200) is used to parse the classification data obtained from the pre-parsing classification according to the mapping relationship, and obtain intermediate data and binary format original files; In addition, the structured conversion processing module (300) is used to take the intermediate data obtained from data parsing and the binary format original file as the data source, call the corresponding structured table template and mapping relationship, and convert and output the structured surveying data.