A method and system for search enhancement based on unstructured data
By constructing a vector knowledge base and implementing retrieval enhancement strategies, the problem of insufficient semantic association in unstructured data retrieval in the natural resources industry has been solved, enabling efficient and accurate data retrieval and review, and improving the level of intelligent approval.
Patent Information
- Application Number
- CN202511525301.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Existing technologies for unstructured data retrieval in the natural resources industry cannot deeply mine the semantic relationships of data, resulting in poor adaptability of retrieval results to business needs, leading to low review efficiency, low automation, and an inability to meet the requirements of intelligent approval.
A vector knowledge base is built based on unstructured natural resource data. Query problems are handled through semantic understanding and vector embedding. Retrieval enhancement strategies are adopted for accurate retrieval, and content completeness, legality, consistency, and standardization are verified.
It enables efficient and accurate retrieval of unstructured data, improves the matching degree between retrieval results and business needs, ensures the completeness, legality and consistency of retrieval results, solves the problem of poor adaptability between retrieval results and business needs, and improves the efficiency and automation of review.
Smart Images

Figure CN120994813B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a retrieval enhancement method and system based on unstructured data. Background Technology
[0002] In the business approval process of the natural resources industry, a large amount of unstructured data is generated, such as approval materials in the form of various documents and images. Traditional approval processes rely on manual review of this unstructured data, resulting in low review efficiency and a low degree of automation in the pre-review of documents, seriously affecting the level of intelligence in natural resources business approval and the effectiveness of resource service guarantees. To address this issue, relevant technical personnel have proposed a data retrieval method for unstructured natural resources data. This method attempts to assist the approval process by processing and retrieving unstructured data.
[0003] However, existing data retrieval methods have obvious shortcomings in practical applications. They can only achieve simple data search, cannot deeply explore the semantic relationships of data, and the retrieval results are poorly adapted to the approval business of the natural resources industry. They are difficult to meet the needs of accurate data extraction, intelligent verification and efficient application in the approval process, and cannot effectively support the intelligent development of natural resources business approval. Summary of the Invention
[0004] This invention provides a retrieval enhancement method and system based on unstructured data, which can achieve efficient and accurate retrieval enhancement of unstructured natural resource report data.
[0005] In a first aspect, the present invention provides a retrieval enhancement method based on unstructured data, comprising:
[0006] Document classification is performed based on unstructured natural resource data to obtain information on various types of documents. Based on the logical relationships between the semantic units in each type of document information, a vector knowledge base is constructed.
[0007] Based on the vector knowledge base, the query question is semantically understood and vector embedded to obtain the semantic vector corresponding to the query question. Based on the semantic vector, the query question is classified using a question classification model to obtain the category information of the query question.
[0008] Based on the category information and the vector knowledge base, a retrieval strategy based on retrieval enhancement is used to perform a retrieval operation and obtain preliminary retrieval results;
[0009] Based on the preliminary search results, content integrity verification, content legality verification, content consistency verification, and submission document standardization verification are performed to obtain enhanced search results.
[0010] In a second aspect, the present invention also provides a retrieval enhancement system based on unstructured data, applied to the retrieval enhancement method based on unstructured data as described in any of the first aspects, wherein the retrieval enhancement system based on unstructured data includes:
[0011] The knowledge base construction module is used to classify documents based on unstructured natural resource data, obtain information on various types of documents, and construct a vector knowledge base based on the logical relationships between the semantic units in each type of document information.
[0012] The question classification module is used to perform semantic understanding and vector embedding processing on the query question based on the vector knowledge base to obtain the semantic vector corresponding to the query question, and to classify the query question based on the semantic vector using a question classification model to obtain the category information of the query question;
[0013] The retrieval operation module is used to perform retrieval operations based on the category information and the vector knowledge base using a retrieval enhancement-based retrieval strategy to obtain preliminary retrieval results;
[0014] The retrieval verification module is used to perform content integrity verification, content legality verification, content consistency verification, and submission document standardization verification based on the preliminary retrieval results, thereby obtaining enhanced retrieval results.
[0015] Thirdly, the present invention also provides an electronic device, comprising: a memory for storing computer software programs; and a processor for reading and executing the computer software programs, thereby realizing the retrieval enhancement method based on unstructured data as described above.
[0016] Fourthly, the present invention also provides a non-transitory computer-readable storage medium storing a computer software program, which, when executed by a processor, implements the retrieval enhancement method based on unstructured data as described above.
[0017] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the retrieval enhancement method based on unstructured data as described above.
[0018] The retrieval enhancement method based on unstructured data provided in this invention constructs a vector knowledge base based on the logical relationships of various semantic units in various types of document information. This solves the problem of not being able to deeply mine the semantic associations of data. Based on the vector knowledge base, semantic understanding and vector embedding are performed on the query question to obtain a semantic vector, clarifying the semantic expression of the query requirement. Based on the semantic vector, query question category information is obtained through question classification, achieving accurate positioning of the query requirement and improving the retrieval targeting. Based on the category information and the vector knowledge base, a retrieval enhancement strategy is adopted to obtain preliminary retrieval results, improving the matching degree between the retrieval results and the query requirement. Based on the preliminary retrieval results, multi-dimensional verification is carried out to obtain retrieval enhancement results, ensuring the completeness, legality, consistency and standardization of the retrieval results. This solves the shortcomings of poor adaptability of retrieval results to natural resource industry approval business, effectively solving the problems of low efficiency and low degree of automation in manual review, and realizing efficient and accurate retrieval enhancement of unstructured natural resource report data. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the retrieval enhancement method based on unstructured data provided by the present invention;
[0020] Figure 2 This is a schematic diagram of the structure of the retrieval enhancement system based on unstructured data provided by the present invention;
[0021] Figure 3 A schematic diagram of an embodiment of the electronic device provided in this invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0024] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0025] Optional, see below Figure 1 , Figure 1 This is a flowchart illustrating the retrieval enhancement method based on unstructured data provided by the present invention. In this embodiment, the execution entity of the retrieval enhancement method based on unstructured data is a data retrieval system. Therefore, the retrieval enhancement method based on unstructured data includes:
[0026] Step 10: Classify documents based on unstructured natural resource data to obtain information on various types of documents, and construct a vector knowledge base based on the logical relationships of the semantic units in each type of document information.
[0027] Optionally, the data retrieval system receives unstructured natural resource data, which may include text, images, audio, etc., covering geological exploration reports, ecological environment monitoring records, forestry resource survey data, etc. It uses text classification algorithms from Natural Language Processing (NLP) technology (such as Support Vector Machines, deep learning models like BERT) to classify the unstructured text data. For non-text data such as images and audio, it is first converted into text using technologies such as Optical Character Recognition (OCR) and speech-to-text before classification. After classification, various document types are obtained, such as geological documents, ecological documents, and forestry documents.
[0028] Furthermore, the data retrieval system performs semantic analysis on each type of document, breaking it down into multiple semantic units. These units can be sentences, paragraphs, or text blocks with independent semantic meaning. By analyzing the logical relationships between these semantic units, such as causal, parallel, and subordinate relationships, word vector technologies (such as Word2Vec and GloVe) are used to transform the semantic units into vector representations. The data retrieval system then organizes these vectors according to logical relationships to construct a vector knowledge base.
[0029] In one embodiment, the data retrieval system receives a batch of unstructured natural resource data, including 10 geological exploration reports, 15 ecological environment monitoring records, and 8 forestry resource survey data. The data retrieval system uses the BERT model to classify these documents, categorizing the 10 geological exploration reports as geological documents, the 15 ecological environment monitoring records as ecological documents, and the 8 forestry resource survey data as forestry documents.
[0030] Taking a geological exploration report as an example, the sentences "A large amount of iron ore reserves were discovered in this area" and "The formation of iron ore is related to paleogeological tectonic movements" are used as semantic units. Semantic analysis determines that these two semantic units have a causal relationship. Then, Word2Vec technology is used to convert these two sentences into vector representations, which are stored in a vector knowledge base according to the causal logic. Similarly, the semantic units of other documents are processed to ultimately construct a complete vector knowledge base.
[0031] Step 20: Based on the vector knowledge base, perform semantic understanding and vector embedding processing on the query question to obtain the semantic vector corresponding to the query question, and classify the query question using a question classification model based on the semantic vector to obtain the category information of the query question.
[0032] Furthermore, when a user inputs a query through the data retrieval system, the system first performs semantic understanding of the query based on a vector knowledge base. Utilizing semantic analysis techniques from natural language processing, the query is broken down into words and phrases, and their semantics are analyzed. Further, the data retrieval system uses vector embedding techniques (such as SentenceTransformer) to transform the query into a corresponding semantic vector, making the query computable in the vector space.
[0033] Furthermore, the data retrieval system employs a pre-trained question classification model (such as the deep learning-based TextCNN model) to classify query questions using semantic vectors as input. Based on existing category information in the vector knowledge base, the question classification model determines which specific category the query question belongs to, such as geology, ecology, or forestry, thus obtaining the category information of the query question.
[0034] In one embodiment, a user inputs the query "What are the air quality indicators in the ecological environment monitoring data of a certain region?" into the data retrieval system. The data retrieval system first uses SentenceTransformer to convert the query into a semantic vector, for example, obtaining a 1024-dimensional vector representation. Then, this semantic vector is input into a pre-trained TextCNN question classification model. Based on existing category information such as ecology, geology, and forestry in the vector knowledge base, the model determines that the query belongs to the ecology category, thus obtaining the category information of the query as ecology.
[0035] Step 30: Based on category information and vector knowledge base, a retrieval-enhanced retrieval strategy is used to perform a retrieval operation to obtain preliminary retrieval results.
[0036] Furthermore, after obtaining the category information of the query question, the data retrieval system, in conjunction with a vector knowledge base, employs a retrieval enhancement strategy. This enhancement strategy, as described in this embodiment, primarily includes two aspects: first, narrowing the search scope using category information, searching only document vectors belonging to the corresponding category in the vector knowledge base; and second, calculating the similarity (e.g., cosine similarity) between the query question's semantic vector and document vectors in the vector knowledge base, and then searching in descending order of similarity. The system retrieves document vectors from the vector knowledge base that have a high similarity to the query question's semantic vector, identifies them as specific documents, and obtains preliminary search results.
[0037] Continuing with the above embodiment, for the query question "What are the air quality indicators in the ecological environment monitoring data of a certain region?" in step 20, its category information is ecology. Only the vectors corresponding to ecology-related documents are retrieved from the vector knowledge base. By calculating the cosine similarity between the semantic vector of the query question and the vector of the ecology-related documents, it is assumed that three ecological environment monitoring record document vectors with high similarity are retrieved. The system returns these three documents as preliminary search results; these three documents may contain content related to air quality indicators.
[0038] Step 40: Based on the preliminary search results, perform content integrity verification, content legality verification, content consistency verification, and submission document standardization verification to obtain enhanced search results.
[0039] Furthermore, the data retrieval system performs content integrity verification, content legality verification, content consistency verification, and submission document standardization verification based on the preliminary retrieval results, obtaining four verification results. If the preliminary retrieval results are confirmed to be without problems after passing the four verification results, enhanced retrieval results are obtained, as detailed in steps 401 to 428.
[0040] Content completeness verification: The data retrieval system checks whether the documents in the initial search results contain all the information required to answer the query. By analyzing the document's content structure and keywords, it determines whether it covers all the key elements of the query. If a document is missing key information, the content is considered incomplete.
[0041] Content legality verification: Based on relevant laws, regulations and industry standards, check whether the content of documents in the preliminary search results is legal and compliant, such as whether the documents contain prohibited information or infringe on intellectual property rights.
[0042] Content consistency verification: Check the consistency of document content in the preliminary search results, determine whether the information in different parts of the document is contradictory, and whether the document is consistent with other relevant information in the vector knowledge base.
[0043] Document submission compliance check: Check whether the document format in the preliminary search results meets the requirements, such as whether the document layout, font, chart format, etc. are standardized, and whether the file naming follows the rules.
[0044] This invention constructs a vector knowledge base based on the logical relationships between semantic units in various types of document information, solving the problem of not being able to deeply mine the semantic relationships of data. Based on the vector knowledge base, semantic understanding and vector embedding are performed on the query question to obtain a semantic vector, clarifying the semantic expression of the query requirement. Based on the semantic vector, query question category information is obtained through question classification, achieving accurate positioning of the query requirement and improving the retrieval targeting. Based on the category information and the vector knowledge base, a retrieval enhancement strategy is adopted to obtain preliminary retrieval results, improving the matching degree between the retrieval results and the query requirement. Based on the preliminary retrieval results, multi-dimensional verification is carried out to obtain retrieval enhancement results, ensuring the completeness, legality, consistency and standardization of the retrieval results, and realizing efficient and accurate retrieval enhancement of unstructured natural resource report data.
[0045] In one embodiment, steps 401 to 405 include:
[0046] Step 401: Extract all structured data fields from the preliminary search results to obtain the initial structured data fields, and obtain the list of pre-set necessary information fields in the natural resources industry approval business based on the initial structured data fields to obtain the business necessary information field list.
[0047] Optionally, the data retrieval system first processes the preliminary retrieval results obtained in step 30. Through data parsing technology, it extracts all structured data fields from the documents of the preliminary retrieval results. These fields store data in a fixed format and structure, such as column headers in tables and field names in databases, thereby obtaining an initial set of structured data fields.
[0048] Furthermore, the data retrieval system searches for the corresponding list of necessary information fields in the preset information database of natural resource industry approval business based on the business areas involved in the initial structured data fields. The list predefines the information fields that must be available to complete the approval business, covering key data items in various business scenarios such as natural resource development and land use approval, and finally obtains the list of necessary information fields for the business.
[0049] Continuing with the above embodiment, the initial search results are three documents related to ecological environment monitoring data for a certain region. The data retrieval system uses data parsing technology to extract structured data fields such as "monitoring time," "monitoring location," and "monitoring indicator name" from these three documents, forming an initial set of structured data fields. Since these documents involve ecological environment monitoring business, the system finds a list of necessary information fields corresponding to ecological environment monitoring approval business in the preset information database of natural resources industry approval business. This list includes fields such as "monitoring time," "monitoring location," "monitoring indicator name," "monitoring data value," and "monitoring unit."
[0050] Step 402: Compare the initial structured data fields with the list of business-required information fields one by one to obtain the field comparison results, and identify the missing required information fields in the initial structured data fields based on the field comparison results to obtain the missing field list.
[0051] Furthermore, the data retrieval system compares the initial structured data fields obtained in step 401 with the list of business-essential information fields one by one. This embodiment of the invention employs an exact matching algorithm to check whether each field in the list of business-essential information fields exists in the initial structured data field set. After the comparison is completed, a field comparison result is obtained, which records the matching status of each business-essential information field in the initial structured data fields. Further, based on the field comparison result, the data retrieval system identifies business-essential information fields that do not exist in the initial structured data fields, compiles these fields into a missing field list, and clarifies the key data fields missing in the document.
[0052] Continuing with the above embodiment, the data retrieval system compares the initial structured data fields "Monitoring Time," "Monitoring Location," and "Monitoring Indicator Name" from step 401 with the list of business-required information fields "Monitoring Time," "Monitoring Location," "Monitoring Indicator Name," "Monitoring Data Value," and "Monitoring Unit." It was found that "Monitoring Data Value" and "Monitoring Unit" were missing from the initial structured data fields, so they were compiled into a list of missing fields.
[0053] Step 403: Check whether there is any incorrectly extracted information content in the preliminary search results that corresponds to each field in the missing field list, obtain the incorrectly extracted information verification results, and locate the incorrectly extracted information in the preliminary search results based on the incorrectly extracted information verification results to obtain the location information of the incorrectly extracted information.
[0054] Furthermore, the data retrieval system performs a deep scan of the document content in the preliminary search results for the missing field list. This embodiment of the invention employs text keyword matching, semantic analysis, and other technologies to search for information in the document that is semantically related to the fields in the missing field list. If relevant information is found, it is determined that this information was not correctly extracted during the previous structured data extraction process, and the relevant information is recorded, resulting in a check result for incorrectly extracted information. Further, based on the check result for incorrectly extracted information, the data retrieval system locates the specific positions of these incorrectly extracted information pieces in the document from the preliminary search results, obtaining their page numbers, paragraphs, sentences, and other positional information to obtain location information.
[0055] Continuing with the above embodiment, after obtaining the list of missing fields as "monitoring data value" and "monitoring unit" in step 402, the data retrieval system performs a deep scan of the three preliminary search result documents. Through keyword matching, the statement "PM2.5 concentration is 35 micrograms per cubic meter" was found in a paragraph of one of the documents. "35 micrograms per cubic meter" corresponds to both "monitoring data value" and "monitoring unit," but it had not been correctly extracted previously. The data retrieval system records this situation, obtains the result of the incorrectly extracted information verification, and determines that the information is located on page 3, paragraph 2 of the document, thus obtaining the location information.
[0056] Step 404: Extract the information content corresponding to the location information and supplement it into the initial structured data field to obtain the supplemented structured data field.
[0057] Furthermore, based on the location information of incorrectly extracted information obtained in step 403, the data retrieval system accurately extracts the information content of the corresponding location from the preliminary retrieval result document. Further, the data retrieval system supplements the extracted information content into the initial structured data field set according to the format and requirements of the business-essential information field list, forming a supplemented structured data field containing more complete information.
[0058] Continuing with the above embodiment, based on the location information obtained in step 403 (page 3, paragraph 2 of the document), the data retrieval system extracts the information content "35 micrograms / cubic meter" from that location. Following the format of the business-required information field list, "35" is added to the "monitoring data value" field, and "micrograms / cubic meter" is added to the "monitoring unit" field. These are then added to the initial structured data field set, resulting in supplemented structured data fields that include "monitoring time", "monitoring location", "monitoring indicator name", "35", and "micrograms / cubic meter".
[0059] Step 405: Based on the supplemented structured data fields, compare them again with the list of business-required information fields until it is confirmed that all required information fields are included, and obtain the content integrity verification result.
[0060] Furthermore, the data retrieval system performs a new round of comparison between the supplemented structured data fields obtained in step 404 and the list of business-required information fields. Again, an exact matching algorithm is used to check whether each field in the list of business-required information fields exists in the supplemented structured data field set.
[0061] If any fields are missing, repeat steps 403 and 404 to continue checking for any incorrectly extracted information in the document and supplementing it, then compare again. Repeat this process until all fields in the list of essential business information fields can be found in the supplemented structured data field set. At this point, the document content is confirmed to be complete, and the content integrity check result is "complete." If missing fields still exist after multiple iterations, the content integrity check result is "incomplete."
[0062] Continuing with the above embodiment, the supplemented structured data fields "monitoring time", "monitoring location", "monitoring indicator name", "35" and "micrograms / cubic meter" obtained in step 404 are compared again with the list of business-required information fields "monitoring time", "monitoring location", "monitoring indicator name", "monitoring data value" and "monitoring unit". It is found that all required information fields are included, and the content integrity verification result is "complete".
[0063] This invention, through its embodiments, proceeds from extracting initial structured data fields, to comparing them with a list of necessary business information fields to identify missing fields, then to verifying and supplementing incorrectly extracted information, and finally repeatedly comparing and confirming completeness. This ensures that the search results documents contain all the key information required for approval processes in the natural resources industry, preventing approval processes from being disrupted due to missing data and improving the accuracy of the search results data.
[0064] In one embodiment, steps 406 to 409 include:
[0065] Step 406: Extract the information content corresponding to all structured data in the preliminary search results to obtain the search information content, and obtain the legality and regulatory requirements for various types of information content in the natural resources industry approval business based on the search information content to obtain the information content legality and regulatory list.
[0066] Optionally, the data retrieval system processes the preliminary search results obtained in step 30, extracting corresponding actual information content from all structured data fields. This information content covers various types of data recorded in structured form within the document, resulting in a set of retrieved information content. Further, based on the business domain and data type involved in the retrieved information content, the data retrieval system searches for corresponding legality regulations in the legality norms database for natural resource industry approval processes. This database pre-stores legality regulations for various types of information content in natural resource industry approval processes, including data formats, value ranges, and sensitive information processing rules, ultimately obtaining a list of information content legality regulations.
[0067] Continuing with the above embodiment, the preliminary search results in step 30 are three documents related to ecological environment monitoring data for a certain region. The data retrieval system extracts the corresponding information content from the structured data fields "monitoring time," "monitoring data value," and "monitoring unit" in the documents, such as "2024-01-01," "35," and "micrograms per cubic meter," forming a set of search information content. Since this information involves ecological environment monitoring operations, the system finds the relevant legal requirements for ecological environment monitoring data in the natural resources industry approval business legality norms database, forming a list of information content legality norms. These norms stipulate that "monitoring time" must be in "YYYY-MM-DD" format, "monitoring data value" must be a value greater than 0, and "monitoring unit" must conform to national standard units of measurement, etc.
[0068] Step 407: Match each piece of information in the retrieved information content with the corresponding normative requirements in the list of legality norms for information content to obtain the matching results between the information content and the normative requirements. Based on the matching results, identify the information content in the retrieved information content that does not meet the legality normative requirements to obtain a list of illegal information content.
[0069] Furthermore, the data retrieval system matches each specific piece of information in the retrieved information content obtained in step 406 with the corresponding regulatory requirements in the information content legality specification list. This embodiment of the invention employs precise matching and rule-based judgment algorithms to check whether each piece of information content meets the corresponding legality specification requirements. After matching is completed, a matching result between the information content and the regulatory requirements is obtained, recording the compliance status of each piece of information content with the regulatory requirements. Further, based on the matching results, the data retrieval system identifies information content in the retrieved information content that does not meet the legality specification requirements, compiles these illegal information contents into an illegal information content list, and clarifies the data items in the document that have legality issues.
[0070] Continuing with the above embodiment, for the search information content "2024-01-01", "35", "micrograms / cubic meter" in step 406 and the list of legality standards for information content ("monitoring time" must be in "YYYY-MM-DD" format, "monitoring data value" must be a value greater than 0, and "monitoring unit" must conform to national standard units of measurement), a match is made one by one. It is found that all information content meets the requirements, and the list of illegal information content is empty at this time; if there is a record with "monitoring time" as "January 1, 2024", which does not meet the "YYYY-MM-DD" format requirement, then "January 1, 2024" is added to the list of illegal information content.
[0071] Step 408: Analyze the specific reasons why each item in the list of illegal information content does not meet the requirements of the regulations, obtain the analysis results of the illegality reasons, and based on the analysis results of the illegality reasons, confirm whether the illegal information content belongs to the special circumstances allowed in the approval business of the natural resources industry, and obtain the special circumstances confirmation results.
[0072] Furthermore, the data retrieval system conducts in-depth analysis of each illegal information item in the list of illegal information content. By comparing the information content with the legality requirements, and combining business knowledge and relevant regulations, it determines the specific reasons why the illegal information content does not meet the requirements, and forms an analysis result of the illegality reasons.
[0073] Furthermore, based on the special provisions for approval business in the natural resources industry, the data retrieval system determines whether each of the reasons in the analysis of illegal reasons belongs to a permitted special circumstance. These special circumstances may be due to historical data legacy issues, special business scenario requirements, or other reasons that do not meet the requirements of the regulations but are recognized by the business, and finally obtain the special circumstance confirmation result.
[0074] Continuing with the above example, the list of illegal information contains a record with a "monitoring time" of "January 1, 2024". The data retrieval system analyzes that this is illegal because the format does not conform to the "YYYY-MM-DD" requirement. Next, the system queries the special regulations for approval procedures in the natural resources industry and finds that for some historical data, inconsistent time formats are allowed. Since this data falls under the category of historical data, it is confirmed that this illegal information falls under a special circumstance, and the special circumstance confirmation result is "yes".
[0075] Step 409: Based on the special circumstances, confirm the specific location and reason for the illegal information content, obtain a detailed record of the illegal information, and summarize the detailed record of the illegal information into a content legality verification report to obtain the content legality verification result.
[0076] Furthermore, based on the special case confirmation results obtained in step 408, the data retrieval system records the specific location (such as page number, paragraph, sentence, etc.) of the illegal information content in the preliminary search result document, as well as the reason for its illegality, forming a detailed record of illegal information.
[0077] If any illegal information is found and it does not fall under special circumstances, it is considered a serious legality issue; if it falls under special circumstances, it is noted in the record. Finally, the system summarizes all illegal information records in detail, generates a content legality verification report, and derives the content legality verification result based on the circumstances of the illegal information. If the list of illegal information is empty or all illegal information falls under special circumstances, the content legality verification result is "legal"; if there is any illegal information that does not fall under special circumstances, the content legality verification result is "illegal".
[0078] Continuing with the above embodiment, for the illegal information content recorded as "January 1, 2024" as the "monitoring time," the data retrieval system records its location as the first paragraph on page 2 of a document. The reason for its illegality is that the format does not conform to the "YYYY-MM-DD" requirement, and it is noted that this is a special case. This record is summarized with other illegal information records (if any) to generate a content legality verification report. Since this illegal information is a special case and there is no other seriously illegal information, the content legality verification result is "legal."
[0079] This invention, through its embodiments, proceeds from extracting search information and obtaining legality requirements, to matching and identifying illegal information, then analyzing the causes and judging special cases, and finally recording and summarizing to generate a verification report. This ensures that the information content in the search results meets the legality requirements of natural resource industry approval business, effectively identifies and processes illegal information, and reasonably distinguishes special cases, providing reliable data basis for natural resource industry approval business and avoiding business risks caused by data legality issues.
[0080] In one embodiment, steps 410 to 414 include:
[0081] Step 410: Extract the approval document identification information contained in the preliminary search results to obtain the target approval document identification, and retrieve the structured information of other approval documents associated with the target approval document identification from the natural resources industry approval business database to obtain the associated approval document structured information.
[0082] Optionally, the data retrieval system processes the preliminary retrieval results obtained in step 30. Using data parsing technology, it extracts the approval document identification information contained in the documents of the preliminary retrieval results. The approval document identification information is key information used to uniquely identify an approval document, such as the document number and project name. After extraction, the target approval document identification is obtained. Further, the data retrieval system uses the target approval document identification as an index to query the natural resources industry approval business database. This database stores relevant information for all approval documents. The system retrieves the structured information of other approval documents that are related to the target approval document identification (such as belonging to different stages of the same project, related projects, etc.). This structured information is stored in a fixed format and structure, ultimately yielding the structured information of the associated approval documents.
[0083] Continuing with the above example, the initial search results are three approval reports concerning ecological environment monitoring data for a certain region. The data retrieval system extracts the report number "HJJC-2024-001" from the documents as the target approval report identifier. Using the report number as an index, a query is performed in the natural resources industry approval business database to retrieve the structured information of the other two approval reports associated with that report number. This includes structured fields such as basic information of the reports and approval process information, as well as corresponding content, thus obtaining the structured information of the associated approval reports.
[0084] Step 411: Extract the same information category fields from the structured information of the associated approval documents that correspond to the preliminary search results, obtain the same category information fields, and compare the content of the same category information fields in the preliminary search results with the content of the corresponding fields in the structured information of the associated approval documents one by one to obtain the comparison results of the same category field content.
[0085] Furthermore, the data retrieval system analyzes the structured information of related approval documents, identifying fields with the same information categories as those in the initial search results. For example, if the initial search results are related to ecological environment monitoring data, and the structured information of the related approval documents also contains fields related to ecological environment monitoring, these fields with the same category are extracted. The data retrieval system then compares the specific content of the fields with the corresponding fields in the structured information of the related approval documents, one by one. Using an exact matching algorithm, it checks the consistency of the corresponding field content word by word. After the comparison is completed, the system obtains the comparison results of the same category of fields, recording the consistency of each field with the content in the initial search results and the structured information of the related approval documents.
[0086] Continuing with the above embodiment, for fields such as "monitoring time" and "monitoring data value" in the preliminary search results of step 410, as well as the same "monitoring time" and "monitoring data value" fields in the structured information of the associated approval report, the data retrieval system extracts these information fields of the same category. The content of the "monitoring time" field "2024-01-01" in the preliminary search results is compared with the content of the "monitoring time" field "2024-01-02" in the structured information of the associated approval report. Similarly, other information fields of the same category are compared one by one to obtain the comparison results of the content of the same category fields, and the consistency of each field content is recorded.
[0087] Step 412: Based on the comparison results of fields of the same category, identify fields that are inconsistent with the content in the preliminary search results and the structured information of the associated approval documents, obtain a list of inconsistent fields, and retrieve the original data source of the inconsistent fields in the preliminary search results and the structured information of the associated approval documents based on the list of inconsistent fields, to obtain the original data source information of the inconsistent fields.
[0088] Furthermore, based on the comparison results of fields of the same category, the data retrieval system filters out fields whose content differs from that in the preliminary retrieval results and the structured information of the associated approval documents, and compiles these fields into a list of inconsistent fields.
[0089] Furthermore, based on the list of inconsistent fields, the data retrieval system retrieves the original data source for each inconsistent field from the storage locations of the preliminary search results documents and the structured information of the associated approval reports. The original data source may be a specific paragraph in a document, a specific record in a database, etc. The system obtains relevant information about these original data sources, including storage path and data collection time, to form the original data source information for the inconsistent fields.
[0090] Continuing with the above embodiment, assuming that the comparison results of the same category fields in step 411 show that the "monitoring time" and "monitoring data value" fields are inconsistent, the data retrieval system will compile the "monitoring time" and "monitoring data value" into a list of inconsistent fields. Then, the system determines from the preliminary retrieval results document that the data for the "monitoring time" field originates from the third paragraph on page 1, and determines from the storage database of structured information of associated approval reports that its corresponding data originates from the fifth record of a certain data table; similarly, it obtains the original data source of the "monitoring data value" field, thus obtaining the original data source information of the inconsistent fields.
[0091] Step 413: Based on the original data source information of the inconsistent fields, verify the actual content of the inconsistent fields in the original data source, confirm the correct content of the inconsistent fields, and obtain the confirmation result of the correct content of the inconsistent fields.
[0092] Furthermore, the data retrieval system traces the original data source information of the inconsistent fields back to their actual location and conducts a detailed verification of the content of the inconsistent fields. Through manual review (if necessary), reference to relevant business regulations, and consideration of the data collection context, the correct content of the inconsistent fields is determined. For example, checking the original data collection records and reviewing relevant project documents, the correct content of each inconsistent field is ultimately determined, forming a confirmation result for the correct content of the inconsistent fields.
[0093] Continuing with the above embodiment, for the "monitoring time" field in step 412, the data retrieval system traces back to the third paragraph on page 1 of the preliminary retrieval result document and the fifth record in the data table corresponding to the structured information of the associated approval report. By checking the monitoring log of the project, it is found that the actual monitoring time is "2024-01-01", and the correct content of the "monitoring time" field is determined to be "2024-01-01". Similarly, the "monitoring data value" field is checked, and the correct content of the inconsistent field is finally confirmed.
[0094] Step 414: Based on the confirmation result of the correct content in the inconsistent field, record the information in the preliminary search results that is inconsistent with the correct content and the verification process to obtain the content consistency verification result.
[0095] Furthermore, the data retrieval system confirms the results based on the correct content of the inconsistent fields, marking information that is inconsistent with the correct content in the preliminary retrieval result document. Simultaneously, it records the verification process for these inconsistent information, including the original data source, verification method, and the basis for determining the correct content.
[0096] Furthermore, the data retrieval system derives a content consistency verification result based on the inconsistencies. If there are no inconsistencies or all inconsistencies have been identified as correct and corrected, the content consistency verification result is "consistent"; if there are inconsistencies that are not identified as correct or cannot be corrected, the content consistency verification result is "inconsistent".
[0097] Following the above embodiment, for the "monitoring time" and "monitoring data value" fields, the data retrieval system marks the original error information in the preliminary retrieval result document and records the verification process: by reviewing the monitoring log, the correct content of "monitoring time" is determined to be "2024-01-01", and by verifying the original data of the monitoring instrument, the correct content of "monitoring data value" is determined. Since all inconsistent information has been confirmed to be correct, the content consistency verification result is "consistent".
[0098] This invention, through its embodiments, extracts the target approval document identifier and retrieves the structured information of related documents, then compares the content of fields of the same category and verifies inconsistent fields, ultimately determining the correct content and recording the verification results. This ensures the consistency between the preliminary search results and the information of related approval documents, avoids data contradictions affecting the accuracy and fairness of natural resource industry approval processes, and improves data quality and the reliability of business processing.
[0099] In one embodiment, steps 415 to 419 include:
[0100] Step 415: Extract the original approval report documents corresponding to the preliminary search results to obtain the target original approval report documents, and obtain the file format standards stipulated by the natural resources industry approval business system based on the target original approval report documents to obtain the system's list of allowed file formats.
[0101] Optionally, the data retrieval system processes the preliminary search results, extracting the original approval documents corresponding to the preliminary search results from the storage system using file indexing and storage path information, and identifying them as the target original approval documents. These documents may be in various forms such as documents, tables, and images.
[0102] Furthermore, the data retrieval system searches for the corresponding file format standard in the format standard library of the natural resources industry approval business system based on the business type (such as ecological environment monitoring, geological exploration, etc.) of the original approval report document. The standard library pre-stores the format requirements for various types of approval business documents, extracts all allowed file format types, and compiles them into a list of allowed file formats of the system.
[0103] Continuing with the above example, the initial search results are three approval reports concerning ecological and environmental monitoring data for a certain region. Based on the storage path, these three original approval reports are extracted from the file server as the target original approval reports. Since these files belong to ecological and environmental monitoring business, the file format standards corresponding to ecological and environmental monitoring business are searched in the format standard library of the natural resources industry approval business system to obtain a list of allowed file formats, including formats such as ".doc", ".pdf", and ".xls".
[0104] Step 416: Based on the system's list of allowed file formats, detect the format type of the target original approval report file, determine whether it belongs to the format type in the system's list of allowed file formats, and obtain the file format detection result.
[0105] Furthermore, the data retrieval system employs file format recognition technology (such as based on file extensions, file header identifiers, etc.) to detect the format type of the target original approval document. The system then compares the detected file format types one by one with the format types in the system's list of allowed file formats. If the format type of the target original approval document exists in the system's list of allowed file formats, the document is deemed compliant; otherwise, it is deemed non-compliant, ultimately yielding the file format detection result and determining whether the file format is compliant.
[0106] Continuing with the above embodiment, for the three target original approval documents in step 415, the data retrieval system identifies them by file extension, finding that two are in ".pdf" format and one is in ".txt" format. These format types are compared with the system's list of allowed file formats (".doc", ".pdf", ".xls"), resulting in the following file format detection results: two ".pdf" files meet the requirements, while one ".txt" file does not.
[0107] Step 417: If the target original approval document file format is determined to be not a system-allowed format based on the file format detection result, then the actual format type of the target original approval document file is identified to obtain the actual file format information. Based on the actual file format information, the system allows a query to see if there is a compatible format type in the list of file formats that corresponds to the actual file format information, and the format compatibility query result is obtained.
[0108] Furthermore, when the file format detection results determine that the target original approval document file format does not belong to the format type in the system's allowed file format list, the data retrieval system employs more detailed file format analysis techniques (such as file content feature analysis) to accurately identify the actual format type of the target original approval document file and obtain the actual file format information. Further, based on the actual file format information, the data retrieval system searches the system's allowed file format list to determine if a format type compatible with the actual file format exists. If it exists, the compatible format type is recorded; if it does not exist, it indicates that there is no compatible format, ultimately yielding the format compatibility query result.
[0109] Continuing with the above embodiments, for the ".txt" file whose format does not meet the requirements in step 416, the data retrieval system confirms that its actual file format information is plain text format through file content feature analysis. Then, it searches in the system's list of allowed file formats (".doc", ".pdf", ".xls") and finds that the ".doc" format can achieve content compatibility with the ".txt" format through copying and pasting, thus determining that ".doc" is a compatible format type, and obtaining the format compatibility query result that a compatible format type (".doc") exists.
[0110] Step 418: If a compatible format type is determined based on the format compatibility query results, the compatible format type and format conversion requirements are recorded to obtain format conversion reference information. Based on the format conversion reference information, it is checked whether the target original approval report document can be converted into a compatible format according to the format conversion requirements to obtain the format conversion feasibility result.
[0111] Furthermore, when the data retrieval system determines the existence of a compatible format type based on the format compatibility query results, the system records the compatible format type and the specific requirements (such as conversion tools, operation steps, etc.) needed to convert the target original approval document into that compatible format, forming format conversion reference information. Further, the data retrieval system checks the target original approval document based on the format conversion reference information. By simulating the conversion process and checking the integrity of the document content and format compatibility, it determines whether the document can be successfully converted into a compatible format according to the format conversion requirements. If the conversion is successful, the format conversion feasibility result is "feasible"; if the conversion is unsuccessful, it is "infeasible," ultimately obtaining the format conversion feasibility result.
[0112] Continuing with the above embodiment, for the ".txt" file whose compatible format type is determined to be ".doc" in step 417, ".doc" is recorded as a compatible format type, and the format conversion requirement is specified as opening the ".txt" file directly using word processing software (such as Microsoft Word) and saving it as ".doc" format to obtain format conversion reference information. Next, the system simulates the conversion process and checks to find that the content of the ".txt" file is not lost and the format is normal after conversion to ".doc" format, thus determining the format conversion feasibility result as "feasible".
[0113] Step 419: Based on the format conversion feasibility results, summarize the format detection information, compatibility query information, and conversion feasibility information of the original target approval documents to obtain the document format standardization inspection results.
[0114] Furthermore, the data retrieval system summarizes file format detection information (whether the file format conforms to the system's allowed formats), format compatibility query information (whether there are compatible format types), and format conversion feasibility information (whether it can be converted to a compatible format). If the file format detection result is compliant, or although it does not meet the requirements but a compatible format exists and conversion is feasible, the file format standardization check result is "compliant"; if the file format does not meet the requirements and there is no compatible format or conversion is not feasible, the file format standardization check result is "non-compliant," and the final file format standardization check result is obtained.
[0115] Continuing with the above example, for the three target original approval documents, two ".pdf" files passed the format test and met the requirements. One ".txt" file, although initially not meeting the requirements, had a compatible ".doc" format that could be converted. The data retrieval system aggregated this information and concluded that the file format compliance test result was "compliant".
[0116] This invention, from extracting files and obtaining format standards to detecting file formats, judging compatibility and conversion feasibility, and finally summarizing the test results, ensures that the format of submitted files meets the requirements of the natural resources industry approval business system. This effectively avoids the approval process being blocked due to non-standard file formats, improves the efficiency of approval business processing, and ensures the smooth progress of approval work and the standardization of data storage and processing.
[0117] In one embodiment, steps 420 to 424 include:
[0118] Step 420: Extract the image content contained in the original approval report file corresponding to the preliminary search results to obtain the target approval report image, and obtain the image resolution and clarity standards specified by the natural resources industry approval business system based on the target approval report image to obtain the image quality standard requirements.
[0119] Optionally, the data retrieval system processes the original approval report files corresponding to the preliminary retrieval results. Through file parsing technology, it identifies and extracts all image content contained in the original approval report files and identifies these images as target approval report images. These images may exist in different formats (such as JPEG, PNG, etc.) in the files.
[0120] Furthermore, the data retrieval system, based on the business type of the target approval document image (such as on-site photos in ecological environment monitoring, schematic diagrams in geological exploration, etc.), searches for the corresponding image resolution and clarity standards in the standard and specification library of the natural resources industry approval business system. This standard library pre-stores quality requirements for images of various approval business types. The system extracts the relevant standards, organizes them into image quality standard requirements, and clearly specifies the resolution values, clarity details, and other indicators that the images should achieve.
[0121] Continuing with the above example, the initial search results were three approval reports concerning ecological environment monitoring data for a certain region. The data retrieval system extracted five on-site monitoring photos and two ecological environment schematic diagrams from these documents through file parsing, which were then used as the target approval report images. Since these images pertain to ecological environment monitoring, the system searched the standard and specification library of the natural resources industry approval system for the corresponding image quality standards. The image quality standards required were: resolution no less than 300 dpi, and text, data labels, and other content in the image must be clearly distinguishable, without blurring or ghosting.
[0122] Step 421: Based on the image quality standard requirements, detect the resolution parameters of each image in the target approval document image, determine whether it meets the resolution standard in the image quality standard requirements, obtain the image resolution detection result, and identify whether the text and graphic details in the image whose resolution meets the standard based on the image resolution detection result are clear and distinguishable, obtain the image clarity detection result.
[0123] Furthermore, the data retrieval system uses image information reading technology to extract resolution parameters (such as horizontal and vertical resolution) from each image in the target approval document. Then, the extracted resolution parameters are compared one by one with the resolution standards in the image quality requirements.
[0124] If the resolution parameters of an image meet or exceed the standard requirements, the image resolution is deemed to meet the standard; otherwise, it is deemed not to meet the standard. The final image resolution detection result is obtained, clarifying the compliance status of each image's resolution.
[0125] Furthermore, for images whose resolution meets the standard in the image resolution detection results, the data retrieval system uses image recognition and analysis technology to identify text and graphic details in the image and determine whether they are clear and legible, such as whether the text can be accurately identified and whether the graphic lines are clear. Based on the judgment results, the image sharpness detection results are obtained, clarifying whether the sharpness of these images meets the requirements.
[0126] Continuing with the above embodiment, for the 7 target approval document images in step 420, the resolution parameters of each image are read. It is found that 3 images have a resolution of 200 dpi, and 4 images have a resolution of 350 dpi. Comparing these resolution parameters with the image quality standard requirement (not less than 300 dpi), the image resolution detection results are as follows: 3 images do not meet the standard, and 4 images do meet the standard.
[0127] For the four images that meet the resolution standard, image recognition technology was used to check and found that the data label in one image was slightly blurry, while the text and graphic details in the other three images were clearly distinguishable. The image clarity test results were: three images meet the clarity standard, and one image does not meet the clarity standard.
[0128] Step 422: Based on the image clarity detection results, identify images in the target approval document images that do not meet the resolution or clarity standards, obtain a list of unqualified images, and analyze the specific degree of non-compliance of each unqualified image based on the list of unqualified images to obtain the degree of non-compliance analysis results.
[0129] Furthermore, based on the image resolution and image sharpness detection results, the data retrieval system filters out images that do not meet the image quality standards in terms of resolution or sharpness, and compiles these images into a list of unqualified images.
[0130] Furthermore, the data retrieval system performs a detailed analysis of each image in the list of substandard images. By comparing the actual parameters of the images with the standard requirements, it assesses the specific degree to which each substandard image fails to meet the standard. For example, it calculates the numerical difference in resolution below the standard, analyzes the range and degree of blurriness, and ultimately obtains the substandard degree analysis results, quantifying the severity of the problem for each substandard image.
[0131] Continuing with the above embodiments, and based on the results of step 421, the data retrieval system identifies 3 images with resolution that do not meet the standard and 1 image with clarity that does not meet the standard, for a total of 4 images as unqualified images, which are compiled into a list of unqualified images. Analysis of the images in the list of unqualified images reveals that 3 images with a resolution of 200 dpi are each 100 dpi lower than the standard; and in the image with clarity that does not meet the standard, the blurred area of the data marker occupies 15% of the image area. This analysis of the degree of unqualification clarifies the specific circumstances under which each unqualified image does not meet the standard.
[0132] Step 423: Based on the results of the non-compliance analysis, confirm whether the non-compliance images can be technically processed to meet the image quality standards, and obtain the feasibility results of image processing.
[0133] Furthermore, based on the analysis results of the degree of non-compliance, the data retrieval system, for each image in the list of non-compliant images, combines existing image processing technologies (such as resolution enhancement algorithms, image sharpening technologies, etc.) to determine whether the image can be processed through technical means to meet the image quality standards.
[0134] If the assessment determines that existing technology can be used to process the image to meet the standards, the image processing feasibility result is "feasible"; if existing technology cannot meet the requirements, it is "infeasible". The final image processing feasibility result clarifies the processing possibility of each unqualified image.
[0135] Continuing with the above embodiment, the data retrieval system evaluates the four unqualified images from step 422. It is found that the three low-resolution images can be upgraded to 300 dpi using the resolution enhancement function of professional image processing software; however, the image whose clarity does not meet the standard has a large blurred area, and existing image sharpening techniques cannot fully restore the clarity of the data identifier. The final image processing feasibility result is: processing three images is feasible, and processing one image is not feasible.
[0136] Step 424: Based on the feasibility results of image processing, summarize the quality inspection information, non-compliance analysis information and processing feasibility information of the target approval report images to obtain the image quality standardization inspection results.
[0137] Furthermore, the data retrieval system summarizes the image resolution detection results and image sharpness detection results (i.e., quality detection information) from step 421, the non-compliance analysis results (non-compliance analysis information) from step 422, and the image processing feasibility results from step 423.
[0138] If all images meet the standards in terms of resolution and clarity, or if there are substandard images that can be technically processed to meet the standards, the image quality standardization inspection result is "standard"; if there are substandard images that cannot be technically processed to meet the standards, the image quality standardization inspection result is "non-standard", and the final image quality standardization inspection result is obtained.
[0139] Continuing with the above example, for the 7 target approval document images, the relevant information is summarized as follows: 4 images passed the quality inspection, 3 images with substandard resolution can be processed, and 1 image with substandard clarity cannot be processed. Since there is a substandard image that cannot be processed, the data retrieval system determines the image quality compliance inspection result to be "non-compliant".
[0140] This invention, from extracting images and obtaining quality standards to detecting resolution and clarity, analyzing non-compliance, judging processing feasibility, and finally summarizing the inspection results, ensures that the quality of images in approval documents meets the requirements of the natural resources industry approval business system. This avoids affecting the accuracy and efficiency of the approval process due to image quality issues and ensures that the image information in the approval documents can clearly and accurately convey the content.
[0141] In one embodiment, steps 425 to 428 include:
[0142] Step 425: Extract the document structure information of the original approval report file corresponding to the preliminary search results to obtain the target approval report document structure, and obtain the standard document structure framework of the approval report document stipulated in the natural resources industry approval business based on the target approval report document structure to obtain the standard document structure framework.
[0143] Optionally, the data retrieval system processes the original approval report documents corresponding to the preliminary retrieval results. Using document parsing technology, it identifies the chapter divisions, paragraph levels, and heading settings in the documents, extracts the structural information of the entire document, and thus obtains the structure of the target approval report document. The structural information describes the compositional relationship and hierarchical order of the various parts of the document.
[0144] Furthermore, based on the business type involved in the target approval document (such as ecological environment monitoring report, land resource approval document, etc.), the data retrieval system searches for the corresponding standard structure framework of the approval document in the standard specification library of natural resource industry approval business. The standard specification library stores the standard structure templates that various approval business documents should follow. The system extracts the corresponding standard structure framework, which clearly specifies the structural modules that the document should contain, as well as the hierarchical relationship and content requirements of each module.
[0145] Continuing with the above example, the initial search results were three approval reports concerning ecological and environmental monitoring data for a certain region. The data retrieval system, using document parsing technology, analyzed the document's chapter structure and paragraph formatting to obtain the structure of the target approval report document. One document was found to contain three main chapters: "Monitoring Background," "Monitoring Methods," and "Monitoring Results." Because these documents pertain to ecological and environmental monitoring operations, the system located the standard structural framework for ecological and environmental monitoring approval reports in the natural resources industry's standard and specification library. This framework stipulates that the document should include five structural modules: "Monitoring Background," "Monitoring Methods," "Monitoring Results," "Results Analysis," and "Conclusions and Recommendations."
[0146] Step 426: Based on the standard document structure framework, compare the target approval report document structure with the standard document structure framework, identify the missing structural modules in the target approval report document structure, obtain a list of missing structural modules, and check whether there are any incorrectly identified structural module contents in the target original approval report file based on the list of missing structural modules, and obtain the unidentified structural module verification results.
[0147] Furthermore, the data retrieval system compares the target approval document structure with the standard document structure framework item by item. By comparing the names, hierarchical relationships, and content requirements of each structural module, it determines which standard-specified structural modules are missing from the target approval document structure and compiles these missing structural modules into a list of missing structural modules.
[0148] Furthermore, the data retrieval system performs a deep scan of the original approval documents for each module in the missing structural module list. Using techniques such as text keyword matching and semantic analysis, it checks for content related to the missing structural modules that was not correctly identified during the previous structural extraction process. The system records the check results, obtaining the verification results for unidentified structural modules, and clarifies whether there are any potential supplementary structural modules in the document.
[0149] Continuing with the above embodiments, regarding the target approval report document structure (containing three main sections: "Monitoring Background," "Monitoring Methods," and "Monitoring Results") and the standard document structure framework (containing five structural modules: "Monitoring Background," "Monitoring Methods," "Monitoring Results," "Results Analysis," and "Conclusions and Recommendations") in step 425, the data retrieval system found that the "Results Analysis" and "Conclusions and Recommendations" structural modules were missing, forming a list of missing structural modules. Further, the data retrieval system scanned the content of the original approval report document and found a brief analysis of the monitoring results at the end of the document, as well as several recommendations regarding subsequent ecological protection. This content had not previously been identified as an independent structural module, resulting in an unidentified structural module verification result, confirming the existence of incorrectly identified structural module content.
[0150] Step 427: If the unidentified structural module verification results determine that there are incorrectly identified structural module contents, then the structural attributes of the incorrectly identified structural module contents are re-parsed to obtain the re-parsed structural module attribute information. Based on the re-parsed structural module attribute information, the incorrectly identified structural module contents are classified into the modules corresponding to the standard document structure framework, and the target approval report document structure is supplemented and improved to obtain the supplemented document structure.
[0151] Furthermore, when the original approval document contains incorrectly identified structural modules based on the results of the unidentified structural module verification, the data retrieval system uses text analysis and semantic understanding technologies to re-parse this incorrectly identified content. It analyzes the content's theme, logical relationships, and expressive intent to determine its structural attributes, such as whether the content belongs to the analysis, suggestion, or other categories, thus obtaining the re-parsed structural module attribute information.
[0152] Furthermore, based on the re-parsed structural module attribute information, the data retrieval system matches the content of incorrectly identified structural modules with the various modules in the standard document structure framework, classifying them into the corresponding structural modules. By adding or adjusting relevant content in the target approval report document structure, the document structure is supplemented and improved, ultimately resulting in the supplemented document structure.
[0153] Continuing with the above embodiments, the data retrieval system re-parses the incorrectly identified structural module content (the result analysis text and ecological protection recommendations at the end of the document) found in step 426. Analysis determines that the result analysis text belongs to the "Result Analysis" module, and the ecological protection recommendations belong to the "Conclusions and Recommendations" module, resulting in re-parsed structural module attribute information. Next, the data retrieval system categorizes the result analysis text into the "Result Analysis" module and the ecological protection recommendations into the "Conclusions and Recommendations" module, supplementing and improving the target approval report document structure to obtain a supplemented document structure containing five structural modules: "Monitoring Background," "Monitoring Methods," "Monitoring Results," "Result Analysis," and "Conclusions and Recommendations."
[0154] Step 428: Based on the supplemented document structure, compare it again with the standard document structure framework until it is confirmed that all standard structure modules are included, and obtain the document structure integrity confirmation result. Based on the document structure integrity confirmation result, summarize the detection information, supplementary information and confirmation information of the target approval report document structure to obtain the document structure standardization inspection result.
[0155] Furthermore, the data retrieval system performs a new round of detailed comparison between the supplemented document structure and the standard document structure framework. Using precise matching and logical judgment algorithms, it checks whether each structural module in the standard document structure framework exists in the supplemented document structure, and whether the hierarchical relationships and content of each module meet the standard requirements. If any missing or non-compliant structural modules are found, the system returns to steps 426 and 427 to continue checking for incorrectly identified content in the file and supplementing it, then comparing again. This process is repeated until all structural modules in the standard document structure framework can be found in the supplemented document structure with corresponding and compliant content. At this point, the document structure is confirmed to be complete, and a document structure integrity confirmation result is obtained.
[0156] Furthermore, the data retrieval system summarizes the detection information from step 426 (list of missing structural modules, etc.), the supplementary information from step 427 (re-parsing and classification status, etc.), and the document structure integrity confirmation results. If the document structure is complete and conforms to the standard, the document structure standardization check result is "standard"; if there are non-standard situations, it is "non-standard," and the final document structure standardization check result is obtained.
[0157] Continuing with the above embodiments, the supplemented document structure was compared again with the standard document structure framework. It was found that the supplemented document structure fully included the five structural modules: "Monitoring Background," "Monitoring Methods," "Monitoring Results," "Result Analysis," and "Conclusions and Recommendations." Furthermore, the hierarchical relationships and content of each module met the standard requirements, resulting in a "complete" document structure integrity confirmation. Summarizing the previous detection information (finding the missing "Result Analysis" and "Conclusions and Recommendations" modules), the supplementary information (re-analysis and classification of relevant content), and the integrity confirmation result, the document structure standardization inspection result was "standardized."
[0158] This invention, through its embodiments, extracts document structure information and obtains a standard structural framework, then compares and identifies missing modules, checks and supplements unidentified content, and repeatedly compares and confirms completeness. This ensures that the structure of approval documents conforms to the regulations and requirements of the natural resources industry's approval business, avoids the impact on the processing flow and efficiency of approval business due to chaotic document structure or missing key modules, and gives approval documents a unified and standardized structure, facilitating information retrieval, review, and management.
[0159] Furthermore, the retrieval enhancement system based on unstructured data provided by the present invention will be described below. The retrieval enhancement system based on unstructured data described below can be referred to in correspondence with the retrieval enhancement method based on unstructured data described above.
[0160] Optional, refer to Figure 2 , Figure 2This is a structural diagram of the retrieval enhancement system based on unstructured data provided by the present invention. The retrieval enhancement system based on unstructured data includes:
[0161] The knowledge base construction module 210 is used to classify documents based on unstructured natural resource data, obtain information on various types of documents, and construct a vector knowledge base based on the logical relationships of various semantic units in the information on various types of documents.
[0162] The question classification module 220 is used to perform semantic understanding and vector embedding processing on the query question based on the vector knowledge base, to obtain the semantic vector corresponding to the query question, and to classify the query question based on the semantic vector using a question classification model to obtain the category information of the query question;
[0163] The retrieval operation module 230 is used to perform retrieval operations based on category information and vector knowledge base using a retrieval enhancement-based retrieval strategy to obtain preliminary retrieval results;
[0164] The retrieval verification module 240 is used to perform content integrity verification, content legality verification, content consistency verification, and submission document standardization verification based on the preliminary retrieval results, thereby obtaining enhanced retrieval results.
[0165] The embodiments of the present invention achieve efficient and accurate retrieval enhancement of unstructured natural resource report data.
[0166] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating an embodiment of the electronic device provided in this invention. For example... Figure 3 As shown, this embodiment of the invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it performs the following steps:
[0167] Document classification is performed based on unstructured natural resource data to obtain information on various types of documents. Based on the logical relationships between the semantic units in each type of document information, a vector knowledge base is constructed.
[0168] Based on a vector knowledge base, semantic understanding and vector embedding are performed on the query questions to obtain the semantic vectors corresponding to the query questions. Based on the semantic vectors, a question classification model is used to classify the query questions to obtain the category information of the query questions.
[0169] Based on category information and a vector knowledge base, a retrieval-enhanced retrieval strategy is employed to perform retrieval operations and obtain preliminary retrieval results.
[0170] Based on the preliminary search results, the search results were enhanced by performing content integrity checks, content legality checks, content consistency checks, and submission document standardization checks.
[0171] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0172] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.
[0174] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0175] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0176] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
Claims
1. A retrieval enhancement method based on unstructured data, characterized in that, include: Document classification is performed based on unstructured natural resource data to obtain various types of document information. Based on the logical relationships of the semantic units in each type of document information, a vector knowledge base is constructed. The logical relationships include causal relationships, parallel relationships, and subordinate relationships. Based on the vector knowledge base, the query question is semantically understood and vector embedded to obtain the semantic vector corresponding to the query question. Based on the semantic vector, the query question is classified using a question classification model to obtain the category information of the query question. The category information includes geological, ecological and forestry categories. Based on the category information and the vector knowledge base, a retrieval strategy based on retrieval enhancement is used to perform a retrieval operation and obtain preliminary retrieval results; Based on the preliminary search results, content integrity verification, content legality verification, content consistency verification, and submission document standardization verification are performed respectively to obtain enhanced search results; The steps for verifying the conformity of submitted documents based on the preliminary search results include: Extract the image content contained in the original approval report file corresponding to the preliminary search results to obtain the target approval report image, and obtain the image resolution and clarity standards specified by the natural resources industry approval business system based on the target approval report image to obtain the image quality standard requirements; Based on the image quality standard requirements, the resolution parameters of each image in the target approval document image are detected to determine whether they meet the resolution standard in the image quality standard requirements, and the image resolution detection result is obtained. Then, the text and graphic details in the image whose resolution meets the standard based on the image resolution detection result are identified to obtain the image clarity detection result. Based on the image clarity detection results, images in the target approval document images that do not meet the resolution or clarity standards are identified, resulting in a list of unqualified images. Based on the list of unqualified images, the specific degree of non-compliance of each unqualified image is analyzed to obtain the degree of non-compliance analysis results. Based on the analysis results of the degree of non-compliance, it is confirmed whether the non-compliant images can be processed to meet the image quality standards, and the feasibility results of image processing are obtained. Based on the image processing feasibility results, the quality inspection information, non-compliance analysis information, and processing feasibility information of the target approval document images are summarized to obtain the image quality standardization inspection results.
2. The retrieval enhancement method based on unstructured data according to claim 1, characterized in that, The steps for verifying content integrity based on the preliminary search results include: Extract all structured data fields from the preliminary search results to obtain initial structured data fields, and obtain a list of pre-set necessary information fields in the natural resources industry approval business based on the initial structured data fields to obtain a list of necessary information fields for the business. The initial structured data fields are compared one by one with the list of business-required information fields to obtain the field comparison results. Based on the field comparison results, the missing required information fields in the initial structured data fields are identified to obtain a list of missing fields. Check whether there is any incorrectly extracted information content in the preliminary search results that corresponds to each field in the missing field list, obtain the incorrectly extracted information verification results, and locate the incorrectly extracted information in the preliminary search results based on the incorrectly extracted information verification results to obtain the location information of the incorrectly extracted information; Extract the information content corresponding to the location information and supplement it into the initial structured data field to obtain the supplemented structured data field; Based on the supplemented structured data fields, they are compared again with the list of business-required information fields until it is confirmed that all required information fields are included, thus obtaining the content integrity verification result.
3. The retrieval enhancement method based on unstructured data according to claim 1, characterized in that, The steps for verifying the legality of the content based on the preliminary search results include: Extract the information content corresponding to all structured data in the preliminary search results to obtain the search information content, and obtain the legality and regulatory requirements for various types of information content in the natural resources industry approval business based on the search information content to obtain the information content legality and regulatory list. Each piece of information in the retrieved information content is matched with the corresponding specification requirements in the information content legality specification list to obtain the matching result of information content and specification requirements. Based on the matching result, information content in the retrieved information content that does not meet the legality specification requirements is identified to obtain a list of illegal information content. Analyze the specific reasons why each item in the list of illegal information content does not meet the requirements of the regulations, obtain the analysis results of the illegality reasons, and based on the analysis results of the illegality reasons, confirm whether the illegal information content belongs to the special circumstances allowed in the approval business of the natural resources industry, and obtain the special circumstances confirmation results; Based on the special circumstances, the specific location and reason for the illegal information content are recorded to obtain a detailed record of the illegal information. The detailed record of the illegal information is then summarized into a content legality verification report to obtain the content legality verification result.
4. The retrieval enhancement method based on unstructured data according to claim 1, characterized in that, The steps for verifying content consistency based on the preliminary search results include: Extract the approval report identification information contained in the preliminary search results to obtain the target approval report identification, and retrieve the structured information of other approval reports associated with the target approval report identification from the natural resources industry approval business database to obtain the associated approval report structured information; Extract the same information category field from the structured information of the associated approval report that corresponds to the preliminary search result to obtain the same category information field, and compare the content of the same category information field in the preliminary search result with the content of the corresponding field in the structured information of the associated approval report one by one to obtain the same category field content comparison result; Based on the comparison results of the same category fields, identify the fields that are inconsistent with the content in the preliminary search results and the structured information of the associated approval documents, obtain a list of inconsistent fields, and retrieve the original data source of the inconsistent fields in the preliminary search results and the structured information of the associated approval documents based on the list of inconsistent fields to obtain the original data source information of the inconsistent fields. Based on the original data source information of the inconsistent fields, the actual content of the inconsistent fields in the original data source is checked, the correct content of the inconsistent fields is confirmed, and the correct content confirmation result of the inconsistent fields is obtained. Based on the confirmation results of the correct content in the inconsistent fields, record the information in the preliminary search results that is inconsistent with the correct content and the verification process to obtain the content consistency verification results.
5. The retrieval enhancement method based on unstructured data according to claim 1, characterized in that, The steps for verifying the conformity of submitted documents based on the preliminary search results include: Extract the original approval report file corresponding to the preliminary search results to obtain the target original approval report file, and obtain the file format standard specified by the natural resources industry approval business system based on the target original approval report file to obtain the system's allowed file format list; Based on the system's list of allowed file formats, the system detects the format type of the target original approval document and determines whether it belongs to the format type in the system's list of allowed file formats, thus obtaining the file format detection result. If the target original approval document file format is determined to be not a system-allowed format based on the file format detection result, the actual format type of the target original approval document file is identified to obtain the actual file format information. Based on the actual file format information, the system's allowed file format list is queried to see if there is a compatible format type corresponding to the actual file format information, and the format compatibility query result is obtained. If a compatible format type is determined based on the format compatibility query results, the compatible format type and format conversion requirements are recorded to obtain format conversion reference information. Based on the format conversion reference information, it is checked whether the target original approval document can be converted into a compatible format according to the format conversion requirements to obtain a format conversion feasibility result. Based on the format conversion feasibility results, the format detection information, compatibility query information, and conversion feasibility information of the target original approval document are summarized to obtain the document format standardization test results.
6. The retrieval enhancement method based on unstructured data according to claim 1, characterized in that, The steps for verifying the conformity of submitted documents based on the preliminary search results include: Extract the document structure information of the original approval report file corresponding to the preliminary search results to obtain the target approval report document structure, and obtain the standard document structure framework of the approval report document stipulated in the natural resources industry approval business based on the target approval report document structure to obtain the standard document structure framework. Based on the standard document structure framework, the target approval report document structure is compared with the standard document structure framework to identify the missing structural modules in the target approval report document structure, obtain a list of missing structural modules, and check whether there are any incorrectly identified structural module contents in the target original approval report file based on the list of missing structural modules to obtain the unidentified structural module verification result; If, based on the verification results of the unidentified structural modules, it is determined that there are structural modules that have not been correctly identified, then the structural attributes of the structural modules that have not been correctly identified are re-parsed to obtain the re-parsed structural module attribute information. Based on the re-parsed structural module attribute information, the structural modules that have not been correctly identified are classified into the modules corresponding to the standard document structure framework, thereby supplementing and improving the target approval report document structure to obtain the supplemented document structure. Based on the supplemented document structure, it is compared with the standard document structure framework again until it is confirmed that all standard structure modules are included, and the document structure integrity confirmation result is obtained. Based on the document structure integrity confirmation result, the detection information, supplementary information and confirmation information of the target approval report document structure are summarized to obtain the document structure standardization inspection result.
7. A retrieval enhancement system based on unstructured data, characterized in that, Applied to the retrieval enhancement method based on unstructured data as described in any one of claims 1 to 6; The retrieval enhancement system based on unstructured data includes: The knowledge base construction module is used to classify documents based on unstructured natural resource data, obtain various types of document information, and construct a vector knowledge base based on the logical relationships of various semantic units in the various types of document information. The logical relationships include causal relationships, parallel relationships, and subordinate relationships. The question classification module is used to perform semantic understanding and vector embedding processing on the query question based on the vector knowledge base to obtain the semantic vector corresponding to the query question, and to classify the query question based on the semantic vector using a question classification model to obtain the category information of the query question. The category information includes geological, ecological and forestry categories. The retrieval operation module is used to perform retrieval operations based on the category information and the vector knowledge base using a retrieval enhancement-based retrieval strategy to obtain preliminary retrieval results; The retrieval verification module is used to perform content integrity verification, content legality verification, content consistency verification, and submission document standardization verification based on the preliminary retrieval results, so as to obtain enhanced retrieval results; The steps for verifying the conformity of submitted documents based on the preliminary search results include: Extract the image content contained in the original approval report file corresponding to the preliminary search results to obtain the target approval report image, and obtain the image resolution and clarity standards specified by the natural resources industry approval business system based on the target approval report image to obtain the image quality standard requirements; Based on the image quality standard requirements, the resolution parameters of each image in the target approval document image are detected to determine whether they meet the resolution standard in the image quality standard requirements, and the image resolution detection result is obtained. Then, the text and graphic details in the image whose resolution meets the standard based on the image resolution detection result are identified to obtain the image clarity detection result. Based on the image clarity detection results, images in the target approval document images that do not meet the resolution or clarity standards are identified, resulting in a list of unqualified images. Based on the list of unqualified images, the specific degree of non-compliance of each unqualified image is analyzed to obtain the degree of non-compliance analysis results. Based on the analysis results of the degree of non-compliance, it is confirmed whether the non-compliant images can be processed to meet the image quality standards, and the feasibility results of image processing are obtained. Based on the image processing feasibility results, the quality inspection information, non-compliance analysis information, and processing feasibility information of the target approval document images are summarized to obtain the image quality standardization inspection results.
8. An electronic device, comprising: Memory, used to store computer software programs; A processor for reading and executing the computer software program, characterized in that, when the processor executes the computer software program, it implements the retrieval enhancement method based on unstructured data as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium, wherein a computer software program is stored therein, characterized in that, When the computer software program is executed by the processor, it implements the retrieval enhancement method based on unstructured data as described in any one of claims 1 to 6.
Citation Information
Patent Citations
ES retrieval knowledge base method based on BERT enhancement
CN118885565A
Operation and maintenance question-answering system and method based on large language model workflow
CN119441410A