Semantic understanding-based environmental impact report auxiliary auditing system

Through the environmental impact report review system based on semantic understanding, the problems of low review efficiency and poor consistency in existing technologies have been solved, accurate review and anomaly identification of environmental impact reports have been achieved, and the comprehensiveness and timeliness of the review have been improved.

CN120805925APending Publication Date: 2025-10-17SHANGHAI RUIDUN INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510966769.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively handling non-uniform format chapters and multi-level semantically nested content in the review of environmental impact reports, resulting in incomplete extraction of key review elements, the rule base being unable to respond to policy updates in a timely manner, and low review efficiency and poor consistency.

Method used

An environmental impact report auxiliary review system based on semantic understanding is adopted, including text cleaning, chapter segmentation, rule generation and content review modules. Through semantic analysis and standard database comparison, it can achieve accurate extraction and difference analysis of monitoring items, and mark anomalies and logical deviations.

Benefits of technology

It significantly improves the purity and consistency of audit data, enhances the accuracy of semantic content positioning, achieves content granularity refinement and targeted audit judgment, and improves the comprehensiveness and timeliness of audit results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805925A_ABST
    Figure CN120805925A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semantic understanding, in particular to an environmental impact report auxiliary auditing system based on semantic understanding, which comprises a text cleaning module, a chapter segmentation module, a rule generation module, a content auditing module and a result output module. According to the method, the paragraph structure and the layout identification are analyzed in a unified mode, key paragraph types in environmental influence chapters are efficiently recognized, the text integration accuracy is improved by combining inter-paragraph similarity calculation and repetition rate screening, chapter boundary accurate positioning and affiliation adjustment are achieved on the basis of format feature comparison of serial numbers and title styles, and the text integration efficiency is improved. A multi-layer parameter alignment rule set is constructed to support comprehensive verification of standard numbers, monitoring frequencies and periodic elements, an exception labeling mechanism is combined to complete difference item logic judgment and consistency evaluation, an auditing basis with rule driving and data comparison capabilities is formed, chapter positioning precision, element extraction comprehensiveness and logic exception recognition efficiency are improved, and the method is suitable for popularization and application. And the pertinence and the automatic processing depth in the auditing process are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of semantic understanding, in particular to an environmental impact report auxiliary auditing system based on semantic understanding. BACKGROUND

[0002] The technical field of semantic understanding relates to the identification, analysis and modeling of the semantics contained in text, speech and other information in the natural language processing process, including semantic role labeling, contextual semantic disambiguation, entity relationship identification, semantic dependency analysis, etc. The purpose is to enable the computing system to identify the real meaning implied in the language, thereby supporting information extraction, question and answer systems, text generation, sentiment analysis and other intelligent applications. Semantic understanding is often achieved through means such as word vector generation, context modeling, semantic graph construction and semantic reasoning. It relies on a large amount of corpus and machine learning methods to model and deduce the connections between words, syntax and context in language, forming a semantic structure representation that can be called by downstream tasks. Among them, the environmental impact report auxiliary auditing system refers to the semantic analysis of the text content in the environmental impact report in the environmental impact assessment process to assist the auditing personnel in identifying whether the project meets the key auditing elements such as policy regulations, completeness and standardization. In the traditional way, the report book is usually searched and compared by means of keyword matching, template comparison and manually preset sentence structure recognition based on rules to form the basis for auditing and auxiliary judgment. The small class information such as construction project profile, pollution source analysis, environmental status description, evaluation factor selection, etc. in the report book is searched and compared.

[0003] The prior art uses a static keyword library and manually preset templates for text comparison. The rule system lacks dynamic expansion capability. When analyzing non-uniform format chapters and multi-level semantic nested content, it is limited by the preset sentence pattern recognition mechanism and is difficult to effectively handle title number variations and cross-paragraph semantic echo relationships, resulting in insufficient completeness of key auditing elements such as evaluation factor selection and pollution source analysis. The fixed rule library cannot respond to policy clause updates in time, and rule lag bias may occur in the matching of environmental monitoring projects and standard parameters. The auditing process dominated by human experience cannot guarantee the consistency of large-scale text processing, and the auditing efficiency is subject to the static nature of the rule library. In complex semantic scenarios, it lacks automated reasoning capability, affecting the comprehensiveness and timeliness of the auditing results. SUMMARY

[0004] To solve the technical problems existing in the prior art, the embodiments of the present application provide an environmental impact report auxiliary auditing system based on semantic understanding. The technical solution is as follows:

[0005] On the one hand, an environmental impact report auxiliary auditing system based on semantic understanding is provided, which includes:

[0006] The text cleaning module obtains the original text by receiving the uploaded environmental impact report, analyzes the paragraph structure and layout identification, identifies the environmental impact chapter characteristics, screens the literature and repeated paragraphs, adjusts the structure to screen the associated paragraphs, outputs the preprocessed text, and passes to the chapter segmentation module;

[0007] The chapter segmentation module obtains the preprocessed text, compares the title style and number according to the chapter format of the environmental impact report, screens the title group, calculates the section boundary, adjusts the division chapter, outputs the chapter division set, and passes to the rule generation module and the content audit module;

[0008] The rule generation module obtains the chapter division set, extracts the monitoring item title and paragraph, calls the environmental assessment standard database, screens the monitoring parameter and standard number, constructs the comparison rule, outputs the comparison rule set, and passes to the content audit module;

[0009] The content audit module obtains the chapter division set and the comparison rule set, analyzes the data consistency, screens the repeated and difference items, analyzes the unit, frequency and cycle difference according to the project requirements, calculates the consistency, analyzes the logical contradiction and deviation degree, and forms the abnormal mark set, and passes to the result output module.

[0010] As a further scheme of the present application, the preprocessed text includes a screened literature set, a deduplicated paragraph group, and an associated paragraph structure, the chapter division set includes title group data, section boundary parameters, and chapter attribution indexes, the comparison rule set includes a monitoring parameter set, a standard number index, and a comparison condition matrix, and the abnormal mark set includes a difference data set, an abnormal type label, and a logical contradiction index.

[0011] As a further scheme of the present application, the text cleaning module includes:

[0012] The paragraph recognition submodule obtains the original text by receiving the uploaded environmental impact report, analyzes the starting position, paragraph length and paragraph number of the text paragraph, detects the paragraph beginning punctuation, number style and paragraph indentation form, identifies the environmental impact report chapter title corresponding paragraph, screens the paragraph type, and obtains the paragraph category identification data;

[0013] The content screening submodule screens the literature feature paragraph and the repeated paragraph based on the paragraph category identification data, compares the paragraph beginning identifier, the literature reference symbol and the paragraph tail mark, analyzes the repetition degree of the paragraph in the chapter, screens the retained paragraph content, and obtains the associated paragraph screening data;

[0014] The page structure adjusting sub-module adjusts the order and logical attribution of the paragraphs, identifies the page areas of the cover, table of contents and references, reclassifies the associated paragraphs into the chapter logical structure, establishes paragraph attribution blocks, and generates the preprocessed text.

[0015] As a further scheme of the present application, the chapter segmentation module comprises:

[0016] The title filtering sub-module acquires the preprocessed text, collects paragraph title data according to the chapter format of the environmental impact report, compares the numbering style, punctuation format and keyword structure, filters the titles in the paragraphs with chapter characteristics, calculates the title line length and the number of adjacent paragraphs, and generates chapter title position distribution values.

[0017] The structure positioning sub-module collects the number of lines between titles, the average number of words in a paragraph and the indentation format data based on the chapter title position distribution values, compares the physical distance between titles and the paragraph structure distribution rules, calculates the paragraph attribution section range according to the distance and structure relationship, and generates chapter paragraph structure interval values.

[0018] The chapter attribution sub-module collects the attribution position data of each paragraph according to the chapter paragraph structure interval values, adjusts the paragraph attribution level, associates the chapter title number and name, summarizes the chapter structure information, integrates the uncovered paragraph blocks, and obtains the chapter division set.

[0019] As a further scheme of the present application, the rule generation module comprises:

[0020] The information extraction sub-module extracts the monitoring item title and paragraph by acquiring the chapter division set, extracts the monitoring object name, index item name and measurement unit content in the paragraph, and filters the matching items that meet the monitoring factor name, unit type and object quantity, and generates monitoring index extraction data.

[0021] The parameter filtering sub-module calls the environmental evaluation standard database based on the monitoring index extraction data, collects the standard number, index name and applicable conditions, compares the matching degree between the monitoring factor name and the standard index item, filters the standard index parameters that meet the requirements of the monitoring item type, and generates standard parameter filtering data.

[0022] The standard index parameter refers to a data set composed of the technical index name, index limit range, measurement unit type, applicable item classification and corresponding standard number associated with the monitoring object according to the monitoring item entries recorded in the environmental evaluation standard database.

[0023] The standard construction submodule filters data according to the standard parameters, integrates standard numbers, monitoring item names and standard index parameters, associates standard items under applicable conditions, combines to form corresponding comparison relationship items, establishes a standard comparison index relationship table, and obtains a comparison rule set.

[0024] As a further scheme of the present application, the specific formula for matching degree between the comparison monitoring factor name and the standard index item is:

[0025] The index matching offset value is calculated.

[0026] Wherein, Z u represents the matching offset value between the u-th monitoring factor and the standard index item, η uv represents the v-th semantic correlation factor adjustment weight of the u-th monitoring factor, X uv represents the similarity score between the v-th comparison field of the u-th monitoring factor and the standard index item, represents the arithmetic mean of the similarity scores of all comparison fields of the u-th monitoring factor, κ uv represents the structure difference correction offset value of the v-th field of the u-th monitoring factor, γ uv represents the normalized basic weight of the v-th field of the u-th monitoring factor, ζ u represents the semantic distribution density coefficient of the u-th monitoring factor, m represents the total number of fields participating in comparison, u is the monitoring factor serial number index, and v is the corresponding comparison field serial number index in the monitoring factor.

[0027] As a further scheme of the present application, the content review module comprises:

[0028] The data grouping submodule acquires the record number, sampling position, sampling time and measurement value of monitoring and calculation data by acquiring the chapter division set and the comparison rule set, compares the sampling position and time combination data, filters repeated and different items, establishes a corresponding relationship of monitoring data, and generates a data grouping matching value.

[0029] The index difference submodule extracts the unit category, sampling frequency and monitoring period value of each monitoring item based on the data grouping matching value, calculates the unit conversion coefficient, frequency ratio and period interval length, and generates an index difference value by comprehensively analyzing the difference index data.

[0030] The abnormality labeling submodule acquires the item standard index parameters in the comparison rule set according to the index difference value, analyzes the measurement value and the standard interval overlap length, calculates the logical offset degree, filters the data records whose logical offset degree exceeds the logical offset threshold, integrates an abnormality identification data set, and generates an abnormality marking set.

[0031] The logic offset threshold is set by weighted accumulation of unit conversion error value, sampling frequency change value and monitoring period offset value, and the calculation method is: logic offset threshold = unit conversion error value × 0.4 + sampling frequency change value × 0.3 + monitoring period offset value × 0.3.

[0032] As a further scheme of the present application:

[0033] The result output module obtains the abnormal mark set, analyzes the abnormal record structure, chapter position, paragraph number and corresponding standard content in the environmental impact report, calculates the logical order of the abnormal record, adjusts the abnormal data summary and establishes the error type index, integrates the chapter number, paragraph number, monitoring item, abnormal type and standard number, and outputs the report auditing data.

[0034] The report auditing data includes chapter positioning index, abnormal type classification and standard comparison table.

[0035] As a further scheme of the present application, the result output module includes:

[0036] The position arrangement submodule obtains the abnormal mark set, collects the chapter position, paragraph number, paragraph start and end position and chapter title corresponding to the abnormal record in the environmental impact report, integrates the chapter position and paragraph number index data, establishes the abnormal record position information, and generates the abnormal position index value.

[0037] The logic sorting submodule obtains the paragraph number sequence, chapter title number sequence and paragraph logic sequence code based on the abnormal position index value, calculates the logical order value of the abnormal record in the document, integrates the logical arrangement data, summarizes the abnormal data record, extracts the abnormal type identification and establishes the error type index table, and generates the abnormal order index data.

[0038] The result integration submodule collects the monitoring item and standard number in the comparison rule set corresponding to the abnormal record according to the abnormal order index data, integrates the chapter number, paragraph number, monitoring item, abnormal type and standard number, establishes the abnormal type and position correspondence table, and obtains the report auditing data.

[0039] As a further scheme of the present application, the specific formula for calculating the logical order value of the abnormal record in the document is:

[0040]

[0041] Calculate the logical order characteristic value.

[0042] Wherein, L s represents the logical order characteristic value corresponding to the current abnormal record, n s represents the number of paragraph numbers involved in the abnormal record, and pi represents the original paragraph number of the i-th paragraph in the document, q i Represents the chapter title number associated with the i-th paragraph, a i Represents the index value of the position of the i-th paragraph in the chapter, l i Represents the logical order code value of the i-th paragraph, n r represents the number of paragraphs in the standard paragraph set, p j represents the paragraph number of the jth standard paragraph, q j Represents the chapter title number of the jth standard paragraph, l j Represents the logical order code value of the jth standard paragraph, i is the index variable of the paragraph number in the current exception record, j is the index variable of the paragraph number in the standard paragraph set, s is the index number of the current exception record, and r is the index number of the standard reference paragraph set.

[0043] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0044] By performing structured recognition and eliminating duplicate information on the original text, the data purity and context consistency of subsequent processing can be significantly improved. With the help of title style comparison and section boundary analysis, the chapter content can be accurately divided and reasonably attributed, thereby enhancing the positioning accuracy of semantic content. By accurately extracting monitoring project titles and parameters, and combining with standard databases to build rule sets, the comparison process has clear standards and upper and lower limits. Combined with factors such as unit frequency period, difference analysis and consistency calculation are carried out, and then anomalies are marked and logical deviations are identified. In the entire review process, content granularity refinement, element consistency identification and pre-prompt of abnormal risks are achieved, effectively improving the meticulousness of semantic analysis and the pertinence of review judgments. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0046] Figure 1 is a system flow chart of the present invention;

[0047] Figure 2 Schematic diagram of the system framework of the present invention. DETAILED DESCRIPTION

[0048] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0049] In the embodiments of the present application, the words such as "example", "for example" are used to represent an example, illustration, or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.

[0050] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. "Of", "corresponding" and "corresponding" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.

[0051] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1, and the meanings expressed are consistent when the distinction is not emphasized.

[0052] In order to make the technical problems, technical schemes and advantages to be solved by the present application more clear, the following will be described in detail in combination with the drawings and specific embodiments.

[0053] The embodiments of the present application provide an environmental impact report aided auditing system based on semantic understanding, please refer to Figure 1 to Figure 2 The present application provides a technical scheme, an environmental impact report aided auditing system based on semantic understanding, which comprises:

[0054] A text cleaning module acquires an original text by receiving an uploaded environmental impact report, analyzes paragraph structure and layout identification, identifies environmental impact chapter features, screens literature and repeated paragraphs, adjusts structure to screen associated paragraphs, outputs preprocessed text, and passes to a chapter segmentation module;

[0055] The chapter segmentation module acquires preprocessed text, compares title styles and numbers according to the chapter format of the environmental impact report, screens title groups, calculates section boundaries, adjusts attribution to divide chapters, outputs chapter division sets, and passes to a rule generation module and a content auditing module;

[0056] The rule generation module acquires chapter division sets, extracts monitoring item titles and paragraphs, calls an environmental assessment standard database, screens monitoring parameters and standard numbers, constructs comparison rules, outputs comparison rule sets, and passes to the content auditing module;

[0057] The content review module, by obtaining the chapter division set and the comparison rule set, analyzes data consistency, filters repeated and different items, analyzes units, frequency, and period differences according to project requirements, calculates consistency, analyzes logical contradictions and deviation degrees, marks to form an abnormal marker set, and passes to the result output module;

[0058] The result output module, by obtaining the abnormal marker set, analyzes the abnormal record structure, chapter position, paragraph number, and corresponding standard content in the environmental impact report, calculates the logical order of the abnormal record, adjusts the abnormal data summary and establishes the error type index, integrates the chapter number, paragraph number, monitoring item, abnormal type, and standard number, and outputs the report review data.

[0059] The pre-processing text includes screening the literature set, removing the paragraph group, and associating the paragraph structure. The chapter division set includes title group data, section boundary parameters, and chapter attribution index. The comparison rule set includes monitoring parameter set, standard number index, and comparison condition matrix. The abnormal marker set includes difference data set, abnormal type label, and logical contradiction index. The report review data includes chapter positioning index, abnormal type classification, and standard comparison table.

[0060] Please refer to Figure 1 to Figure 2 , the text cleaning module includes:

[0061] The paragraph recognition submodule obtains the original text by receiving the uploaded environmental impact report, analyzes the starting position, paragraph length, and paragraph number of the text paragraph, detects the paragraph beginning punctuation, numbering style, and paragraph indentation form, identifies the paragraph corresponding to the chapter title of the environmental impact report, filters the paragraph type, and obtains the paragraph category identification data.

[0062] The original paragraph start position is based on the content of the uploaded environmental impact report, which consists of multiple chapters with inconsistent typesetting features and text styles. These chapters include numerically numbered sections, Chinese title sections, and other text sections. Paragraphs in such documents are not necessarily separated by clear line breaks, and often have different first line indents, number of spaces, and numbering styles, posing challenges to paragraph recognition. During the execution process, the original document structure is first parsed to extract the starting position index, line break mark status and prefix character sequence corresponding to each line of text, and these character sequences are matched with the preset numbering templates. The numbering templates cover forms such as "1.", "(I)", "1.1", "Chapter X", etc. The character comparison method is used to determine whether it belongs to the chapter number. After identification, it is used as a potential paragraph starting point; secondly, for the paragraph indentation format, the number of spaces or tabs at the beginning of each line is recorded. Through traversal statistics, the indentation threshold is set to the width of two Chinese characters (that is, 4 half-width characters) to determine whether the line belongs to the beginning of a new paragraph. If consecutive lines have the same indentation form and no number matching, they are marked as the same paragraph; thirdly, the paragraph length is limited. If the number of paragraph characters is less than 30, it is marked as a title paragraph. If it is greater than 200 characters, it is marked as a new paragraph. The middle range is further judged by semantic merging context; then the first punctuation of the paragraph is detected to determine whether it starts with symbols such as "", ":", ")", ".", etc. If such symbols exist, the content of the previous line is traced back and merged to eliminate the phenomenon of mis-segmentation; in addition, by counting the numbering patterns in consecutive similar paragraphs, the correspondence between the paragraph number and the line is extracted, and a paragraph number sequence is constructed to identify paragraph continuity and numbering jumps, and the numbering jump points are used to assist in paragraph boundary judgment; finally, matching rules are used to extract typical chapter titles such as "Introduction", "Project Overview", "Environmental Impact Analysis", etc., and a chapter-paragraph mapping table is constructed. A paragraph number dictionary is established according to the chapter hierarchy relationship to generate paragraph recognition results and output the paragraph number, starting position index and length information. The following is some sample data;

[0063] Table 1 Paragraph identification information table

[0064]

[0065] As shown in Table 1, different paragraphs are marked according to numbering, indentation, and title features, providing basic paragraph information support for subsequent content screening and structural adjustment. The final result is to obtain paragraph category identification data.

[0066] The content screening submodule screens document feature paragraphs and repeated paragraphs based on paragraph category identification data, compares paragraph start identifiers, document citation symbols, and paragraph end annotations, analyzes the degree of paragraph repetition in chapters, screens and retains paragraph content, and obtains related paragraph screening data;

[0067] The content starting part is referenced by paragraph category identification data, and the identified paragraph categories of "literature feature paragraph" and "repeated paragraph" are extracted and processed. During the execution process, first, the paragraph sequence is traversed, and the prefix character is extracted from the beginning of each text. It is judged whether it belongs to the defined identifier set, including "[1]", "(Smith, 2020)", "①" or "reference" and other features. If the matching is successful, the paragraph is marked as a reference segment. Then, the endnote identification is performed on the tail character of the paragraph, and it is detected whether it ends with "data source:", "from", "source" and other words. If it exists, it is also marked as a reference segment, and the tail feature string is recorded. For whether there are "according to xx literature", "see xx" and other statements in the middle of the paragraph, it is judged whether it has literature properties through string matching and keyword extraction. Then, the identified paragraph and other paragraphs in the chapter to which it belongs are compared in terms of repetition. The Jaccard similarity calculation method is used. Assuming that there are 10 paragraphs in a chapter, the feature word set is extracted after the text vectorization, for example, P003 contains "pollution source, monitoring factor, concentration, limit value", P005 contains "pollution source, concentration, concentration limit value", the intersection of the two is 3 words, the union set is 5 words, and the similarity is 3 / 5=0.6. If the value exceeds the threshold value of 0.5, it is marked as a repeated paragraph. Further, the paragraphs with repetition higher than the threshold value are compared by comparing the length and complexity of the beginning identifier, and the paragraph containing more information is preferentially retained, for example, the paragraph containing time, unit and reference details is retained, and the repeated part is deleted. Finally, the paragraph retention and rejection list is formed, the associated paragraph filtering structure body is constructed, and the paragraph ID and the binary identification content of whether it is retained are output.

[0068] The layout structure adjustment submodule adjusts the paragraph order and logical attribution according to the associated paragraph filtering data, identifies the cover, table of contents and reference literature layout area, reclassifies the associated paragraphs into chapter logical structures, establishes paragraph attribution blocks, and generates preprocessed text.

[0069] The processing starts with the content of the associated paragraph screening data, which records whether each paragraph is retained and its logical relationship in the document and chapter attribution. During execution, first read the paragraph list with the retention flag "1", and perform reordering operation according to its original text position and chapter mapping information. In the sorting process, prefer to retain chapter title segments (such as "Chapter 4 Environmental Protection Measures"), and arrange the text segments belonging to the chapter in ascending order of paragraph number. At the same time, identify and extract the paragraphs containing "cover", "table of contents", "references" and mark them as page number segments, and set up independent nodes in the structure tree for classification. In the specific adjustment, if a paragraph originally belongs to chapter "3.2 Water Environment Impact Analysis", its position is offset between "4.1 Environmental Protection Measures", then it needs to be renumbered as "3.2.3" and be classified under the original chapter. When updating the paragraph attribution, the logical continuity of the previous and next paragraphs needs to be judged. For example, the previous paragraph of P009 is "P008: Analysis of river water quality trend", and the next paragraph is "P010: Propose treatment suggestions". If P009 still discusses trends or data review, it should be attributed to P008, otherwise it can be a new logical paragraph. After sorting, build paragraph attribution block structure, including fields such as "chapter number", "paragraph number", "text start position", "whether it is a title segment", etc. Finally, output the structure to form the new paragraph hierarchical relationship, which provides the structure basis for subsequent pre-processing text generation, and finally get the paragraph structure adjustment result data.

[0070] Please refer to Figure 1 to Figure 2 The chapter segmentation module includes:

[0071] The title screening submodule acquires the pre-processed text, collects paragraph title data according to the chapter format of the environmental impact report, compares the number style, punctuation format and keyword structure, screens the titles in the paragraphs with chapter characteristics, calculates the length of the title line and the number of adjacent paragraphs, and generates the chapter title position distribution value.

[0072] In the execution process in the title screening submodule, first, the obtained preprocessed text needs to be analyzed, and the chapter format of the environmental impact report is combined with the title data in the paragraph to perform screening. In the execution process, the system compares the numbering style, punctuation format, and keyword structure of the title, and the specific steps include: extracting the chapter title in the text through regular expressions or other text processing tools and performing format checking. First, for the analysis of the numbering style, it checks whether the numbering of each title conforms to the standard of the report, such as "Chapter 1", "1.1 Section" and the like; then, for the checking of the punctuation format, it ensures that the position and use of punctuation marks in the title conform to the document specification, and no incorrect punctuation usage can occur, such as commas should be avoided at the end of the title; finally, the comparison of the keyword structure confirms whether there are valid keywords in the title, such as "environment", "pollution", "impact" and other keywords related to the content of the report. If the title meets the above checking standards, the system will further calculate the length of the title line and the number of adjacent paragraphs to evaluate the distribution position of the title, and generate the chapter title position distribution value. In the execution process, for example, if "Chapter 2" appears before a title, the length of the title line is 15 characters, and there are 2 lines of spacing between the adjacent paragraphs, then the generated title position distribution value is the coordinates and spacing of the title.

[0073] The structural positioning submodule collects the title-to-title line spacing paragraph number, the average number of words in the paragraph, and the indentation format data based on the chapter title position distribution value, compares the physical distance between the titles and the paragraph structure distribution rule, calculates the paragraph belonging section range according to the distance and structure relationship, and generates the chapter paragraph structure interval value;

[0074] In the structural positioning sub-module, based on the chapter title position distribution value, the physical distance between titles, the average number of words in the section, and the indentation format need to be analyzed during the process. These parameters will be derived through the calculation of the number of lines between titles, the average number of words per section, and the indentation format of each section to obtain the specific structural distribution rule. First, the number of lines between titles is obtained by measuring the number of sections between adjacent chapter titles. For example, if there are 5 sections of text between the title of Chapter 2 and the title of Chapter 3, and the average number of words per section is 120 words, then the number of lines between titles in this section is 5. Next, the average number of words in the section needs to calculate the average number of words in each paragraph. Assuming that a paragraph has 3 lines with word counts of 100, 120, and 110 respectively, the average number of words in the section is 110 words. In the indentation format analysis, the system will confirm the indentation of each section. If the first line of a paragraph is indented by 2 spaces, it will be marked as "2 spaces". By comparing this information, the system can determine the range of the section to which the paragraph belongs, and generate the chapter paragraph structure interval value according to the physical distance and structural relationship. For example, if there are 5 lines between Chapter 2 and Chapter 3, the average number of words in the section is 120 words, and the indentation format is consistent, then the structure interval value can be calculated as the structure interval of the chapter.

[0075] The chapter attribution sub-module collects the attribution position data of each section, adjusts the attribution level of the paragraph, associates the chapter title number and name, summarizes the chapter structure information, integrates the uncovered paragraph blocks, and obtains the chapter division set according to the chapter paragraph structure interval value.

[0076] During the execution process of the chapter attribution sub-module, the attribution position data of each paragraph needs to be collected according to the chapter paragraph structure interval value, the attribution level of the paragraph needs to be adjusted, and the chapter title number and name need to be associated. First, for the attribution position data of each section, the system will determine which chapter the paragraph belongs to according to the chapter paragraph structure interval value obtained earlier. For example, assuming that the chapter paragraph structure interval value indicates that the structure interval of Chapter 2 contains paragraphs 1 to 10, the system will attribute these paragraphs to Chapter 2. Then, the system will adjust the level of the paragraph according to the calculation result to ensure that the paragraph meets the level specification in the report. If a paragraph belongs to a section, it will be attributed to the section level, and if it belongs to a chapter, it will be attributed to the chapter level. Finally, by associating the number and name of the chapter title, the system can form a complete chapter structure information table and integrate the uncovered paragraph blocks to ensure that each paragraph is accurately attributed to the appropriate chapter. For example, if paragraph 12 is attributed to Chapter 3, the attribution information of paragraph 12 will include the number and name of Chapter 3;

[0077] Table 2 Chapter title position distribution table

[0078] Title Number Title Name Line Length (Characters) Number of Adjacent Paragraphs Title Position Distribution Value Chapter 1 Chapter Title 1 15 2 (1,2) Chapter 2 Chapter Title 2 20 3 (2,3) Chapter 3 Chapter Title 3 25 4 (3,4)

[0079] As shown in Table 2, the table lists the numbering, name, title line length, adjacent paragraph number and position distribution value of three chapter titles.

[0080] Please refer to Figure 1 to Figure 2 , the rule generation module comprises:

[0081] The information extraction submodule extracts the monitoring item title and paragraph by obtaining the chapter division set, extracts the monitoring object name, index item name and measurement unit content in the paragraph, and screens the matching items conforming to the monitoring factor name, unit type and object quantity to generate monitoring index extraction data;

[0082] The information extraction submodule obtains the text data after chapter segmentation, first reads the starting position and ending position of each paragraph, judges whether the paragraph contains environmental monitoring information according to the keyword field (such as "project name", "monitoring item", "environmental index") appearing at the beginning of the paragraph, and extracts the first character string (such as the first i=5 to j=10 characters) after the preliminary screening of each paragraph. Determine whether it is a title type sentence, for example, if the first k paragraph starts with "air pollution monitoring item", it can be initially confirmed as a title, and the next step is processed. Perform word segmentation analysis on each sentence in the paragraph, and search for keywords "air", "surface water", "groundwater", "soil" in turn. If it matches successfully, mark the paragraph as a monitoring information paragraph. Then read each sentence in it, identify the word group containing the unit field "μg / m 3 ", "mg / L", "mg / m 3 ", and extract the prepositional phrase as the monitoring factor name, such as the short sentence "PM2.5 concentration is 35 μg / m 3 ", which can be parsed to obtain the factor name "PM2.5", the value x=35, and the unit μg / m 3 . Then determine whether the factor is a legal monitoring item. The list method can be used to determine x i ∈L, where L={PM2.5, SO2, NO2, O3, CO}. If it matches, record it, match the unit field, and handle it according to the conversion rule when the unit is inconsistent. μg / m 3 →mg / m 3 Use the transformation If there are multiple objects in the paragraph (such as "PM2.5 in the air is 35 μg / m 3 , COD in surface water is 2.3 mg / L"), it is necessary to construct data structures (O1, P1, U1, V1) and (O2, P2, U2, V2) respectively, that is, object, factor, unit, and value four-tuple structure. After each four-tuple is constructed, uniform format verification is performed. If the unit field U i is illegal or the value V i is missing, it is discarded, and finally an array is formed Each is a legal monitoring structure item, and intersects with the factor list S in the standard library, and outputs structured monitoring index information;

[0083] Table 3 Ambient air monitoring index table

[0084] Monitoring Items Monitoring Objects Index Unit Sampling Values Reference Threshold Value PM2.5 Air pg / m 3 ]] 35.0 75.0 SO2 Air pg / m 3 ]] 50.0 60.0 NO2 Air pg / m 3 ]] 45.0 80.0 <![CDATA[O3]]> Air pg / m 3 ]] 120.0 160.0 CO Air mg / m 3 ]]> 1.2 4.0

[0085] As shown in Table 3, each monitoring data includes factor name, object, unit, sampling value and reference standard value, which can be used by the comparison module.

[0086] The parameter screening submodule extracts data based on the monitoring index, calls the environmental evaluation standard database, collects standard numbers, index names and applicable conditions, compares the matching degree between the monitoring factor name and the standard index item, screens the standard index parameters that meet the type requirements of the monitoring project, and generates standard parameter screening data;

[0087] The specific formula for comparing the matching degree between the monitoring factor name and the standard index item is:

[0088]

[0089] Calculate the index matching offset value;

[0090] Wherein, Z u represents the matching offset value between the u-th monitoring factor and the standard index item, η uv represents the v-th semantic correlation factor adjustment weight of the u-th monitoring factor, X uv represents the similarity score between the v-th comparison field of the u-th monitoring factor and the standard index item, represents the arithmetic mean of the similarity scores of all comparison fields of the u-th monitoring factor, κ uv represents the structure difference correction offset value of the v-th field of the u-th monitoring factor, γ uv represents the normalized basic weight of the v-th field of the u-th monitoring factor, ζ u represents the semantic distribution density coefficient of the u-th monitoring factor, m represents the total number of fields participating in comparison, and u is the monitoring factor serial number index, and v is the corresponding comparison field serial number index in the monitoring factor;

[0091] Formula:

[0092]

[0093] Formula details and formula calculation derivation process:

[0094] The formula is used to calculate the matching offset value between the monitoring factor and the standard index item, and the result is used to screen the parameters with higher matching degree in the monitoring project, to support the formation of standard index parameter screening data;

[0095] Parameter meaning and setting value:

[0096] η uv is the semantic correlation factor adjustment weight, obtained by matching the level evaluation of the domain expert semantic library, set to 0.25;

[0097] X uv is the similarity score, calculated by a text similarity model such as SimHash algorithm, set X 11 = 0.82, X 13 = 0.68;

[0098] is the average similarity score of the monitoring factor u, calculated by the arithmetic mean of field X uv ,

[0099]

[0100] κ uv is the structure difference correction offset value, derived from the field format structure matching degree calculation, the structure consistency level is extracted by label sequence comparison, set κ 11 = 0.08, κ 13 = 0.05;

[0101] γ uv is the basic weight, determined by the frequency ratio of the field in the monitoring task, the frequency ratio is based on the statistical results of the monitoring history database, the field with the highest frequency is assigned a weight of 1, and the frequency is linearly reduced and normalized, set γ 11 = 0.9, γ 13 = 0.7;

[0102] ζ u is the semantic distribution density coefficient, representing the semantic coverage degree of the monitoring factor u in the index parameter knowledge base, calculated and normalized according to the number of independent terms of the monitoring factor in the standard index term, set ζ1= 0.15;

[0103] m is the total number of fields participating in comparison, set to 3;

[0104] Substitute the parameters into the formula for calculation:

[0105]

[0106] |0.0383|+|0.0283|+|-0.0042|=0.0383+0.0283+0.0042=0.0708;

[0107] γ 11 + ζ1 = 0.9 + 0.15 = 1.05;

[0108] (1.05) 2 = 1.1025;

[0109] γ 12 + ζ1= 0.8 + 0.15 = 0.95;

[0110] (0.95) 2 = 0.9025;

[0111] γ 13 + ζ1= 0.7 + 0.15 = 0.85;

[0112] (0.85) 2 = 0.7225;

[0113]

[0114] Substitute into the formula to calculate the matching offset value:

[0115]

[0116] The result Z1=0.0429 indicates the matching offset degree of the first monitoring factor after the structural difference, semantic offset and basic weight weighting calculation with the standard index item. The closer the result is to 0, the more sufficient the matching is, and the higher the matching accuracy is. The value is used to drive the output generation of the standard index parameter screening step;

[0117] The standard index parameter refers to the data set composed of technical index name, index limit range, measurement unit category, applicable project classification and corresponding standard number associated with the monitoring object according to the monitoring project item recorded in the environmental assessment standard database.

[0118] The standard construction submodule integrates the standard number, monitoring project name and standard index parameter according to the standard parameter screening data, associates the standard items under the applicable conditions, combines to form the corresponding comparison relationship items, establishes the standard comparison index relationship table, and obtains the comparison rule set;

[0119] The standard construction submodule reads the standard number set S={s1, s2, …, s n}, each standard item contains number, applicable range, project name and limit value information, and the number field s i such as "GB3095-2012" identifies the standard version, and then reads the field "applicable medium". Through set judgment M∈{air, surface water, groundwater, soil}, if M=air, only the items with "air" as the object in paragraph 1 are allowed to match. Each monitoring project name p j in the standard item is combined with the unit u j and the limit value l j to form a judgment condition, and the judgment condition is used to determine whether the monitoring item in the monitoring project item set matches the standard item.j = P i ) ∧ (u j = U i ) holds, a comparison operation is entered, if the units are different, conversion rules are adopted: if u j = mg / m 3 , then a difference calculation is performed: Δ i = x i - l j , a judgment rule is set: if Δ i > 0, it is an over-limit item; if Δ i ≤ 0, it is a qualified item. Taking PM2.5 as an example, the sampling value is x = 35, the standard value is l = 75, then Δ = 35 - 75 = -40, it is judged to be qualified, and the result is written into the index table structure R = {(s i , p j , l j , x i , Δ i , state)}, where the state is "qualified" or "over-limit", and each index table record establishes a one-to-many relationship: mapping: p j → {s1, s2, …}, which is convenient for subsequent standard updating and tracing. The comparison index result is directly called and compared with the corresponding fields of the sampling value and the reference threshold value in Table 3 above, to form a complete standard comparison structure.

[0120] Please refer to Figure 1 to Figure 2 , the content review module includes:

[0121] The data grouping submodule collects the record number, sampling location, sampling time and measurement value of the monitoring and calculation data by obtaining the chapter division set and comparison rule set, compares the sampling location and time combination data, filters repeated and different items, establishes the corresponding relationship of the monitoring data, and generates data grouping matching values;

[0122] The data grouping submodule obtains the chapter division set as input, first extracts the field values in the monitoring record one by one, including the record number (such as R001, R002, …), the sampling location (such as "point A"), the sampling time (such as "2025-01-01 08:00") and the measurement value (such as 2.5 mg / L), and performs string consistency check on the record number and the sampling location. If the location i = location j and |time i - time j | < 60s, it is marked as a repeated sampling group. For example, R001 and R003 are both sampled from point A, and the time is completely consistent, so they are judged as repeated records. Then the measurement values of the two records are judged. If the value i = value j , it is set as a completely repeated group and is merged, if |value i- value j | 0, still record as difference group, such as R005 value is 2.6 mg / L and R001 value is 2.5 mg / L, the difference is 0.1, which needs to be set as a difference data group, the sampling time difference is 5 minutes, which is less than 1 hour, and can be classified as a time approximate group; by comparing each group of field position, time, and value triplets, a hash structure or sequence list is used to build a group identification, and each group is assigned a number G k , for example, G1 = {R001, R003}, G2 = {R005}, a group matching value set containing group number, record number array, sampling consistency mark (completely consistent / position consistent and time different), and measurement value difference index is generated, which is used for downstream difference calculation and abnormality judgment processing;

[0123] Table 4 monitoring data record table

[0124] Record Number Sampling Location Sampling Time Measurement Value (mg / L) R001 Point A 2025-01-0108:00 2.5 R002 Point B 2025-01-0108:10 3.1 R003 Point A 2025-01-0108:00 2.5 R004 Point C 2025-01-0108:20 3.3 R005 Point A 2025-01-0108:05 2.6

[0125] As shown in Table 4, record numbers R001 and R003 belong to a completely repeated group, R005 has the same sampling position as the previous two but has a slight difference in time, and therefore is classified as a difference data group.

[0126] The index difference submodule extracts the unit category, sampling frequency, and monitoring period values of each monitoring item based on the data group matching values, calculates the unit conversion coefficient, frequency ratio, and period interval length, and generates the index difference value by integrating the difference index data;

[0127] The index difference submodule takes the data group matching values generated as input, extracts the unit category (such as mg / L, μg / L), sampling frequency (number of daily samples), and monitoring period (start and end time interval) values of each monitoring item, and calculates the unit conversion coefficient through known proportional conversion, for example, the conversion coefficient of μg / L→mg / L is k=0.001, and if the data A unit is 450 μg / L, it is converted to 450×0.001=0.45 mg / L; the sampling frequency ratio β is calculated by taking the ratio of the two frequencies, for example, if frequency 1=3 times / day and frequency 2=1 time / day, then β=3, and the monitoring period difference ΔT is the difference between the monitoring periods of the two groups, for example, group 1: 10 days, group The final difference value δ is calculated as follows: δ=∈×0.4+β×0.3+ΔT×0.3, where: ∈: the numerical difference after unit conversion, such as 0.05, β: the sampling frequency difference (such as 3-1=2), ΔT: the period difference (such as 3 days), and the example values are substituted as follows: δ=0.05×0.4+2×0.3+3×0.3=0.02+0.6+0.9=1.52, and the index difference value δ=1.52 is used for subsequent abnormality deviation degree judgment.

[0128] The anomaly labeling submodule collects the project standard index parameters in the comparison rule set according to the index difference value, analyzes the overlap length of the measured value and the standard interval, calculates the logical offset degree, screens the data records whose logical offset degree exceeds the logical offset threshold, integrates the anomaly identification data set, and generates the anomaly marking set;

[0129] The anomaly labeling submodule performs overlap analysis by comparing the measured value with the project standard parameter interval in the comparison rule set (for example, the standard range of ammonia nitrogen is [0.5, 3.0] mg / L). Assuming that the current measured value is x, the overlap length with the standard interval is defined as: L = max (0, min (x, u) - max (x, l)), where x: measured value, l = 0.5, u = 3.0: lower and upper limits, and if x = 2.5, then L = min (2.5, 3.0) - max (2.5, 0.5) = 2.5 - 0.5 = 2.0. The logical offset degree λ is calculated: If the standard interval is [0.5, 3.0] and the length is 2.5, then The calculation method of the logical offset threshold θ is: θ = ∈ × 0.4 + β × 0.3 + ΔT × 0.3. Substituting the aforementioned numerical values: θ = 0.05 × 0.4 + 2 × 0.3 + 3 × 0.3 = 0.02 + 0.6 + 0.9 = 1.52. If λ > θ, the record is marked as abnormal. For example, x = 6.2 mg / L, which is not within the standard interval, the logical offset degree λ = 1.0. Although λ = 1.0 < θ = 1.52, the logical offset does not exceed the threshold, and it is not marked as abnormal. However, if θ is reduced or the measured value deviates seriously resulting in λ > θ, an abnormal flag is needed to be marked, and the record number and the logical offset result are combined to form an anomaly marking record set for downstream audit process, forming a traceable anomaly detection chain;

[0130] The logical offset threshold is set by weighted accumulation of the unit conversion error value, the sampling frequency change value, and the monitoring cycle offset value. The calculation method is: logical offset threshold = unit conversion error value × 0.4 + sampling frequency change value × 0.3 + monitoring cycle offset value × 0.3.

[0131] Please refer to Figure 1 to Figure 2 The result output module includes:

[0132] The position arrangement submodule acquires the anomaly marking set, collects the chapter position, paragraph number, paragraph start and end position, and chapter title corresponding to the abnormal record in the environmental impact report, integrates the chapter position and paragraph number index data, establishes the abnormal record position information, and generates the abnormal position index value.

[0133] The position arrangement submodule acquires the anomaly marking set, which contains an abnormal identifier ∈ i First, all abnormal items are numbered and indexed. Assuming that the anomaly set is {∈1, ∈2, …, ∈n read the index information of the chapter where the anomaly is located, extract the chapter number field C i , the paragraph number field P i , match the chapter title through the full-text location index file of the document, obtain the starting position of the label "Chapter x" and "chapter name" through keyword matching, parse the chapter title and its position (S i , E i ), map the text paragraph number P i to the actual character starting position and ending position, perform start and end position cutting on each paragraph number, read the data within the character index range, and confirm from the original text that it is the paragraph corresponding to the anomaly record, for example, "paragraph 12" is positioned at character positions 4523 to 4576, which appears in the corresponding chapter "Current Monitoring Results", then record (C i , P i , S i , E i , T i ), where T i is the chapter title content. In the processing process, it needs to be judged whether multiple anomaly items are concentrated in the same chapter. If there are multiple anomaly items corresponding to the same chapter number C i , rearrange them in ascending order of paragraph number, and perform continuity judgment on the starting value S i . If the starting values of any two adjacent anomalies differ by less than 200 characters, it is judged to be the same position cluster, and the anomaly cluster number field G i is set. The table is appended with a field to record the anomaly belonging group. Finally, a structured data table is established through the seven fields of anomaly item number, chapter number, paragraph number, starting and ending position, chapter title, cluster number, which serves as the basis data for subsequent anomaly analysis and visualization marking. The structure is shown in Table 5.

[0134] Table 5: Chapter position information table of anomaly record

[0135] Chapter Number Paragraph Number Starting Position Ending Position Chapter Title Chapter 3 12 4523 4576 Current Monitoring Results Chapter 4 8 7921 7999 Environmental Impact Analysis Chapter 5 21 11230 11312 Comparison Analysis

[0136] Referring to Table 5, the chapter number and start and end character positions corresponding to each anomaly record in the document have been determined, and the position information can be used for subsequent comparison and tracking processing.

[0137] The logical sorting submodule obtains the paragraph number sequence, chapter title number sequence, and paragraph logical order code based on the anomaly position index value, calculates the logical order value of the anomaly record in the document, integrates the logical arrangement data, summarizes the anomaly data record, extracts the anomaly type identifier and establishes the error type index table, and generates the anomaly order index data.

[0138] The specific formula for calculating the logical order value of the anomaly record in the document is:

[0139]

[0140] The logical order characteristic value of the calculation logic sequence;

[0141] Wherein, L s represents the logical order characteristic value corresponding to the current abnormal record, n s represents the number of paragraph numbers involved in the abnormal record, p i represents the original paragraph number of the i-th paragraph in the document, q i represents the chapter title number associated with the i-th paragraph, a i represents the position index value of the i-th paragraph in the chapter, l i represents the logical order encoding value of the i-th paragraph, n r represents the number of paragraphs in the standard paragraph set, p j represents the paragraph number of the j-th standard paragraph, q j represents the chapter title number of the j-th standard paragraph, l j represents the logical order encoding value of the j-th standard paragraph, i is the index variable of the paragraph number in the current abnormal record, j is the index variable of the paragraph number in the standard paragraph set, s is the index number of the current abnormal record, and r is the index number of the standard reference paragraph set;

[0142] Formula:

[0143]

[0144] Formula details and formula calculation derivation process:

[0145] The formula is used to calculate the logical order characteristic value of the abnormal record in the document, and the result is used to integrate the logical arrangement data, summarize the abnormal data records, extract the abnormal type identification and establish the error type index table, and generate the abnormal order index data;

[0146] Parameter meaning and setting value:

[0147] n s : the number of paragraph numbers involved in the abnormal record, obtained by counting the number of paragraphs contained in the abnormal record, and the setting value is 3;

[0148] p i : the original paragraph number of the i-th abnormal paragraph in the document, obtained by document structure analysis, and the setting value is 12, 15, and 18;

[0149] q i : the chapter title number of the i-th abnormal paragraph, obtained by document structure analysis, and the setting value is 2, 2, and 3;

[0150] a i: The position index value of the i-th abnormal paragraph in the chapter, obtained by document structure analysis, and set to 3, 5, 2;

[0151] l i : The logical order coding value of the i-th abnormal paragraph, obtained by document structure analysis, and set to 4, 5, 6;

[0152] n r : The total number of paragraphs contained in the standard paragraph set, obtained by counting the number of paragraphs in the standard paragraph set, and set to 3;

[0153] p j : The original paragraph number of the j-th standard paragraph, obtained by document structure analysis, and set to 10, 13, 16;

[0154] q j : The chapter title number of the j-th standard paragraph, obtained by document structure analysis, and set to 2, 2, 3;

[0155] l j : The logical order coding value of the j-th standard paragraph, obtained by document structure analysis, and set to 3, 4, 5;

[0156] Substitute the parameters into the formula to calculate:

[0157] Calculate the average value of the abnormal record part:

[0158]

[0159]

[0160] Calculate the average value of the standard paragraph set part:

[0161]

[0162] Calculate the value of L s :

[0163] L s = |10.0033-6.0667| = 3.9366;

[0164] The result 3.9366 indicates the degree of difference between the logical order characteristic value of the current abnormal record and the average value of the standard paragraph set. This result is used to integrate the logical arrangement data, aggregate the abnormal data records, extract the abnormal type identifier, and establish an error type index table, and generate abnormal order index data.

[0165] The result integration submodule indexes data according to the abnormal sequence, collects the monitoring items corresponding to the abnormal records and the standard numbers in the comparison rule set, integrates the chapter number, paragraph number, monitoring item, abnormality type and standard number, establishes a corresponding table between abnormality type and location, and obtains report review data;

[0166] The result integration submodule further reads the monitoring item p corresponding to each abnormal record based on the aforementioned location information. i , extract the project name, paragraph number and chapter number, and match the standard number s corresponding to the project through the standard comparison structure generated by the comparison module i ,The execution process is as follows: First, locate the abnormal type field, which is obtained by comparing it with the standard limit field, that is, when the sampling value x i With limit l i Satisfy x i >l i Set the exception type to "out of limit" when the exception occurs, otherwise it is "qualified", and set the exception type field e in the structure i , then call the chapter number field C in the paragraph position information table i , paragraph number P i , combined with the project field p i , abnormal field e i and Standard Number fields i Constructing a quintuple structure (C i ,P i ,p i ,e i ,s i ), write the abnormality type and location comparison table in sequence. During the construction process, it is necessary to determine whether there are multiple abnormal items of the same monitoring project under the same chapter number. If so, keep all records and append the sequence number field k i For example, if the paragraph numbers 8, 9, and 11 of "SO2" in "Chapter 4" all exceed the standard, the record is k = 1, 2, 3, and a composite primary key index (C i ,P i ,p i ,k i ) to ensure that data is not repeated or lost. In the output structure, each abnormal type field can be set according to enumeration such as "overlimit", "low value", "no value", etc. By comparing the field value classification, the abnormal type and position comparison structure is finally output as a data interface format for report review. In the comparison structure, for example, the PM2.5 monitoring value is x = 94μg / m 3 The corresponding standard limit is l=75μg / m 3 , the judgment result is "out of limit", and write the type field e i = Exceeded, create a complete record structure by combining the paragraph number "12" with the chapter number "Chapter 3".

[0167] The above-described embodiments can be implemented in part or in whole through software, hardware (e.g., circuitry), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When loaded and executed by a computer, the computer instructions or computer programs cause the computer to perform the processes or functions described in the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, such as from a website site, a computer, a server, or a data center to another website site, a computer, a server, or a data center, through a wired (e.g., infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. that includes one or more collections of available media. The available media can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.

[0168] It should be understood that the term "and / or" used herein is merely to describe an associated relationship between associated objects, and can represent three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but can also represent an "and / or" relationship, which can be understood in the context before and after it.

[0169] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be singular or plural.

[0170] It should be understood that in various embodiments of the present application, the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0171] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0172] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0173] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is merely logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0174] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0175] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.

[0176] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0177] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. The environmental impact report auxiliary review system based on semantic understanding is characterized by: The system comprises: The text cleaning module receives the original text of the uploaded environmental impact report, analyzes the paragraph structure and layout identification, identifies the characteristics of the environmental impact chapter, selects references and repeated paragraphs, adjusts the structure and selects related paragraphs, outputs the preprocessed text, and passes it to the chapter segmentation module; A chapter segmentation module obtains the preprocessed text, compares the title style and numbering according to the chapter format of the environmental impact report, filters the title group, calculates the section boundary, adjusts the attribution division chapter, outputs the chapter division set, and passes it to the rule generation module and the content review module; A rule generation module obtains the chapter division set, extracts monitoring project titles and paragraphs, calls the environmental assessment standard database, filters monitoring parameters and standard numbers, constructs comparison rules, outputs the comparison rule set, and passes it to the content review module; The content review module obtains the chapter division set and the comparison rule set, analyzes data consistency, filters out duplicate and different items, analyzes unit, frequency, and period differences according to project requirements, calculates consistency, analyzes logical contradictions and deviation levels, annotates to form an abnormal tag set, and passes it to the result output module.

2. The environmental impact report auxiliary review system based on semantic understanding according to claim 1 is characterized in that: The preprocessed text includes a screening document set, a deduplicated paragraph group, and an associated paragraph structure; the chapter division set includes title group data, section boundary parameters, and a chapter attribution index; the comparison rule set includes a monitoring parameter set, a standard number index, and a comparison condition matrix; and the abnormal marking set includes a difference data set, an abnormal type label, and a logical contradiction indicator.

3. The environmental impact report auxiliary review system based on semantic understanding according to claim 1 is characterized in that: The text cleaning module includes: The paragraph recognition submodule receives the original text of the uploaded environmental impact report, analyzes the starting position, length and number of the text paragraphs, detects the punctuation at the beginning of the paragraph, the numbering style and the indentation form of the paragraph, identifies the paragraph corresponding to the chapter title of the environmental impact report, selects the paragraph type, and obtains the paragraph category identification data; A content screening submodule, based on the paragraph category identification data, screens document feature paragraphs and repeated paragraphs, compares paragraph start identifiers, document citation symbols, and paragraph end annotations, analyzes the degree of paragraph repetition in the chapter, screens and retains paragraph content, and obtains related paragraph screening data; The layout structure adjustment submodule filters data according to the related paragraphs, adjusts the paragraph order and logical attribution, identifies the cover, table of contents, and reference layout areas, reclassifies the related paragraphs into the chapter logical structure, establishes paragraph attribution blocks, and generates preprocessed text.

4. The environmental impact report auxiliary review system based on semantic understanding according to claim 3 is characterized in that: The chapter segmentation module includes: The title screening submodule obtains the pre-processed text, collects paragraph title data according to the chapter format of the environmental impact report, compares the numbering style, punctuation format and keyword structure, screens the titles with chapter characteristics in the paragraphs, calculates the length of the title line and the number of adjacent paragraphs, and generates a chapter title position distribution value; The structure positioning submodule collects the number of paragraphs with line spacing between titles, the average number of words within a paragraph, and the indentation format data based on the chapter title position distribution value, compares the physical distance between titles with the paragraph structure distribution pattern, calculates the paragraph belonging segment range based on the distance and structure relationship, and generates the chapter paragraph structure interval value; The chapter attribution submodule collects the attribution position data of each paragraph according to the chapter paragraph structure interval value, adjusts the paragraph attribution level, associates the chapter title number and name, summarizes the chapter structure information, integrates the uncovered paragraph blocks, and obtains the chapter division set.

5. The environmental impact report auxiliary review system based on semantic understanding according to claim 4 is characterized in that: The rule generation module includes: The information extraction submodule obtains the chapter division set, extracts the monitoring item titles and paragraphs, extracts the monitoring object names, indicator item names and measurement unit contents in the paragraphs, and screens the matching items that meet the monitoring factor names, unit categories and object quantities to generate monitoring indicator extraction data; The parameter screening submodule extracts data based on the monitoring indicators, calls the environmental assessment standard database, collects the standard number, indicator name and applicable conditions, compares the matching degree between the monitoring factor name and the standard indicator item, screens the standard indicator parameters that meet the requirements of the monitoring project type, and generates standard parameter screening data; The standard indicator parameters refer to a data set consisting of the technical indicator name, indicator limit range, measurement unit category, applicable project classification and corresponding standard number associated with the monitoring object, obtained based on the monitoring project items recorded in the environmental assessment standard database; The standard construction submodule filters data according to the standard parameters, integrates the standard number, monitoring project name and standard indicator parameters, associates the standard items under the applicable conditions, combines them to form corresponding comparison relationship items, establishes a standard comparison index relationship table, and obtains a comparison rule set.

6. The environmental impact report auxiliary review system based on semantic understanding according to claim 5 is characterized in that: The specific formula for comparing the matching degree between the monitoring factor name and the standard indicator item is: Calculate the indicator matching deviation value; Among them, Z u Represents the matching deviation value between the u-th monitoring factor and the standard indicator item, η uv represents the weight adjustment of the vth semantic relevance factor of the uth monitoring factor, X uv Represents the similarity score between the vth comparison field of the uth monitoring factor and the standard indicator item, represents the arithmetic mean of the similarity scores of all comparison fields of the u-th monitoring factor, κ uv represents the structural difference correction offset value of the vth field of the uth monitoring factor, γ uv represents the normalized basic weight of the vth field of the uth monitoring factor, ζ u Represents the semantic distribution density coefficient of the u-th monitoring factor, m represents the total number of fields involved in the comparison, u is the monitoring factor serial number index, and v is the corresponding comparison field serial number index in the monitoring factor.

7. The environmental impact report auxiliary review system based on semantic understanding according to claim 5 is characterized in that: The content review module includes: The data grouping submodule acquires the chapter division set and the comparison rule set, collects the record number, sampling location, sampling time and measurement value of the monitoring and calculation data, compares the sampling location and time combination data, filters out duplicate and difference items, establishes a corresponding relationship between the monitoring data, and generates a data grouping matching value; The indicator difference submodule extracts the unit category, sampling frequency and monitoring period value of each monitoring item based on the data group matching value, calculates the unit conversion coefficient, frequency ratio and period interval length, and integrates the difference indicator data to generate the indicator difference value; The anomaly marking submodule collects the standard indicator parameters of the project in the comparison rule set according to the indicator difference value, analyzes the overlapping length of the measurement value and the standard interval, calculates the logical deviation, filters the data records whose logical deviation exceeds the logical deviation threshold, integrates the anomaly identification data set, and generates an anomaly marking set; The logic offset threshold is set by weighted accumulation of the comprehensive unit conversion error value, the sampling frequency change value and the monitoring period offset value. The calculation method is: logic offset threshold = unit conversion error value × 0.4 + sampling frequency change value × 0.3 + monitoring period offset value × 0.

3.

8. The environmental impact report auxiliary review system based on semantic understanding according to claim 1 is characterized by: The result output module obtains the abnormality mark set, analyzes the abnormality record structure, chapter position, paragraph number and corresponding standard content in the environmental impact report, calculates the logical order of the abnormality record, adjusts the abnormality data summary and establishes an error type index, integrates the chapter number, paragraph number, monitoring item, abnormality type and standard number, and outputs the report review data; The report review data includes chapter location index, abnormality type classification, and standard comparison table.

9. The environmental impact report auxiliary review system based on semantic understanding according to claim 8 is characterized in that: The result output module includes: The position arrangement submodule obtains the abnormality mark set, collects the chapter position, paragraph number, paragraph start and end position and chapter title corresponding to the abnormal record in the environmental impact report, integrates the chapter position and paragraph number index data, establishes the abnormal record position information, and generates the abnormal position index value; A logic sorting submodule, based on the abnormal position index value, obtains the paragraph number sequence, chapter title number sequence and paragraph logical order code, calculates the logical order value of the abnormal record in the document, integrates the logical arrangement data, summarizes the abnormal data records, extracts the abnormal type identifier and establishes the error type index table, and generates abnormal sequence index data; The result integration submodule collects the monitoring items and standard numbers in the comparison rule set corresponding to the abnormal records according to the abnormal sequence index data, integrates the chapter number, paragraph number, monitoring item, abnormal type and standard number, establishes a corresponding table between abnormal type and location, and obtains report review data.

10. The environmental impact report auxiliary review system based on semantic understanding according to claim 9 is characterized in that: The specific formula for calculating the logical order value of the abnormal record in the document is: Calculate logical sequence eigenvalues; Among them, L s Represents the logical sequence characteristic value corresponding to the current abnormal record, n s Represents the number of paragraphs involved in the abnormal record, p i represents the original paragraph number of the i-th paragraph in the document, q i Represents the chapter title number associated with the i-th paragraph, a i Represents the index value of the position of the i-th paragraph in the chapter, l i Represents the logical order code value of the i-th paragraph, n r represents the number of paragraphs in the standard paragraph set, p j represents the paragraph number of the jth standard paragraph, q j Represents the chapter title number of the jth standard paragraph, l j Represents the logical order code value of the jth standard paragraph, i is the index variable of the paragraph number in the current exception record, j is the index variable of the paragraph number in the standard paragraph set, s is the index number of the current exception record, and r is the index number of the standard reference paragraph set.

Citation Information

Cited By

  • Authentication whole process efficiency improving method

    CN121365986A

  • Environmental assessment report essential factor intelligent review method for solving large model context limitation

    CN121681837A

  • Multi-dimensional agent review method and system for aviation evaluation permission report

    CN122335233A