A method and system for intelligent recognition of environmental evaluation forms
Through optical character recognition and attribute labeling technology, image acquisition, preprocessing, segmentation and recognition of environmental evaluation forms is solved, and the problem of low recognition efficiency and accuracy in traditional methods is achieved, and efficient and accurate processing of environmental evaluation forms is achieved.
Patent Information
- Application Number
- CN202411393243.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-10-08
AI Technical Summary
In traditional environmental evaluation form processing methods, the limitations of algorithms and models lead to low recognition efficiency and accuracy, and insufficient image data recognition and segmentation accuracy.
The image segmentation results are identified through optical character recognition technology, and the recognition results are evaluated and corrected based on attribute annotation. Combined with image acquisition, preprocessing, image segmentation and natural language processing technologies, key information is extracted and stored in a structured form.
It realizes efficient and accurate processing of environmental assessment forms, ensures the integrity and consistency of data, and provides technical support for environmental protection work.
Smart Images

Figure CN119360399B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, specifically to the field of image recognition and optical character recognition technology, and in particular to a method and system for intelligent recognition of environmental assessment forms. Background Art
[0002] In environmental protection work, intelligent processing of environmental assessment forms is crucial for data analysis and decision-making. Traditional environmental assessment form processing relies primarily on computer algorithms and pre-set programs. While this approach avoids extensive manual work, in practice, computer processing of environmental assessment forms still presents challenges that affect form recognition efficiency and accuracy, such as algorithm and model limitations and insufficient image data recognition and segmentation accuracy.
[0003] In this context, an intelligent recognition method and system for environmental assessment forms came into being. The system aims to identify image segmentation results through optical character recognition technology, and evaluate and correct the recognition results based on attribute annotation, so as to solve the technical problems of low recognition efficiency and accuracy caused by the diversity of form content. Summary of the Invention
[0004] This application provides an intelligent recognition method and system for environmental evaluation forms, aiming to identify image segmentation results through optical character recognition technology, and evaluate and correct the recognition results based on attribute labeling, so as to solve the technical problem of low recognition efficiency and accuracy caused by the diversity of form content.
[0005] In view of the above problems, the present application provides an environmental evaluation form intelligent recognition method and system.
[0006] The first aspect disclosed in the present application provides an intelligent recognition method for environmental assessment forms, the method comprising: using an image acquisition device to acquire image information of an environmental assessment form according to preset clarity requirements, and constructing an environmental assessment form image library; preprocessing the images in the environmental assessment form image library, including denoising and edge detection; using image segmentation technology to segment the preprocessed environmental assessment form image according to preset segmentation requirements, and annotating according to attributes corresponding to the segmentation results; using optical character recognition technology to sequentially identify the segmentation results of the environmental assessment form to obtain recognition results; based on the attribute annotation, evaluating and correcting the recognition results to obtain a corrected recognition result; according to the structural content of the environmental assessment form, extracting key information from the corrected recognition result through natural language processing technology, and storing it in a structured form based on the structural content.
[0007] Another aspect disclosed in the present application provides an intelligent recognition system for environmental assessment forms, the system comprising: an information acquisition module, the information acquisition module being used to use an image acquisition device to collect image information of an environmental assessment form according to preset clarity requirements and construct an environmental assessment form image library; a preprocessing module, the preprocessing module being used to preprocess images in the environmental assessment form image library, including denoising and edge detection; an image segmentation module, the image segmentation module being used to use image segmentation technology to segment the preprocessed environmental assessment form image according to preset segmentation requirements and annotate according to attributes corresponding to the segmentation results; a segmentation and recognition module, the segmentation and recognition module being used to use optical character recognition technology to sequentially identify the segmentation results of the environmental assessment form to obtain recognition results; an evaluation and correction module, the evaluation and correction module being used to evaluate and correct the recognition results based on attribute annotation to obtain corrected recognition results; an information extraction module, the information extraction module being used to extract key information from the corrected recognition results according to the structural content of the environmental assessment form through natural language processing technology, and store the key information in a structured form based on the structural content.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0009] The aforementioned intelligent recognition method for environmental assessment forms uses image acquisition equipment to accurately capture images of environmental assessment forms according to established clarity standards and systematically organizes these images into an image library. Subsequently, the form images in the image library undergo preprocessing. Denoising effectively removes any interfering factors from the image. Edge detection technology precisely defines the various components of the form, laying a solid foundation for subsequent processing. Image segmentation technology then divides the preprocessed form image into distinct logical regions based on pre-set segmentation rules. Each region is precisely labeled based on its content attributes, providing clear guidance for subsequent recognition. Furthermore, optical character recognition technology accurately recognizes text within the segmented form regions, converting the text in the image into computer-processable text data, providing raw material for subsequent data analysis. To ensure the accuracy of the recognition results, the recognition results are rigorously evaluated and corrected based on previously annotated attribute information. By comparing the annotated information with the recognition results, potential errors are promptly identified and corrected, ensuring data accuracy. Finally, natural language processing technology is used to extract key information from the corrected recognition results based on the structural characteristics of the environmental assessment form. This information is stored in a structured format to facilitate subsequent data analysis and processing. The entire process enables efficient and accurate processing of environmental assessment forms, providing strong technical support for the smooth implementation of environmental protection work.
[0010] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 A schematic flow chart of an intelligent identification method for an environmental evaluation form in one embodiment;
[0013] Figure 2 This is an architecture diagram of an intelligent recognition system for environmental evaluation forms in one embodiment.
[0014] Explanation of the accompanying symbols: information acquisition module 1, preprocessing module 2, image segmentation module 3, segmentation and recognition module 4, evaluation and correction module 5, information extraction module 6. DETAILED DESCRIPTION
[0015] The embodiments of the present application provide a method and system for intelligently identifying environmental evaluation forms to solve the technical problem of insufficient recognition accuracy and efficiency due to the diversity of form content.
[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0018] Example 1:
[0019] like Figure 1 As shown, the present application provides an intelligent recognition method for environmental evaluation forms, the method comprising:
[0020] Use image acquisition equipment to collect image information of environmental assessment forms according to preset clarity requirements and build an environmental assessment form image library;
[0021] The Environmental Assessment Form is an important tool for systematically evaluating and predicting the potential environmental impacts of a project. It records in detail various environmental information about the project, including its impacts on the atmosphere, water, soil, biodiversity, and other aspects, as well as its impacts on human health and the socio-economic situation.
[0022] In this embodiment of the present application, the system terminal uses an image acquisition device to capture image information from the environmental assessment form according to preset clarity standards. During this process, the device ensures that the captured image is clear and detailed for subsequent processing and analysis. The captured image information is organized and stored in a dedicated image library, forming a complete image collection of the environmental assessment form. This step provides the necessary image material and data foundation for subsequent form recognition and processing.
[0023] Preprocessing the images in the environmental assessment form image library, including denoising and edge detection;
[0024] In one embodiment, in order to improve the quality of the environmental assessment form image, the system terminal performs a series of preprocessing operations on each form image in the image library. First, denoising technology is used to eliminate interference factors such as noise and spots in the image, making the image clearer and purer. Subsequently, edge detection is performed to determine whether there are any problems with the collected image information. If there are problems, the system terminal will store the problematic images and generate re-collection instructions for these problematic images based on the image numbers, allowing the image acquisition device to re-collect these images. After that, the edge detection steps are repeated until all images are correct, and the detected images are stored in the environmental assessment form image library. After these preprocessing steps, the quality of the form image will be significantly improved, laying a solid foundation for subsequent recognition work.
[0025] Furthermore, the present application provides preprocessing of images in the environmental assessment form image library, including denoising and edge detection, and the method further includes:
[0026] Performing denoising processing on the images in the environmental assessment form image library;
[0027] Use the denoised image to perform edge detection, extract the images that do not meet the edge detection requirements, and build an abnormal image library;
[0028] Preferably, to optimize the recognition of environmental assessment form images, the system terminal first performs denoising on each form image in the image library. This step aims to remove noise and interference from the image, making the image content clearer and more accurate. The denoising significantly improves image quality, laying a good foundation for subsequent edge detection. Subsequently, edge detection is performed on the denoised image. Edge detection is a key step in identifying the boundaries of different regions in an image. The system terminal uses a preset edge detection algorithm to perform edge location and refinement, non-maximum suppression and edge tracking, and post-processing to complete edge detection on the processed image. The edge detection results are then visually and algorithmically evaluated to identify images with blurred or broken edges, or where the detected edges do not match the actual form content. Images that do not meet the required edge detection results are then extracted from the image library and stored in a dedicated abnormality library. The establishment of this abnormality library facilitates further analysis and processing of these images, as well as re-acquisition and pre-processing, to improve the overall recognition accuracy of environmental assessment form images. In summary, by denoising and edge detection of environmental assessment form images, combined with the establishment of an abnormal image library, the image recognition effect can be better optimized and the accuracy and efficiency of environmental assessment form data processing can be improved.
[0029] Confirming the number of the acquired image based on the abnormal image library generates a re-acquisition instruction, sends the instruction to the image acquisition device, and establishes a mapping relationship between the re-acquisition image and the abnormal image;
[0030] The edge detection of the re-collected images is performed based on the abnormal image library until the preset requirements are met, and all images that meet the preset requirements are stored in the environmental assessment form image library.
[0031] Preferably, the system terminal confirms the image number that needs to be recaptured based on the image information in the abnormal image library and generates a corresponding recapture instruction. This instruction will be sent to the image acquisition device to instruct the device to recapture these images. At the same time, in order to ensure the consistency and traceability of the data, a mapping relationship between the recaptured images and the abnormal images is established. In this way, in subsequent processing and analysis, it is clear which images are recaptured and their corresponding original abnormal images. After the recapture is completed, the system terminal will perform edge detection on these images again. This process will be repeated until all recaptured images meet the preset edge detection requirements. Once the images meet the requirements, they will be stored in the environmental assessment form image library for subsequent use. Through this process, it can be ensured that the image quality in the environmental assessment form image library is effectively improved, providing accurate and reliable data support for subsequent environmental assessment work.
[0032] Use image segmentation technology to segment the pre-processed environmental assessment form image according to the preset segmentation requirements, and mark it according to the attributes corresponding to the segmentation results;
[0033] In one embodiment, the quality and clarity of the pre-processed environmental assessment form image are significantly improved, laying a good foundation for subsequent segmentation and annotation work. The system terminal will use image segmentation technology to further process these images. During the segmentation process of the environmental assessment form image, the system terminal accurately divides the different parts of the form according to preset segmentation requirements, such as titles, fields and other features, and annotates the relevant attributes. After the segmentation is completed, each area is annotated according to the attributes corresponding to the segmentation results. This helps the system terminal better understand the structure and content of the form. Through image segmentation and annotation, the environmental assessment form image can be converted into structured data, which is convenient for subsequent automatic recognition and extraction. This can not only improve the efficiency of environmental assessment work, but also improve the accuracy of the data.
[0034] Furthermore, the present application provides the method of using the image segmentation technology to segment the pre-processed environmental assessment form image according to preset segmentation requirements, and labeling according to the attributes corresponding to the segmentation results, and further comprising:
[0035] According to the recognition task in the form recognition request, the preset segmentation requirements are set, including title, field, and data;
[0036] Based on the preset segmentation requirements, generating the attribute annotation;
[0037] The preset segmentation requirements are added to the image segmentation technology to generate a segmentation execution module, the pre-processed environmental evaluation form image is segmented, and the corresponding segmentation results are annotated based on the attribute annotation.
[0038] Preferably, image segmentation technology is used to segment the preprocessed environmental assessment form image according to preset segmentation requirements and annotate the image according to the attributes corresponding to the segmentation results. Attribute annotation refers to the classification and meaning of the segmented image regions, including title annotation, field annotation, and data annotation. Based on the specific task requirements in the form recognition request, the system terminal first sets preset segmentation requirements. These requirements primarily cover key components of the form, such as the title, fields, and data. These requirements provide clear guidance for subsequent image segmentation, ensuring accurate segmentation of each important region within the form. Subsequently, based on these preset segmentation requirements, corresponding attribute annotations are generated. These annotations effectively label the segmented image regions with their categories and meanings, facilitating understanding of the specific content represented by each region. The system terminal then uses the integrated preset segmentation requirements as input parameters and embeds them into a preset image segmentation algorithm, forming a segmentation execution module. The preprocessed environmental assessment form image is then input into the segmentation execution module. The segmentation execution module automatically segments the image according to the preset segmentation requirements, dividing the form into multiple regions or objects. Then, based on the previously generated attribute annotations, each region in the segmentation result is assigned a corresponding label to identify the function and content of the different regions. Finally, the segmentation result with attribute annotations is output as a data structure. This output not only contains the segmented image regions but also the attribute annotation information for each region, facilitating subsequent information extraction and processing. Through this series of operations, the system terminal can achieve accurate segmentation and annotation of the environmental assessment form image, laying a solid foundation for subsequent data extraction and analysis. This not only improves the efficiency and accuracy of form recognition, but also significantly enhances overall work efficiency.
[0039] Using optical character recognition technology to sequentially identify the segmentation results of the environmental assessment form to obtain recognition results;
[0040] In one embodiment, optical character recognition (OCR) technology primarily converts text information within an image into a text format that can be edited and processed by a computer. When processing an environmental assessment form, the system terminal first uses OCR technology to scan and recognize the form image, extracting the text and converting it into text. After converting the text into text, the system terminal then applies operations based on the segmentation results of the form image. This segmentation result provides a demarcation of different areas within the form, such as the title, fields, and data. Based on these segmentation results, more precise and targeted operations can be performed on the converted text information. Specifically, the system terminal matches the recognized text information with the corresponding form areas based on the labels in the segmentation results. This allows accurate extraction of the value of each form field, or further analysis and processing of the text within a specific area, generating a recognition result. In summary, this method allows the system terminal to efficiently convert the text information within the environmental assessment form into a computer-processable text format and, based on the segmentation results, perform precise operations on the converted text information, providing a convenient and accurate foundation for subsequent data analysis and processing.
[0041] Furthermore, the present application provides the method of sequentially recognizing the segmentation results of the environmental evaluation form using optical character recognition technology to obtain recognition results, and the method also includes:
[0042] Converting the environmental assessment form into text information;
[0043] Performing content recognition on the text information according to the segmentation result to obtain a recognition result;
[0044] Based on the attribute annotation of the environmental evaluation form, the character structure of each segmentation result is determined, and the recognition result is converted according to the character structure to obtain the recognition result.
[0045] Preferably, the system terminal first performs optical character recognition on the entire form image. The optical character engine scans the entire image, analyzes the pixel patterns, and identifies all visible text. After the optical character engine recognizes the text, it converts it into computer-editable text information, providing a basis for subsequent operations. Subsequently, the system terminal performs content recognition on the converted text information based on the results of the previous segmentation of the form image. Because the segmentation results have already divided the form into different regions, the system terminal can perform targeted content recognition on each region, ensuring accurate extraction of information for each field or title. During the content recognition process, the system terminal compares the identifiers in the annotations with the corresponding information in the segmentation results to determine the attribute annotations corresponding to each segmentation result. Based on the matched attribute annotations, the system terminal determines the character structure of each segmentation result. The character structure describes the arrangement, font, size, and other characteristics of the text in the region. For example, some fields may require fixed-length numbers, while others may allow text containing special characters or spaces. With this character structure information, the system terminal performs conversion on the recognition results. The purpose of the conversion is to ensure that the recognition results conform to the character structure requirements of the form. This includes adjusting the length, format, and font of the text. For example, if a field requires a fixed-length number, but the length of the number in the recognition result does not match, padding or truncation is required. Finally, the converted recognition result is output to obtain the final recognition result. This result exists in text form and contains key information in the environmental assessment form, which facilitates subsequent data analysis and processing. In summary, after converting the environmental assessment form into text information, content recognition is performed by combining the segmentation results and attribute annotations, and the recognition results are converted according to the character structure, ultimately obtaining accurate recognition results. This process improves the automation level of form processing and ensures the accuracy and completeness of information.
[0046] Based on the attribute annotation, the recognition result is evaluated and corrected to obtain a corrected recognition result;
[0047] In one embodiment, the process of evaluating and correcting recognition results based on attribute annotation primarily utilizes known form attribute annotation information to check and correct the text results recognized by optical character recognition technology. This step aims to improve the accuracy and completeness of recognition results and ensure that the converted text is consistent with the original form content. First, the system terminal evaluates the recognition results by comparing the content format requirements, necessary parameters, and other information defined in the attribute annotations. If discrepancies are found between the recognition results and the annotation information, such as formatting errors, missing data, or garbled characters, correction is required. This correction process includes various methods. Simple formatting errors or spelling errors can be addressed through automatic replacement. Complex cases, such as missing data or garbled characters, require correction through the use of correction algorithms. After the evaluation and correction are complete, corrected recognition results are obtained. These results have been rigorously checked and adjusted to better conform to the attribute requirements of the original form, providing a reliable foundation for subsequent data processing and analysis. In summary, evaluating and correcting recognition results based on attribute annotation is a key step in ensuring conversion quality. This step improves the accuracy of recognition results, reduces errors and omissions, and provides high-quality text data for subsequent processing of environmental assessment forms.
[0048] Furthermore, the present application provides a method for evaluating and correcting the recognition result based on attribute annotation to obtain a corrected recognition result, and the method further includes:
[0049] Determining content format requirements and necessary parameter information of the segmentation result based on the attribute annotations;
[0050] Traversing and matching the recognition results according to the content format requirements and necessary parameter information to determine abnormal matching results;
[0051] Preferably, based on attribute annotation, the system terminal first determines the content format requirements and necessary parameter information for each segmented result in the environmental assessment form. This typically includes the data type, format specifications, length restrictions, and specific identifiers that each segmented region should contain. This information is crucial for subsequent verification of the recognition results. Once these format requirements and parameter information are in place, a traversal matching process can be performed on the recognition results. This means that the system terminal will examine each data item in the recognition result one by one and compare it with the corresponding content format requirements and parameter information. The purpose of this process is to identify abnormal matching results that do not meet the preset requirements. Abnormal matching results include data with incorrect format, data missing necessary parameters, or data containing invalid characters. Through traversal matching, these abnormal results can be accurately located, providing a basis for subsequent data correction. In summary, determining content format requirements and necessary parameter information based on attribute annotation, and then performing traversal matching on the recognition results, is an effective data verification process. This helps the system terminal promptly detect and correct abnormal data in the recognition results, ensuring data accuracy and completeness, and providing reliable data support for subsequent environmental assessment work.
[0052] When there is no abnormal matching result, a mapping relationship between the recognition result and the form is established and exported to the recognition database;
[0053] When the abnormal matching result exists, deviation information is obtained, and the deviation information is supplemented by using the attribute tolerance to obtain a corrected recognition result; for the deviation information that exceeds the attribute tolerance, correction recognition constraints are generated, the constraints are added to the optical character recognition technology, a re-recognition module is generated for recognition, and the recognition results that meet the requirements are stored in the recognition database.
[0054] Preferably, attribute tolerance refers to the extent to which data can deviate from preset standards and still be considered valid. This is a preset value. When the optical character recognition results fully match the content format requirements and necessary parameter information of the form attribute annotations, with no exceptions, the system terminal successfully establishes a mapping relationship between the recognition results and the form. This means that each recognized data item is accurately mapped to the corresponding location in the form, ensuring data accuracy and integrity. These mapping relationships and their corresponding recognition results are then exported and stored in the recognition database for subsequent data analysis and processing. However, if an anomaly is detected during the matching process—that is, the recognition result does not meet the preset format requirements or parameter information—the system terminal must take further action. First, the deviation information is obtained and supplemented or adjusted using the attribute tolerance. By utilizing this tolerance, minor deviations can be corrected, resulting in a corrected recognition result. However, if some deviation information exceeds the attribute tolerance range, the system terminal cannot simply modify it to correct it. In this case, correction recognition constraints are generated, which specifically describe which types of deviations are unacceptable and require re-recognition using the optical character recognition technology. These constraints are incorporated into the optical character recognition technology, generating a re-recognition module specifically designed to handle data items that failed initial recognition. This re-recognition process yields more accurate, satisfying results and stores them in the recognition database. In short, this process continuously optimizes and refines the recognition results. By establishing mapping relationships, applying attribute tolerances for corrections, and generating the re-recognition module, we ensure that the data ultimately stored in the recognition database is accurate, complete, and reliable.
[0055] Furthermore, the present application provides a method for adding the constraint condition to the optical character recognition technology to generate a re-recognition module for recognition, and the method further includes:
[0056] Determine the image number of the environmental evaluation form based on the deviation information, and traverse from the abnormal image;
[0057] Determine whether there is an abnormal graph based on the traversal result, and if so, locate the deviation information of the abnormal graph;
[0058] Optionally, when the system terminal detects discrepancies in the results of identifying an environmental assessment form, it first determines the specific form image number based on the discrepancy information. This is because forms are typically composed of multiple images, each of which may contain different information areas. Determining the image number allows for more precise location of the problematic image. Subsequently, a traversal process begins, starting with the identified anomalous image. The purpose of this traversal process is to carefully examine every part of the image to identify the specific location or area that may have caused the recognition discrepancy. During the traversal process, the system terminal determines whether any anomalies exist. Anomalies are images that contain obvious recognition errors or formatting issues. Once an anomaly is found, further analysis and comparison of image elements such as text, lines, and symbols is performed to determine the specific location of the discrepancy. In summary, this process involves determining the image number, traversing the anomalous image, determining the presence of the anomaly, and ultimately locating the discrepancy. This helps to more accurately identify the problem, providing important information for subsequent data correction and image optimization.
[0059] Adding the located abnormal image to the recognition list, performing recognition through the re-recognition module, and obtaining an auxiliary recognition result;
[0060] The auxiliary recognition result is fused with the recognition result, and the recognition result is corrected using the fusion information that meets the content format requirements and necessary parameter information to obtain the corrected recognition result.
[0061] Optionally, once the system terminal locates anomalous images in the environmental assessment form, it adds them to the recognition list. This allows for further recognition processing of these specific images to correct any errors that may have occurred during the initial recognition process. Subsequently, these anomalous images are recognized using the previously generated re-recognition module. This re-recognition module is built based on the deviation information discovered during the previous recognition process and the corrected recognition constraints, thus focusing on addressing situations that may have led to recognition errors. Through the re-recognition module, auxiliary recognition results are obtained for these anomalous images. After obtaining the auxiliary recognition results, the system terminal compares and matches the auxiliary recognition results with the data items in the previous recognition results one by one. For data items in the same location, if their content is identical or very similar, one of them is selected as the final result; if there are significant differences, the auxiliary recognition result is used as the final result. During the fusion process, the fusion results are filtered according to preset content format requirements and necessary parameter information. This means that the system terminal checks whether each data item conforms to the expected format and contains all necessary parameter information. Data items that do not meet the requirements are marked as anomalous. After information filtering, the original recognition results are corrected using the fusion information that meets the requirements. This includes replacing incorrect data items and supplementing missing information. The goal of correction is to ensure that the final recognition result is both accurate and meets the form requirements. In summary, this process is a cycle of recognition, fusion, and correction. The re-recognition module obtains auxiliary recognition results, which are then fused with the original results. Finally, the fused information is used to correct the recognition results, resulting in a more accurate and reliable corrected recognition result.
[0062] Furthermore, the present application provides that before storing the recognition results that meet the requirements in the recognition database, the method further includes:
[0063] When the recognition result cannot meet the content format requirements and necessary parameter information, the deviation information is marked and an abnormal reminder message is generated.
[0064] Optionally, when the identification results fail to meet the content format requirements of the environmental assessment form or are missing necessary parameter information, deviations are identified and an exception alert is generated. This means that during the identification process, if the system terminal discovers that certain data items do not conform to the preset format specifications or parameter requirements, or that certain key information is missing, it will specifically mark these deviations. These identifiers help quickly locate the problem and identify which areas require special attention and correction. The system terminal also generates exception alerts based on these identifiers. These alerts are presented to the user in an easy-to-understand manner, clearly indicating which data items have issues and possible causes or impacts. This allows users to quickly understand the issues in the identification results and take appropriate measures to correct or supplement them. In summary, identifying deviations and generating exception alerts are important steps to ensure the accuracy and completeness of identification results. They help to promptly identify and resolve problems, improving the efficiency and accuracy of data processing.
[0065] Further; extracting key information from the corrected recognition result by natural language processing technology, the method also includes;
[0066] Use natural language processing techniques (such as BERT) to convert text content into feature vectors;
[0067] Randomly select eigenvectors as initial cluster centers;
[0068] Calculate the distance between each eigenvector and all cluster centers and assign it to the nearest cluster center;
[0069] Calculate the mean of all eigenvectors in each cluster and update the cluster center position;
[0070] Repeat the assignment of data points and updating of cluster centers until the cluster centers no longer change;
[0071] The sum of squared errors within the cluster is calculated using the objective function through the eigenvectors and cluster centers to accurately extract the evaluation items in the environmental evaluation form and cluster similar environmental evaluation forms together. The objective function is as follows:
[0072]
[0073] Among them, υ(T ί ) is the eigenvector, μκ is the cluster center, J is the sum of squared errors within the cluster.
[0074] That is, this application obtains the evaluation content of each form;
[0075] Use natural language processing technology (such as BERT) to convert the text content of each form into a feature vector. Specifically, extract the text content T of each form. ί , and then use the BERT model to transform the text T ί Convert to vector representation υ(T ί ); where υ(T ί ) is a characteristic vector of the form, T ί The text content of the form;
[0076] Select K eigenvectors as initial cluster centers;
[0077] Assign each form feature vector to the nearest cluster center;
[0078] Calculate each eigenvector υ(T ί ) to each cluster center.
[0079] Assign each eigenvector to the closest μκ Cluster center;
[0080] Furthermore, K eigenvectors may be randomly selected as the eigenvectors of the initial cluster centers.
[0081] In summary, the embodiments of the present application have at least the following technical effects:
[0082] In the embodiments of the present application, an image acquisition device is used to capture image information of an environmental assessment form according to preset clarity requirements and to construct an image library. Subsequently, the images in the image library are preprocessed. Image segmentation technology is then used to segment the preprocessed images according to preset segmentation requirements, and the images are annotated according to the attributes corresponding to the segmentation results. Optical character recognition technology is then used to sequentially recognize the segmented results of the environmental assessment form, converting the form content into text information. The recognition results are then formatted according to the attribute annotations to obtain a final recognition result. The recognition results are then evaluated and corrected based on the attribute annotations. Anomalous matching results are identified by comparing the recognition results with preset format requirements and necessary parameter information. In cases where no anomalies exist, a mapping relationship between the recognition results and the form is established and exported to a recognition database. In cases where anomalies exist, attribute tolerance is used to supplement deviation information or generate corrective recognition constraints, which are then incorporated into optical character recognition technology for re-recognition to obtain a corrected recognition result. Finally, natural language processing technology is used to extract key information from the corrected recognition result and store it in a structured format according to the structural content of the environmental assessment form. Before storage, if the recognition result still fails to meet format requirements or lacks necessary parameters, deviations will be identified and an abnormality alert will be generated. These technical effects jointly solve the technical problems of insufficient recognition accuracy and efficiency caused by the diversity of form content, ensuring that environmental assessment work is carried out efficiently and accurately.
[0083] Example 2:
[0084] Based on the same inventive concept as the method for intelligently identifying an environmental evaluation form in the aforementioned embodiment, Figure 2 As shown, the present application provides an environmental assessment form intelligent recognition system, the system comprising:
[0085] Information collection module 1: The information collection module 1 is used to use an image collection device to collect image information of the environmental evaluation form according to preset clarity requirements and build an environmental evaluation form image library;
[0086] Preprocessing module 2: The preprocessing module 2 is used to preprocess the images in the environmental assessment form image library, including denoising and edge detection;
[0087] Image segmentation module 3: The image segmentation module 3 is used to segment the pre-processed environmental assessment form image according to preset segmentation requirements using image segmentation technology, and mark the attributes corresponding to the segmentation results;
[0088] Segmentation and recognition module 4: the segmentation and recognition module 4 is used to sequentially recognize the segmentation results of the environmental assessment form using optical character recognition technology to obtain recognition results;
[0089] Evaluation and correction module 5: the evaluation and correction module 5 is used to evaluate and correct the recognition result based on the attribute annotation to obtain a corrected recognition result;
[0090] Information extraction module 6: The information extraction module 6 is used to extract key information from the modified recognition result according to the structural content of the environmental evaluation form through natural language processing technology, and store it in a structured form based on the structural content.
[0091] Furthermore, the preprocessing module 2 is used to perform the following method:
[0092] Performing denoising processing on the images in the environmental assessment form image library;
[0093] Use the denoised image to perform edge detection, extract the images that do not meet the edge detection requirements, and build an abnormal image library;
[0094] Confirming the number of the acquired image based on the abnormal image library generates a re-acquisition instruction, sends the instruction to the image acquisition device, and establishes a mapping relationship between the re-acquisition image and the abnormal image;
[0095] The edge detection of the re-collected images is performed based on the abnormal image library until the preset requirements are met, and all images that meet the preset requirements are stored in the environmental assessment form image library.
[0096] Furthermore, the image segmentation module 3 is used to perform the following method:
[0097] According to the recognition task in the form recognition request, the preset segmentation requirements are set, including title, field, and data;
[0098] Based on the preset segmentation requirements, generating the attribute annotation;
[0099] The preset segmentation requirements are added to the image segmentation technology to generate a segmentation execution module, the pre-processed environmental evaluation form image is segmented, and the corresponding segmentation results are annotated based on the attribute annotation.
[0100] Furthermore, the segmentation and recognition module 4 is used to perform the following method:
[0101] Converting the environmental assessment form into text information;
[0102] Performing content recognition on the text information according to the segmentation result to obtain a recognition result;
[0103] Based on the attribute annotation of the environmental evaluation form, the character structure of each segmentation result is determined, and the recognition result is converted according to the character structure to obtain the recognition result.
[0104] Furthermore, the evaluation and correction module 5 is used to perform the following method:
[0105] Determining content format requirements and necessary parameter information of the segmentation result based on the attribute annotations;
[0106] Traversing and matching the recognition results according to the content format requirements and necessary parameter information to determine abnormal matching results;
[0107] When there is no abnormal matching result, a mapping relationship between the recognition result and the form is established and exported to the recognition database;
[0108] When the abnormal matching result exists, deviation information is obtained, and the deviation information is supplemented by using the attribute tolerance to obtain a corrected recognition result; for the deviation information that exceeds the attribute tolerance, correction recognition constraints are generated, the constraints are added to the optical character recognition technology, a re-recognition module is generated for recognition, and the recognition results that meet the requirements are stored in the recognition database.
[0109] Furthermore, the evaluation and correction module 5 is used to perform the following method:
[0110] Determine the image number of the environmental evaluation form based on the deviation information, and traverse from the abnormal image;
[0111] Determine whether there is an abnormal graph based on the traversal result, and if so, locate the deviation information of the abnormal graph;
[0112] Adding the located abnormal image to the recognition list, performing recognition through the re-recognition module, and obtaining an auxiliary recognition result;
[0113] The auxiliary recognition result is fused with the recognition result, and the recognition result is corrected using the fusion information that meets the content format requirements and necessary parameter information to obtain the corrected recognition result.
[0114] Furthermore, the evaluation and correction module 5 is used to perform the following method:
[0115] When the recognition result cannot meet the content format requirements and necessary parameter information, the deviation information is marked and an abnormal reminder message is generated.
[0116] It should be noted that the above-mentioned order of the embodiments of the present application is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0117] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0118] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.
Claims
1. A method for intelligent recognition of environmental evaluation forms, characterized in that: include: Use image acquisition equipment to collect image information of environmental assessment forms according to preset clarity requirements and build an environmental assessment form image library; Preprocessing the images in the environmental assessment form image library, including denoising and edge detection; Use image segmentation technology to segment the pre-processed environmental assessment form image according to the preset segmentation requirements, and mark it according to the attributes corresponding to the segmentation results; Using optical character recognition technology to sequentially identify the segmentation results of the environmental assessment form to obtain recognition results; Based on the attribute annotation, the recognition result is evaluated and corrected to obtain a corrected recognition result; According to the structural content of the environmental assessment form, key information in the revised recognition result is extracted by natural language processing technology, and stored in a structured form based on the structural content; The step of evaluating and correcting the recognition result based on attribute annotation to obtain a corrected recognition result includes: Determining content format requirements and necessary parameter information of the segmentation result based on the attribute annotations; Traversing and matching the recognition results according to the content format requirements and necessary parameter information to determine abnormal matching results; When there is no abnormal matching result, a mapping relationship between the recognition result and the form is established and exported to the recognition database; When an abnormal matching result exists, deviation information is obtained, and the deviation information is supplemented by using the attribute tolerance to obtain a corrected recognition result; for deviation information that exceeds the attribute tolerance, correction recognition constraints are generated, the constraints are added to the optical character recognition technology, a re-recognition module is generated for recognition, and recognition results that meet the requirements are stored in the recognition database; The step of adding the constraint condition to the optical character recognition technology to generate a re-recognition module for recognition includes: When deviation information is found in the result of identifying the environmental evaluation form, the image number of the environmental evaluation form is determined based on the deviation information, and traversal is performed from the abnormal image; Determine whether there is an abnormal image based on the traversal results. If so, locate the deviation information of the abnormal image, that is, accurately analyze and compare the elements in the image to determine the specific location where the deviation information appears; Adding the located abnormal image to the recognition list, performing recognition through the re-recognition module, and obtaining an auxiliary recognition result; The auxiliary recognition result is fused with the recognition result, and the recognition result is corrected using the fusion information that meets the content format requirements and necessary parameter information to obtain the corrected recognition result.
2. The method according to claim 1, wherein Preprocess the images in the environmental assessment form image library, including denoising and edge detection, including: Performing denoising processing on the images in the environmental assessment form image library; Use the denoised image to perform edge detection, extract the images that do not meet the edge detection requirements, and build an abnormal image library; Confirming the number of the acquired image based on the abnormal image library generates a re-acquisition instruction, sends the instruction to the image acquisition device, and establishes a mapping relationship between the re-acquisition image and the abnormal image; The edge detection of the re-collected images is performed based on the abnormal image library until the preset requirements are met, and all images that meet the preset requirements are stored in the environmental assessment form image library.
3. The method according to claim 1, wherein The method of using image segmentation technology to segment the pre-processed environmental assessment form image according to preset segmentation requirements and marking the image according to the attributes corresponding to the segmentation results includes: According to the recognition task in the form recognition request, the preset segmentation requirements are set, including title, field, and data; Based on the preset segmentation requirements, generating the attribute annotation; The preset segmentation requirements are added to the image segmentation technology to generate a segmentation execution module, the pre-processed environmental evaluation form image is segmented, and the corresponding segmentation results are annotated based on the attribute annotation.
4. The method according to claim 1, wherein The step of sequentially recognizing the segmentation results of the environmental assessment form using optical character recognition technology to obtain recognition results includes: Converting the environmental assessment form into text information; Performing content recognition on the text information according to the segmentation result to obtain a recognition result; Based on the attribute annotation of the environmental evaluation form, the character structure of each segmentation result is determined, and the recognition result is converted according to the character structure to obtain the recognition result.
5. The method according to claim 1, wherein Before storing the recognition results that meet the requirements in the recognition database, the method further includes: When the recognition result cannot meet the content format requirements and necessary parameter information, the deviation information is marked and an abnormal reminder message is generated.
6. The method according to claim 1, wherein The method further comprises extracting key information from the modified recognition result by natural language processing technology according to the structural content of the environmental evaluation form; Use natural language processing technology to convert text content into feature vectors; Randomly select eigenvectors as initial cluster centers; Calculate the distance between each eigenvector and all cluster centers and assign it to the nearest cluster center; Calculate the mean of all eigenvectors in each cluster and update the cluster center position; Repeat the assignment of data points and updating of cluster centers until the cluster centers no longer change; The sum of squared errors within the cluster is calculated using the objective function through the eigenvectors and cluster centers to accurately extract the evaluation items in the environmental evaluation form and cluster similar environmental evaluation forms together. The objective function is as follows: ; Among them, υ(T ί ) is the eigenvector, μκ is the cluster center, J is the sum of squared errors within the cluster.
7. An intelligent recognition system for environmental evaluation forms, characterized in that: A method for intelligently identifying an environmental assessment form according to any one of claims 1 to 6, comprising: Information collection module: Use image acquisition equipment to collect image information of environmental assessment forms according to preset clarity requirements and build an environmental assessment form image library; Preprocessing module: preprocessing the images in the environmental assessment form image library, including denoising and edge detection; Image segmentation module: Use image segmentation technology to segment the pre-processed environmental assessment form image according to the preset segmentation requirements, and mark it according to the attributes corresponding to the segmentation results; Segmentation and recognition module: Use optical character recognition technology to sequentially recognize the segmentation results of the environmental assessment form to obtain recognition results; Evaluation and correction module: Based on the attribute annotation, the recognition result is evaluated and corrected to obtain a corrected recognition result; Information extraction module: according to the structural content of the environmental evaluation form, the key information in the modified recognition result is extracted by natural language processing technology, and stored in a structured form based on the structural content.
Citation Information
Patent Citations
Community resident event identification method and device based on text clustering
CN114328812A
Symmetric table character data structured extraction method and system based on semantic analysis
CN115147857A
Financial form identification method and device, electronic equipment and storage medium
CN117831052A