Enterprise environment report file analysis method and system
By designing enterprise environmental report file analysis methods and systems, using layout analysis models and OCR models to extract key information in report files, building a database for risk assessment, solving unstructured data processing problems, and improving risk assessment efficiency and information integration capabilities.
Patent Information
- Application Number
- CN202510144744.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is difficult to effectively process and analyze unstructured data, especially in corporate environmental risk assessment, resulting in inefficient assessment and difficulty in integrating information.
Design an enterprise environmental report file analysis method and system, convert report files in different formats into a unified document format, build a layout analysis model and OCR model, extract file description elements and text data, and build a database to achieve automated risk assessment.
It realizes automated processing and risk assessment of unstructured data, improves evaluation efficiency, enhances information integration and data correlation analysis capabilities, and fully taps the potential value of unstructured data.
Smart Images

Figure CN120067194A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data structure analysis, and particularly relates to a method and system for analyzing enterprise environmental report files. Background Art
[0002] Currently, the assessment of enterprise sudden environmental accident risks mainly relies on two major types of data sources: structured data and unstructured data.
[0003] Structured data: Structured data includes the online monitoring data of enterprises, valid certificate registration information, ledger data, administrative penalty records, law enforcement records, etc. This type of data is stored in the form of fixed-format tables or databases, and has the characteristics of strong structure and unified format. Therefore, the analysis of structured data is relatively mature, and usually combines technical means such as big data analysis, statistical analysis, mechanism modeling, and machine learning for processing to achieve a quantitative assessment of enterprise environmental risks.
[0004] Unstructured data: Unstructured data refers to texts and reports that have not been structurally processed, such as environmental impact assessment reports provided by environmental assessment agencies, risk assessment reports self-evaluated by enterprises, and pollutant discharge permit implementation reports. These reports have complex content, contain a large amount of natural language text, tables, and charts, and the information is distributed dispersedly, lacking a unified structure. Currently, the processing of unstructured data mainly relies on manual methods, that is, manually extracting and reviewing text information and entering relevant content into the system. This method has a large workload, low efficiency, and limitations in information integration and data correlation analysis, and it is difficult to fully explore the potential value in unstructured data.
[0005] In the enterprise risk assessment system, currently, the public enterprise sudden environmental event risk grading standard is mostly used to assess risks, but mainly relies on manual registration and review methods. In order to adapt to the environmental protection characteristics and data resource conditions of different regions, the risk grading assessment method needs to flexibly adjust indicators and assessment methods. Due to the increasing importance of unstructured data in risk assessment, there is an urgent need for a new technical means to effectively combine unstructured data and grading assessment methods to achieve systematic and automated assessment. Summary of the Invention
[0006] The first object of the present invention is to provide a method for analyzing enterprise environmental report files.
[0007] The second object of the present invention is to provide a system for analyzing enterprise environmental report files.
[0008] The present invention designs and implements an enterprise environmental report file analysis method and system, which can process unstructured file data such as enterprise environmental report files, facilitate the extraction of index data affecting environmental risk assessment, and provide an analysis means for environmental risk assessment.
[0009] To further illustrate the content of the present invention, on the one hand, the present invention provides an enterprise environmental report file analysis method, including the following steps:
[0010] S1. Convert enterprise environmental report files in different formats into a unified document format;
[0011] S2. Construct a layout analysis model, and use the layout analysis model to distinguish the content in the enterprise environmental report file to obtain file description elements, where the file description elements include text, title, illustration, table, header, footer, and formula;
[0012] S3. Construct an OCR model, use the OCR model to identify the file description elements, and extract the file description elements to obtain the text data corresponding to the file description elements;
[0013] S3. Use the text data identified by the OCR model to construct a database, and generate a new database content layout by inputting the retrieval text to associate with the data content tags in the database.
[0014] Further, in the step S1, converting enterprise environmental report files in different formats into a unified document format specifically includes:
[0015] Obtain enterprise environmental report files in an external database through an external API interface, and convert them into a unified pdf format document through a document conversion tool.
[0016] Further, in the step S2, constructing a layout analysis model, and using the layout analysis model to distinguish the content in the enterprise environmental report file to obtain file description elements, where the file description elements include text, title, illustration, table, header, footer, and formula, specifically includes:
[0017] Construct a layout analysis coordinate system, perform layout interception on the coordinate starting point and coordinate ending point of each content in the enterprise environmental report file to obtain an image area, and perform block division on the obtained image area to obtain file description elements, where the file description elements include text, title, illustration, table, header, footer, and formula.
[0018] Further, based on the layout analysis coordinate system, for the table in the description elements, identify the intersection points of the horizontal and vertical lines in the table and split them into cells.
[0019] Further, in the step S3, an OCR model is constructed, and the OCR model is used to identify the file description elements, and the file description elements are extracted to obtain the text data corresponding to the file description elements, specifically including:
[0020] Use the OCR model to identify the file description elements, extract the text information in the file description elements and output it as text data, and store it in the text data format of the txt format.
[0021] On the other hand, the present invention also provides an enterprise environmental report file analysis system, including the following steps:
[0022] A file data source acquisition and conversion module, which is used to convert enterprise environmental report files in different formats into a unified document format;
[0023] A file layout analysis module, which is used to construct a layout analysis model, and use the layout analysis model to distinguish the content in the enterprise environmental report file to obtain file description elements, and the file description elements include text, title, illustration, table, header, footer and formula;
[0024] A file data content extraction module, which is used to construct an OCR model, use the OCR model to identify the file description elements, and extract the file description elements to obtain the text data corresponding to the file description elements;
[0025] A recombination output module, which is used to construct a database by using the text data recognized by the OCR model, and generate a new database content layout by inputting a retrieval text to associate with the data content tags in the database.
[0026] Further, the file data source acquisition and conversion module is specifically used to obtain the enterprise environmental report file in the external database through an external API interface, and convert it into a unified pdf format document through a document conversion tool.
[0027] Further, the file layout analysis module is specifically used to construct a layout analysis coordinate system, perform layout interception on the coordinate starting point and coordinate ending point of each content in the enterprise environmental report file to obtain an image area, and divide the obtained image area to obtain file description elements, and the file description elements include text, title, illustration, table, header, footer and formula.
[0028] Further, the file data content extraction module is specifically used to use the OCR model to identify the file description elements, extract the text information in the file description elements and output it as text data, and store it in the text data format of the txt format.
[0029] In a third aspect, a computer-readable storage medium stores a computer program, characterized in that: when the computer program is executed by a processor, the steps of the method as described above are implemented.
[0030] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings
[0031] Figure 1 is a flowchart showing the actions of an embodiment for implementing a method for analyzing an enterprise environmental report file of the present invention;
[0032] Figure 2 is a flowchart showing the file processing actions in an embodiment of the present invention;
[0033] Figure 3 is a flowchart showing the database acquisition actions in an embodiment of the present invention. Detailed Embodiments
[0034] To better elaborate the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0035] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the embodiments of the present application.
[0036] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0037] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0038] In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0039] As an exemplary example of the present invention, as Figures 1 to 3 shown, as one aspect of this embodiment, this embodiment provides a method for analyzing enterprise environmental report files, including the following steps:
[0040] S1. Convert enterprise environmental report files in different formats into a unified document format;
[0041] Obtain enterprise environmental report files in an external database through an external API interface, and convert them into a unified pdf format document through a document conversion tool.
[0042] For example, the system obtains a list of enterprise statuses to be processed and a corresponding list of reports to be processed from the business system database through an interface. For example, through an API interface with a login authentication mechanism, it pulls to the local path from the business system. The content includes unstructured documents such as the pulled environmental risk assessment reports, as well as JSON files saved with information such as enterprise names, USCCs, file types, and business IDs. The file data sources can include formats such as PDF and Word.
[0043] S2. Build a layout analysis model, and use the layout analysis model to distinguish the content in the enterprise environmental report file to obtain file description elements. The file description elements include the main text, title, illustrations, tables, headers, footers, and formulas;
[0044] Build a layout analysis coordinate system, perform layout interception on the coordinate starting point and coordinate ending point of each content in the enterprise environmental report file to obtain an image area, and perform block differentiation on the obtained image area to obtain file description elements. The file description elements include the main text, title, illustrations, tables, headers, footers, and formulas.
[0045] In addition, based on the layout analysis coordinate system, for the tables in the description elements, mark the intersection points of the horizontal and vertical lines in the table and split them into cells.
[0046] Enterprise environmental report documents contain a large number of tables, which play a crucial role in enterprise information assessment. Therefore, it is necessary to organize the recognized text content of the table part. By analyzing the coordinates of the layout, the image area corresponding to the table part is intercepted. Binary + morphological erosion and dilation operations are used to obtain the table lines, and the intersection points of the horizontal and vertical lines of the table are identified. Next, the table is split into cell form, which is conducive to subsequent use of the OCR model to separately identify the content of the cells. Finally, the corresponding markdown table structure is generated according to the table format, and the OCR recognition results are filled in.
[0047] S3. Build an OCR model, use the OCR model to identify the file description elements, and extract the file description elements to obtain the text data corresponding to the file description elements;
[0048] Specifically, use the OCR model to identify the file description elements, extract the text information in the file description elements and output it as text data, and store it in the text data format of txt format.
[0049] Based on the above steps S2 and S3, in the processing example of the OCR model and the layout analysis model, the layout analysis recognition result is a text block with layout structure information (distinguishing titles, text, tables, etc.). For the convenience of semantic analysis and retrieval, it is necessary to clean the text content and split the text block into text sentences.
[0050] Data cleaning is performed on the text content extracted by the OCR model, including removing noise characters, special symbols, etc., to ensure the standardization and consistency of the text.
[0051] The cleaned text is further processed into sentences, and the long paragraphs are refined into independent sentences for subsequent semantic analysis. The sentence splitting rule uses a regular expression method, with the end-of-sentence symbols as the splitting points, etc., to ensure the accuracy of the sentence boundaries. The splitting symbols include the full stop '.。', semicolon ';;', question mark '??', exclamation mark '!!', line break symbol '\n', etc., including Chinese and English symbols, which can be adjusted according to the actual business scenario.
[0052] For the table, the table is converted into markdown format code. Separated from the text, it is retained as a markdown code block with layout structure information;
[0053] For structures such as illustrations, page numbers, and headers, they are directly removed. After sentence splitting, the stored text changes from a text block with layout structure information (distinguishing titles, text, tables, etc.) to text sentences with layout structure information.
[0054] According to the content of the sentence splitting result, with the chapter as the organizational structure, the text sentences, table codes, etc. are sorted and arranged.
[0055] Before organizing the structural information, confirm the level and number of layers of all chapters in the document according to the layout analysis results. (For example: This document has three levels of chapters, corresponding to the first-level headings, second-level headings, and third-level headings.) Use the page number as the basic traversal variable, and extract all text sentences on the current page for processing during each traversal.
[0056] If a text sentence of the chapter title type is traversed, update the currently corresponding chapter according to the chapter title level. (For example: Update from section 1.1.2 to section 1.1.3)
[0057] If a text sentence or markdown table code is traversed, add the text content to the text list of the current chapter.
[0058] Process like this until all page numbers are traversed. Each chapter and all text sentences or markdown table codes under the corresponding chapter can be obtained, organized in the order of the smallest level of all chapters in the document (for example: organized in the order of 1.1.1, 1.1.2, 2.1.1, 2.1.2...).
[0059] In the previous structural information organization process, each text sentence already has information such as text content, the chapter it belongs to, and the page number. After completing the chapter organization, update the in-paragraph position (which sentence of the chapter the current sentence belongs to), the total number of sentences in the chapter, and the total number of words in the chapter. Add this part of information to the structural information to ensure the accuracy of context retrieval.
[0060] S4. Use the text data recognized by the OCR model to build a database, and generate a new database content layout by inputting the retrieval text to associate with the data content tags in the database.
[0061] A text embedding model (such as BERT, RoBERTa, etc.) can be used to vectorize each clause, converting it into a high-dimensional semantic vector for similarity calculation in the vector space.
[0062] The vectorized clauses are stored in a vector database (such as FAISS) to form a semantic-based document index library. The vector database supports fast similarity retrieval to ensure that the system can efficiently match the text content related to the evaluation metrics.
[0063] Combined with the above, the method in this embodiment can then analyze each part of the enterprise unstructured report by sequentially calling the large model according to each item of the risk assessment metrics, and can summarize the scores and analysis bases of each risk metric.
[0064] As the second aspect of this embodiment, the present invention also provides an enterprise environmental report file analysis system, including the following steps:
[0065] File data source acquisition and conversion module, which is used to convert enterprise environmental report files in different formats into a unified document format;
[0066] File layout analysis module, which is used to construct a layout analysis model, and use the layout analysis model to distinguish the content in the enterprise environmental report file to obtain file description elements, and the file description elements include text, title, illustration, table, header, footer and formula;
[0067] File data content extraction module, which is used to construct an OCR model, use the OCR model to identify the file description elements, and extract the file description elements to obtain the text data corresponding to the file description elements;
[0068] Reorganization and output module, which is used to construct a database by using the text data obtained by the OCR model, and generate a new database content layout by inputting the retrieval text to associate with the data content tags in the database.
[0069] Further, the file data source acquisition and conversion module is specifically used to obtain the enterprise environmental report file in the external database through an external API interface, and convert it into a unified pdf format document through a document conversion tool.
[0070] Further, the file layout analysis module is specifically used to construct a layout analysis coordinate system, perform layout interception on the coordinate starting point and coordinate ending point of each content in the enterprise environmental report file to obtain an image area, and perform block division on the obtained image area to obtain file description elements, and the file description elements include text, title, illustration, table, header, footer and formula.
[0071] Further, the file data content extraction module is specifically used to use the OCR model to identify the file description elements, extract the text information in the file description elements and output it as text data, and store it in the text data format of txt.
[0072] As the third aspect of this embodiment, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program, and is characterized in that: when the computer program is executed by a processor, the steps of the method as described above are implemented.
[0073] When the enterprise environmental report file analysis system is implemented in the form of software functional units and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described embodiments of various pharmaceutical document structured content analysis methods can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0074] According to the disclosure and teachings of the above specification, those skilled in the art to which the present invention pertains can also make changes and modifications to the above-described embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for convenience of description and do not constitute any limitation to the present invention.
Claims
1. A method for analyzing corporate environmental report documents, characterized in that: The steps include: S1. Convert corporate environmental report files in different formats into a unified document format; S2. constructing a layout analysis model, and using the layout analysis model to distinguish the contents of the corporate environmental report file to obtain file description elements, wherein the file description elements include text, title, illustration, table, header, footer and formula; S3, constructing an OCR model, using the OCR model to identify the file description elements, and extracting the file description elements to obtain text data corresponding to the file description elements; S4. Use the text data obtained by OCR model recognition to build a database, and generate a new database content layout by inputting the search text to associate the data content tags in the database.
2. The method for analyzing corporate environmental report files according to claim 1, characterized in that: In the step S1, the enterprise environmental report files in different formats are converted into a unified document format, specifically including: obtaining the enterprise environmental report files in an external database through an external API interface, and converting them into a unified PDF format document through a document conversion tool.
3. The method for analyzing corporate environmental report files according to claim 1, characterized in that: In step S2, a layout analysis model is constructed, and the layout analysis model is used to distinguish the contents in the corporate environmental report file to obtain file description elements, wherein the file description elements include text, title, illustration, table, header, footer and formula, specifically including: A layout analysis coordinate system is constructed, and the coordinate starting points and end points of each content in the corporate environmental report file are intercepted to obtain the image area, and the acquired image area is divided into blocks to obtain file description elements, which include text, title, illustrations, tables, headers, footers and formulas.
4. The method for analyzing corporate environmental report files according to claim 3, characterized in that: Based on the layout analysis coordinate system, for the table in the description element, the intersections of the horizontal and vertical lines in the table are identified and split into cells.
5. The method for analyzing corporate environmental report files according to claim 1, characterized in that: In the step S3, an OCR model is constructed, the OCR model is used to identify the file description elements, and the file description elements are extracted to obtain text data corresponding to the file description elements, which specifically includes: The OCR model is used to identify the file description elements, and the text information in the file description elements is extracted and output as text data, which is stored in the text data format of txt.
6. An enterprise environmental report document analysis system, characterized in that: The steps include: A file data source acquisition and conversion module, wherein the file data source acquisition and conversion module is used to convert enterprise environmental report files of different formats into a unified document format; A document layout analysis module, which is used to construct a layout analysis model, and uses the layout analysis model to distinguish the contents of the corporate environmental report file to obtain document description elements, wherein the document description elements include text, title, illustrations, tables, headers, footers and formulas; A file data content extraction module, which is used to construct an OCR model, use the OCR model to identify file description elements, and extract the file description elements to obtain text data corresponding to the file description elements; The reorganization output module is used to construct a database using text data obtained by OCR model recognition, and generate a new database content layout by inputting the search text to associate the data content tags in the database.
7. The method for analyzing corporate environmental report files according to claim 1, characterized in that: The file data source acquisition and conversion module is specifically used to obtain the enterprise environment report file in the external database through an external API interface, and convert it into a unified PDF format document through a document conversion tool.
8. The method for analyzing corporate environmental report files according to claim 1, characterized in that: The document layout analysis module is specifically used to construct a layout analysis coordinate system, perform layout interception on the coordinate starting point and coordinate end point of each content in the enterprise environmental report file to obtain an image area, and divide the obtained image area into blocks to obtain file description elements, which include text, title, illustrations, tables, headers, footers and formulas.
9. The method for analyzing corporate environmental report files according to claim 1, characterized in that: The file data content extraction module is specifically used to use the OCR model to identify the file description elements, extract the text information in the file description elements and output it as text data, and store it in the text data format of txt format.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.