A file processing method and apparatus
By extracting and generating associated features in complex document processing, the problems of fuzzy file classification, dispersed information extraction and insufficient abstract flexibility in the prior art are solved, and precise processing of complex documents and high-quality abstract generation are achieved.
Patent Information
- Application Number
- CN202510407038.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The prior art is difficult to identify semantic associations across data types in complex document processing, resulting in fuzzy file classification, dispersed information extraction and insufficient flexibility in digests.
By obtaining the target file, semantic extraction and generating correlation features between multiple content units, classifying files based on these features, extracting domain information, and generating dynamic summary.
It realizes accurate classification and information extraction of complex documents. The generated summary not only retains the accuracy of key data, but also meets user needs and is readable, solving the shortcomings of traditional methods in file processing.
Smart Images

Figure CN119917464B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and particularly to a file processing method and apparatus. Background Art
[0002] In complex document processing, documents such as technical specifications and contract agreements usually contain a mixture of structured and unstructured data. For example, technical documents contain both technical parameters in tabular form and functional descriptions in text paragraphs; business contracts coexist with clause lists and supplementary descriptions of free text. Existing technologies have certain limitations.
[0003] Traditional methods rely on keyword matching or fixed templates and are difficult to identify semantic associations across data types. For example, the technical parameters in a table need to be analyzed in association with the acceptance terms in the text, but existing systems cannot effectively model such logical relationships. Multi-dimensional information such as technology, business, and cost in the document is scattered in different content units. For example, in a bidding scenario, the equipment performance parameters (technical attributes) and itemized quotes (cost attributes) need to be associated across tables and text, and traditional methods lack the ability to handle them uniformly. Users need to dynamically adjust the summary content according to their needs. For example, technical reviews require a detailed parameter summary, while business decisions require an overview of the key points of the terms. Existing technologies cannot be flexibly adapted.
[0004] Therefore, there is an urgent need for a file processing method that can overcome the deficiencies of traditional methods in document classification, information extraction, and summary flexibility. Summary of the Invention
[0005] In view of this, the present invention provides a file processing method and apparatus to solve the problems of fuzzy document classification, scattered information extraction, and insufficient summary flexibility of traditional methods. The technical solution is as follows:
[0006] In a first aspect, the present invention provides a file processing method, which includes:
[0007] Obtain a target file; the target file includes a plurality of content units; the types of content units include structured data and unstructured data;
[0008] Perform semantic extraction on the plurality of content units respectively, and generate association features between the plurality of content units correspondingly;
[0009] Classify the target file based on the association features to generate at least one classification label;
[0010] Extract domain information from the association features according to the classification label;
[0011] Generate a summary of the target file based on the domain information, classification label, and association features between the plurality of content units.
[0012] A document processing method provided by an embodiment of the present invention systematically solves the core problems in document processing, especially tendering documents, through the synergistic effect of multi-dimensional technical features. First, by dividing content units into structured and unstructured data and generating associated features, the limitation of semantic fragmentation between tables and texts in traditional methods is broken, enabling the complete capture of the logical dependency relationships of cross-modal data, thereby improving classification accuracy. Secondly, a multi-level classification mechanism based on associated features is used to achieve fine-grained identification of document attributes, avoiding misjudgments caused by semantic isolation in traditional single classification models. Finally, a dynamic summary is generated by combining domain information and associated features. Through the dual mechanisms of logical association and domain adaptation, both the accuracy of key data is retained and a readable expression that meets the user's needs is generated, thereby reducing the manual processing cost while meeting the differentiated needs of multiple roles such as bid evaluation experts and project managers for the summary content.
[0013] In an alternative embodiment, the structured data includes at least one of a table, a list of terms, and a technical parameter table; the unstructured data includes at least one of a free text paragraph and descriptive text in an image.
[0014] A document processing method provided by an embodiment of the present invention, where the structured data includes a table, a list of terms, and a technical parameter table (such as an itemized quotation sheet), and the unstructured data includes a free text paragraph (such as a technical solution description) and descriptive text in an image (such as drawing annotations in a scanned document). For tables, a structured parsing algorithm (such as table restoration based on OCR) is used, and for free text, a BERT model is used to extract semantics, resulting in a significant improvement in efficiency compared to the hybrid processing mode. Clearly defining the scope of structured data (such as a technical parameter table) can accurately locate key fields (such as unit price and total price), avoiding missed extraction caused by fuzzy definitions in traditional methods (for example, misjudging the "auxiliary machine quotation" in a nested table as ordinary text).
[0015] In an alternative embodiment, semantic extraction is performed on several content units respectively, and associated features between multiple content units are generated correspondingly, including:
[0016] Jointly encode the structured data and the unstructured data to generate an initial semantic vector;
[0017] Obtain the logical association relationships between different content units, and the logical association relationships include at least one of a parameter reference relationship, a clause dependency relationship, and a context semantic connection;
[0018] Based on the logical association relationships, correct the initial semantic vector to obtain the associated features between multiple content units.
[0019] A file processing method provided by an embodiment of the present invention maps tabular values and text descriptions to a unified vector space, and modifies semantic vectors based on parameter reference relationships and clause dependency relationships. Joint encoding solves the problem of semantic fragmentation between tables and texts. Logical association correction can identify the dependency relationships between clauses, avoid missing key constraints in manual review, make the associated features more conform to the business logic of bidding documents, and improve the accuracy of subsequent classification and summarization.
[0020] In an alternative embodiment, before performing semantic extraction on several content units, it further includes:
[0021] Perform OCR parsing, table structure restoration, and text cleaning on the target file to generate standardized structured data and unstructured data.
[0022] A file processing method provided by an embodiment of the present invention processes nested tables in scanned documents / PDFs through OCR parsing and table restoration, and removes garbled characters and irrelevant formats through text cleaning. OCR parsing supports character recognition of scanned files (such as scanned copies of paper tender documents), solves the limitation that traditional methods only support electronic documents, and expands the coverage of the bidding scenario. Table structure restoration ensures the complete parsing of the technical parameter table (such as a quotation form with nested multi-level headers) by identifying the row and column relationships of the table (such as merging cell restoration), and avoids field misalignment caused by complex structures in traditional table parsing tools. Text cleaning removes irrelevant characters (such as headers and page numbers) and abnormal spaces, improves the input quality of the semantic extraction model, and reduces misjudgments caused by noise interference (such as misidentifying "Page 3" as the main text content).
[0023] In an alternative embodiment, classifying the target file based on the associated features to generate at least one classification label, including:
[0024] Analyze the semantic information of the associated features through a pre-trained language model to identify the attribute labels of the target file;
[0025] According to the attribute labels, use a multi-label classification model to map the attribute labels to domain labels.
[0026] A file processing method provided by an embodiment of the present invention utilizes the context understanding ability of the language model to identify implicit business clauses in technical attributes, and solves the misclassification caused by semantic ambiguity in traditional classification models. Through multi-label classification, it supports accurate annotation of multi-domain cross-content in bidding documents, meets the needs of multi-dimensional information retrieval in bid evaluation. Multi-label classification supports multiple label annotation of the same file, and solves the problem of insufficient label capacity of a single classification model. Domain labels drive subsequent processing logic.
[0027] In an alternative embodiment, the target file includes a bidding document or a tender document, and the method further includes:
[0028] Analyze the semantic information of the associated features through a pre-trained language model, distinguish the target document as a tender document or a bid document, and generate a first-level classification label; and label the attributes of the target document to generate a second-level classification label; the second-level classification label includes technical attributes, commercial attributes, or cost attributes; the pre-trained language model includes a BERT model;
[0029] According to the attribute labels, use a multi-label classification model to map the attribute labels to domain labels; the domain labels include materials, auxiliary machines, equipment, desulfurization, denitration, main engines, and network and information systems; the multi-label classification model includes a RoBERTa model;
[0030] Among them, if the semantic information of the associated features contains tender content and the technical parameters are range values, then confirm that the target document is a tender document;
[0031] If the semantic information of the associated features contains bid content and the technical parameters are definite values, then confirm that the target document is a bid document.
[0032] A document processing method provided by an embodiment of the present invention quickly screens target documents through first-level classification, automatically identifies tender documents based on range value features, and avoids confusion in processing logic caused by manual annotation errors. Quickly locate bid response content through definite values, assist the bid evaluation experts in comparing tender requirements and bid commitments, and improve the bid evaluation efficiency.
[0033] In an alternative embodiment, extracting domain information from the associated features includes:
[0034] Generate an attention weight matrix according to the classification label;
[0035] Perform weighted calculation on the associated features through a multi-head self-attention mechanism to obtain the weights of each feature vector;
[0036] Compare the weights of each feature vector with a weight threshold, and extract the target feature vector from each feature vector as the domain information according to the comparison result.
[0037] A document processing method provided by an embodiment of the present invention dynamically adjusts the attention weights through classification labels, focuses on key content units, and suppresses the interference of irrelevant information. Capture semantic associations in different dimensions through a multi-head self-attention mechanism, accurately extract the core domain information, and support the generation of high-quality abstracts.
[0038] In an alternative embodiment, generating a summary of the target document includes:
[0039] Determine a summary template according to the domain information and classification label;
[0040] Fill the template content based on the associated features of multiple content units to generate an abstract of the target file.
[0041] A file processing method provided by an embodiment of the present invention automatically matches a template based on domain tags to ensure that the abstract content is highly adapted to user needs (such as the technical details concerned by bid evaluation experts). Integrate information across data units (such as parameters in a table + clauses in text) into the same abstract to avoid the problem of information fragmentation in traditional abstract generation modules (such as only extracting table fields and ignoring related text).
[0042] In an optional implementation manner, filling the template content based on the associated features of multiple content units includes:
[0043] Screen the associated features between multiple content units through the TextRank algorithm, screen out high-weight text paragraphs and table fields, and generate an extractive abstract;
[0044] Input the extractive abstract into a Seq2Seq model to generate a natural language text abstract.
[0045] A file processing method provided by an embodiment of the present invention ensures that the abstract contains the core information of bidding documents (such as key performance indicators and quoted amounts) by calculating the importance of content units (such as the weight of a technical parameter table is higher than that of conventional descriptive text), and solves the problem of missing key points caused by traditional random extraction. The Seq2Seq generation converts the extractive abstract into coherent natural language, improves the readability of the abstract, and meets the fast reading needs of non-technical users.
[0046] In summary, for the file processing method provided by the present invention, first, through data preprocessing, multi-format files are standardized to provide high-quality input for subsequent processing; second, joint encoding and logical association correction break through the semantic barriers between structured and unstructured data, enabling the multi-level classification model to achieve accurate classification based on fine-grained semantics; further, the combination of the attention weight mechanism and domain label mapping ensures that information extraction focuses on core content; finally, hybrid abstract generation through templated filling and natural language conversion not only retains the accuracy of key data but also generates a readable expression that conforms to the user scenario. The synergistic effect of the above technical features enables the system to achieve a breakthrough improvement in classification accuracy, information extraction integrity, and abstract generation efficiency, while being compatible with multi-format files and large-scale real-time processing requirements.
[0047] In a second aspect, the present invention provides a file processing device, and the device includes:
[0048] An acquisition module, configured to acquire a target file; the target file includes a plurality of content units; the types of the content units include structured data and unstructured data;
[0049] A parsing module, configured to perform semantic extraction on the plurality of content units respectively, and correspondingly generate association features between multiple content units;
[0050] A classification module, configured to classify the target file based on the association features, and generate at least one classification label;
[0051] An extraction module, configured to extract domain information from the association features according to the classification label;
[0052] A generation module, configured to generate an abstract of the target file based on the domain information, the classification label, and the association features between multiple content units.
[0053] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, wherein the memory stores computer instructions, and the processor executes the computer instructions to execute the file processing method according to the first aspect or any corresponding embodiment thereof.
[0054] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the file processing method according to the first aspect or any corresponding embodiment thereof.
[0055] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the file processing method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings
[0056] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 is a schematic flowchart of a file processing method according to an embodiment of the present invention;
[0058] Figure 2 is a schematic flowchart of a bidding document processing method according to an embodiment of the present invention;
[0059] Figure 3 is a structural block diagram of a file processing device according to an embodiment of the present invention;
[0060] Figure 4 is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0061] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] In the field of complex document processing technology, traditional file classification and information extraction methods have long faced multi-dimensional technical bottlenecks. Existing systems usually process data based on a single modality (such as only processing text or tables), resulting in the breakage of cross-modal logical associations. Specifically, the core problems of the existing technologies are that most solutions only support first-level classification (such as distinguishing tender / bid documents) and cannot achieve multi-level mapping of "file type - attribute category - domain label". Entity recognition methods based on keyword matching (such as regular expressions) are difficult to capture implicit relationships in nested structures. Existing abstract algorithms (such as those based on template filling) cannot dynamically adapt to the needs of user roles. Evaluation experts need to focus on technical parameter comparison, while financial personnel are concerned about cost composition, but traditional systems can only generate fixed-format abstracts, resulting in information overload or omission of key data.
[0063] In the typical application scenario of tendering and bidding document processing, the above defects are particularly acute. Tendering documents and bidding documents usually contain a large amount of complex structured and unstructured information. Tendering documents usually include technical requirements, commercial terms and other instructions, while bidding documents contain commercial quotations, technical solutions and price lists. These documents are not only long in length but also diverse in structure, and the manual review and classification process is time-consuming and error-prone. Especially in large-scale projects, the quantity of documents submitted by multiple bidders is extremely large, further increasing the burden of manual processing.
[0064] The limitations of the existing technologies in such tendering and bidding document processing scenarios are particularly significant and usually cannot meet the specific requirements of tendering and bidding documents. Specifically, existing methods are difficult to accurately classify documents according to the context semantics of the document content, especially when distinguishing technical, commercial and price documents, confusion often occurs. There are many important entities and relationships in tender and bidding documents, such as supplier names, key technical parameters, quoted amounts, etc. These information are often nested in complex sentences or tables, and simple rule-based methods are difficult to extract complete and accurate data. Traditional methods cannot effectively generate document abstracts for specific requirements, such as only extracting technical details related to a specific field (such as desulfurization equipment parameters).
[0065] Therefore, to address the deficiencies of existing technologies in complex document processing, embodiments of the present invention provide a file processing method that can solve the problems of fuzzy file classification, scattered information extraction, and insufficient flexibility in abstracts in traditional methods.
[0066] The method flow of a file processing method provided by embodiments of the present invention is as Figure 1 shown and includes the following steps.
[0067] S101. Obtain a target file; the target file includes a number of content units; the types of content units include structured data and unstructured data.
[0068] Specifically, in S101, the target file is an electronic document to be processed, including but not limited to formats such as PDF, Word, scanned images, etc. A content unit is the smallest logical unit in the document that is segmented and processed, and is divided into two categories: structured data and unstructured data. Structured data is data with a fixed format, such as a table (a set of data aligned in rows and columns), a list of terms (a set of entries with numbers / symbols), a technical parameter table (a combination of parameter name - value - unit). Unstructured data is a free text paragraph without a fixed format or descriptive text in an image (such as the text description in a scanned document).
[0069] S102. Perform semantic extraction on each of the several content units and generate association features between multiple content units accordingly.
[0070] Specifically, in S102, semantic extraction is to perform semantic analysis on each content unit and extract its core meaning. For example, the table field "Unit Price: ¥100" is extracted as "Price Entity - Value 100 - Currency Unit ¥". The text paragraph "The supplier promises to deliver within 10 days" is extracted as "Delivery Clause - Time 10 days - Responsible Party Supplier". Association features refer to the logical or semantic association relationships between different content units. For example, "Total Price = Unit Price × Quantity" in the table and "The quotation includes taxes" in the text form a numerical dependency relationship. "Payment Terms: 30% Advance Payment" in the list of terms and "Performance Bond Template" in the attachment form a clause reference relationship.
[0071] S103. Classify the target file based on the association features and generate at least one classification label.
[0072] Specifically, in S103, the classification label is a category identifier that describes the attributes of a document. For example, at the first level of classification: file type labels (such as "bidding documents", "tender documents"). At the second level of classification: file attribute labels (such as "technical documents", "commercial documents", "price documents"). At the third level of classification: field labels (such as "equipment field", "materials field"). The classification basis determines the classification logic according to the type and distribution of associated features. For example, if the relationship between technical parameters and clause references in the associated features is dominant, it is classified as a "technical document".
[0073] S104. Extract domain information from the associated features according to the classification label.
[0074] Specifically, in S104, the domain information is the professional field or topic information to which the document content belongs. For example, if the classification label is a "technical document" and "desulfurization efficiency" and "sulfur capacity" frequently appear in the associated features, the domain information is "desulfurization equipment". The extraction method of domain information is to screen the associated features related to the classification label.
[0075] S105. Generate an abstract of the target document based on the domain information, classification label, and the associated features between multiple content units.
[0076] Specifically, in S105, integrate the domain information, classification label, and associated features, and output a concise document summary. Determine the focus of the abstract according to the classification label (such as technical parameters, commercial terms), extract the key content through the associated features, and organize the sentences according to the natural language rules.
[0077] A file processing method provided by an embodiment of the present invention achieves a systematic improvement in the efficiency of complex document processing through the synergistic effect of progressive technical features. Firstly, by splitting the target file into structured data and unstructured data, a differentiated parsing logic is designed to solve the format conflict and information omission problems caused by mixed processing in traditional methods. Secondly, semantic extraction is performed on each content unit and cross-unit correlation features are constructed, which breaks the limitation of data islands in traditional methods and forms a semantic network for key information scattered in tables and texts, providing a multi-dimensional contextual basis for subsequent processing. On this basis, multi-level classification labels are generated based on the distribution law of correlation features, and fine-grained classification is achieved by dynamically perceiving contextual semantics to avoid the risk of misjudgment in traditional single-label classification. Furthermore, the classification labels are combined to filter the core information of the domain in the correlation features, and redundant data interference is accurately eliminated to ensure the targeting of information extraction. Finally, a dynamic summary is generated by integrating domain information, classification labels and correlation features, and the focus of the summary is adaptively adjusted according to the classification results. The correlation features are used to ensure logical coherence, so as to output a highly readable summary that meets the needs of multiple roles while retaining data accuracy, greatly reducing the time spent on manual review, and systematically solving the core problems of information fragmentation, extensive classification, and rigid summary in traditional document processing.
[0078] Optionally, in the above steps, structured data includes tables, clause lists, and technical parameter tables (such as itemized quotations), and unstructured data includes free text paragraphs (such as technical solution descriptions) and descriptive text in images (such as drawing annotations in scanned copies). Using structured parsing algorithms (such as OCR-based table restoration) for tables and using the BERT model to extract semantics for free text greatly improves efficiency compared to the hybrid processing mode. Clearly defining the scope of structured data (such as technical parameter tables) can accurately locate key fields (such as unit price and total price), avoiding omissions caused by fuzzy definitions in traditional methods (for example, misjudging the "auxiliary machine quotation" in a nested table as ordinary text).
[0079] Optionally, in the above step S102, semantic extraction is performed on several content units respectively, and correlation features between multiple content units are generated correspondingly. Specifically, it includes: jointly encoding structured data and unstructured data to generate an initial semantic vector; obtaining the logical correlation relationships between different content units, where the logical correlation relationships include at least one of parameter reference relationship, clause dependency relationship, and context semantic connection; based on the logical correlation relationships, correcting the initial semantic vector to obtain the correlation features between multiple content units. Mapping table values and text descriptions to a unified vector space, and correcting the semantic vector based on the parameter reference relationship and clause dependency relationship. Joint encoding solves the problem of semantic fragmentation between tables and texts. Logical correlation correction can identify the dependency relationships between clauses, avoid missing key constraints in manual review, make the correlation features more conform to the business logic of bidding documents, and improve the accuracy of subsequent classification and summarization.
[0080] Before performing semantic extraction on several content units respectively, a preprocessing process is also included, which performs OCR parsing, table structure restoration, and text cleaning on the target file to generate standardized structured data and unstructured data. The nested tables in scanned documents / PDFs are processed through OCR parsing and table restoration, and garbled characters and irrelevant formats are removed through text cleaning. OCR parsing supports character recognition of scanned files (such as scanned paper tender documents), solves the limitation that traditional methods only support electronic documents, and expands the coverage of the bidding scenario. Table structure restoration ensures the complete parsing of the technical parameter table (such as a quotation form with nested multi-level headers) by identifying the row and column relationships of the table (such as restoring merged cells), and avoids field misalignment caused by complex structures in traditional table parsing tools. Text cleaning removes irrelevant characters (such as headers and page numbers) and abnormal spaces, improves the input quality of the semantic extraction model, and reduces misjudgments caused by noise interference (such as misidentifying "Page 3" as the main text content).
[0081] Optionally, in the above step S103, based on the correlation features, the target file is classified to generate at least one classification label, including: analyzing the semantic information of the correlation features through a pre-trained language model to identify the attribute labels of the target file; according to the attribute labels, using a multi-label classification model to map the attribute labels to domain labels.
[0082] Taking the bidding documents including the tender documents and bid documents as an example of the target documents, the semantic information of the associated features is analyzed through a pre-trained language model to distinguish whether the target document is a tender document or a bid document, and a first-level classification label is generated; and the attributes of the target document are marked to generate a second-level classification label; the second-level classification label includes technical attributes, commercial attributes or cost attributes; the pre-trained language model includes the BERT model; according to the attribute labels, a multi-label classification model is used to map the attribute labels to domain labels; the domain labels include materials, auxiliary machines, equipment, desulfurization, denitration, main engines, and network and information systems; the multi-label classification model includes the RoBERTa model; among them, if the semantic information of the associated features contains tender content and the technical parameters are range values, it is confirmed that the target document is a tender document; if the semantic information of the associated features contains bid content and the technical parameters are definite values, it is confirmed that the target document is a bid document. By using the context understanding ability of the language model, the implicit commercial terms in the technical attributes are identified, which solves the misclassification caused by semantic ambiguity in the traditional classification model. Through multi-label classification, accurate annotation of multi-domain cross-content in the bidding documents is supported, meeting the requirements for multi-dimensional information retrieval in bid evaluation. Multi-label classification supports multiple label annotations for the same document, solving the problem of insufficient label capacity of a single classification model. The target documents are quickly screened through the first-level classification, and the tender documents are automatically identified based on the range value features, avoiding the confusion of processing logic caused by manual annotation errors. The bid response content is quickly located through the definite value, assisting the bid evaluation experts to compare the tender requirements with the bid commitments and improving the bid evaluation efficiency.
[0083] Optionally, in the above step S104, extracting domain information from the associated features specifically includes: generating an attention weight matrix according to the classification label; performing weighted calculation on the associated features through the multi-head self-attention mechanism to obtain the weights of each feature vector; comparing the weights of each feature vector with a weight threshold, and extracting the target feature vector from each feature vector as the domain information according to the comparison result. Dynamically adjust the attention weights through the classification label, focus on the key content units, and suppress the interference of irrelevant information. Capture the semantic associations in different dimensions through the multi-head self-attention mechanism, accurately extract the core domain information, and support the generation of high-quality abstracts.
[0084] Optionally, in the above step S105, determine an abstract template according to the field information and classification tags; fill the template content based on the association features of multiple content units to generate an abstract of the target file. Automatically match the template based on the field tags to ensure that the abstract content is highly adapted to the user requirements (such as the technical details concerned by the bid evaluation experts). Integrate the information across data units (such as the parameters in the table + the terms in the text) into the same abstract to avoid the information fragmentation problem of the traditional abstract generation module (such as only extracting the table fields and ignoring the associated text). Specifically, screen the association features between multiple content units through the TextRank algorithm, screen out the high-weight text paragraphs and table fields, and generate an extractive abstract; input the extractive abstract into the Seq2Seq model to generate a natural language text abstract. By calculating the importance of the content units (such as the weight of the technical parameter table is higher than that of the regular descriptive text), ensure that the abstract contains the core information of the bidding documents (such as key performance indicators and quoted amounts), and solve the problem of missing key points caused by traditional random extraction. The Seq2Seq generation transforms the extractive abstract into coherent natural language, improves the readability of the abstract, and meets the fast reading needs of non-technical users.
[0085] In summary, for the file processing method provided by the embodiments of the present invention, first, standardize multi-format files through data preprocessing to provide high-quality input for subsequent processing; second, combine encoding and logical association correction to break through the semantic barriers between structured and unstructured data, enabling the multi-level classification model to achieve accurate classification based on fine-grained semantics; further, combine the attention weight mechanism and the field tag mapping to ensure that information extraction focuses on the core content; finally, the hybrid abstract generation through templated filling and natural language transformation not only retains the accuracy of key data but also generates a readable expression that conforms to the user scenario. The synergistic effect of the above technical features enables the system to achieve a breakthrough improvement in classification accuracy, information extraction integrity, and abstract generation efficiency, while being compatible with multi-format files and large-scale real-time processing requirements.
[0086] Exemplarily, the file processing method provided by the above embodiments will be described in detail below using the typical application scenario of bidding document processing. Based on the file processing method provided by the above embodiments, a processing system based on bidding documents is constructed for classifying and extracting abstracts of bidding documents.
[0087] The bidding document processing system includes: a data preprocessing module, a classification module, an information extraction module, an abstract generation module, a field mapping module, and a user interaction and visualization module.
[0088] The data preprocessing module is used for file parsing, text cleaning and generating structured data, supporting the parsing of files in formats such as PDF and Word, and processing scanned documents in combination with OCR technology. It deletes irrelevant characters, spaces and abnormal formats, and extracts pure text content. It identifies tables and nested format content and converts them into parsable structured data.
[0089] The classification module is used for the initial classification of files for secondary fine classification. It uses deep learning models (such as BERT or RoBERTa) to classify documents, distinguishing between tender documents and bid documents. On the basis of the initial classification, it is further subdivided into technical, commercial and price documents, and the corresponding subdivision modules (such as materials, auxiliary machines, etc.) are marked.
[0090] The information extraction module is used to extract key information such as supplier names, bid amounts, and technical indicators through the BERT-CRF model. Combining with the BiLSTM+Attention model, it extracts complex relationships such as payment terms and delivery times. It logically associates multi-segment information to generate a complete set of key information.
[0091] The abstract generation module is used to calculate the importance of sentences using the TextRank algorithm and generate extractive abstracts. It uses pre-trained generation models (such as T5, GPT) to generate smooth and coherent abstract abstracts. At the same time, it supports adjusting the abstract length and granularity based on user needs to meet the requirements of different scenarios.
[0092] The domain mapping module uses a multi-label classification model to map the document content to seven major domains (materials, auxiliary machines, equipment, desulfurization, denitration, main engines, network and information systems). It generates specific content abstracts for the domains and supports customized processing for bid evaluation requirements.
[0093] The user interaction and visualization module is used to provide an intuitive graphical interface, supporting file upload, viewing of classification results and abstracts. It supports exporting the results in formats such as JSON and Excel for subsequent evaluation and recording. It achieves efficient file classification and abstract extraction through asynchronous processing and provides dynamic result display.
[0094] The processing flow of this tender and bid document processing system is as Figure 2 shown, including the following processes:
[0095] Users upload tender documents or bid documents; the data preprocessing module parses the document format and cleans the text; the classification module identifies the document category and further subdivides it; the information extraction module extracts key information and aggregates relevant content; the abstract generation module generates high-quality abstracts and supports user customization; the domain mapping module completes content classification and domain label marking; users view the results through the interface and export or adjust the abstract according to their needs.
[0096] The functional implementation methods of each module will be described in detail below.
[0097] The functions of the data preprocessing module include OCR parsing, table parsing, and text cleaning.
[0098] For OCR parsing, the Tesseract OCR model is used for character recognition, and combined with a deep learning model (such as CRAFT) to improve the character recognition accuracy of scanned documents. The final output of OCR parsing is structured plain text data or table images.
[0099] Character recognition can be regarded as a classification problem. Assuming a given image segment I, the OCR model outputs a probability distribution, and the expression is:
[0100] ;
[0101] In the formula, c is the possible character category, I is the image to be recognized, the softmax function converts the result after linear transformation into a probability distribution, f(I) is the image feature extractor (such as the feature vector extracted by CNN), and W and b are the weight matrix and bias of the model.
[0102] The CRAFT (Character Region Awareness for Text Detection) model realizes more accurate text extraction by detecting the boundaries of text regions. The loss function of the detection region is:
[0103] ;
[0104] In the formula, Lregion is the text region detection loss, Laffinity is the character association loss, and λ1 and λ2 are weight hyperparameters.
[0105] For table parsing, deep learning models such as TabularNet are used to recognize complex table structures and generate parsable structured data. Recursive parsing is performed on nested tables to extract the logical relationship between the table header and the data.
[0106] Based on the TabularNet model, the rows, columns, and cell boundaries of the table are detected. Its prediction process is as follows:
[0107] ;
[0108] In the formula, S represents the boundary of the table cell (such as the upper left and lower right coordinates), Ws and bs are the weights and biases of the table detection layer respectively, and f(I) is the image feature.
[0109] For nested tables, a recursive parsing strategy is adopted. The nested area R is detected from the main table, and R is used as the input for the sub-table. The above table detection process is repeated for the sub-table until the parsing is completed. The recursive formula for recursive parsing:
[0110] ;
[0111] where R k represents the nested area at the k-th layer, and f detect is the table detection function.
[0112] Parse the mapping relationship between the header and data cells based on a rule engine or neural network:
[0113] ;
[0114] where M(h, d) represents the mapping relationship between the header h and the data cell d. The argmax function returns the independent variable value corresponding to the maximum value of the sim function. H and D are the sets of headers and data cells respectively, and sim(h, d) is the similarity function (such as cosine similarity or deep matching model).
[0115] Text cleaning is a cleaning module based on regular expressions and a rule engine, which deletes irrelevant characters and unifies the text format. Use the regular expression R1 to delete invalid characters, and the expression is:
[0116] ;
[0117] where R1 is the regular expression for matching invalid characters, and Re.sub is the regular matching replacement function.
[0118] Unify the date format, and the expression is:
[0119] ;
[0120] where R2 is the regular expression matching rule for the date format, and YYYY-MM-DD is the general date format for year, month, and day.
[0121] The classification module uses a pre-trained language model (such as BERT) for fine-tuning and then classification. The first-level classification is to distinguish between tender documents and bid documents. The second-level classification is further subdivided into technical, commercial, and price documents according to the content.
[0122] The text input generates feature representations through the BERT model, and the expression is:
[0123] ;
[0124] where xi Denote the i-th sentence or paragraph of the input as h i Denote the context feature vector output by BERT as
[0125] Use a classifier to predict the feature vector, and the expression is:
[0126] ;
[0127] In the formula, W and b are the weight matrix and bias vector parameters of the classifier, and P(y∣x) is the classification probability distribution.
[0128] Use the multi-task learning method to jointly train the first-level classification and second-level classification models, and the expression is:
[0129] ;
[0130] In the formula, α and β are the task weights, and L 一级分类 and L 二级分类 are the cross-entropy losses of the classification tasks.
[0131] Introduce the attention mechanism to optimize the classification results, especially for documents containing multiple topics. Extract important contexts through the attention mechanism, and the expression is:
[0132] ;
[0133] In the formula, is the weight matrix, which performs a linear transformation on the feature vector
[0134] The extracted feature vector can be expressed as:
[0135] ;
[0136] In the formula, is the attention weight, and H is the weighted feature representation.
[0137] The specific classification steps are as follows:
[0138] First, the weighted feature representation H obtained through the attention mechanism has integrated the important information of each part of the document. Input the weighted feature representation H into the fully connected layer of the classifier, and the expression is:
[0139] ;
[0140] In the formula, is the weight matrix of the first-level / second-level classifier, is the bias vector.
[0141] Subsequently, the obtained result vector is converted into probabilities using the softmax function, with the expression:
[0142] ;
[0143] Finally, according to the probability distribution output by the softmax function, the category with the highest probability is selected as the result of the final classification, with the expression:
[0144] ;
[0145] This classification process can be used for quickly classifying bidding documents. For example, after classifying the technical solutions and quotation documents of the bidding documents, they are respectively handed over to the evaluation experts. In the scenario of multiple themes in complex documents (such as those containing both technology and business), the classification accuracy is improved.
[0146] The specific extraction process of the information extraction module is as follows:
[0147] First, the BERT-CRF model is used to perform entity recognition in combination with the domain dictionary. The recognized categories include: supplier name, amount, date, technical parameters, etc.
[0148] Specifically, after using BERT to extract context features, the CRF layer is used to predict the entity label sequence, with the expression:
[0149] ;
[0150] In the formula, s t (y t ) is the score of BERT for the label y t , and A[y t−1 ,y t is the transition score between labels.
[0151] Subsequently, relation extraction is performed based on the relation extraction model with a two-tower architecture (BiLSTM+Attention).
[0152] Specifically, assume that the input entity pair is: (e1, e2); the BiLSTM is used to obtain the context features of the entities, with the expression:
[0153] ;
[0154] In the formula, is the feature representation of the text words.
[0155] The context between entities is weighted through the attention mechanism, with the expression:
[0156] ;
[0157] Predict the relationship using a classifier, and the expression is:
[0158] ;
[0159] In the formula, and are the weight matrix and bias vector of the classifier.
[0160] Finally, use the graph neural network (GNN) to establish associations for the cross-segment scattered information to form a complete context.
[0161] Specifically, establish the relationship between paragraph nodes and The expression is:
[0162] ;
[0163] In the formula, N(i) is the set of neighbor nodes of node , represents the feature vector of node after k+1 iterations of update, is the weight matrix. is the activation function, which performs a non-linear transformation on the obtained result.
[0164] The abstract generation module includes extractive abstract generation and abstractive abstract generation, and supports dynamic abstract adjustment.
[0165] Extractive abstract generation uses the TextRank algorithm, combined with sentence vector similarity calculation, to screen high-importance sentences. The expression is:
[0166] ;
[0167] In the formula, In( ) is the set of all nodes pointing to node , Out( ) is the set of all nodes pointed to by node , is the weight of the edge from node to node , is the weight of all edges pointed to by node , and d is the damping coefficient.
[0168] Abstractive abstract generation is based on the Seq2Seq model (such as T5) to generate natural language abstracts.
[0169] Specifically, it optimizes the generation quality through reinforcement learning and trains in combination with the ROUGE metric. The expression is:
[0170] ;
[0171] In the formula, T represents the length of the generated sequence, and t is a time step index, taking values from 1 to T in sequence. is the true word (or token) generated at time step t. P represents the probability of generating under the condition of the given input sequence and the generated word sequence.
[0172] The dynamic summary adjusts the summary based on user needs and can generate long summaries or short summaries.
[0173] The domain mapping module uses the RoBERTa model for multi-label classification and maps the document to seven domains. The expression of the document is:
[0174] ;
[0175] The domain mapping expression is:
[0176] ;
[0177] The RoBERTa model can also be fine-tuned with domain data to improve the classification accuracy of the few-shot domain. The expression of fine-tuning is:
[0178] ;
[0179] In the formula, N represents the number of samples, represents the true label of the i-th sample. represents the predicted probability of the model for the i-th sample.
[0180] The specific process of fine-tuning is as follows:
[0181] Collect a large number of bidding documents from the public resource trading platform channels, and conduct screening and annotation.
[0182] Select the RoBERTa model, load the weights of the pre-trained model, set relevant hyperparameters, select a suitable learning rate range and dynamically monitor and adjust, and reasonably set the batch size according to the computing resources and data volume, and determine the appropriate number of training epochs, etc. through methods such as cross-validation.
[0183] Introduce L1 or L2 regularization to constrain the parameters of the model, prevent overfitting, improve the generalization ability of the model, and adopt corresponding data augmentation methods for different types of data.
[0184] For small-sample data, usually some layers are frozen first, and only the last few layers are fine-tuned. After the model is stable, more layers are gradually unfrozen for fine-tuning. The prepared domain data is input into the model for training, and the performance metrics of the model, such as accuracy, recall, F1 value, etc., are monitored, and the training strategy is adjusted in a timely manner.
[0185] The trained model is evaluated using the test set, and various performance metrics are calculated to comprehensively understand the classification ability of the model in the small-sample domain. The evaluation results are analyzed to find out the problems existing in the model, and methods such as adjusting hyperparameters, improving data augmentation methods, and replacing the model structure are used for optimization. The optimized model is deployed to the actual application environment to ensure the stable operation of the model.
[0186] The user interaction and visualization module displays the classification, information extraction, and summary generation results through a visual interface and supports export in multiple formats (such as JSON, Excel). Based on the asynchronous processing framework, after the user uploads a file, they can view the classification results and the progress of summary generation in real time.
[0187] In addition, the model parameters in each module can be adjusted in combination with user evaluations, especially the model used in the classification module, to continuously improve the classification and generation effects.
[0188] Specifically, if the user feedback shows that the classification effect of the model fluctuates greatly on new data, or the generated content is unstable, it may be that the learning rate is too large and needs to be appropriately reduced, such as adjusting from the current 0.001 to 0.0001. This will make the model take smaller steps during parameter updates and converge more robustly, improving the stability of classification and generation. On the contrary, if the model has a slow learning speed and is difficult to quickly adapt and accurately classify or generate when facing new complex types of bidding documents, the learning rate can be appropriately increased, such as increasing from 0.0001 to 0.0005, to speed up the model's learning speed for new features.
[0189] When the user reports that there are concentrated errors in the classification of bidding documents for specific large-scale projects and the training time is long, the batch size can be considered to be increased. For example, from 32 to 64, which can enable the model to calculate based on more data during each parameter update, reduce the variance of gradient estimation, and improve the training efficiency. If the classification results vary significantly between different batches, it indicates that the model is sensitive to data batches, and the batch size can be reduced to 16 to increase the randomness of training, enabling the model to better learn data features and improve generalization ability.
[0190] If the user feedback indicates that the model performs well on known data but has poor classification or generation results for new bidding documents, there may be an overfitting problem. In this case, the regularization parameter should be increased. For example, the L2 regularization coefficient is increased from 0.001 to 0.01 to impose stronger constraints on the model parameters and reduce the risk of overfitting. If the model performs poorly on all types of data, it may be underfitting, and the regularization parameter needs to be decreased, such as to 0.0001, to allow the model more freedom to learn data features.
[0191] If the user feedback shows that the model is inaccurate in classifying bidding documents with complex structures and numerous clauses, or the generated abstract fails to comprehensively cover key information, one can try increasing the number of neurons in the hidden layer to endow the model with stronger feature learning and expression capabilities. For example, the number of neurons in a certain hidden layer is increased from 128 to 256. If the model training time is too long and the improvement in classification and generation results is not obvious, the number of neurons in the hidden layer can be appropriately reduced, such as from 256 to 192, to simplify the model structure and improve the training efficiency.
[0192] In summary, the bidding document processing system provided in this example realizes the automatic classification of bidding documents and tender documents through deep learning technology. Compared with the traditional method that relies on manual screening, it significantly improves the efficiency and reduces the error rate. The system can conduct detailed classification of technical, commercial, and price documents according to the content of the documents, providing accurate support for users to quickly locate key information. At the same time, the system further refines the classification to specific fields (such as materials, auxiliary machines, equipment, etc.), making the classification results more targeted and facilitating the division of labor among different review experts. This function provides a reliable solution for the batch processing of large-scale bidding documents, greatly reducing the time cost. It realizes the accurate extraction of key information in the documents (such as supplier names, amounts, dates, technical parameters, etc.). With the cross-segment information aggregation technology, the system can integrate relevant data scattered in different paragraphs into a logically consistent structured content, ensuring the comprehensiveness and consistency of information extraction. Compared with traditional manual extraction, this invention significantly reduces human omissions, provides high-quality data support for review experts, and demonstrates its superiority especially in the parsing of complex documents and tables. By combining extractive summaries with abstractive summaries, this invention provides users with flexible and diverse summary generation capabilities. The extractive summary uses algorithms to automatically screen important sentences in the document and generate a condensed version of high-weight content, which is suitable for quick browsing. The abstractive summary, on the other hand, generates smooth and natural text through a generative model, facilitating the interpretation of complex technical solutions or commercial terms. In addition, users can adjust the length and granularity of the summary according to their needs and obtain content that fits the usage scenario in real time, thereby greatly improving the decision-making efficiency and user experience. Through multi-label classification technology and domain mapping models, the document content is accurately divided into seven key fields (such as materials, auxiliary machines, equipment, etc.). This field-specific support enables the system to automatically adapt to the needs of different industries and specialties, enhancing the practicality of the classification and summary results. The domain mapping module also combines adaptive optimization technology to improve the confidence and interpretability of the classification, providing an efficient auxiliary tool for domain experts to allocate documents and significantly accelerating the review process. The system also has an intuitive user interface, enabling users to complete tasks such as file upload, viewing classification results, and generating summaries through simple operations. The system supports the export of results in multiple formats (such as JSON, Excel), facilitating subsequent data processing. At the same time, the real-time feedback function allows users to dynamically adjust the classification and summary strategies during the operation to ensure that the output results match the actual needs. This user-friendly design significantly reduces the usage threshold of the system, and both technical experts and business personnel can easily get started, greatly enhancing the overall user experience. It supports the parsing of multiple file formats (such as PDF, Word, pictures, etc.) and can be compatible with diverse document types submitted by different industries and suppliers, including complex tables and nested structures.The modular design enables the system to be flexibly expanded according to user needs. For example, new parsing rule libraries can be added or the generation model can be adjusted to adapt to future diversified business requirements. This feature not only ensures the stability of the system in the current application scenario but also provides technical support for future industry expansion.
[0193] In this embodiment, a file processing device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0194] This embodiment provides a file processing device, as Figure 3 shown, including:
[0195] An acquisition module 301, configured to acquire a target file; the target file includes a plurality of content units; the types of the content units include structured data and unstructured data;
[0196] A parsing module 302, configured to perform semantic extraction on the plurality of content units respectively, and generate association features between the plurality of content units correspondingly;
[0197] A classification module 303, configured to classify the target file based on the association features to generate at least one classification label;
[0198] An extraction module 304, configured to extract domain information from the association features according to the classification label;
[0199] A generation module 305, configured to generate an abstract of the target file based on the domain information, the classification label, and the association features between the plurality of content units.
[0200] The further function descriptions of the above-mentioned modules and units are the same as those in the corresponding above embodiments, and will not be repeated here.
[0201] The file processing device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0202] An embodiment of the present invention also provides a computer device having the above Figure 3 shown file processing device.
[0203] Please refer to Figure 4 ,Figure 4 FIG. Figure 4 is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 4 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 4 In
[0204] FIG. Figure 4 , one processor 10 is taken as an example.
[0205] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device may be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0206] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0207] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device. In addition, the memory 20 may include a high-speed random access memory, and may further include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0208] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected by a bus or other means. Figure 4 Taking the connection by bus as an example.
[0209] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0210] The embodiments of the present invention also provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0211] A part of the present invention can be applied as a computer program product, such as computer program instructions, which when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0212] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A file processing method, characterized in that: The method comprises: Acquire a target file; the target file includes a plurality of content units; the types of the content units include structured data and unstructured data; Performing semantic extraction on the plurality of content units respectively, and correspondingly generating association features between the plurality of content units; Classify the target file based on the associated features and generate at least one classification label; Extracting domain information from the associated features according to the classification labels; Generate a summary of the target file based on the domain information, the classification label, and the association features between the plurality of content units; The semantic extraction of the plurality of content units is respectively performed, and the associated features between the plurality of content units are correspondingly generated, including: Jointly encode structured data and unstructured data to generate an initial semantic vector; Acquire a logical association relationship between different content units, wherein the logical association relationship includes at least one of a parameter reference relationship, a clause dependency relationship, and a context semantic connection; Based on the logical association relationship, the initial semantic vector is modified to obtain association features between multiple content units.
2. The method according to claim 1, characterized in that The structured data includes at least one of a table, a list of terms and a technical parameter table; the unstructured data includes at least one of a free text paragraph and a descriptive text in an image.
3. The method according to claim 1, characterized in that Before extracting the semantics of the plurality of content units respectively, the method further includes: The target file is subjected to OCR analysis, table structure restoration and text cleaning to generate standardized structured data and unstructured data.
4. The method according to claim 1, characterized in that: The step of classifying the target file based on the associated features to generate at least one classification label includes: Analyzing semantic information of the associated features through a pre-trained language model to identify attribute labels of the target file; According to the attribute labels, a multi-label classification model is used to map the attribute labels to domain labels.
5. The method according to claim 4, characterized in that The target document includes a bidding document or a tender document, and the method further includes: Analyze the semantic information of the associated features through a pre-trained language model, distinguish the target file as a bidding document or a tender document, and generate a primary classification label; and annotate the attributes of the target file to generate a secondary classification label; the secondary classification label includes a technical attribute, a business attribute or a cost attribute; the pre-trained language model includes a BERT model; According to the attribute labels, a multi-label classification model is used to map the attribute labels to domain labels; the domain labels include materials, auxiliary machines, equipment, desulfurization, denitrification, host, and network and information systems; the multi-label classification model includes a RoBERTa model; If the semantic information of the associated feature includes the bidding content and the technical parameter is a range value, then the target document is confirmed to be a bidding document; If the semantic information of the associated feature includes bidding content and the technical parameter is a definite value, the target document is confirmed to be a bidding document.
6. The method according to claim 5, characterized in that The extracting domain information from the associated features includes: Generate an attention weight matrix according to the classification label; The associated features are weighted by a multi-head self-attention mechanism to obtain the weight of each feature vector; The weight of each feature vector is compared with the weight threshold, and the target feature vector is extracted from each feature vector as the domain information according to the comparison result.
7. The method according to claim 6, characterized in that The generating a summary of the target file comprises: Determine a summary template according to the field information and classification labels; The template content is filled based on the associated features of the plurality of content units to generate a summary of the target file.
8. The method according to claim 7, characterized in that The filling of the template content based on the associated features of the plurality of content units includes: The TextRank algorithm is used to filter the correlation features between multiple content units, select high-weight text paragraphs and table fields, and generate extractive summaries; The extractive summary is fed into a Seq2Seq model to generate a natural language text summary.
9. A file processing device, characterized in that: The device comprises: An acquisition module, used for acquiring a target file; the target file includes a plurality of content units; the types of the content units include structured data and unstructured data; A parsing module, used to perform semantic extraction on the plurality of content units respectively, and correspondingly generate correlation features between the plurality of content units; A classification module, used to classify the target file based on the associated features and generate at least one classification label; An extraction module, used to extract domain information from the associated features according to the classification labels; A generation module, configured to generate a summary of the target file based on the domain information, the classification label, and the association features between the plurality of content units; Wherein, the analysis module is specifically used for: Jointly encode structured data and unstructured data to generate an initial semantic vector; Acquire a logical association relationship between different content units, wherein the logical association relationship includes at least one of a parameter reference relationship, a clause dependency relationship, and a context semantic connection; Based on the logical association relationship, the initial semantic vector is modified to obtain association features between multiple content units.
Citation Information
Patent Citations
Transform structure-based intelligent evaluation report generation method and system
CN118261163A
Large model multi-modal data semantic representation alignment method
CN119380341A
Cited By
Document content logic review system based on large language model
CN121278111A