File processing method, device and equipment based on AI, RPA, LLM and AI Agent

By combining multiple visual recognition models and LLM with an AI Agent to generate a set of structural and semantic feature tags for electronic documents, the problem of low document processing efficiency and high cost in existing technologies is solved, and efficient and accurate document attribute tag generation is achieved.

CN121789231APending Publication Date: 2026-04-03BEIJING BENYING NETWORK TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

When dealing with massive amounts of documents, existing technologies are inefficient and costly when using purely manual methods, while semi-automatic methods using single models have low recognition accuracy and poor system scalability, making it difficult to meet timeliness and consistency requirements.

Method used

Multiple target visual recognition models and large language models (LLM) are combined with an AI agent to generate attribute label sets of structural features and semantic features respectively. Through logical reasoning and confidence filtering, the target attribute label set is fused to generate a target attribute label set, reducing human intervention.

Benefits of technology

It significantly improves the accuracy and automation of tag generation, reduces the risk of misjudgment, shortens processing time, reduces labor costs, and improves the accuracy and efficiency of document processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789231A_ABST
    Figure CN121789231A_ABST
Patent Text Reader

Abstract

The invention provides a file processing method, device and equipment based on AI, RPA, LLM and AI Agent, and relates to the field of AI, RPA and AI Agent. The method comprises the steps of obtaining a to-be-processed electronic file; performing visual identification on the electronic file by adopting a plurality of target visual identification models to generate a first attribute tag set; performing semantic understanding and logical reasoning on the text content extracted by the AI Agent of the electronic file by adopting LLM to obtain a second attribute tag set; and generating a target attribute tag set of the electronic file according to the first attribute tag set and the second attribute tag set. Therefore, the accuracy and the automation level of tag generation of the electronic file are remarkably improved, the misjudgment risk caused by layout complexity, text fuzziness or semantic ambiguity of the electronic file is effectively reduced, the time required by attribute classification is greatly shortened, and the manual intervention and processing cost is remarkably reduced; in addition, the accuracy and efficiency of file processing can be improved based on AI, RPA and AI Agent technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of Artificial Intelligence (AI), Robotic Process Automation (RPA), Large Language Model (LLM), and Artificial Intelligence Agent (AI Agent) for digital employee platforms, and particularly to a document processing method, apparatus, and device based on AI, RPA, LLM, and AI Agent. Background Technology

[0002] Robotic Process Automation (RPA) uses specific "robot software" to simulate human operations on a computer and automatically execute process tasks according to rules.

[0003] Artificial intelligence (AI) is a technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0004] Large Language Models (LLMs) are models trained on massive amounts of text that can recognize human language, perform language-related tasks, and have a large number of parameters.

[0005] Artificial Intelligence Agents (AI Agents) are capable of perceiving their environment, making decisions, and executing actions. Unlike traditional artificial intelligence, they possess the ability to think and act independently, and can utilize tools to achieve given goals. Based on LLM (Limited Learning Model) as its core computing engine, AI Agents can engage in dialogue, perform tasks, reason, and exhibit a degree of autonomy. They possess the ability to autonomously understand, perceive, plan, remember, and use tools, enabling them to automate complex tasks. Specifically, LLM-driven AI Agents, comprised of various AI capabilities, can interact with employees using natural language, understand their instructions and needs, and provide feedback and responses; they can acquire domain-specific knowledge relevant to the business to complete complex professional tasks; they can break down complex tasks into several executable tasks and use data and tools to complete them; and they can collaborate with employees, and AI Agents can also collaborate to complete complex tasks, enabling digital employees to leap from automation to intelligence, helping employees complete their work more efficiently, and fully realizing human-machine collaboration.

[0006] In related technologies, it is usually necessary to manually view, read and understand the content of each document, and then manually add attribute classification tags according to preset rules.

[0007] However, when faced with a massive amount of documents on a document processing platform, this purely manual approach not only fails to meet timeliness requirements but also incurs high labor costs. Summary of the Invention

[0008] This application provides a file processing method, apparatus, and device based on AI, RPA, LLM, and AI Agent to solve one of the technical problems existing in related technologies. The technical solution is as follows: In a first aspect, embodiments of this application provide a file processing method based on AI, RPA, LLM, and AI Agent, including: Obtain the electronic file to be processed; The electronic document is visually recognized using multiple target visual recognition models to generate a first attribute tag set; wherein the first attribute tag set includes multiple first attribute tags describing the structural features of the electronic document; LLM is used to perform semantic understanding and logical reasoning on the text content of the electronic document extracted by the AI ​​Agent to obtain a second attribute tag set; wherein, the second attribute tag set includes multiple second attribute tags describing the semantic features of the electronic document; The target attribute tag set of the electronic document is generated based on the first attribute tag set and the second attribute tag set.

[0009] In one implementation, the first attribute tag set further includes a first confidence level corresponding to each of the first attribute tags, and the second attribute tag set further includes a second confidence level corresponding to each of the second attribute tags; generating the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set includes: From each of the first attribute labels, at least one first valid attribute label is selected; wherein the first confidence level of the first valid attribute label is greater than the corresponding first confidence threshold. From each of the second attribute labels, at least one second valid attribute label is selected; wherein the second confidence level of the second valid attribute label is greater than the corresponding second confidence threshold. Determine whether there is a business logic conflict between each of the first valid attribute tags and each of the second valid attribute tags that have a first set association relationship; wherein, the business logic conflict is used to indicate that the structural feature corresponding to the first valid attribute tag and the semantic feature corresponding to the second valid attribute tag do not conform to the preset business rules; In response to the absence of business logic conflicts between the tag pairs, each of the first valid attribute tags and each of the second valid attribute tags are merged to obtain a target attribute tag set.

[0010] In one implementation, in response to the absence of business logic conflicts between the tag pairs, the first valid attribute tags and the second valid attribute tags are fused to obtain a target attribute tag set, including: In response to the absence of business logic conflicts between the tag pairs, each of the first valid attribute tags and each of the second valid attribute tags are merged to obtain an intermediate attribute tag set; All first attribute tags other than the first valid attribute tag in each of the first attribute tags and all second attribute tags other than the second valid attribute tag in each of the second attribute tags are scheduled to the target object for tag verification, and the first verification attribute tag and the second verification attribute tag fed back by the target object are received. The intermediate attribute tag set, the first verification attribute tag, and the second verification attribute tag are merged to generate the target attribute tag set.

[0011] In one embodiment, the method further includes: In response to the tag pair including a first tag pair with business logic conflict, the tag state of the first valid attribute tag of the first tag pair is updated to an unknown state to obtain the first attribute tag after the state update. The updated first attribute label, the other first attribute labels, and the other second attribute labels are scheduled to the target object for label verification, and the first and second verification attribute labels fed back by the target object are received. The first valid attribute tags, excluding the first valid attribute tags in the first tag pair, the first verification attribute tags, and the second verification attribute tags in each of the first valid attribute tags are merged to generate the target attribute tag set.

[0012] In one implementation, generating the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set includes: For any first valid attribute tag, obtain the target text content associated with the first valid attribute tag from the text content; In response to a business logic conflict between the text attribute tag of the target text content and the second valid attribute tag, additional description information of the first valid attribute tag is generated based on the first valid attribute tag, the text attribute tag, and the second valid attribute tag. The additional description information is used to annotate any of the first valid attribute tags to obtain the annotated first valid attribute tags; The target attribute tag set is generated based on each of the first valid attribute tags, each of the second valid attribute tags, and the additional description information.

[0013] In one embodiment, the method further includes: In response to each of the first confidence scores being less than the first confidence score threshold of the corresponding first attribute label, and each of the second confidence scores being less than the second confidence score threshold of the corresponding second attribute label, the label status of the first attribute label set and the second attribute label set is updated to pending processing; In response to the review operation, the attribute tag set with the tag status of pending processing is reviewed to obtain the target tag set.

[0014] In one implementation, the step of employing a large language model to perform semantic understanding and logical reasoning on the text content extracted from the electronic document by optical character recognition (OCR) to obtain a second attribute tag set includes: Obtain a prompt template; wherein the prompt template is used to indicate the task information to be executed by the large language model; The set semantic analysis instructions and the text content extracted from the electronic document by the AI ​​Agent are used to fill the prompt template to obtain model prompt information; The large language model is invoked to perform semantic understanding and logical reasoning on the text content based on the model prompt information, so as to obtain the second attribute tag set.

[0015] In one implementation, the step of visually recognizing the electronic document using multiple visual recognition models to generate a first attribute tag set includes: The RPA robot sends processing requests to the processing engine associated with multiple target visual models; The processing request includes the electronic document. The processing request is used by the processing engine to call multiple target visual recognition models in parallel to perform visual recognition on the electronic document, and to summarize the attribute tags returned by the multiple visual recognition models to generate a first attribute tag set.

[0016] In one embodiment, before performing visual recognition on the electronic document using multiple target visual recognition models to generate a first attribute tag set, the method further includes: The AI ​​Agent monitors the running status of each candidate visual recognition model; Obtain the recognition results output by each candidate visual recognition model in at least one historical time period; Based on the recognition results of each candidate visual recognition model, determine the evaluation index of each candidate visual recognition model in at least one evaluation dimension. The performance of each candidate visual recognition model is determined based on the evaluation metrics of each candidate recognition model. Based on the performance and operating status of multiple candidate visual recognition models, multiple target visual recognition models are determined from the multiple candidate visual recognition models.

[0017] Secondly, embodiments of this application provide a file processing apparatus based on AI, RPA, LLM, and an AI Agent, comprising: The first acquisition module is used to acquire the electronic files to be processed. The first processing module is used to perform visual recognition on the electronic document using multiple target visual recognition models to generate a first attribute tag set; wherein, the first attribute tag set includes multiple first attribute tags describing the structural features of the electronic document; The second processing module is used to perform semantic understanding and logical reasoning on the text content of the electronic document extracted by the AI ​​Agent using LLM, so as to obtain a second attribute tag set; wherein, the second attribute tag set includes multiple second attribute tags describing the semantic features of the electronic document; The generation module is used to generate the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set.

[0018] In one implementation, the first attribute tag set further includes a first confidence level corresponding to each first attribute tag, and the second attribute tag set further includes a second confidence level corresponding to each second attribute tag. The generation module is used to filter out at least one first valid attribute tag from each of the first attribute tags; wherein the first confidence level of the first valid attribute tag is greater than the corresponding first confidence level threshold. From each of the second attribute labels, at least one second valid attribute label is selected; wherein the second confidence level of the second valid attribute label is greater than the corresponding second confidence threshold. Determine whether there is a business logic conflict between each of the first valid attribute tags and each of the second valid attribute tags that have a first set association relationship; wherein, the business logic conflict is used to indicate that the structural feature corresponding to the first valid attribute tag and the semantic feature corresponding to the second valid attribute tag do not conform to the preset business rules; In response to the absence of business logic conflicts between the tag pairs, each of the first valid attribute tags and each of the second valid attribute tags are merged to obtain a target attribute tag set.

[0019] In one implementation, the generation module is configured to, in response to the absence of business logic conflict between the tag pairs, merge each of the first valid attribute tags and each of the second valid attribute tags to obtain an intermediate attribute tag set; All first attribute tags other than the first valid attribute tag in each of the first attribute tags and all second attribute tags other than the second valid attribute tag in each of the second attribute tags are scheduled to the target object for tag verification, and the first verification attribute tag and the second verification attribute tag fed back by the target object are received. The intermediate attribute tag set, the first verification attribute tag, and the second verification attribute tag are merged to generate the target attribute tag set.

[0020] In one implementation, the generation module is configured to, in response to the tag pair including a first tag pair with business logic conflict, update the tag status of the first valid attribute tag of the first tag pair to an unknown state, so as to obtain the first attribute tag after the status update. The updated first attribute label, the other first attribute labels, and the other second attribute labels are scheduled to the target object for label verification, and the first and second verification attribute labels fed back by the target object are received. The first valid attribute tags, excluding the first valid attribute tags in the first tag pair, the first verification attribute tags, and the second verification attribute tags in each of the first valid attribute tags are merged to generate the target attribute tag set.

[0021] In one implementation, the generation module is configured to obtain target text content associated with any first valid attribute tag from the text content for any first valid attribute tag. In response to a business logic conflict between the text attribute tag of the target text content and the second valid attribute tag, additional description information of the first valid attribute tag is generated based on the first valid attribute tag, the text attribute tag, and the second valid attribute tag. The additional description information is used to annotate any of the first valid attribute tags to obtain the annotated first valid attribute tags; The target attribute tag set is generated based on each of the first valid attribute tags, each of the second valid attribute tags, and the additional description information.

[0022] In one implementation, the generation module is configured to update the label status of the first attribute label set and the second attribute label set to pending processing in response to each of the first confidence scores being less than the first confidence score threshold of the corresponding first attribute label and each of the second confidence scores being less than the second confidence score threshold of the corresponding second attribute label. In response to the review operation, the attribute tag set with the tag status of pending processing is reviewed to obtain the target tag set.

[0023] In one embodiment, the second processing module is used to obtain a prompt template; wherein the prompt template is used to indicate the task information to be executed by the large language model; The set semantic analysis instructions and the text content extracted from the electronic document by the AI ​​Agent are used to fill the prompt template to obtain model prompt information; The large language model is invoked to perform semantic understanding and logical reasoning on the text content based on the model prompt information, so as to obtain the second attribute tag set.

[0024] In one implementation, the first processing module is configured to send a processing request to a processing engine associated with multiple target visual models via an RPA robot. The processing request includes the electronic document. The processing request is used by the processing engine to call multiple target visual recognition models in parallel to perform visual recognition on the electronic document, and to summarize the attribute tags returned by the multiple visual recognition models to generate a first attribute tag set.

[0025] In one embodiment, the file processing device further includes: a determination module, used to monitor the running status of each candidate visual recognition model through an AI Agent; Obtain the recognition results output by each candidate visual recognition model in at least one historical time period; Based on the recognition results of each candidate visual recognition model, determine the evaluation index of each candidate visual recognition model in at least one evaluation dimension. The performance of each candidate visual recognition model is determined based on the evaluation metrics of each candidate recognition model. Based on the performance and operating status of multiple candidate visual recognition models, multiple target visual recognition models are determined from the multiple candidate visual recognition models.

[0026] Thirdly, embodiments of this application provide an electronic device comprising a memory and a processor. The memory and the processor communicate with each other via an internal connection path. The memory stores instructions, and the processor executes the instructions stored in the memory. When the processor executes the instructions stored in the memory, it causes the processor to perform the method described in any of the above embodiments.

[0027] Fourthly, embodiments of this application provide a computer-readable storage medium that stores a computer program, wherein when the computer program is run on a computer, the methods in any of the above-described embodiments are executed.

[0028] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the methods in any of the above-described embodiments.

[0029] The advantages or beneficial effects of the above technical solutions include at least the following: By employing multiple target visual recognition models to perform visual analysis on electronic documents, their structural features can be accurately captured from different angles, thereby generating a comprehensive and robust first attribute tag set. Simultaneously, by utilizing a large language model to perform deep semantic understanding and logical reasoning on the text content extracted by the AI ​​Agent, the semantic intent, key business elements, and contextual relationships of the document can be effectively identified, generating a second attribute tag set rich in semantic information. Furthermore, based on the first and second attribute tag sets, a target attribute tag set is generated. This target attribute tag set integrates multimodal information from both structural and semantic dimensions of the electronic document, significantly improving the accuracy, generalization ability, and automation level of tag generation. It also effectively reduces the risk of misjudgment caused by complex layouts, ambiguous text, or semantic ambiguity, greatly shortening the time required for attribute classification and significantly reducing manual intervention and processing costs. In addition, AI, RPA, and AI Agent technologies can be used to improve the accuracy and efficiency of document processing.

[0030] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of this application will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0031] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.

[0032] Figure 1 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in one embodiment of this application; Figure 2 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in another embodiment of this application; Figure 3 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in another embodiment of this application; Figure 4 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in another embodiment of this application; Figure 5 This is a schematic diagram illustrating the implementation principle of an embodiment of this application; Figure 6 This is a flowchart illustrating the document processing method according to an embodiment of this application; Figure 7 This is a schematic diagram of the file processing system according to an embodiment of this application; Figure 8 This is a structural diagram of a file processing apparatus based on AI, RPA, LLM, and AI Agent provided in one embodiment of this application; Figure 9 A structural block diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation

[0033] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0034] As enterprises deepen their digital transformation, document processing platforms have become core tools for handling various documents (such as contracts, reports, invoices, and forms). These platforms accumulate massive amounts of user data in their daily operations, characterized by multiple sources, multiple formats (such as PDF, JPG, PNG, and DOCX), and multiple modalities (including text, images, tables, and handwriting).

[0035] In related technologies, document processing platforms mainly rely on the following methods for data classification and management: (1) Purely manual classification method: Algorithm engineers or data labelers manually view, read and understand the content of each document, and then manually add classification labels according to preset rules; (2) Semi-automated approach based on a single model: To improve performance, a general-purpose OCR engine is developed or called to perform full-text recognition of the document. Then, the text output by Optical Character Recognition (OCR) is analyzed based on simple rules or keyword matching.

[0036] The above-mentioned purely manual classification method is extremely inefficient and cannot be scaled up. Processing a complex document may take several minutes or even longer. Faced with the massive amounts of data on document processing platforms, the purely manual method cannot meet the timeliness requirements of data preparation for algorithm iteration, and it is also costly. In addition, the purely manual classification method is highly subjective and inconsistent. Different algorithm / annotation personnel may have different understandings of the classification rules, which will lead to inconsistent classification results and uncontrollable data quality, thus affecting the accuracy of subsequent algorithm training and evaluation. The aforementioned semi-automated approach using a single model suffers from low recognition accuracy and poor generalization ability. For example, general OCR models struggle to accurately recognize complex situations such as multilingual text, handwritten text, and tables without clear borders, easily leading to misclassification and weak generalization. A specific model typically excels at handling only one type of task; adding new classification dimensions requires developing, training, and integrating another independent model, resulting in poor system scalability and high development and maintenance costs. Keyword rule systems built on a single OCR output are extremely fragile. Once the document template or wording changes, the existing rules may immediately become invalid, requiring continuous manual maintenance and updates to the rule base, failing to achieve true adaptability and intelligence, and exhibiting insufficient system flexibility.

[0037] To address at least one of the problems existing in related technologies, embodiments of this application propose a document processing method, apparatus, and device based on Artificial Intelligence (AI), Robotic Process Automation (RPA), Large Language Model (LLM), and Artificial Intelligence Agent (AI Agent) for digital employee platforms. These and other aspects of the embodiments of this application will become clear with reference to the following description and accompanying drawings. In these descriptions and drawings, some specific implementations of the embodiments of this application are specifically disclosed to illustrate some ways of carrying out the principles of the embodiments of this application; however, it should be understood that the scope of the embodiments of this application is not limited thereto. Rather, the embodiments of this application include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0038] Before describing the specific embodiments of this application, for ease of understanding, commonly used technical terms will first be introduced: In the description of this application, the term "multiple" means two or more.

[0039] In the description of this application, the term "target visual recognition model" refers to an artificial intelligence (AI) model trained based on computer vision technology, used to identify specific visual elements from images or document pages, such as tables, handwriting, seals, layouts, etc.

[0040] In the description of this application, the term "LLM" refers to a class of natural language processing models based on deep learning. Its main characteristics are a large number of model parameters and a complex neural network structure. It has powerful language understanding, context awareness and language generation capabilities. It can automatically learn useful feature representations from input data and generate relevant text.

[0041] In the description of this application, the term "first attribute tag" refers to a structured tag that describes the structural features of an electronic document, extracted from the image or layout of the electronic document by a target visual recognition model.

[0042] In the description of this application, the term "first attribute tag set" refers to a collection consisting of multiple first attribute tags.

[0043] In the description of this application, the term "second attribute tag" refers to a structured tag generated by performing semantic understanding and logical reasoning on the text content of an electronic document extracted by an AI Agent through LLM, which is used to describe the semantic features of the electronic document.

[0044] In the description of this application, the term "second attribute tag set" refers to a collection consisting of multiple second attribute tags.

[0045] In the description of this application, the term "target attribute tag set" refers to a comprehensive and highly confident set of attribute tags generated by integrating the structural features (i.e., the first attribute tag set) and semantic features (i.e., the second attribute tag set) of an electronic document.

[0046] In the description of this application, the term "first valid attribute label" refers to a label selected from the first attribute label set that has a confidence level higher than a first confidence threshold; In the description of this application, the term "second valid attribute label" refers to a label selected from the set of second attribute labels that has a confidence level higher than the second confidence threshold; In the description of this application, the term "business logic conflict" refers to a situation where the structural feature corresponding to the first valid attribute tag and the semantic feature corresponding to the second valid attribute tag do not conform to preset business rules. For example, if the first valid attribute tag is "contains a table: yes" and the second valid attribute tag is "document type: letter", according to preset business rules, ordinary letters should not contain structured tables used to agree on rights and obligations, and there is a business logic conflict between the first valid attribute tag and the second valid attribute tag.

[0047] In the description of this application, the term "intermediate attribute tag set" refers to the preliminary tag set formed by directly merging the first valid attribute tag and the second valid attribute tag when there is no business logic conflict between the first valid attribute tag and the second valid attribute tag.

[0048] In the description of this application, the term "first review attribute label" refers to a structural attribute label that has been manually reviewed and then confirmed or corrected by the target object (such as a human auditor).

[0049] In the description of this application, the term "second review attribute label" refers to a semantic attribute label that is returned after being confirmed, corrected, or supplemented by the target object (such as a human reviewer) following a manual or intelligent review of a second attribute label that has not reached the second confidence threshold.

[0050] In the description of this application, the term "hint template" refers to a structured, configurable instruction framework used to guide the semantic analysis tasks that a large language model needs to perform.

[0051] These and other aspects of the embodiments of this application will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific implementations of the embodiments of this application are specifically disclosed to illustrate some ways of carrying out the principles of the embodiments of this application; however, it should be understood that the scope of the embodiments of this application is not limited thereto. Rather, the embodiments of this application include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0052] The following describes, with reference to the accompanying drawings, a file processing method, apparatus, and device based on AI, RPA, LLM, and AI Agent according to embodiments of this application.

[0053] Figure 1 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in one embodiment of this application.

[0054] In one possible implementation provided in this application, the document processing method based on AI, RPA, and AI Agent is configured in a document processing device based on AI, RPA, LLM, and AI Agent. This document processing device can be applied to any electronic device with computing capabilities.

[0055] The electronic device can be a personal computer, a mobile terminal, a server (or a server), etc. The mobile terminal can be a mobile phone, a tablet computer, a personal digital assistant, or other hardware device with various operating systems.

[0056] In another possible implementation of this application embodiment, the file processing method based on AI, RPA, LLM and AI Agent can be applied to an AI Agent, wherein the AI ​​Agent can run on any electronic device with computing capabilities.

[0057] In another possible implementation of this application's embodiments, the document processing method based on AI, RPA, LLM, and AI Agent can be applied to a Work Execution Platform (WEP). The implementation of the WEP involves three stages. The first stage is automation: for RPA stages with low business complexity, software automation technology is used to automate rule-based, predefined procedural tasks. The second stage is intelligence: leveraging AI to extend the boundaries of RPA, such as processing unstructured documents and making data-driven decisions. The third stage is human-machine collaboration: utilizing the understanding, planning, and execution capabilities of large models to automate complex tasks end-to-end.

[0058] In the human-machine collaboration phase, the digital employee platform serves as a bridge connecting workers and systems, workers and data, and systems and data. It is capable of: operating complex systems, processing various types of data, and interacting and collaborating with employees. The digital employee platform helps industries build large-scale, model-enabled digital employees (i.e., intelligent agents), achieving automation, intelligence, and human-machine collaboration in business processes.

[0059] The digital employee platform can seamlessly integrate multiple capabilities such as Agentic Process Automation (APA), Agentic Document Processing (ADP), and Agentic Business Insights (ABI). It has five major functions: "business understanding", "process creation", "run anywhere", "centralized management and control", and "human-machine collaboration". It enables enterprises to achieve end-to-end intelligent automation of business processes, replace manual operations, further improve business efficiency, and accelerate digital transformation.

[0060] ADP is a next-generation platform that combines LLM and Vision-Language Model (VLM) with AI Agent technology to achieve end-to-end automated document processing. It represents a new generation of document processing solutions based on LLM and AI Agent technologies. It is no longer a "tool" that requires configuring templates or annotating samples, but rather an "intelligent agent" capable of understanding business needs and autonomously planning and executing. Traditional document processing systems are "tools": users need to explicitly tell the system "how to do it." ADP, on the other hand, is an "intelligent agent": users only need to tell the system "what to do," and the system can autonomously understand, plan, and execute.

[0061] like Figure 1 As shown, this document processing method based on AI, RPA, LLM, and AI Agent may include the following steps S101 to S104: Step S101: Obtain the electronic file to be processed.

[0062] To ensure the consistency of input data, one possible approach is to access the original files to be classified (such as PDFs and DOCXs) from the ADP platform and perform preprocessing operations to obtain the electronic files to be processed. Preprocessing may include, but is not limited to: format standardization (e.g., converting PDFs and DOCXs to JPG images), image deduplication, and image angle correction.

[0063] Step S102: Use multiple target visual recognition models to perform visual recognition on the electronic document to generate a first attribute tag set.

[0064] The first attribute tag set includes multiple first attribute tags that describe the structural characteristics of electronic documents.

[0065] In order to comprehensively capture the structural features of electronic documents, as a possible approach, multiple screened target visual recognition models are used to analyze the electronic documents in parallel, extracting the structural features of the electronic documents, such as whether they contain tables, whether they contain handwritten signatures, whether they contain seals, etc., and converting these features into structured first attribute labels to form a first attribute label set. The target visual recognition models may include table detection models, seal recognition models, general character recognition models, handwritten character detection models, etc.

[0066] It should be noted that, in order to achieve more efficient and unified multimodal document understanding capabilities, multiple target visual recognition models (such as table detection models, seal recognition models, handwriting detection models, etc.) can be replaced by a unified and more powerful VLM. This visual language model can simultaneously perceive structural information and text semantics in images, and complete the joint recognition and reasoning of multiple visual attributes within a single model, thereby simplifying the system architecture and reducing model maintenance costs.

[0067] Step S103: Use LLM to perform semantic understanding and logical reasoning on the text content extracted by the AI ​​Agent from the electronic document to obtain the second attribute tag set.

[0068] The second attribute tag set includes multiple second attribute tags that describe the semantic features of electronic documents.

[0069] To achieve accurate understanding of the semantic content of electronic documents, one possible approach is to have the AIAgent extract processable text content from the electronic document and then call LLM (Limited Language Management) to perform deep semantic analysis on that extracted text content. LLM can accurately identify key semantic features of the electronic document, such as its main language and document type, and structure this information into second attribute tags, thus forming a set of second attribute tags.

[0070] Step S104: Generate the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set.

[0071] To improve the accuracy and comprehensiveness of attribute tags, one possible approach is to generate a target attribute tag set for the electronic document based on a first attribute tag set from visual recognition and a second attribute tag set from a large language model. This target attribute tag set includes the key attributes of the electronic document in both structural and semantic dimensions.

[0072] The document processing method based on AI, RPA, LLM, and AI Agent in this application employs multiple target visual recognition models to perform visual analysis on electronic documents, accurately capturing their structural features from different angles to generate a comprehensive and robust first attribute tag set. Simultaneously, a large language model is used to perform deep semantic understanding and logical reasoning on the text content extracted by the AI ​​Agent, effectively identifying the document's semantic intent, key business elements, and contextual relationships, generating a second attribute tag set rich in semantic information. Furthermore, based on the first and second attribute tag sets, a target attribute tag set is generated. This target attribute tag set integrates multimodal information from both structural and semantic dimensions of the electronic document, significantly improving the accuracy, generalization ability, and automation level of tag generation. It also effectively reduces the risk of misjudgment caused by complex layouts, ambiguous text, or semantic ambiguity, greatly shortening the time required for attribute classification and significantly reducing manual intervention and processing costs. Moreover, AI, RPA, and AI Agent technologies can be used to improve the accuracy and efficiency of document processing.

[0073] To clearly illustrate how the target attribute tag set of an electronic document is generated based on the first attribute tag set and the second attribute tag set in any embodiment of this application, this application also proposes a document processing method based on AI, RPA, LLM, and AI Agent.

[0074] Figure 2 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in another embodiment of this application.

[0075] It should be noted that the document processing method based on AI, RPA, LLM and AI Agent can be executed individually, or it can be executed together with any embodiment or possible implementation in the embodiment of this application, or it can be executed together with any technical solution in related technologies. The embodiments of this application do not limit this.

[0076] like Figure 2 As shown, this document processing method based on AI, RPA, LLM, and AI Agent may include the following steps S201 to S207: Step S201: Obtain the electronic file to be processed.

[0077] Step S202: Use multiple target visual recognition models to perform visual recognition on the electronic document to generate a first attribute tag set.

[0078] The first attribute tag set includes multiple first attribute tags that describe the structural characteristics of electronic documents.

[0079] Step S203: Use LLM to perform semantic understanding and logical reasoning on the text content extracted by the AI ​​Agent from the electronic document to obtain the second attribute tag set.

[0080] The second attribute tag set includes multiple second attribute tags that describe the semantic features of electronic documents.

[0081] The explanation of steps S201 to S203 can be found in the relevant description in any embodiment of this application, and will not be repeated here.

[0082] Step S204: Select at least one first valid attribute tag from each first attribute tag.

[0083] Among them, the first confidence level of the first valid attribute label is greater than the corresponding first confidence threshold.

[0084] To ensure the reliability of visual recognition results, as a possible implementation, the first attribute label set also includes the first confidence level corresponding to each first attribute label. Only the high-confidence structural feature labels (such as "contains tables") output by the target visual recognition model are retained, i.e., the first valid attribute labels, thereby filtering out low-quality recognition results caused by image blurring or model misjudgment.

[0085] Step S205: Select at least one valid second attribute tag from each second attribute tag.

[0086] Among them, the second confidence level of the second valid attribute label is greater than the corresponding second confidence threshold.

[0087] To ensure the accuracy of semantic understanding results, as a possible approach, the second attribute label set also includes a second confidence level corresponding to each second attribute label. The confidence level of the semantic labels (e.g., language: Chinese, document type: contract) generated by LLM based on the text content extracted by the AI ​​Agent is evaluated and filtered to eliminate low-confidence outputs that may have ambiguity or missing information, thereby ensuring the reliability of semantic features.

[0088] Step S206: Determine whether there is a business logic conflict between each first valid attribute tag and each second valid attribute tag that has a first set association relationship.

[0089] Among them, business logic conflict is used to indicate that the structural feature corresponding to the first valid attribute label and the semantic feature corresponding to the second valid attribute label do not conform to the preset business rules.

[0090] To achieve consistency verification between structural and semantic features, one possible approach is to use pre-defined business rules (e.g., ordinary letters should not contain structured tables used to agree on rights and obligations) to perform logical consistency checks on tag pairs with a pre-defined relationship between each first valid attribute tag and each second valid attribute tag. This proactively identifies potential conflicts and prevents erroneous tags from entering the final result. For example, if the first valid attribute tag is "Contains Table: Yes" and the second valid attribute tag is "Document Type: Letter", then the first and second valid attribute tags are a tag pair.

[0091] Step S207: In response to the absence of business logic conflicts between the tag pairs, the first valid attribute tags and the second valid attribute tags are merged to obtain the target attribute tag set.

[0092] To achieve high-confidence, conflict-free multimodal tag fusion, as a possible approach, under the premise of ensuring that both structural and semantic information are reliable and there are no business logic conflicts, each first valid attribute tag and each second valid attribute tag is merged into a target attribute tag set. The target attribute tag set can provide high-quality and reliable decision-making basis for subsequent intelligent document processing tasks such as automatic classification, review, and archiving.

[0093] As an example, in response to the absence of business logic conflicts between tag pairs, each first valid attribute tag and each second valid attribute tag are merged to obtain an intermediate attribute tag set; other first attribute tags in each first attribute tag except for the first valid attribute tag and other second attribute tags in each second attribute tag except for the second valid attribute tag are scheduled to the target object for tag review, and the first and second review attribute tags fed back by the target object are received; the intermediate attribute tag set, the first review attribute tags and the second review attribute tags are merged to generate the target attribute tag set.

[0094] In other words, to improve the accuracy of document processing, after confirming that there are no business logic conflicts between the first and second valid attribute tags, the first and second valid attribute tags are merged to generate an intermediate attribute tag set. At the same time, the remaining first and second attribute tags that do not reach the confidence threshold (i.e., low-confidence tags) are scheduled to the target object (such as a human reviewer or expert system) for targeted review, and the first and second review attribute tags fed back by the target object are received. Finally, the intermediate attribute tag set is merged with the tags returned by the review to generate the target attribute tag set.

[0095] As another possible implementation, in response to a first tag pair including a tag pair with business logic conflict, the tag state of the first valid attribute tag of the first tag pair is updated to an unknown state to obtain the first attribute tag with the updated state; the first attribute tag with the updated state, other first attribute tags, and other second attribute tags are scheduled to the target object for tag review, and the first review attribute tag and the second review attribute tag fed back by the target object are received; the other first valid attribute tags, the first review attribute tag, and the second review attribute tag in each first valid attribute tag are merged except for the first valid attribute tag in the first tag pair to generate a target attribute tag set.

[0096] In other words, to achieve precise intervention and reliable correction of attribute tags with conflicting business logic, when a tag pair containing a first tag pair with conflicting business logic is detected, the tag status of the first valid attribute tag in the first tag pair is updated to an unknown state. That is, the second valid attribute tag (the semantic tag recognized by the large language model) in the first tag pair is used as the standard, and the first valid attribute is removed from the high-confidence tag set to obtain the first attribute tag with the updated status. Then, the first attribute tag with the updated status, the other first attribute tags that were not selected as valid, and all other second attribute tags that were not selected as valid are uniformly scheduled to the target object for centralized review, and the first and second review attribute tags fed back by the target object are received. Finally, in the fusion stage, the remaining first valid attribute tags without conflict are integrated with the first and second review attribute tags returned by the review to generate the final target attribute tag set.

[0097] For example, the target visual recognition model (e.g., a table detection model) might identify "Contains table: Yes (high confidence)", while the LLM semantic judgment might identify "Document type: Letter (high confidence)". The default business rule is that "ordinary letters should not contain structured tables used to agree on rights and obligations". Therefore, the LLM semantic judgment is given priority, as OCR might misidentify well-formatted headers, footers, or lists as tables, while the large language model can more accurately determine the document type based on overall semantics. The final result will adopt "Document type: Letter", and the target visual recognition model's result indicating the presence of a table will be downgraded to unknown for manual review.

[0098] As another possible implementation, for any first valid attribute tag, the target text content associated with any valid first attribute tag is obtained from the text content; in response to the business logic conflict between the text attribute tag and the second valid attribute tag of the target text content, additional description information for any first valid attribute tag is generated based on any first valid attribute tag, the text attribute tag, and the second valid attribute tag; the additional description information is used to annotate any first valid attribute tag to obtain the annotated first valid attribute tag; and a target attribute tag set is generated based on each first valid attribute tag, each second valid attribute tag, and the additional description information.

[0099] For any first valid attribute tag, locate the associated target text content from the text content extracted by OCR (for example, when the first valid attribute tag contains handwriting, the associated target text content is the specific handwritten text content); if there is a business logic conflict between the text attribute tag obtained by analyzing the target text content (e.g., language: English) and the second valid attribute tag (language: Chinese), then combine the first valid attribute tag, the text attribute tag, and the second valid attribute tag to generate additional descriptive information (e.g., the document body is in Chinese, but contains English handwriting); then, attach the additional descriptive information as an annotation to the corresponding first valid attribute tag to form the annotated first valid attribute tag; finally, based on all first valid attribute tags, second valid attribute tags, and corresponding additional descriptive information, generate a target attribute tag set.

[0100] For example, taking a visual recognition model that includes a handwritten text detection model, the semantic judgment result of the large language model is: Language: Chinese (high confidence), but the detection and recognition result of the handwritten text detection model is: Contains handwritten text: Yes (high confidence). However, the handwritten text recognized by OCR is detected as English. In the purely visual feature judgment of "whether it contains handwritten text", the system prioritizes trusting the handwritten text detection model. The final result will record "Contains handwritten text: Yes", and may also add additional descriptive information or notes in the tag (the main body of the document is Chinese, but it contains English handwritten text).

[0101] As another possible implementation, in response to each first confidence level being less than the first confidence threshold of the corresponding first attribute label and each second confidence level being less than the second confidence threshold of the corresponding second attribute label, the label status of the first attribute label set and the second attribute label set is updated to pending processing; in response to the review operation, the attribute label set with the label status pending processing is reviewed to obtain the target label set.

[0102] In other words, in order to effectively process low-confidence recognition results, when it is detected that the first confidence of all first attribute tags is lower than the corresponding first confidence threshold and the second confidence of all second attribute tags is lower than the corresponding second confidence threshold, it is determined that the structural and semantic features of the current document lack sufficient credibility. The tag status of the entire first attribute tag set and the second attribute tag set can be uniformly updated to pending processing. Then, after receiving a review operation instruction (such as manual intervention), the attribute tag set with the tag status pending processing is reviewed as a whole. The target object can re-judge or correct the tag content based on the original electronic document and OCR text, and finally generate an accurate and reliable target attribute tag set.

[0103] In summary, by applying confidence filtering to the first attribute labels generated by visual recognition and the second attribute labels generated by the large language model, only the first and second valid attribute labels with high confidence are retained, effectively eliminating low-quality or uncertain recognition results. Furthermore, for tag pairs with preset associations, it is determined whether there are business logic conflicts between structural and semantic features. Finally, under the premise that there are no business logic conflicts, the two types of valid tags are fused to generate a target attribute label set, which significantly improves the reliability and automation level of the intelligent document understanding system, while greatly reducing manual costs.

[0104] To clearly illustrate how a large language model is used in any embodiment of this application to perform semantic understanding and logical reasoning on the text content of electronic documents extracted by AI Agent in order to obtain a second attribute tag set, this application also proposes a document processing method based on AI, RPA, LLM and AI Agent.

[0105] Figure 3 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in another embodiment of this application.

[0106] It should be noted that the document processing method based on AI, RPA, LLM and AI Agent can be executed individually, or it can be executed together with any embodiment or possible implementation in the embodiment of this application, or it can be executed together with any technical solution in related technologies. The embodiments of this application do not limit this.

[0107] like Figure 3 As shown, the document processing method based on AI, RPA, LLM, and AI Agent may include the following steps S301 to S306: Step S301: Obtain the electronic file to be processed.

[0108] Step S302: The electronic document is visually recognized using multiple target visual recognition models to generate a first attribute tag set.

[0109] The first attribute tag set includes multiple first attribute tags that describe the structural characteristics of electronic documents.

[0110] The explanation of steps S301 to S302 can be found in the relevant description in any embodiment of this application, and will not be repeated here.

[0111] Step S303: Obtain the prompt template.

[0112] The prompt template is used to indicate the task information to be executed by the large language model.

[0113] To achieve precise guidance for reasoning tasks of large language models, one possible approach is to pre-obtain a structured prompt template. This prompt template indicates the specific task information that the large language model needs to perform in the current document processing scenario, such as which fields to extract, what business rules to follow, and what output format to use. This allows the general language model to be adapted into an intelligent parsing engine for specific document types.

[0114] Step S304: The prompt template is filled with the set semantic analysis instructions and the text content extracted from the electronic document by the AI ​​Agent to obtain the model prompt information.

[0115] To achieve precise control over the semantic parsing process of a large language model, one possible approach is to use pre-defined semantic analysis instructions and text content extracted from electronic documents by an AI agent to fill placeholders in a prompt template, generating model prompt information. This prompt information contains both a clear analysis objective and the actual document content, providing the large language model with sufficient input context and task constraints. Semantic analysis instructions could include, for example, carefully analyzing the main languages ​​used in the document text, carefully analyzing the document text, and selecting one and only one best-matching category from a pre-defined list for output.

[0116] Step S305: Call the large language model to perform semantic understanding and logical reasoning on the text content based on the model prompt information to obtain the second attribute tag set.

[0117] The second attribute tag set includes multiple second attribute tags that describe the semantic features of electronic documents.

[0118] To accurately extract the structured semantic content of documents, one possible approach is to call a large language model and perform deep semantic understanding and logical reasoning on the text content extracted by the AI ​​Agent based on the model's prompts, ultimately outputting a set of structured attribute descriptions, namely the second attribute tag set.

[0119] Step S306: Generate the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set.

[0120] The explanation of step S306 can be found in the relevant description in any embodiment of this application, and will not be repeated here.

[0121] In summary, by acquiring prompt templates that indicate the task objectives of the large language model, and dynamically filling the templates with preset semantic analysis instructions and text content extracted by the AI ​​Agent, clear and context-complete model prompt information is generated. This drives the large language model to perform deep semantic understanding and logical reasoning on the text content under clear task constraints, and finally outputs a structured set of second attribute tags, which significantly improves the accuracy and effectiveness of semantic extraction.

[0122] To clearly illustrate how multiple visual recognition models are used to perform visual recognition on the electronic document in any embodiment of this application to generate a first attribute tag set, this application also proposes a document processing method based on AI, RPA, LLM, and AIAgent.

[0123] Figure 4 This is a flowchart of a file processing method based on AI, RPA, LLM, and AI Agent provided in another embodiment of this application.

[0124] It should be noted that the document processing method based on AI, RPA, LLM and AI Agent can be executed individually, or it can be executed together with any embodiment or possible implementation in the embodiment of this application, or it can be executed together with any technical solution in related technologies. The embodiments of this application do not limit this.

[0125] like Figure 4 As shown, this document processing method based on AI, RPA, LLM, and AI Agent may include the following steps S401 to S405: Step S401: Obtain the electronic file to be processed.

[0126] The explanation of step S401 can be found in the relevant description in any embodiment of this application, and will not be repeated here.

[0127] Step S402: The RPA robot sends a processing request to the processing engine associated with multiple target visual models.

[0128] The processing request includes electronic documents. The processing request is used by the processing engine to call multiple target visual recognition models in parallel to perform visual recognition on the electronic documents, and to summarize the attribute tags returned by the multiple visual recognition models to generate the first attribute tag set.

[0129] To achieve efficient and comprehensive extraction of structural features from electronic documents, one possible approach is to use an RPA robot to send processing requests to a processing engine associated with multiple target visual recognition models. The processing request contains the electronic document to be processed, which triggers the processing engine to call multiple dedicated target visual recognition models (such as table detection models, seal recognition models, general character recognition models, and handwriting detection models) in parallel to perform multi-dimensional visual analysis on the same electronic document. The engine then summarizes the structured attribute labels returned by each model in real time, ultimately generating the first attribute label set.

[0130] It should be noted that before using multiple target visual recognition models to visually recognize the electronic document and generate the first attribute tag set, to further improve the accuracy of visual recognition of the electronic document, the AIAgent continuously monitors the running status of each candidate visual recognition model, including service availability, response latency, resource load, etc., and obtains the recognition results output by each candidate visual recognition model in at least one historical period. Based on the recognition results, the evaluation index of each candidate model is calculated under at least one evaluation dimension (such as accuracy, recall, business field coverage, etc.), and then the recognition performance of each candidate model is comprehensively evaluated. Finally, combining the performance of each candidate model and its real-time running status, multiple target visual recognition models with high accuracy and high stability are dynamically selected from the multiple candidate visual recognition models. The candidate visual recognition models include, but are not limited to: table detection models, seal recognition models, general character recognition models, handwritten character detection models, signature detection models, document layout analysis models, QR code / barcode recognition models, image quality assessment models, etc.

[0131] Step S403: Use a large language model to perform semantic understanding and logical reasoning on the text content extracted by the AI ​​Agent from the electronic document in order to obtain the second attribute tag set.

[0132] The second attribute tag set includes multiple second attribute tags that describe the semantic features of electronic documents.

[0133] Step S404: Generate the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set. Explanations of steps S403 to S404 can be found in the relevant descriptions of any embodiment of this application, and will not be repeated here.

[0134] In summary, by sending a processing request containing electronic documents to a processing engine associated with multiple target visual recognition models via an RPA robot, the processing engine is triggered to call multiple target visual recognition models in parallel to perform multi-dimensional visual analysis of the electronic documents. The engine then automatically summarizes the structured attribute labels returned by each model to generate a first attribute label set. This improves the processing efficiency and comprehensiveness of visual recognition, effectively overcoming the blind spots of a single model in scenarios such as complex layouts, low image quality, or overlapping elements, thereby generating a more comprehensive and accurate set of structural features.

[0135] In any embodiment of this application, another document processing method is proposed, such as... Figure 5 As shown, the file processing method may include the following steps: Step S501: Obtain and preprocess the documents to be classified from the ADP platform through the data access and preprocessing module; In this embodiment, the "Chinese Procurement Contract" PDF file is obtained through the ADP platform's Application Programming Interface (API). The preprocessing module first parses the PDF, converting each page into a high-resolution (e.g., 200 DPI) PNG image format. Subsequently, the images undergo automatic angle correction and noise reduction to ensure standardization of subsequent model inputs.

[0136] Step S502: Using a dedicated OCR small model cluster (multiple target visual recognition models), perform preliminary analysis and classification on the preprocessed document in parallel or serial order to generate first-level classification results (such as "Contains handwriting: Yes / No", "Contains tables: Yes / No", "Contains seals: Yes / No"). In this embodiment, the uniformly formatted image of the document is fed in parallel into various dedicated models within the OCR small model cluster, such as... Figure 6 As shown, the specific models may include the following: (1) General text recognition model: Extract all text content in the image, output structured text block information and coordinates, the result is: {text: Purchase contract Party A: ... Party B: ... Product details and specifications ...}; (2) Handwriting detection model: This dedicated model is optimized for features such as pen stroke texture and continuity. It detects handwriting features in the signature area at the end of the document and outputs the following result: {"contains_handwriting": true, "confidence": 0.95}; (3) Table detection and recognition model: This model is based on a deep learning-based object detection architecture (such as YOLOv8) and can recognize tables without borders. It identifies a structured table in the middle of the document and outputs its row and column coordinates and content. The result is: {"contains_table": true, "confidence": 0.98, "table_data": {...}}; (4) Seal / signature detection model: The circular and elliptical red areas in the image are identified by feature recognition. No matching seal features are found. The output result is: {"contains_seal": false, "confidence": 0.99}.

[0137] Step S503: Using the large language model integration module, perform semantic analysis on the text content processed by the AI ​​Agent to generate second-level classification results (such as "Language: Chinese", "Document Type: Contract"). In this embodiment, the system does not directly feed the text extracted by the AI ​​Agent to the large model. Instead, it constructs a highly structured prompt message rich in context and instructions. After receiving this prompt, the large language model (such as GPT-4, Qwen, etc.) can accurately understand the task and perform deep semantic analysis on the text. It identifies core keywords such as "purchase contract," "both parties," and "liability for breach of contract," as well as the contractual context, and finally outputs standardized results.

[0138] Step S504: Through the collaborative decision-making unit, the first-level classification results and the second-level classification results are integrated to generate a comprehensive classification label set containing multiple dimensions; The collaborative decision-making unit receives the output results from all models, and the system calls classification rules and predefined rules in the knowledge base for processing, mainly including: (1) Classification dimension definition: All dimensions of the final output are predefined, namely: ["Language", "Contains handwriting", "Contains tables", "Contains seals / signatures", "Document type"]; (2) Confidence threshold rule: The confidence score of each model output must reach a preset threshold for the result to be considered valid. For example: handwriting detection confidence score ≥ 0.9, table detection confidence score ≥ 0.85, classification results output by the large language model: language confidence score ≥ 0.8, document type confidence score ≥ 0.8; (3) Result fusion and conflict resolution: Map the results of each model to a predefined classification dimension. In this embodiment, the output confidence of all models far exceeds the threshold and there is no logical conflict (for example, the large model judges it as "contract", which is reasonable in business logic to coexist with elements such as tables and handwriting), so it is directly fused; (4) Based on the above rules, the system generates the final comprehensive classification tag set: {Language: Chinese, Includes handwriting: Yes, Includes tables: Yes, Includes seals / signatures: No, Document type: Contract}.

[0139] Step S505: Output and store the comprehensive classification label set through the classification result output and application module for subsequent algorithm development and evaluation.

[0140] The system associates the structured tag set with the original PDF file, preprocessed images, and even the intermediate outputs of each model, storing the data in a database specifically designed for algorithm development. Through the system's API interface, users can instantly filter all documents that meet the criteria using SQL-like queries (e.g., language="Chinese" AND document type="contract" AND contains handwriting="yes"). This data can then be immediately used for specialized evaluation and iterative training of the "handwriting recognition in complex contract formats" algorithm.

[0141] Furthermore, it should be noted that this application also proposes a file processing system for implementing the above-mentioned file processing method, such as... Figure 7 As shown, the file processing system mainly includes the following modules: I. Data Access and Preprocessing Module In this embodiment, the data access and preprocessing module is used to access the raw data to be classified from the ADP platform and perform preprocessing operations on the data. Preprocessing includes, but is not limited to: format unification (converting PDF and DOCX files to JPG images), image deduplication, and image angle correction.

[0142] II. Multi-model Collaborative Classification Engine: This engine is the core of this system and includes: (1) Dedicated OCR Small Model Cluster: Composed of multiple lightweight, specialized OCR models, each focusing on processing a specific classification dimension. It includes at least: (a) General text recognition model: used to recognize text content in documents; (b) Handwritten text detection model: used to identify and determine whether a document contains handwritten text; (c) Table detection and recognition model: used to detect whether a document contains a table, and can further parse the structure and content of the table; (d) Seal / Signature Detection Model: Used to detect whether a document contains a seal or signature area; When analyzing electronic document images, these models can not only identify the location and category of specific structural elements, but also directly output pre-annotated files conforming to industry-standard formats, such as PASCAL VOC (XML format) or COCO (JSON format). These standardized formats fully describe the bounding box coordinates, category labels, and confidence scores of the target objects, and can be directly loaded by mainstream object detection algorithms (such as YOLO, Faster R-CNN, Detectron2, etc.) for model training, fine-tuning, or evaluation without conversion. Thus, the previously arduous task of manual frame-by-frame annotation is transformed into a highly efficient process of "AI-automated initial annotation + rapid manual review," significantly shortening the data preparation cycle, reducing annotation costs, and accelerating the iterative optimization of visual models, thereby greatly improving the efficiency and scalability of the entire intelligent document processing system in the data annotation stage.

[0143] (2) LLM Integration Module: Integrates LLM and is responsible for deep semantic understanding and logical reasoning. It is used for at least: (a) Language recognition: Automatically identify the main language of the document (such as Simplified Chinese, Traditional Chinese, English, Thai, Spanish, etc.) based on the text content; (b) Document type recognition: Determine the category of a document based on its semantic content, for example, identify it as "contract", "invoice", "resume", "order", "business license", etc. (3) Collaborative decision-making unit: It is used to receive and integrate the output results of the OCR small model cluster and the LLM integration module, and generate the final unified, multi-dimensional classification label set according to the preset decision rules.

[0144] III. Classification Rules and Knowledge Base: Stores predefined classification systems, label definitions, model calling logic, and collaborative decision-making rules.

[0145] IV. Classification Result Output and Application Module: This module associates and stores the generated multi-dimensional classification labels with the original data and the data in a standardized format, and provides an interface to output the classification results. The results can be directly applied to: (1) Quickly construct test datasets for specific scenarios (such as "handwritten contract recognition"); (2) Data filtering and aggregation: It enables algorithm engineers to efficiently filter the required data based on multi-dimensional labels (such as "language: Chinese", "contains tables: yes", "document type: contract") for model training.

[0146] To implement the above embodiments, this application also provides a file processing apparatus based on AI, RPA, LLM, and AI Agent.

[0147] Figure 8 This is a structural diagram of a file processing apparatus based on AI, RPA, LLM, and AI Agent provided in one embodiment of this application.

[0148] like Figure 8 As shown, the file processing device 800 based on AI, RPA, LLM and AI Agent includes: a first acquisition module 810, a first processing module 820, a second processing module 830 and a generation module 840.

[0149] The system comprises: a first acquisition module 810 for acquiring an electronic document to be processed; a first processing module 820 for visually recognizing the electronic document using multiple target visual recognition models to generate a first attribute tag set, wherein the first attribute tag set includes multiple first attribute tags describing the structural features of the electronic document; a second processing module 830 for semantic understanding and logical reasoning of the text content extracted from the electronic document by the AI ​​Agent using a large language model to obtain a second attribute tag set, wherein the second attribute tag set includes multiple second attribute tags describing the semantic features of the electronic document; and a generation module 840 for generating a target attribute tag set for the electronic document based on the first and second attribute tag sets.

[0150] In any embodiment of this application, the first attribute tag set further includes a first confidence level corresponding to each first attribute tag, and the second attribute tag set further includes a second confidence level corresponding to each second attribute tag. The generation module 840 is configured to: filter at least one first valid attribute tag from each first attribute tag; wherein the first confidence level of the first valid attribute tag is greater than the corresponding first confidence level threshold; filter at least one second valid attribute tag from each second attribute tag; wherein the second confidence level of the second valid attribute tag is greater than the corresponding second confidence level threshold; determine whether there is a business logic conflict between each first valid attribute tag and a tag pair with a first predetermined association relationship among each second valid attribute tag; wherein the business logic conflict is used to indicate that the structural feature corresponding to the first valid attribute tag and the semantic feature corresponding to the second valid attribute tag do not conform to a preset business rule; and in response to the absence of a business logic conflict between the tag pairs, fuse each first valid attribute tag and each second valid attribute tag to obtain a target attribute tag set.

[0151] In any embodiment of this application, the generation module 840 is configured to merge each first valid attribute tag and each second valid attribute tag to obtain an intermediate attribute tag set; schedule other first attribute tags (excluding the first valid attribute tags) and other second attribute tags (excluding the second valid attribute tags) in each first attribute tag to the target object for tag verification, and receive the first verification attribute tag and the second verification attribute tag fed back by the target object; and merge the intermediate attribute tag set, the first verification attribute tag and the second verification attribute tag to generate a target attribute tag set.

[0152] In any embodiment of this application, the generation module 840 is configured to, in response to a tag pair including a first tag pair with business logic conflict, update the tag status of the first valid attribute tag of the first tag pair to an unknown state to obtain the first attribute tag after the status update; schedule the first attribute tag after the status update, other first attribute tags, and other second attribute tags to a target object for tag review, and receive the first review attribute tag and the second review attribute tag fed back by the target object; and merge the other first valid attribute tags, the first review attribute tag, and the second review attribute tag in each first valid attribute tag except for the first valid attribute tag in the first tag pair to generate a target attribute tag set.

[0153] In any embodiment of this application, the generation module 840 is configured to: obtain target text content associated with any first valid attribute tag from the text content for any first valid attribute tag; in response to a business logic conflict between the text attribute tag and the second valid attribute tag of the target text content, generate additional description information for any first valid attribute tag based on any first valid attribute tag, the text attribute tag, and the second valid attribute tag; annotate any first valid attribute tag with the additional description information to obtain an annotated first valid attribute tag; and generate a target attribute tag set based on each first valid attribute tag, each second valid attribute tag, and the additional description information.

[0154] In any embodiment of this application, the generation module 840 is configured to update the label status of the first attribute label set and the second attribute label set to pending processing in response to the fact that each first confidence level is less than the first confidence level threshold of the corresponding first attribute label and each second confidence level is less than the second confidence level threshold of the corresponding second attribute label; and in response to the review operation, perform label review on the attribute label set whose label status is pending processing to obtain the target label set.

[0155] In any embodiment of this application, the second processing module 830 is used to obtain a prompt template; wherein the prompt template is used to indicate the task information to be executed by the large language model; the prompt template is filled with the set semantic analysis instructions and the text content extracted from the electronic document by the AI ​​Agent to obtain model prompt information; the large language model is called to perform semantic understanding and logical reasoning on the text content based on the model prompt information to obtain a second attribute tag set.

[0156] In any embodiment of this application, the first processing module 820 is used to send a processing request to a processing engine associated with multiple target visual models via an RPA robot; wherein the processing request includes the electronic document, and the processing request is used for the processing engine to call multiple target visual recognition models in parallel to perform visual recognition on the electronic document, and to summarize the attribute tags returned by the multiple visual recognition models to generate a first attribute tag set.

[0157] In any embodiment of this application, the document processing apparatus 800 further includes a determination module.

[0158] The determination module is used to monitor the running status of each candidate visual recognition model through the AI ​​Agent; Obtain the recognition results output by each candidate visual recognition model in at least one historical time period; determine the evaluation index of each candidate visual recognition model in at least one evaluation dimension based on the recognition results of each candidate visual recognition model; determine the performance of each candidate visual recognition model based on the evaluation index of each candidate visual recognition model; and determine multiple target visual recognition models from multiple candidate visual recognition models based on the performance and running status of multiple candidate visual recognition models.

[0159] It should be noted that the file processing apparatus based on AI, RPA, LLM and AI Agent provided in this application embodiment can implement all the method steps implemented in any of the above method embodiments and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail.

[0160] Figure 9 A structural block diagram of an electronic device according to an embodiment of this application is shown. Figure 9 As shown, the electronic device includes a memory 910 and a processor 920. The memory 910 stores a computer program that can run on the processor 920. When the processor 920 executes the computer program, it implements the file processing methods based on AI, RPA, LLM, and AI Agent as described in the above embodiments. The number of memories 910 and processors 920 can be one or more.

[0161] The electronic device also includes: The communication interface 930 is used to communicate with external devices and exchange and transmit data.

[0162] If the memory 910, processor 920, and communication interface 930 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0163] Optionally, in a specific implementation, if the memory 910, processor 920, and communication interface 930 are integrated on a single chip, then the memory 910, processor 920, and communication interface 930 can communicate with each other through an internal interface.

[0164] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the file processing method based on AI, RPA, LLM, and AI Agent provided in any embodiment of this application.

[0165] This application also provides a chip, which includes a processor for calling and executing instructions stored in a memory, causing a communication device equipped with the chip to execute the file processing method based on AI, RPA, LLM, and AI Agent provided in any embodiment of this application.

[0166] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the file processing method based on AI, RPA, LLM, and AIAgent provided in any embodiment of the application.

[0167] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Computing (RISC) machines (ARM) architecture.

[0168] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0169] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0170] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0171] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0172] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.

[0173] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0174] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.

[0175] Furthermore, the functional units in the various embodiments of this application can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0176] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A document processing method based on artificial intelligence (AI), robotic process automation (RPA), large language modeling (LLM), and a digital employee platform intelligent agent (AI Agent), characterized in that: include: Obtain the electronic file to be processed; The electronic document is visually recognized using multiple target visual recognition models to generate a first attribute tag set; wherein the first attribute tag set includes multiple first attribute tags describing the structural features of the electronic document; The LLM is used to perform semantic understanding and logical reasoning on the text content of the electronic document extracted by the AI ​​Agent to obtain a second attribute tag set; wherein, the second attribute tag set includes multiple second attribute tags describing the semantic features of the electronic document; The target attribute tag set of the electronic document is generated based on the first attribute tag set and the second attribute tag set.

2. The method according to claim 1, characterized in that, The first attribute tag set also includes a first confidence level corresponding to each first attribute tag, and the second attribute tag set also includes a second confidence level corresponding to each second attribute tag; The step of generating the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set includes: From each of the first attribute labels, at least one first valid attribute label is selected; wherein the first confidence level of the first valid attribute label is greater than the corresponding first confidence threshold. From each of the second attribute labels, at least one second valid attribute label is selected; wherein the second confidence level of the second valid attribute label is greater than the corresponding second confidence threshold. Determine whether there is a business logic conflict between each of the first valid attribute tags and each of the second valid attribute tags that have a first set association relationship; wherein, the business logic conflict is used to indicate that the structural feature corresponding to the first valid attribute tag and the semantic feature corresponding to the second valid attribute tag do not conform to the preset business rules; In response to the absence of business logic conflicts between the tag pairs, each of the first valid attribute tags and each of the second valid attribute tags are merged to obtain a target attribute tag set.

3. The method according to claim 2, characterized in that, In response to the absence of business logic conflicts between the tag pairs, the first valid attribute tags and the second valid attribute tags are merged to obtain a target attribute tag set, including: In response to the absence of business logic conflicts between the tag pairs, each of the first valid attribute tags and each of the second valid attribute tags are merged to obtain an intermediate attribute tag set; All first attribute tags other than the first valid attribute tag in each of the first attribute tags and all second attribute tags other than the second valid attribute tag in each of the second attribute tags are scheduled to the target object for tag verification, and the first verification attribute tag and the second verification attribute tag fed back by the target object are received. The intermediate attribute tag set, the first verification attribute tag, and the second verification attribute tag are merged to generate the target attribute tag set.

4. The method according to claim 3, characterized in that, The method further includes: In response to the tag pair including a first tag pair with business logic conflict, the tag state of the first valid attribute tag of the first tag pair is updated to an unknown state to obtain the first attribute tag after the state update. The updated first attribute label, the other first attribute labels, and the other second attribute labels are scheduled to the target object for label verification, and the first and second verification attribute labels fed back by the target object are received. The first valid attribute tags, excluding the first valid attribute tags in the first tag pair, the first verification attribute tags, and the second verification attribute tags in each of the first valid attribute tags are merged to generate the target attribute tag set.

5. The method according to claim 2, characterized in that, The step of generating the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set includes: For any first valid attribute tag, obtain the target text content associated with the first valid attribute tag from the text content; In response to a business logic conflict between the text attribute tag of the target text content and the second valid attribute tag, additional description information of the first valid attribute tag is generated based on the first valid attribute tag, the text attribute tag, and the second valid attribute tag. The additional description information is used to annotate any of the first valid attribute tags to obtain the annotated first valid attribute tags; The target attribute tag set is generated based on each of the first valid attribute tags, each of the second valid attribute tags, and the additional description information.

6. The method according to claim 2, characterized in that, The method further includes: In response to each of the first confidence scores being less than the first confidence score threshold of the corresponding first attribute label, and each of the second confidence scores being less than the second confidence score threshold of the corresponding second attribute label, the label status of the first attribute label set and the second attribute label set is updated to pending processing; In response to the review operation, the attribute tag set with the tag status of pending processing is reviewed to obtain the target tag set.

7. The method according to claim 1, characterized in that, The step of using the LLM to perform semantic understanding and logical reasoning on the text content extracted by the AI ​​Agent from the electronic document to obtain a second attribute tag set includes: Obtain a prompt template; wherein the prompt template is used to indicate the task information to be executed by the large language model; The set semantic analysis instructions and the text content extracted from the electronic document by the AI ​​Agent are used to fill the prompt template to obtain model prompt information; The large language model is invoked to perform semantic understanding and logical reasoning on the text content based on the model prompt information, so as to obtain the second attribute tag set.

8. The method according to claim 1, characterized in that, The step of visually recognizing the electronic document using multiple visual recognition models to generate a first attribute tag set includes: The RPA robot sends processing requests to the processing engine associated with multiple target visual models; The processing request includes the electronic document. The processing request is used by the processing engine to call multiple target visual recognition models in parallel to perform visual recognition on the electronic document, and to summarize the attribute tags returned by the multiple visual recognition models to generate a first attribute tag set.

9. The method according to claim 1, characterized in that, Before using multiple target visual recognition models to perform visual recognition on the electronic document to generate the first attribute tag set, the method further includes: The AI ​​Agent monitors the running status of each candidate visual recognition model; Obtain the recognition results output by each candidate visual recognition model in at least one historical time period; Based on the recognition results of each candidate visual recognition model, determine the evaluation index of each candidate visual recognition model in at least one evaluation dimension. The performance of each candidate visual recognition model is determined based on the evaluation metrics of each candidate recognition model. Based on the performance and operating status of multiple candidate visual recognition models, multiple target visual recognition models are determined from the multiple candidate visual recognition models.

10. A document processing device based on AI, RPA, LLM, and AI Agent, characterized in that, include: The first acquisition module is used to acquire the electronic files to be processed. The first processing module is used to perform visual recognition on the electronic document using multiple target visual recognition models to generate a first attribute tag set; wherein, the first attribute tag set includes multiple first attribute tags describing the structural features of the electronic document; The second processing module is used to perform semantic understanding and logical reasoning on the text content of the electronic document extracted by the AI ​​Agent using the LLM, so as to obtain a second attribute tag set; wherein, the second attribute tag set includes multiple second attribute tags describing the semantic features of the electronic document; The generation module is used to generate the target attribute tag set of the electronic document based on the first attribute tag set and the second attribute tag set.

11. The apparatus according to claim 10, characterized in that, The first attribute tag set further includes a first confidence level corresponding to each first attribute tag, and the second attribute tag set further includes a second confidence level corresponding to each second attribute tag. The generation module is configured to: From each of the first attribute labels, at least one first valid attribute label is selected; wherein the first confidence level of the first valid attribute label is greater than the corresponding first confidence threshold. From each of the second attribute labels, at least one second valid attribute label is selected; wherein the second confidence level of the second valid attribute label is greater than the corresponding second confidence threshold. Determine whether there is a business logic conflict between each of the first valid attribute tags and each of the second valid attribute tags that have a first set association relationship; wherein, the business logic conflict is used to indicate that the structural feature corresponding to the first valid attribute tag and the semantic feature corresponding to the second valid attribute tag do not conform to the preset business rules; In response to the absence of business logic conflicts between the tag pairs, each of the first valid attribute tags and each of the second valid attribute tags are merged to obtain a target attribute tag set.

12. The apparatus according to claim 11, characterized in that, The generation module is used for: The first valid attribute tags and the second valid attribute tags are merged to obtain an intermediate attribute tag set; All first attribute tags other than the first valid attribute tag in each of the first attribute tags and all second attribute tags other than the second valid attribute tag in each of the second attribute tags are scheduled to the target object for tag verification, and the first verification attribute tag and the second verification attribute tag fed back by the target object are received. The intermediate attribute tag set, the first verification attribute tag, and the second verification attribute tag are merged to generate the target attribute tag set.

13. The apparatus according to claim 10, characterized in that, The generation module is used for: In response to the tag pair including a first tag pair with business logic conflict, the tag state of the first valid attribute tag of the first tag pair is updated to an unknown state to obtain the first attribute tag after the state update. The updated first attribute label, the other first attribute labels, and the other second attribute labels are scheduled to the target object for label verification, and the first and second verification attribute labels fed back by the target object are received. The first valid attribute tags, excluding the first valid attribute tags in the first tag pair, the first verification attribute tags, and the second verification attribute tags in each of the first valid attribute tags are merged to generate the target attribute tag set.

14. An electronic device, characterized in that, include: A processor and a memory, wherein instructions are stored in the memory and loaded and executed by the processor to implement the method as described in any one of claims 1 to 9.

15. A computer-readable storage medium storing a computer program therein, the computer program implementing the method as described in any one of claims 1-9 when executed by a processor.

Citation Information

Patent Citations

  • Digital processing method, system and equipment for paper archives based on artificial intelligence

    CN112800949A

  • Intelligent document identification method and device, electronic equipment and storage medium

    CN118194842A

  • Image file information automatic extraction method and system based on agent workflow

    CN119917576A

  • Label data generation method and device and storage medium

    CN120451981A

  • Image annotation method and apparatus

    WO2024045641A1

Cited By

  • Document field extraction method and device, computer device and readable storage medium

    CN122223737A