File analysis method, electronic equipment, storage medium and product
By analyzing and analyzing different types of documents, and using big models to generate Q&A information and description information, the existing document analysis methods are solved, and efficient and accurate document analysis is achieved.
Patent Information
- Application Number
- CN202510110544.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing document analysis methods have a narrow application range and need to be set up in advance, which is time-consuming, inefficient and heavily dependent on personal experience, affecting the accuracy of the analysis results.
By obtaining the target file to be analyzed, calling the corresponding parser according to the file format for parsing, extracting key information, and using the big model to generate and answer mining, generating question-and-answer information and description information.
Effectively respond to different types of documents, expand the scope of document analysis, reduce setting time, improve analysis efficiency and accuracy, and reduce dependence on personal experience.
Smart Images

Figure CN120030145A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information retrieval technology, and in particular to a file analysis method, electronic equipment, storage medium and product. Background Art
[0002] In the current information age, document data processing and document analysis have become an important research field. Existing document analysis technologies mainly include rule-based analysis methods and template-based analysis methods.
[0003] However, due to the diversity of existing document formats, rule-based analysis methods and template-based analysis methods are difficult to handle all types of documents, which to a certain extent limits the widespread application of document analysis and the data source for large model training. In addition, before analyzing a document using the above methods, it is necessary to set the corresponding analysis rules or templates for the document in advance. This setting process is time-consuming, inefficient, and heavily dependent on the personal experience of the person who sets it, affecting the accuracy of the analysis results. Summary of the invention
[0004] In view of the shortcomings of existing methods, the present application proposes a file analysis method, electronic device, storage medium and product, which can solve the problems that existing document analysis methods have a narrow application range, require pre-settings, are time-consuming, inefficient and rely on personal experience, affecting the accuracy of analysis results.
[0005] According to one aspect of an embodiment of the present application, an embodiment of the present application provides a file analysis method, the method comprising:
[0006] Acquire a target file to be analyzed, call a corresponding parser according to the format of the target file to parse the target file, and obtain a parsing result, wherein the target file includes a policy file;
[0007] Extracting key information from the parsing result according to the prompt words corresponding to the target file;
[0008] Question generation and answer mining are performed according to the key information and the analysis results to obtain question and answer information corresponding to the target file, and description information of the target file is generated based on the question and answer information and the key information.
[0009] In a possible implementation, the format includes at least one of docx, PDF, and picture, and calling a corresponding parser according to the format of the target file to parse the target file includes:
[0010] Classifying the target files according to the formats of the target files;
[0011] If it is determined that the target file includes an image or the format is a preset format, extracting text from the target file using an OCR model, and generating text according to the extraction result, wherein the OCR model is trained using images, and the preset format includes a PDF generated by scanning;
[0012] The parser is called to parse the text and obtain a parsing result.
[0013] In a possible implementation, the training of the OCR model includes:
[0014] Acquire pictures for training according to the field corresponding to the target file, and annotate the pictures, wherein the annotations include annotations of text content and annotations of structural information;
[0015] Generate a training set and a test set based on the annotated images, and use the training set and the test set to perform model training, wherein the base model for the model training includes PaddleOCR;
[0016] If it is determined that the trained model does not meet the preset requirements, incremental training is performed on the model to obtain the OCR model.
[0017] In a possible implementation, the parsing result includes document content, and extracting key information from the parsing result according to the prompt word corresponding to the target file includes:
[0018] The key information is extracted from the document content using the key information extraction module of the large model, and the key information corresponds to the prompt words set in the key information module.
[0019] In a possible implementation, the large model includes a question generation module and an answer mining module, and the question generation based on the key information and the parsing result includes:
[0020] The question generation module of the large model is used to generate questions corresponding to the target file. The questions are generated by the question generation module based on the answers output by the answer mining module and the key information and the document content. The question generation module includes a prompt combination composed of multiple question prompt words, and the question prompt words are used to guide the question generation module to generate the questions.
[0021] In a possible implementation, the answer mining includes:
[0022] The answer mining module is used to extract the answer corresponding to the question from a reference text, the reference text includes the key information and the document content, and the answer mining module includes a prompt combination composed of multiple answer prompt words, and the answer prompt words are used to guide the answer mining module to generate the answer.
[0023] In a possible implementation, the generating the description information of the target file based on the question-answer information and the key information includes:
[0024] Integrate the question and answer to obtain question and answer information, and integrate the question and answer information with the key information to obtain description information;
[0025] The description information is output in a structured form.
[0026] According to one aspect of an embodiment of the present application, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described above.
[0027] According to one aspect of an embodiment of the present application, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed, the steps of the method described above are implemented.
[0028] According to one aspect of an embodiment of the present application, an embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0029] The beneficial technical effects brought about by the technical solution provided by the embodiment of the present application include:
[0030] The present application provides a file analysis method, which has the beneficial effects of obtaining a target file to be analyzed, calling a corresponding parser to parse the target file according to the format of the target file, and obtaining a parsing result; extracting key information from the parsing result according to the prompt words corresponding to the target file; generating questions and mining answers according to the key information and the parsing result, obtaining the question and answer information corresponding to the target file, and generating description information of the target file based on the question and answer information and the key information. The present application can perform corresponding parsing according to the type of document and analyze the content of the document using a large model, thereby effectively responding to different types of documents, expanding the scope of document analysis without pre-setting analysis rules and templates, effectively shortening the time spent on setup, improving analysis efficiency and reliance on personal experience, and improving the accuracy and stability of analysis results.
[0031] Additional aspects and advantages of the present application will be partially given in the following description, which will become apparent from the following description, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0033] Figure 1 A flowchart of a file analysis method provided in an embodiment of the present application;
[0034] Figure 2 A flowchart of OCR model training provided in an embodiment of the present application;
[0035] Figure 3 A workflow diagram for file analysis provided in an embodiment of the present application;
[0036] Figure 4 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0038] It will be understood by those skilled in the art that, unless specifically stated, the "said" and "the" used herein may also include plural forms. It should be further understood that the wording "including" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the implementation of other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element may be directly connected or coupled to the other element, or it may refer to the connection relationship between the element and the other element through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein refers to at least one of the items defined by the term, for example, "A and / or B" may be implemented as "A", or as "B", or as "A and B".
[0039] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0040] The embodiment of the present application provides a file analysis method, which can be used in mobile phones, tablet computers, servers, cloud platforms and other smart terminals. The smart terminal obtains a target file to be analyzed, obtains a parsing result of the target file, uses a large model to extract key information and perform question and answer on the parsing result, and obtains description information of the target file.
[0041] like Figure 1-Figure 3 As shown, the file analysis method of the present application includes:
[0042] S101: Obtain a target file to be analyzed, call a corresponding parser according to the format of the target file to parse the target file, and obtain a parsing result.
[0043] Optionally, the target file can be various policy files, which can be scanned copies, PDFs, pictures, Excel, docx and other formats, or URL network links. The smart terminal executing the method of the present application can obtain the target file to be analyzed from a specified website, external storage device or local storage according to the received instructions, or obtain files uploaded by the user, and determine the template file to be analyzed from these files according to the user's selection instructions.
[0044] Optionally, a corresponding parser is called to parse the target file according to the format of the target file, including: classifying the target file according to the format of the target file; if it is determined that the target file includes a picture or the format is a preset format, using an OCR model to extract text from the target file, and generating text according to the extraction result, the OCR model is trained using pictures, and the preset format includes generating PDF by scanning; calling the parser to parse the text and obtain the parsing result.
[0045] Optionally, to improve the accuracy of the parsing results, different parsers may be used to parse target files of different formats. Moreover, to facilitate the parser in obtaining information in the image, an OCR model may be used to identify images or image-like objects (such as files or pages generated by scanning) in the target file before parsing with the parser.
[0046] Optionally, if it is determined that the target file does not need to use the OCR model for text extraction, a parser is called to perform parsing according to the classification of the target file to obtain a parsing result of the target file.
[0047] In one embodiment, a document file to be analyzed is received according to a user's instruction, and the document file is determined as a target file. The format of the target file is obtained, and the target file is classified according to the format. After classification, it is detected whether the target file meets the conditions for using the OCR model for recognition, and the conditions include the existence of an image file, the existence of an image in the file, the file is a PDF obtained by scanning, and other conditions that require the use of the OCR model for text extraction. If so, the file or image that needs to be recognized by the OCR model in the target file is obtained, and the text information in the file or image is extracted by the OCR model, and the text information is converted into editable text. The text is placed in the category corresponding to the original file or image for subsequent parsing.
[0048] Optionally, the training of the OCR model includes: obtaining pictures for training according to the field corresponding to the target file, annotating the pictures, and the annotations include annotations of text content and structural information; generating a training set and a test set based on the annotated pictures, and using the training set and the test set to perform model training, wherein the base model for model training includes PaddleOCR; if it is determined that the trained model does not meet the preset requirements, performing incremental training on the model to obtain an OCR model.
[0049] Optionally, the field corresponding to the target file can be determined, a large number of files related to the field can be obtained, the large number of files can be converted into pictures, and the pictures can be labeled to obtain samples for model training.
[0050] Optionally, when annotating text content, manual annotation or tool annotation can be used. The corresponding relationship between different regions and text in the image can be annotated by this annotation method, thereby providing learning answers for the model to be trained.
[0051] Optionally, when annotating structural information, various structural information in the image can be annotated, and the structural information includes structural information such as title, text, date, etc. This structural information helps the model understand the file structure and improve recognition accuracy.
[0052] Optionally, the base model may be a PP-OCR series model of PaddleOCR, or other models that can be used for OCR recognition.
[0053] Optionally, the preset requirement may be that the accuracy of the model is greater than a preset threshold. Incremental training may use new images to form a new training set, and use the new training set to train the model again, or incremental training may be performed by adjusting training parameters, such as adjusting the number of training set samples or learning rounds.
[0054] In one embodiment, the target file may be a policy document, and the training process of the OCR model may include: preparing a source file data set, recording text labels on the file data, recording labels on the document's structural information on the file data, dividing the training set and the test set, using the training set to train the OCR model, evaluating the trained OCR model on the test set, performing incremental training on the model based on the evaluation results, detecting whether the model has achieved the required effect, and obtaining an OCR model that can accurately identify various parts of the policy document. Specifically, the training process may be: preparing the original file data set: collecting a large number of policy documents, including scans, PDFs, pictures, etc. in various formats, as basic data for training and testing the OCR model. And converting all these files into pictures. This is the starting point of the entire process, and the data quality directly affects the model effect. Marking the file data with text labels: manually or using tools to annotate the text content in the file, that is, marking which areas in the picture converted from the collected policy documents correspond to which text, and providing the model with the "correct answer" for learning. This step is the key to training the model to recognize text. Label all file data with the document's structural information: In addition to the text content, you also need to label all file structural information, such as title, body, issuing agency, date, etc., to help the model understand the file layout and organizational structure. This helps improve the model's understanding of the file structure and recognition accuracy. Split training set and test set: Divide the labeled data set into training set and test set. The training set is used to train the model, and the test set is used to evaluate the model performance to prevent overfitting. Use the training set to train the OCR model: The base model of the OCR model selects PaddleOCR's PP-OCR series model. Because this series of models performs well in Chinese scenarios. And the lightweight PP-OCR series model is suitable for quickly building an OCR system. Install the PaddleOCR operating environment and download the base model. Use the collected training set data to train the OCR model, so that the model learns the correspondence between images and text and the rules of file structure to meet the requirements of use. This is the core process of model learning and optimization. Evaluate the trained OCR model on the test set: Use the test set to evaluate the performance of the trained OCR model, such as character recognition rate, layout analysis accuracy, etc., to test the model's performance on unseen data and whether the model meets the preset requirements. You can judge whether the model performance meets expectations based on the accuracy of the recognition results. If not, improvements and incremental training are needed. This is an important judgment node that determines whether iterative optimization is needed. Incremental training of the model: If the model does not meet the preset requirements, use new data to form a new training set or adjust training parameters, such as the number of training set samples or learning rounds, to train again to improve model performance. This is an iterative optimization process.Improve the model by using the recreated training set on the existing model for further training instead of retraining the model. Continuously improve the model performance through incremental training. Obtain an OCR model that can accurately identify each part of the policy document: When the performance of the model on the test set meets the expected requirements, the final usable OCR model is obtained and can be applied to the actual scenario. This is the ultimate goal of the process, to produce a usable model.
[0055] Optionally, for different types of target files, corresponding parsers are used to extract text content and structural information. The parser can extract the content and structural information of the target file to obtain a parsing result, and the parsing result can be converted into a specific structured format. The type of the structured format can be determined according to the type of the large model used subsequently and the data requirements of the large model.
[0056] Optionally, after obtaining the parsing result, the parsing result may be stored, and the parsing result may be displayed according to the received query request.
[0057] S102: Extract key information from the parsing result according to the prompt words corresponding to the target file.
[0058] Optionally, the parsing result includes document content, and key information in the parsing result is extracted according to the prompt words corresponding to the target file, including: using the key information extraction module of the large model to extract key information from the document content, and the key information corresponds to the prompt words set in the key information module.
[0059] Optionally, the key information extraction module can be an Agent, which is composed of a combination of multiple prompt words, and each combination has multiple pre-edited guiding sentences related to the key information extraction task. When the target file is a policy document, the key information extraction module can extract key information such as the issuing agency, issuing date, effective time, and file level from the parsing results. The type of key information can be determined based on the content, type, and other information of the policy document. Among them, different prompt words can be set for different types of target files to improve the accuracy of key information and its fit with the file content.
[0060] In one embodiment, examples of prompt words for the key information extraction module may be:
[0061] Role-guided sentence: You are an assistant for extracting the issuing agency of a policy document. You can accurately extract the issuing agency from the text entered by the user.
[0062] Requirement guiding statement: Requirements: 1. If there are typos in the text entered by the user, your output needs to correct them. 2. If there are multiple issuing agencies, separate them with a comma. 3. Directly output the issuing agencies without any extra content.
[0063] Example prompting statement: User input: **Municipal Civil Affairs Bureau** **Municipal Finance Bureau** **Municipal Human Resources and Social Security Bureau** Document No. Fa 〔2022〕18 **Municipal Civil Affairs Bureau** **Municipal Finance Bureau** Zhanjiang Municipal Human Resources and Social Security Bureau Notice on Forwarding the Work of Promoting the Issuance of Social Assistance and Subsidy Funds in the Civil Affairs Field through Social Security Cards To the Civil Affairs Bureaus (Social Affairs Bureaus), Finance Bureaus, and Human Resources and Social Security Bureaus of each county (city, district): Hereby forward the Notice of **Provincial Department of Civil Affairs, Provincial Department of Finance, Provincial Department of Human Resources and Social Security** on Promoting the Issuance of Social Assistance and Subsidy Funds in the Civil Affairs Field through Social Security Cards. Model output: **Municipal Civil Affairs Bureau,** **Municipal Finance Bureau,** **Municipal Human Resources and Social Security Bureau**
[0064] Model output result: [Extracted issuing agencies]
[0065] After obtaining the parsing result output by the parser, call the key information extraction module. Based on the above prompt words, this key information extraction module obtains the key information corresponding to the parsing result. Among them, this key information can be output in a structured form so that other modules or objects can use this structured information for other processing.
[0066] S103: Generate questions and mine answers based on the key information and the parsing result to obtain the Q&A information corresponding to the target document, and generate the description information of the target document based on the Q&A information and the key information.
[0067] Optionally, the large model also includes a question generation module and an answer mining module. Generating questions based on the key information and the parsing result includes: using the question generation module of the large model to generate questions corresponding to the target document. The questions are generated by the question generation module based on the answers output by the answer mining module, the key information, and the document content. The question generation module includes a prompt combination composed of multiple question prompt words, and the question prompt words are used to guide the question generation module to generate questions.
[0068] Optionally, both the question generation module and the answer mining module can be Agents. Multiple combinations composed of prompt words can be stored in these two modules. After obtaining the parsing result and the key information, select one or more combinations, and generate questions and mine answers according to the selected combinations.
[0069] Optionally, answer mining includes: using an answer mining module to extract answers corresponding to questions from a reference text, the reference text includes key information and document content, and the answer mining module includes a prompt combination consisting of multiple answer prompt words, and the answer prompt words are used to guide the answer mining module to generate answers. Through the mutual cooperation of the question generation module and the answer mining module, the quality of questions and answers is continuously improved.
[0070] Optionally, in order to avoid the situation where the results output by the large model include content that does not belong to the target file during key information extraction, question generation, and answer mining, the prompt words used for guidance may also include vocabulary that limits the scope of the output results, and guide the large model to use the received data to output the results through the vocabulary.
[0071] In one embodiment, the question generation module is a question generation agent, which is composed of a combination of multiple prompt words, each of which contains multiple pre-edited guide sentences related to the question generation task. The module performs reverse question answering based on the document content and key information supplemented by the answers mined by the answer mining module to generate relevant questions. The combination of prompt words used for question generation can be:
[0072] Role-guided sentence: You are a question generation assistant. The user's input contains reference text and generated answers, and your task is to generate questions that can be answered by the content of the reference text.
[0073] Requirements: 1. The questions you output must strictly exist in the reference text. You cannot output content that does not exist in the reference text. 2. Generate appropriate questions for the generated answers. 3. You must generate questions for all answers.
[0074] Model output: [generated question].
[0075] In another embodiment, the answer mining module can be an answer mining agent, which is composed of multiple prompt word combinations, each of which has multiple pre-edited guide sentences related to the answer mining task. The module is responsible for combining key information, parsing structure, and questions generated by the question generation agent to find answers corresponding to the questions. Examples of prompt words used by the model can be:
[0076] Role-guided sentence: You are an answer extraction assistant. The user's input includes reference text and questions. Your task is to accurately extract the answer to the question from the reference text.
[0077] Requirements: 1. The answer you output must strictly exist in the reference text. You cannot output content that does not exist in the reference text. 2. When the answer to the question does not exist in the reference text, directly output "empty".
[0078] Model output: [extracted answer].
[0079] Optionally, generating description information of the target file based on the question and answer information and the key information includes: integrating the question and answer to obtain the question and answer information, integrating the question and answer information and the key information to obtain description information; and outputting the description information in a structured form. The structured form may be in JSON format.
[0080] Optionally, after obtaining the question and answer information and key information, the question and answer information and key information may be converted into description information described in natural language for the convenience of user viewing.
[0081] Optionally, when displaying the description information, the key information may be displayed first, and then the questions and answers in the question and answer information may be displayed in the form of question and answer pairs.
[0082] When the target file is a policy document, the file analysis method of the present application performs file parsing on policy documents of different formats. Unlike ordinary OCR, which only outputs the text on the image, it can organize natural language descriptions of specific fields, such as the red-headed title: xxx, the seal date xxx under xxx, etc. And it can infer key information from the text, such as the effective or expiration date of the policy, other files mentioned in the text, etc. The original files input into the OCR model and parser are various policy documents. The OCR and file parsing services provide the big model with parsing results, which contain the file content and the structured information of the file, such as the position and size of elements such as text blocks, paragraphs, titles, tables, and pictures. The output content is the text content and structured information of the document, and is expressed in JSON format. Through the guiding statements for different tasks in the key information extraction agent, question generation agent, and answer mining agent, the big model is guided to complete the corresponding target tasks. The output (key information) of the extracted key information and the question and answer information are combined to form the final description information of the natural language description. Let the big model convert the combined information into fluent and natural text. The final output is natural language description information composed of key information extraction results and question-answer mining.
[0083] The file analysis method of this application has the following advantages:
[0084] 1. A pre-trained OCR model can read policy documents and other different types of documents in batches, and can be trained specifically for the typesetting of policy documents. Targeted training is performed on the red header with large red font, the specific specifications used for the document number, and the issuing agency and the date of writing at the red stamp at the end of the document. 2. Corresponding parsing method codes are specially constructed for files of various formats, and there is a most suitable parsing method for each file. 3. A group of key information combinations extracted from policy documents. In addition to the main content of the text, policy documents also contain a lot of important information, such as the issuing agency, document number, effective date, date of writing, etc. The correct and rapid extraction of this information is also the advantage of the present invention over other general methods. 4. A model specifically for judging the information extracted from the policy. The model can determine whether the information in the file is direct information that can be directly used for output, or whether the required information is not clearly stated in the file, or whether the model is based on the information conditions given in the text. For example: effective from xx time, valid for x years. In this way, the final expiration time can be inferred, and then the natural language output is organized into the final answer composed of the key information extraction results and question-answering mining.
[0085] Based on the same inventive concept, the embodiment of the present application provides an electronic device, such as Figure 4 As shown, Figure 4 The electronic device 2000 shown includes: a processor 2001 and a memory 2003. The processor 2001 and the memory 2003 are connected in communication with each other, for example, via a bus 2002.
[0086] Processor 2001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0087] The bus 2002 may include a path to transmit information between the above components. The bus 2002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0088] The memory 2003 can be a ROM (Read-Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (random access memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read-Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.
[0089] Optionally, the electronic device 2000 may further include a communication unit 2004. The communication unit 2004 may be used for receiving and sending signals. The communication unit 2004 may allow the electronic device 2000 to communicate with other devices wirelessly or by wire to exchange data. It should be noted that in actual applications, the communication unit 2004 is not limited to one.
[0090] Optionally, the electronic device 2000 may further include an input unit 2005. The input unit 2005 may be used to receive input digital, character, image and / or sound information, or generate key signal input related to user settings and function control of the electronic device 2000. The input unit 2005 may include, but is not limited to, one or more of a touch screen, a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, a camera, a microphone, etc.
[0091] Optionally, the electronic device 2000 may further include an output unit 2006. The output unit 2006 may be used to output or display information processed by the processor 2001. The output unit 2006 may include but is not limited to one or more of a display device, a speaker, a vibration device, and the like.
[0092] Although the electronic device 2000 having various devices is shown in the figure, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0093] Optionally, the memory 2003 is used to store a computer program for executing the solution of the present application, and the execution is controlled by the processor 2001. The processor 2001 is used to execute the computer program stored in the memory 2003 to implement the steps of any method provided in the embodiments of the present application.
[0094] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided in the present application / implements the steps of various optional implementation methods of the method provided in the present application.
[0095] Those skilled in the art will appreciate that the various operations, methods, steps, measures, and schemes in the processes discussed in this application may be alternated, changed, combined, or deleted. Furthermore, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be alternated, changed, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and schemes in the related art that are similar to those disclosed in this application may also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0096] In the description of the present application, the directions or positional relationships indicated by words such as "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside" and "outside" are based on the exemplary directions or positional relationships shown in the accompanying drawings. They are for the convenience of describing or simplifying the description of the embodiments of the present application, and do not indicate or imply that the referred device or component must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be understood as limitations on the present application.
[0097] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0098] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0099] In the description of this specification, specific features, structures, materials or characteristics may be combined in an appropriate manner in any one or more embodiments or examples.
[0100] The above is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the scheme of the present application, other similar implementation methods based on the technical ideas of the present application are also within the protection scope of the embodiments of the present application.
Claims
1. A file analysis method, characterized in that: The method comprises: Acquire a target file to be analyzed, call a corresponding parser according to the format of the target file to parse the target file, and obtain a parsing result, wherein the target file includes a policy file; Extracting key information from the parsing result according to the prompt words corresponding to the target file; Question generation and answer mining are performed according to the key information and the analysis results to obtain question and answer information corresponding to the target file, and description information of the target file is generated based on the question and answer information and the key information.
2. The file analysis method according to claim 1, characterized in that: The format includes at least one of docx, PDF, and picture, and calling a corresponding parser to parse the target file according to the format of the target file includes: Classifying the target files according to the formats of the target files; If it is determined that the target file includes an image or the format is a preset format, extracting text from the target file using an OCR model, and generating text according to the extraction result, wherein the OCR model is trained using images, and the preset format includes a PDF generated by scanning; The parser is called to parse the text and obtain a parsing result.
3. The file analysis method according to claim 2, characterized in that: The training of the OCR model includes: Acquire pictures for training according to the field corresponding to the target file, and annotate the pictures, wherein the annotations include annotations of text content and annotations of structural information; Generate a training set and a test set based on the annotated images, and use the training set and the test set to perform model training, wherein the base model for the model training includes PaddleOCR; If it is determined that the trained model does not meet the preset requirements, incremental training is performed on the model to obtain the OCR model.
4. The file analysis method according to claim 1, characterized in that: The parsing result includes document content, and extracting key information from the parsing result according to the prompt word corresponding to the target file includes: The key information is extracted from the document content using the key information extraction module of the large model, and the key information corresponds to the prompt words set in the key information module.
5. The file analysis method according to claim 4, characterized in that: The large model includes a question generation module and an answer mining module. The question generation based on the key information and the analysis result includes: The question generation module of the large model is used to generate questions corresponding to the target file. The questions are generated by the question generation module based on the answers output by the answer mining module and the key information and the document content. The question generation module includes a prompt combination composed of multiple question prompt words, and the question prompt words are used to guide the question generation module to generate the questions.
6. The file analysis method according to claim 5, characterized in that: The answer mining includes: The answer mining module is used to extract the answer corresponding to the question from a reference text, the reference text includes the key information and the document content, and the answer mining module includes a prompt combination composed of multiple answer prompt words, and the answer prompt words are used to guide the answer mining module to generate the answer.
7. The file analysis method according to claim 6, characterized in that: The generating the description information of the target file based on the question-answer information and the key information includes: Integrate the question and answer to obtain question and answer information, and integrate the question and answer information with the key information to obtain description information; The description information is output in a structured form.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, which implements the steps of the method according to any one of claims 1 to 7 when executed by a processor.