Information extraction method, device, equipment, medium and product

By using a multimodal information extraction model in information extraction, the inspection report files are analyzed and information extracted, and the problems of low and inaccurate information extraction efficiency in the prior art are solved, achieving a more efficient and robust information extraction effect.

CN120015220APending Publication Date: 2025-05-16GUANGZHOU KINGMED CENTER FOR CLINICAL LABORATORY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093219.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing information extraction methods are inefficient and inaccurate when processing inspection and detection results reports, especially when processing multimodal data.

Method used

Using a method based on the multimodal information extraction model, information extraction results of multimodal file analysis are extracted. By obtaining the analysis results of the pending inspection report file and the corresponding information extraction prompt words, it is input into the pre-trained multimodal information extraction model to generate information extraction results.

Benefits of technology

It improves the robustness and efficiency of information extraction, can process multimodal file input information more accurately, and solves the problems of inefficiency and inaccuracy in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015220A_ABST
    Figure CN120015220A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information extraction method and device, equipment, a medium and a product, and the method comprises the steps: carrying out the analysis processing of a to-be-processed inspection report file, and obtaining a file analysis result; according to the inspection type corresponding to the to-be-processed inspection report file and the file analysis result, information extraction cue words corresponding to the file analysis result are determined, and information extraction cue word information comprises a task instruction, the file analysis result, an extraction task sample and an output format sample; and inputting the file analysis result and the information extraction cue word into a pre-trained multi-modal information extraction model to obtain an information extraction result. According to the technical scheme of the embodiment of the invention, the problems that the efficiency is low and the accuracy is not enough when text recognition is carried out firstly and then information extraction is carried out in current information extraction are solved, information extraction can be carried out on the multi-modal file analysis result based on the multi-modal information extraction model, multi-modal model input information can be processed, and the information extraction efficiency is improved. And the robustness and efficiency of information extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to an information extraction method, device, equipment, medium and product. Background Art

[0002] The test result report contains a wealth of test result information. By extracting the useful information from the report, related services such as secondary diagnosis, health management, and archiving records can be provided.

[0003] In the existing information extraction process, most of them first perform text analysis on the entire document, extract all the text, and then use the language model to extract information; or, first perform document layout recognition, divide the layout area into different categories, such as pure text area, image area and table area, and then perform text recognition separately, and finally use the language model to extract information and output it in a summary.

[0004] However, the information extraction efficiency of the above detection information extraction process is not high, and the robustness of processing different types of data needs to be improved. Summary of the invention

[0005] The embodiments of the present invention provide an information extraction method, apparatus, device, medium and product, which can extract information from multimodal file parsing results based on a multimodal information extraction model, can process multimodal model input information, and improve the robustness and efficiency of information extraction.

[0006] In a first aspect, an embodiment of the present invention provides an information extraction method, the method comprising:

[0007] Obtaining the inspection report file to be processed, and parsing the inspection report file to be processed to obtain a file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed;

[0008] According to the inspection type and file parsing result corresponding to the inspection report file to be processed, determine the information extraction prompt word corresponding to the file parsing result; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples;

[0009] The file parsing results and information extraction prompt words are input into the pre-trained multimodal information extraction model to obtain the information extraction results corresponding to the inspection report file to be processed.

[0010] In a second aspect, an embodiment of the present invention provides an information extraction device, the device comprising:

[0011] A file parsing module is used to obtain the inspection report file to be processed, and parse the inspection report file to be processed to obtain a file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed;

[0012] A prompt word determination module is used to determine the information extraction prompt word corresponding to the file parsing result according to the inspection type corresponding to the inspection report file to be processed and the file parsing result; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples;

[0013] The information extraction module is used to input the file parsing results and information extraction prompt words into the pre-trained multimodal information extraction model to obtain the information extraction results corresponding to the inspection report file to be processed.

[0014] In a third aspect, an embodiment of the present invention further provides a computer device, the computer device comprising:

[0015] one or more processors;

[0016] A memory for storing one or more programs;

[0017] When the one or more programs are executed by one or more processors, the one or more processors implement the information extraction method provided by any embodiment of the present invention.

[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an information extraction method as provided in any embodiment of the present invention.

[0019] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the information extraction method provided by any embodiment of the present invention.

[0020] The embodiments of the above invention have the following advantages or beneficial effects:

[0021] The embodiment of the present invention obtains the inspection report file to be processed and parses the inspection report file to be processed to obtain the file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed; according to the inspection type and file parsing result corresponding to the inspection report file to be processed, the information extraction prompt word corresponding to the file parsing result is determined; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples; the file parsing results and the information extraction prompt word are input into a pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed. The technical solution of the embodiment of the present invention solves the problem that the current information extraction first performs text recognition and then extracts information, which is inefficient and inaccurate. It can extract information from the multimodal file parsing results based on the multimodal information extraction model, and can process the multimodal model input information, thereby improving the robustness and efficiency of information extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flow chart of an information extraction method provided by an embodiment of the present invention;

[0023] Figure 2 is a flow chart of an information extraction method provided by an embodiment of the present invention;

[0024] Figure 3 is a schematic diagram of information extraction provided by an embodiment of the present invention;

[0025] Figure 4 is a flow chart of an information extraction method provided by an embodiment of the present invention;

[0026] Figure 5 is a structural schematic diagram of an information extraction device provided by an embodiment of the present invention;

[0027] Figure 6 It is a structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0029] Figure 1The present invention provides a flowchart of an information extraction method, which can be applied to information extraction scenarios. The method can be performed by an information extraction device, which can be implemented by software and / or hardware and integrated into a computer device with application development functions.

[0030] like Figure 1 As shown, the information extraction method of this embodiment includes the following steps:

[0031] S110, obtaining the inspection report file to be processed, and parsing the inspection report file to be processed to obtain a file parsing result.

[0032] The file parsing result includes the location structure information of the file content in the inspection report file to be processed.

[0033] The file format of the inspection report file to be processed can be a file format such as picture, Word or PDF format (Portable Document Format, portable document format), including structured and unstructured content, structured content such as table, unstructured content such as one or more lines of text content, or one or more columns of text content. Usually the inspection report file to be processed has a certain layout or typesetting, and the texts in different positions have content relevance. For example, the test result of detection item A will be displayed in the row or column corresponding to detection item A. The file parsing result may include file content, i.e., the text content of the file, and the position structure information of the text content. The position structure information may include the coordinate information of the text content in the inspection report file to be processed, i.e., the coordinate information of the text block. The present embodiment can improve the information extraction effect by considering the layout relationship between texts.

[0034] S120: Determine, according to the inspection type corresponding to the inspection report file to be processed and the file parsing result, the information extraction prompt word corresponding to the file parsing result.

[0035] The information extraction prompt word can be an extraction instruction prompt, including information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples. The inspection type can be at least one inspection type in the biological and medical fields, such as a microbiological inspection type. It can be understood that the information content that needs to be extracted in the inspection report to be processed is different for different inspection types. Therefore, the information extraction task instructions, preset information extraction task samples and preset information extraction result output format samples that match the inspection type can be determined according to the inspection type, and then the information extraction prompt word corresponding to the file parsing result is obtained according to the file parsing results, the information extraction task instructions, the preset information extraction task samples and the preset information extraction result output format samples.

[0036] The information extraction task instructions are used to define the information that the model is expected to extract, such as the information content corresponding to the information item that matches the inspection type. The preset information extraction task examples are used to help the model better understand the extraction requirements of the corresponding inspection type. The preset information extraction result output format examples can be the structure of the expected model output, so that the downstream can perform other data processing and data analysis based on the information extraction results.

[0037] S130, inputting the file parsing result and the information extraction prompt word into a pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed.

[0038] The pre-trained multimodal information extraction model can be obtained by training based on a multimodal large model, such as Qwen2-VL, InternVL and other models. By taking inspection report files in different file formats such as images and PDF formats and the corresponding information extraction content as samples, the multimodal large model is trained to obtain a pre-trained multimodal information extraction model. The pre-trained multimodal information extraction model extracts information from the inspection report file to be processed according to the information extraction prompt words and the sample features learned in the training stage, and obtains the information extraction result corresponding to the inspection report file to be processed.

[0039] The technical solution of this embodiment obtains the inspection report file to be processed, and parses the inspection report file to be processed to obtain the file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed; according to the inspection type and file parsing result corresponding to the inspection report file to be processed, the information extraction prompt word corresponding to the file parsing result is determined; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples; the file parsing results and the information extraction prompt word are input into the pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed. The technical solution of the embodiment of the present invention solves the problem that the current information extraction first performs text recognition and then extracts information, which is inefficient and inaccurate. It can extract information from the multimodal file parsing results based on the multimodal information extraction model, and can process the multimodal model input information, thereby improving the robustness and efficiency of information extraction.

[0040] Figure 2 This is a flowchart of an information extraction method provided by an embodiment of the present invention. This embodiment and the information extraction method in the above embodiment belong to the same inventive concept, and further describes the process of determining the information extraction prompt word. The method can be executed by an information extraction device, which can be implemented by software and / or hardware and integrated into a computer device with application development function.

[0041] like Figure 2 As shown, the information extraction method of this embodiment includes the following steps:

[0042] S210, obtaining the inspection report file to be processed, and parsing the inspection report file to be processed to obtain a file parsing result.

[0043] The file parsing result includes the location structure information of the file content in the inspection report file to be processed.

[0044] In an optional implementation, the process of parsing the inspection report file to be processed and obtaining the file parsing result may be as follows:

[0045] Determine the file format of the inspection report file to be processed; if the file format of the inspection report file to be processed is not an image and can be parsed, parse the inspection report file to be processed through a preset file parsing component to obtain corresponding text and tables, and use the text and tables as file parsing results. If the file format of the inspection report file to be processed is not an image and cannot be parsed, perform file format conversion on the inspection report file to be processed to obtain a corresponding file image, and use the file image as the file parsing result. If the file format of the inspection report file to be processed is an image, the inspection report file to be processed can be directly used as the file parsing result.

[0046] This embodiment can set different parsing processes according to different file types and whether the file types can be directly parsed by preset tools, so as to realize multi-modal information extraction. Figure 3 As shown, when the file format of the inspection report file to be processed is PDF and can be parsed, for example, for the inspection report file to be processed in PDF file format that has not been encrypted or permission set, the inspection report file to be processed is parsed through preset file parsing components, such as PyPDF2 and PyMuPDF, to obtain the text and table in PDF, as well as the coordinate information of the text and table. When the file format of the inspection report file to be processed is PDF and cannot be parsed, for example, for the inspection report file to be processed in PDF file format that has been encrypted or permission set, the inspection report file to be processed is converted into a file format that can be processed by the machine, and the corresponding file image is obtained, and the file image is used as the file parsing result.

[0047] In this embodiment, during the information extraction process, the picture is encoded, the prompt word is encoded as text, the text encoding and the picture encoding are integrated, and the picture encoding is driven by the text encoding to achieve the effect of extracting the specified information.

[0048] When the file format of the inspection report file to be processed is an image, the inspection report file to be processed itself can be directly used as the file parsing result. The inspection report file to be processed in image format can be obtained by taking a photo or screenshot of the original electronic or paper report file.

[0049] S220. According to the inspection type corresponding to the inspection report file to be processed, determine the information extraction items, preset information extraction task samples and preset information extraction result output format samples that match the inspection type.

[0050] For example, when the test type is clinical testing, for the test report, the information extraction items matching the test type may include "name", "gender", "age", "department", "collection date", "receiving date", "clinical diagnosis", "examination items", "examination / sending doctor", "project name", "results", "identification", "unit", "reference value / range", "explanation (remarks)", etc., and the information can be assembled in the format of key-value pairs.

[0051] In step S210, the file parsing result includes the position structure information of the file content in the inspection report file to be processed, so that the matching accuracy of the information extraction items is improved by considering the position structure information of the file content in the information extraction process.

[0052] S230: Determine information extraction task instructions according to the information extraction items.

[0053] The information extraction task instruction defines the information extraction task that the model needs to perform. The information extraction task instruction can be determined based on the information extraction items and the role definition.

[0054] S240, splicing the information extraction task instruction, the file parsing result, the preset information extraction task sample and the preset information extraction result output format sample to obtain the information extraction prompt word.

[0055] S250, inputting the file parsing result and the information extraction prompt word into a pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed.

[0056] like Figure 3 As shown, if the file parsing result is an image, the file parsing result in image format and the information extraction prompt words are input into the pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed in image format or unparseable PDF format.

[0057] For the unprocessed inspection report file in a parsable PDF format, the parsed text block text, table text, and information extraction prompt words are input into a pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the unprocessed inspection report file in a parsable PDF format.

[0058] The technical solution of this embodiment is to obtain the inspection report file to be processed, and parse the inspection report file to be processed to obtain the file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed; according to the inspection type corresponding to the inspection report file to be processed, determine the information extraction items, preset information extraction task samples and preset information extraction result output format samples that match the inspection type; determine the information extraction task instructions according to the information extraction items; splice the information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples to obtain information extraction prompt words; wherein the information extraction prompt word information includes the information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples; input the file parsing results and the information extraction prompt words into the pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed. The technical solution of the embodiment of the present invention solves the problem that the current information extraction method of first performing text recognition and then extracting information is inefficient and inaccurate. It can extract information from multimodal file parsing results based on a multimodal information extraction model, and can process multimodal model input information to improve the robustness and efficiency of information extraction. It can also improve the accuracy of model output results by including structured prompt words for information extraction items of expected model output.

[0059] Figure 4 A flowchart of an information extraction method provided in an embodiment of the present invention, which belongs to the same inventive concept as the information extraction method in the above embodiment, further describes the training process of the multimodal information extraction model. The method can be executed by an information extraction device, which can be implemented by software and / or hardware and integrated into a computer device with application development function.

[0060] like Figure 4 As shown, the training process of the multimodal information extraction model in the information extraction method of this embodiment includes the following steps:

[0061] S310, obtaining a sample of the inspection report file and the corresponding file information extraction result.

[0062] Obtain inspection report files of at least two file formats for at least one inspection item as inspection report file samples. The file formats may include image formats and PDF formats, and the PDF formats include parsable PDF and unparsable PDF. Determine the file information extraction results corresponding to the inspection report file samples, and the file information extraction results of the samples can guide the parameter adjustment of the model during the training process. The file information extraction results may include the results of the expected model extraction, which may specifically be the field information in the expected preset format.

[0063] S320: construct a text information extraction sample set and a graphic information extraction sample set respectively according to the inspection report file sample and the file information extraction result.

[0064] According to the parsing results of the inspection report samples in the parseable PDF format and the inspection type of the samples, the extraction prompt words of the text information extraction sample set are determined, and the extraction prompt words and the file information extraction results are combined to obtain the text information extraction sample set.

[0065] According to the image conversion results of the inspection report samples in the unparseable PDF format, or the inspection report samples in the image format, and the inspection type of the samples, the extraction prompt words of the image and text information extraction sample set are determined, and the extraction prompt words and the file information extraction results are combined to obtain the image and text information extraction sample set.

[0066] In an optional embodiment, a text information extraction sample set and a graphic information extraction sample set are respectively constructed according to the inspection report file sample and the file information extraction result. The parseable file sample in the inspection report file sample may be parsed to obtain the corresponding files and tables, and corresponding text information extraction prompt words are constructed based on the parsed files and tables and the corresponding file information extraction results to obtain the text information extraction sample set; the text format of the unparseable file sample in the inspection report file sample is converted to obtain the corresponding file image, and corresponding graphic information extraction prompt words are constructed based on the file image obtained by the format conversion and the corresponding file information extraction result to obtain the graphic information extraction sample set.

[0067] The text information extraction prompt words and the graphic information extraction prompt words are information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples corresponding to the inspection report file samples.

[0068] For parsable file samples, such as parsable PDF, the model's extraction tasks and field information that need to be extracted are determined based on the parsed text files and tables and the corresponding file information extraction results. The corresponding text information extraction prompt words are constructed according to the extraction tasks and the field information that need to be extracted. The text information extraction prompt words and the corresponding file information extraction results are used to create a data set in the data format of a question-answer pair, where the question is the text information extraction prompt word and the answer is the file information extraction result.

[0069] For unparseable file samples, such as file samples whose original samples are in picture format or file samples in image format obtained through format conversion of unparseable PDF files, the model's extraction task and field information to be extracted are determined based on the file image obtained through format conversion and the corresponding file information extraction results, and the corresponding graphic and text information extraction prompt words are constructed. The graphic and text information extraction prompt words and the corresponding file information extraction results are used to create a data set in the data format of question-answer pairs, where the question is the graphic and text information extraction prompt word and the answer is the file information extraction result.

[0070] The method for determining the file information extraction result can be to parse the PDF file using open source tools such as PyPDF2 and PyMuPDF, and then manually complete the semi-automatic annotation method through tools such as large language models to obtain the file information extraction result corresponding to the parseable file sample. Convert the unparseable PDF file into an image, and form image data with data such as photos and screenshots. Use OCR to extract text, and then manually complete the semi-automatic annotation method through tools such as large language models to obtain the file information extraction result corresponding to the unparseable file sample.

[0071] S330 , training the multimodal information extraction model to be trained based on the text information extraction sample set and the graphic information extraction sample set to obtain a multimodal information extraction model.

[0072] Among them, the multimodal information extraction model to be trained is a multimodal large language model.

[0073] The text information extraction sample set and the graphic information extraction sample set are divided into training set, validation set and test set according to the preset ratio, and a multimodal information extraction model for inspection and testing reports is built. The multimodal large model can be used as the multimodal information extraction model to be trained, such as Qwen2-VL, InternVL and other series of models.

[0074] The dataset is used for training. The training methods can be full parameter fine-tuning, Lora fine-tuning, etc. The model is verified with the validation set. Finally, the extraction effect is evaluated on the test set to obtain a multimodal information extraction model.

[0075] In an optional implementation, the original image samples in the graphic information extraction sample set are subjected to at least one of random rotation, Gaussian noise addition, salt and pepper noise addition, image brightness adjustment, and hue adjustment to obtain new image samples; the new image samples and the corresponding original image sample file information extraction results are combined to form new graphic information extraction samples to achieve sample enhancement.

[0076] The original image samples of the training set are subjected to data enhancement, including at least one of random rotation, adding Gaussian or salt and pepper noise, brightness change, hue change, etc., to obtain new image samples and combine the new image samples with the file information extraction results of the corresponding original image samples to form new graphic information extraction samples to achieve sample enhancement, so that the model can learn to ignore Gaussian noise, salt and pepper noise, etc., focus on the essential features of the image for information extraction, and enhance the robustness of the model.

[0077] The technical solution of this embodiment is to obtain the inspection report file sample and the corresponding file information extraction result; construct a text information extraction sample set and a graphic information extraction sample set according to the inspection report file sample and the file information extraction result; train the multimodal information extraction model to be trained based on the text information extraction sample set and the graphic information extraction sample set to obtain a multimodal information extraction model; wherein the multimodal information extraction model to be trained is a multimodal large language model. The technical solution of the embodiment of the present invention solves the problem that the current information extraction first performs text recognition and then extracts information, which is inefficient and inaccurate. It can extract information from the file parsing results in picture or text format based on the multimodal information extraction model, and can process model input information in picture or text format, thereby improving the robustness and efficiency of information extraction.

[0078] Figure 5 This is a schematic diagram of the structure of an information extraction device provided by an embodiment of the present invention, and this embodiment is applicable to the scene of information extraction. The information extraction device can be implemented by software and / or hardware, and integrated into a computer terminal device with application development function.

[0079] like Figure 5 As shown, the information extraction device includes: a file parsing module 410, a prompt word determination module 420 and an information extraction module 430.

[0080] Among them, the file parsing module 410 is used to obtain the inspection report file to be processed, and parse the inspection report file to be processed to obtain the file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed; the prompt word determination module 420 is used to determine the information extraction prompt word corresponding to the file parsing result according to the inspection type and file parsing result corresponding to the inspection report file to be processed; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples; the information extraction module 430 is used to input the file parsing results and the information extraction prompt word into a pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed.

[0081] The technical solution of this embodiment obtains the inspection report file to be processed, and parses the inspection report file to be processed to obtain the file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed; according to the inspection type and file parsing result corresponding to the inspection report file to be processed, the information extraction prompt word corresponding to the file parsing result is determined; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples; the file parsing results and the information extraction prompt word are input into the pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed. The technical solution of the embodiment of the present invention solves the problem that the current information extraction first performs text recognition and then extracts information, which is inefficient and inaccurate. It can extract information from the multimodal file parsing results based on the multimodal information extraction model, and can process the multimodal model input information, thereby improving the robustness and efficiency of information extraction.

[0082] In an optional implementation, the file parsing module 410 is specifically used for:

[0083] Determine the file format of the inspection report file to be processed; when the file format of the inspection report file to be processed is not an image and can be parsed, parse the inspection report file to be processed through a preset file parsing component to obtain corresponding text and tables, and use the text and tables as file parsing results; when the file format of the inspection report file to be processed is not an image and cannot be parsed, convert the file format of the inspection report file to be processed to obtain the corresponding file image, and use the file image as the file parsing result; when the file format of the inspection report file to be processed is an image, use the inspection report file to be processed as the file parsing result.

[0084] In an optional implementation, the prompt word determination module 420 is specifically used to:

[0085] According to the inspection type corresponding to the inspection report file to be processed, determine the information extraction items, preset information extraction task samples and preset information extraction result output format samples that match the inspection type; determine the information extraction task instructions according to the information extraction items; splice the information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples to obtain information extraction prompt words.

[0086] In an optional embodiment, the device further comprises:

[0087] The model training module is used to obtain inspection report file samples and corresponding file information extraction results; construct text information extraction sample sets and graphic information extraction sample sets according to the inspection report file samples and the file information extraction results respectively; train the multimodal information extraction model to be trained based on the text information extraction sample sets and the graphic information extraction sample sets to obtain a multimodal information extraction model; wherein the multimodal information extraction model to be trained is a multimodal large language model.

[0088] In an optional implementation, the model training module is specifically used to:

[0089] Parse the parseable file samples in the inspection report file samples to obtain corresponding files and tables, and construct corresponding text information extraction prompt words based on the parsed files and tables and the corresponding file information extraction results to obtain a text information extraction sample set; perform text format conversion on the unparseable file samples in the inspection report file samples to obtain corresponding file images, and construct corresponding graphic and text information extraction prompt words based on the file images obtained by the format conversion and the corresponding file information extraction results to obtain a graphic and text information extraction sample set.

[0090] In an optional implementation, the model training module is further used to:

[0091] The original image samples in the image and text information extraction sample set are subjected to at least one of random rotation, Gaussian noise addition, salt and pepper noise addition, image brightness adjustment, and hue adjustment to obtain new image samples; the new image samples and the file information extraction results of the corresponding original image samples are combined into new image and text information extraction samples to achieve sample enhancement.

[0092] The information extraction device provided in the embodiment of the present invention can execute the information extraction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0093] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 6 A block diagram of an exemplary computer device 12 suitable for use in implementing embodiments of the present invention is shown. Figure 6 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as an intelligent controller and server, a mobile phone and other terminal devices.

[0094] like Figure 6 As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).

[0095] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0096] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0097] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 6 not shown, usually called a "hard drive"). Although Figure 6 Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The system memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.

[0098] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.

[0099] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the computer device 12 via the bus 18. It should be understood that although Figure 6 Not shown, other hardware and / or software modules may be used in conjunction with computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RFID systems, tape drives, and data backup storage systems.

[0100] The processing unit 16 executes various functional applications and data processing by running the program stored in the system memory 28, for example, implementing the information extraction method provided in the embodiment of the present invention, which includes:

[0101] Obtaining the inspection report file to be processed, and parsing the inspection report file to be processed to obtain a file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed;

[0102] According to the inspection type and file parsing result corresponding to the inspection report file to be processed, determine the information extraction prompt word corresponding to the file parsing result; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples;

[0103] The file parsing results and information extraction prompt words are input into the pre-trained multimodal information extraction model to obtain the information extraction results corresponding to the inspection report file to be processed.

[0104] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the information extraction method provided by any embodiment of the present invention is implemented. The method includes:

[0105] Obtaining the inspection report file to be processed, and parsing the inspection report file to be processed to obtain a file parsing result; wherein the file parsing result includes the location structure information of the file content in the inspection report file to be processed;

[0106] According to the inspection type and file parsing result corresponding to the inspection report file to be processed, determine the information extraction prompt word corresponding to the file parsing result; wherein the information extraction prompt word information includes information extraction task instructions, file parsing results, preset information extraction task samples and preset information extraction result output format samples;

[0107] The file parsing results and information extraction prompt words are input into the pre-trained multimodal information extraction model to obtain the information extraction results corresponding to the inspection report file to be processed.

[0108] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0109] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0110] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0111] Computer program code for performing the operation of the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, Python, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0112] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the information extraction method provided in any embodiment of the present application.

[0113] In the process of implementation, the computer program product can be written in one or more programming languages ​​or a combination thereof to perform the computer program code of the present invention, including object-oriented programming languages, such as Java, Smalltalk, Python, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0114] It should be understood by those skilled in the art that the modules or steps of the present invention described above can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, optionally, they can be implemented by a program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0115] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An information extraction method, characterized in that: include: Obtaining a test report file to be processed, and parsing the test report file to be processed to obtain a file parsing result; wherein the file parsing result includes location structure information of the file content in the test report file to be processed; According to the inspection type corresponding to the inspection report file to be processed and the file parsing result, determine the information extraction prompt word corresponding to the file parsing result; wherein the information extraction prompt word information includes information extraction task instructions, the file parsing result, a preset information extraction task sample and a preset information extraction result output format sample; The file parsing result and the information extraction prompt word are input into a pre-trained multimodal information extraction model to obtain an information extraction result corresponding to the inspection report file to be processed.

2. The method according to claim 1, characterized in that The parsing of the inspection report file to be processed to obtain a file parsing result includes: Determine the file format of the inspection report file to be processed; In the case where the file format of the inspection report file to be processed is not an image and can be parsed, the inspection report file to be processed is parsed by a preset file parsing component to obtain corresponding text and table, and the text and table are used as the file parsing result; When the file format of the inspection report file to be processed is not an image and cannot be parsed, convert the file format of the inspection report file to be processed to obtain a corresponding file image, and use the file image as the file parsing result; In the case that the file format of the inspection report file to be processed is an image, the inspection report file to be processed is used as the file parsing result.

3. The method according to claim 1 or 2, characterized in that: The step of determining, according to the inspection type corresponding to the inspection report file to be processed and the file parsing result, an information extraction prompt word corresponding to the file parsing result comprises: According to the inspection type corresponding to the inspection report file to be processed, determine the information extraction items matching the inspection type, the preset information extraction task sample and the preset information extraction result output format sample; Determine the information extraction task instruction according to the information extraction item; The information extraction task instruction, the file parsing result, the preset information extraction task sample and the preset information extraction result output format sample are spliced ​​to obtain the information extraction prompt word.

4. The method according to claim 1, characterized in that: The training process of the multimodal information extraction model includes: Obtain inspection report file samples and corresponding file information extraction results; Constructing a text information extraction sample set and a graphic information extraction sample set respectively according to the inspection report file sample and the file information extraction result; Training the multimodal information extraction model to be trained based on the text information extraction sample set and the graphic information extraction sample set to obtain the multimodal information extraction model; Wherein, the multimodal information extraction model to be trained is a multimodal large language model.

5. The method according to claim 4, characterized in that The step of constructing a text information extraction sample set and a graphic information extraction sample set according to the inspection report file sample and the file information extraction result respectively includes: Parse the parseable file samples in the inspection report file samples to obtain corresponding files and tables, and construct corresponding text information extraction prompt words based on the parsed files and tables and the corresponding file information extraction results to obtain the text information extraction sample set; perform text format conversion on the unparseable file samples in the inspection report file samples to obtain corresponding file images, and construct corresponding graphic information extraction prompt words based on the file images obtained by format conversion and the corresponding file information extraction results to obtain the graphic information extraction sample set.

6. The method according to claim 4, characterized in that The method further comprises: Perform at least one of random rotation, Gaussian noise addition, salt and pepper noise addition, image brightness adjustment, and hue adjustment on the original image sample in the image and text information extraction sample set to obtain a new image sample; The new image sample and the corresponding file information extraction result of the original image sample are combined into a new image and text information extraction sample to achieve sample enhancement.

7. An information extraction device, characterized in that: include: A file parsing module, used to obtain a test report file to be processed, and parse the test report file to be processed to obtain a file parsing result; wherein the file parsing result includes location structure information of the file content in the test report file to be processed; A prompt word determination module is used to determine the information extraction prompt word corresponding to the file parsing result according to the inspection type corresponding to the inspection report file to be processed and the file parsing result; wherein the information extraction prompt word information includes information extraction task instructions, the file parsing result, a preset information extraction task sample and a preset information extraction result output format sample; The information extraction module is used to input the file parsing result and the information extraction prompt word into a pre-trained multimodal information extraction model to obtain the information extraction result corresponding to the inspection report file to be processed.

8. A computer device, characterized in that: The computer device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the information extraction method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the information extraction method as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the information extraction method according to any one of claims 1 to 6.