Image Document Structured Analysis Method, System, Electronic Device, Storage Medium

By splicing, covering and integrating image documents, and using the LX-BioLayoutLM model for structured analysis, the problem of difficulty in accurately identifying and structured extraction of image documents in the existing technology is solved, and efficient structured analysis of image documents is achieved.

CN114708595BActive Publication Date: 2025-06-20LINGXI QUANTUM (BEIJING) MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210255581.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-06-20
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and structure the information of various parts of the image literature, and it is especially impossible to effectively process non-text content such as charts.

Method used

By splicing each page of the image document into a synthetic diagram, occluding and integrating the layout, then using the LX-BioLayoutLM model to structure the analytical diagram to generate a parsed document with structured labels.

Benefits of technology

The batch structured analysis of image documents is realized, and various parts of information in image documents can be effectively extracted, including text and non-text content, improving the structured extraction accuracy and efficiency of document data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708595B_ABST
    Figure CN114708595B_ABST
Patent Text Reader

Abstract

The present invention provides an image document structured parsing method, system, electronic device, and storage medium. The method includes: splicing each page image in the image document in sequence to obtain a composite image; performing a covering process on the corresponding part in the composite image based on the predetermined text content information to be excluded; performing layout integration on the composite image to obtain a to-be-parsed image; inputting the to-be-parsed image into the LX-BioLayoutLM model to obtain a parsed document with structured tags; wherein the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and text information in the to-be-parsed image. The present invention realizes batch structured parsing of image documents, facilitating structured extraction of document data in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document processing, and in particular, to an image document structured parsing method, system, electronic device, and storage medium. Background Art

[0002] There is a wide demand for the extraction and parsing of image text content in many actual production scenarios. With the booming development of related algorithm technologies for computer vision tasks based on artificial intelligence, using AI technology to parse image text content is a very valuable and meaningful direction.

[0003] An image document usually contains several parts of information such as a title, keywords, abstract, text, charts, references, etc. How to accurately identify and extract each part of the information in the document during the process of parsing an image document, that is, the page layout analysis of the image document, is a very important research topic.

[0004] Currently, there are various different ways to perform text parsing on image documents. For example, the built-in toolkit of the Python language can be used to directly parse and output the text in the image document. However, this parsing method only reads the text in the image and cannot classify and identify each paragraph of text according to category attributes, and this type of method can only be applied to text content and cannot solve non-text content such as charts. There are also some image parsing tools that split and identify the content of an image by uploading an image file. This method is often only effective for specific types of images and cannot achieve structured output of the image content either.

[0005] Therefore, how to structurally extract the content in an image document according to a certain fixed rule has become a solution worthy of exploration. Summary of the Invention

[0006] In view of the problems existing in the prior art, the present invention provides an image document structured parsing method, system, electronic device, and storage medium.

[0007] An image document structured parsing method provided by the present invention includes:

[0008] Sequentially splicing the images of each page in the image document to obtain a composite image;

[0009] Based on the predetermined text content information to be excluded, covering the corresponding part in the composite image;

[0010] Integrating the layout of the composite image to obtain an image to be parsed;

[0011] Input the graph to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured tags;

[0012] Among them, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of image information and text information in the graph to be parsed.

[0013] According to an image document structured parsing method provided by the present invention, after inputting the graph to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured tags, it includes:

[0014] Compare the length of the text in the parsed document with the length of the text in the graph to be parsed to obtain the integrity of text structuring;

[0015] If the integrity exceeds a predetermined threshold, confirm the parsed document as the document after structured parsing of the image document.

[0016] According to an image document structured parsing method provided by the present invention, after the step of, if the integrity exceeds a predetermined threshold, confirm the parsed document as the document after structured parsing of the image document, it includes:

[0017] Based on the structured tags in the parsed document, select several required tags and batch extract the paragraph text in the image document corresponding to the tags.

[0018] According to an image document structured parsing method provided by the present invention, the training data set of the LX-BioLayoutLM model is structured labeled image documents.

[0019] According to an image document structured parsing method provided by the present invention, in the BERT model part of the LX-BioLayoutLM model, the text in the graph to be parsed and the corresponding position information are used as inputs, and the text vector reflecting text semantic understanding and the position embedding vector representing the mapping relationship between the text paragraph and the image are used as outputs.

[0020] According to an image document structured parsing method provided by the present invention, in the LayoutLM model part of the LX-BioLayoutLM model, the graph to be parsed, the text in the graph to be parsed and the corresponding position information are used as inputs, and the character-level 2D position embedding vector and the image embedding vector reflecting the image feature information are used as outputs.

[0021] An image document structured parsing method provided by the present invention, wherein the LX-BioLayoutLM model includes an image alignment layer, and the image alignment layer takes the graph to be parsed, a text vector reflecting text semantic understanding, a position embedding vector representing the mapping relationship between the text paragraph and the image, a character-level 2D position embedding vector, and an image embedding vector reflecting image feature information as inputs, and outputs a parsed document with structured tags.

[0022] The present invention also provides an image document structured parsing system, and the system includes:

[0023] An image synthesis module, which splices the images of each page in the image document in sequence to obtain a synthesized graph;

[0024] A custom information removal module, which performs a covering process on the corresponding part in the synthesized graph based on the predetermined text content information to be removed;

[0025] An image layout reset module, which performs layout integration on the synthesized graph to obtain a graph to be parsed;

[0026] A parsing module, which inputs the graph to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured tags;

[0027] Among them, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and text information in the graph to be parsed.

[0028] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the image document structured parsing method described in any one of the above are implemented.

[0029] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image document structured parsing method described in any one of the above are implemented.

[0030] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the image document structured parsing method described in any one of the above are implemented.

[0031] The image document structured parsing method, system, electronic device, and storage medium provided by the present invention realize batch structured parsing of image documents, facilitating structured extraction of document data in complex scenarios. Description of the Drawings

[0032] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0033] Figure 1 It is a schematic flowchart of a method for structured analysis of image documents provided by the present invention;

[0034] Figure 2 It is a schematic structural diagram of a system for structured analysis of image documents also provided by the present invention;

[0035] Figure 3 It is a schematic physical structure diagram of an electronic device provided by the present invention. Detailed implementation manners

[0036] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0037] The following will, in conjunction with the drawings, through specific embodiments and their application scenarios, provide a detailed description of the method for structured analysis of image documents provided in the embodiments of the present application.

[0038] Figure 1 It is a schematic flowchart of a method for structured analysis of image documents provided by the present invention. As Figure 1 shown, a method for structured analysis of image documents provided by the present invention includes:

[0039] Step 100: Stitch the images of each page in the image document in sequence to obtain a composite image.

[0040] It should be noted that the image document refers to a PDF document in the medical field in this embodiment.

[0041] Preferably, by stitching the images of each page of the image document in sequence to form a large image, this process is jointly completed by combining the built-in PIL package and cv2 package of the python language with the image stitching and synthesis logic.

[0042] Step 200: Based on the predetermined text content information to be excluded, perform a covering process on the corresponding part in the composite image.

[0043] During the process of image content extraction, it is often the case that image documents in the same batch may have some identical non-essential text content information, and there is a risk that this text content information may interfere with subsequent recognition. Therefore, by providing the non-essential text content information that can be ignored in the current batch of image documents, the relevant content of the synthesized image is automatically covered; when it is not necessary to eliminate it, the non-essential text content information is set to a null value.

[0044] Step 300: Integrate the layout of the synthesized image to obtain the image to be parsed.

[0045] After the above-mentioned synthesis and covering processes are performed on the input image, redundant layout structures may be generated. By cropping the redundant information structures, they are integrated into a picture form that is closer to the training data of the subsequent model.

[0046] Step 400: Input the image to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured labels;

[0047] Among them, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and text information in the image to be parsed.

[0048] Preferably, the LX-BioLayoutLM model is a simple and effective pre-trained model trained for medical image understanding tasks. It can be mainly divided into two parts: the BERT model and the LayoutLM model. Through the BERT model, the semantics of the image document are understood, and through the LayoutLM model, information such as the visual features and relative positions of the text in the image document are captured to solve problems at the visual information level.

[0049] This embodiment realizes batch structured parsing of image documents, facilitating structured extraction of document data in complex scenarios.

[0050] Furthermore, in another embodiment, this embodiment provides a method for structured parsing of image documents. After inputting the image to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured labels, it includes:

[0051] Compare the length of the text in the parsed document with the length of the text in the image to be parsed to obtain the integrity of text structuring;

[0052] If the integrity exceeds a predetermined threshold, confirm the parsed document as the document after structured parsing of the image document.

[0053] It should be noted that the LX-BioLayoutLM model is usually used to batch process multiple image documents. The LX-BioLayoutLM model finally generates files with multiple class html code segments, and each file corresponds to an original image document.

[0054] Based on the length of the image text input into the LX-BioLayoutLM model and the length of the text in the final output file, the integrity of the recognition result is judged. Based on experimental data, it is found that it is more appropriate to set the threshold to 75%.

[0055] This embodiment discloses how to determine whether the parsing of the image document meets the predetermined requirements. For the parsing that fails to meet the standard, the structured parsing text generated by the parsing and the original image document are returned simultaneously.

[0056] Furthermore, in another embodiment, this embodiment provides a method for structured parsing of image documents. If the integrity exceeds the predetermined threshold, the parsed document is confirmed as the document after the structured parsing of the image document, and then it includes:

[0057] Based on the structured tags in the parsed document, several required tags are selected, and the paragraph text in the image document corresponding to the tags is batch-extracted.

[0058] It should be noted that the general task is that a batch of image documents are specified to extract the content of specific modules (such as the title and keywords) therein. The parsed document is obtained through the model, the structured tags corresponding to the specific modules are found, and then the content corresponding to the tags in the parsed document is extracted to complete the task.

[0059] This embodiment realizes the extraction and output of information such as the title, abstract, keywords, text, icons, and references in the image document according to the module attributes as required, greatly reducing the difficulty of people's parsing of the image document and improving the accuracy and efficiency of image data processing.

[0060] Furthermore, in another embodiment, this embodiment provides a method for structured parsing of image documents, and the training data set of the LX-BioLayoutLM model is structured-marked image documents.

[0061] Preferably, 8000 medical image documents are structured-marked. By respectively using <title>< / title> 、 <abstract>< / abstract> 、 <k>< / k> 、 <text>< / text> 、 、 <r>< / r> and other class html tags for marking, a training data set is constructed.

[0062] The training data is obtained by manually labeling the biomedical literature database data, and the trained model has better effects on biomedical image literature.

[0063] This embodiment discloses a construction method for the training data set of the LX-BioLayoutLM model.

[0064] Furthermore, in another embodiment, this embodiment provides an image literature structured parsing method. The LayoutLM model part in the LX-BioLayoutLM model takes the to-be-parsed graph, the text in the to-be-parsed graph, and the position information corresponding to the text as inputs, and outputs a character-level 2D position embedding vector and an image embedding vector reflecting the image feature information; the LayoutLM model part in the LX-BioLayoutLM model takes the to-be-parsed graph, the text in the to-be-parsed graph, and the position information corresponding to the text as inputs, and outputs a character-level 2D position embedding vector and an image embedding vector reflecting the image feature information; the LX-BioLayoutLM model includes an image alignment layer. The image alignment layer takes the to-be-parsed graph, the text vector reflecting text semantic understanding, the position embedding vector representing the mapping relationship between the text paragraph and the image, the character-level 2D position embedding vector, and the image embedding vector reflecting the image feature information as inputs, and outputs a parsed document with structured labels.

[0065] It should be noted that the LX-BioLayoutLM model itself has an OCR function. When obtaining the to-be-parsed graph, the text in the graph and the position information corresponding to the text can be obtained accordingly. The 2D position embedding vector is used to represent the relative position markers in the document and capture the relationships between the symbols in the document; the image embedding vector is used to capture some presentation features, such as the direction, type, and color of the characters. In order to align the image feature information and the text information of the document, an image embedding vector layer is added to the model to represent the image features in the language presentation. The alignment of the image feature information and the text information of the document is the process of corresponding the image features such as font, color, etc. with the corresponding text content information.

[0066] LX-BioLayoutLM can be divided into two parts: the BERT model and the LayoutLM model. Among them, the BERT model is a model in the field of natural language processing and is also used here to process text information to solve the problem at the semantic understanding level in this task; the LayoutLM model is used to capture visual features such as the relative position of the document image and the text, and solve the problem at the visual information level. Through the combination of the two parts, the LX-BioLayoutLM model is pre-trained using medical image and text data to realize the layout recognition of the information in the document image, and finally realize the distinction of information such as the title, abstract, text, charts, etc. in the image document, and achieve a structured and labeled output.

[0067] The image alignment layer takes the image to be parsed, the text vector reflecting text semantic understanding, the position embedding vector representing the mapping relationship between the text paragraph and the image, the character-level 2D position embedding vector, and the image embedding vector reflecting the image feature information as inputs. By aligning these four vectors with the image to be parsed itself, it normalizes the semantic understanding of the content and the captured image features, and finally synthesizes and obtains the html structured file.

[0068] The image alignment layer is a convolutional neural network layer. By using the image to be parsed, it converts the image into a matrix composed of pixel values, and then scans the matrix row by row (multiplying the matrix by the matrix corresponding to the scanning area, adding them up, and then taking the average) to obtain a new feature matrix. The new feature matrix is weighted and summed with the matrix formed by splicing the four vectors, and then normalized using an activation function. Finally, the normalized matrix is decoded to obtain the html structured file. In this embodiment, the vectors are obtained through the encoding of the transformer network.

[0069] In addition, when training the LX-BioLayoutLM model, multi-task learning objectives are added, including the loss of the Masked Visual-Language Model (MVLM) and the loss of Multi-label Document Classification (MDC). This will further drive the joint pre-training of text and layout.

[0070] This embodiment discloses the training and application process of the LX-BioLayoutLM model.

[0071] Next, the image document structured parsing system provided by the present invention will be described. The image document structured parsing system described below can be mutually referred to with the image document structured parsing method described above.

[0072] Figure 2 Shown in the figure is a schematic structural diagram of an image document structured parsing system further provided by the present invention. Figure 2 As shown, an image document structured parsing system further provided by the present invention includes:

[0073] An image synthesis module that splices the images on each page of the image document in sequence to obtain a synthesized image;

[0074] A custom information removal module that covers the corresponding part of the synthesized image based on the predetermined text content information to be removed;

[0075] An image layout reset module that integrates the layout of the synthesized image to obtain the image to be parsed;

[0076] The parsing module inputs the graph to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured tags;

[0077] Among them, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and text information in the graph to be parsed.

[0078] This embodiment realizes batch structured parsing of image documents, which is convenient for structured extraction of document data in complex scenarios.

[0079] Through training and fine-tuning, the system accuracy can reach 96.5%. The actual inference result is verified by manual inspection and the accuracy is 93%. Structured parsing of image documents through the system provided by the present invention can relatively perfectly solve the problem that manual data extraction is time-consuming and expensive.

[0080] Figure 3 It is a schematic physical structure diagram of an electronic device provided by the present invention. As Figure 3 shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the method for structured parsing of image documents, and the method includes:

[0081] Stitch the images of each page in the image document in sequence to obtain a composite graph;

[0082] Based on the predetermined text content information to be excluded, cover the corresponding part in the composite graph;

[0083] Integrate the layout of the composite graph to obtain a graph to be parsed;

[0084] Input the graph to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured tags;

[0085] Among them, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and text information in the graph to be parsed.

[0086] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0087] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the image document structured parsing method provided by the above-mentioned various methods. The method includes:

[0088] Stitch the images of each page in the image document in sequence to obtain a composite image;

[0089] Based on the predetermined text content information to be excluded, cover the corresponding part in the composite image;

[0090] Integrate the layout of the composite image to obtain an image to be parsed;

[0091] Input the image to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured tags;

[0092] Among them, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and text information in the image to be parsed.

[0093] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the image document structured parsing method provided by the above-mentioned various methods. The method includes:

[0094] Stitch the images of each page in the image document in sequence to obtain a composite image;

[0095] Based on the predetermined text content information to be excluded, cover the corresponding part in the composite image;

[0096] Perform page layout integration on the composite image to obtain the image to be parsed;

[0097] Input the image to be parsed into the LX - BioLayoutLM model to obtain a parsed document with structured tags;

[0098] Among them, the LX - BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of image information and text information in the image to be parsed.

[0099] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0100] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general - purpose hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above - mentioned technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer - readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0101] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. An image document structured analysis method, characterized in that, The method comprises: Sequentially stitch the images on each page of the image document to obtain a composite image; Based on the predetermined text content information to be removed, the corresponding part in the composite image is covered; Integrate the composite image on the page to obtain the image to be analyzed; Input the graph to be parsed into the LX-BioLayoutLM model to obtain a parsed document with structured tags; The LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of image information and text information in the image to be parsed; The LX-BioLayoutLM model refers to a pre-trained model based on medical image understanding task training; The BERT model part in the LX-BioLayoutLM model takes the text in the image to be parsed and the position information corresponding to the text as input, and outputs a text vector reflecting the semantic understanding of the text and a position embedding vector representing the mapping relationship between the text paragraph and the image; The LayoutLM model part in the LX-BioLayoutLM model takes the graph to be parsed, the text in the graph to be parsed and the position information corresponding to the text as input, and takes the character-level 2D position embedding vector and the image embedding vector reflecting the image feature information as output; The LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of image information and text information in the image to be parsed, including: The BERT model is used to understand the semantics of the image document, and the LayoutLM model is used to capture the visual features of the image document and the relative position information of the text.

2. The image document structured analysis method according to claim 1, characterized in that, The graph to be parsed is input into the LX-BioLayoutLM model to obtain a parsed document with structured tags, and then includes: Compare the length of the text in the parsed document with the length of the text in the to-be-parsed graph to obtain the completeness of the text structure; If the completeness exceeds a predetermined threshold, the parsed document is confirmed as a document after structural parsing of the image document.

3. The image document structured analysis method according to claim 2, characterized in that, If the completeness exceeds a predetermined threshold, the parsed document is confirmed as a document after structural parsing of the image document, and then includes: Based on the structured tags in the parsed document, several tags required are selected, and paragraph texts in the image documents corresponding to the tags are extracted in batches.

4. The image document structured analysis method according to claim 1, characterized in that, The training data set of the LX-BioLayoutLM model is structured labeled image documents.

5. The image document structured analysis method according to claim 1, characterized in that, The LX-BioLayoutLM model includes an image alignment layer, which takes the graph to be parsed, a text vector reflecting text semantic understanding, a position embedding vector representing the mapping relationship between a text paragraph and an image, a character-level 2D position embedding vector, and an image embedding vector reflecting image feature information as input, and outputs a parsed document with structured tags.

6. An image document structured analysis system, characterized in that, The system comprises: An image synthesis module, wherein the image synthesis module sequentially splices the images of each page in the image document to obtain a synthetic image; A custom information elimination module, which based on the predetermined text content information to be eliminated, covers the corresponding part in the composite image; An image layout reset module, which integrates the layout of the composite image to obtain an image to be parsed; An analysis module, which inputs the image to be parsed into the LX-BioLayoutLM model to obtain an analysis document with structured labels; Wherein, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and the text information in the image to be parsed; Wherein, the LX-BioLayoutLM model refers to a pre-trained model trained based on medical image understanding tasks; Wherein, for the BERT model part in the LX-BioLayoutLM model, the text in the image to be parsed and the position information corresponding to the text are used as inputs, and the text vector reflecting the text semantic understanding and the position embedding vector representing the mapping relationship between the text paragraph and the image are used as outputs; Wherein, for the LayoutLM model part in the LX-BioLayoutLM model, the image to be parsed, the text in the image to be parsed and the position information corresponding to the text are used as inputs, and the character-level 2D position embedding vector and the image embedding vector reflecting the image feature information are used as outputs; Wherein, the LX-BioLayoutLM model is based on the BERT model and the LayoutLM model to complete the alignment of the image information and the text information in the image to be parsed, including: Completing the understanding of the semantics of the image document through the BERT model, and capturing the visual features and text relative position information of the image document through the LayoutLM model.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the image document structured analysis method described in any one of claims 1-5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the image document structured analysis method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Medical event information extraction method and device, computer equipment and storage medium

    CN113241138A

  • Medical document analysis method and device, equipment and medium

    CN113420116A