Structured analysis method and device for physical examination report, equipment and storage medium

By converting physical examination reports into image sets and using a multimodal large language model for entity recognition and relation extraction, the problem of differences in report formats among different physical examination institutions is solved, achieving efficient structured parsing and data consistency.

CN120877322APending Publication Date: 2025-10-31HANGZHOU WANGDAO HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511049572.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Differences in report formats among different medical examination institutions make it difficult to summarize and analyze examination results across institutions and time periods, affecting the efficiency of health record establishment and intelligent diagnosis.

Method used

The medical examination report is converted into an image set. The text blocks and their coordinate positions of each image are extracted. A multimodal large language model is used to fuse the text and layout information to perform entity recognition and relation extraction, generating structured data.

Benefits of technology

It improves the accuracy and versatility of structured parsing of physical examination reports, can adapt to reports of different formats, and supports health record keeping and intelligent diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877322A_ABST
    Figure CN120877322A_ABST
Patent Text Reader

Abstract

The invention discloses a structural analysis method and device for a physical examination report, equipment and a storage medium. The method comprises the following steps: converting an obtained physical examination report into a picture set with a uniform format; respectively extracting a plurality of text blocks corresponding to each picture in the picture set and a picture coordinate position of each text block, and constructing a relationship between a text and a page layout; inputting the picture set, the plurality of text blocks corresponding to each picture and the picture coordinate position of each text block into a trained multi-modal large language model for entity recognition, so that entity words corresponding to each text block can be more accurately determined by comprehensively combining texts, page layout and picture features; and based on the entity word corresponding to each text block, generating the structured data corresponding to the physical examination report, so that the text analysis capability of the physical examination report is enhanced, and the accuracy of a structured analysis result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text processing technology, and in particular to a structured parsing method, apparatus, device, and storage medium for medical examination reports. Background Technology

[0002] As people become more health-conscious, regular health checkups have become a common need. However, the report formats of different medical examination institutions vary significantly, and reports exist in various forms such as paper reports, PDF files, and photos. This diversity makes it difficult to summarize and analyze examination results across institutions and time periods, which seriously hinders the efficient implementation of subsequent health management work such as health record establishment and intelligent diagnosis.

[0003] In existing technologies, entity relationships are usually inferred by recognizing the text in a medical examination report. However, since medical examination reports from different medical examination institutions vary greatly in document layout and style, structural parsing of the text content alone can lead to misjudgments due to text fragments or formatting changes, affecting the accuracy of the structured parsing results. Summary of the Invention

[0004] To address the aforementioned issues, this application provides a structured analysis method, apparatus, device, and storage medium for medical examination reports, with the aim of improving the accuracy of the structured analysis results.

[0005] The embodiments of this application disclose the following technical solutions:

[0006] Firstly, this application provides a structured parsing method for medical examination reports, including:

[0007] The acquired medical examination report is converted into an image set; each page of the medical examination report corresponds to one image in the image set.

[0008] Extract the multiple text blocks corresponding to each image in the image set, as well as the image coordinates of each text block;

[0009] The image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block are input into the trained multimodal large language model for entity recognition to determine the entity words corresponding to each text block.

[0010] Based on the entity words corresponding to each text block, the structured data corresponding to the physical examination report is generated.

[0011] Optionally, as described above, for each image, the multiple text blocks corresponding to the image and the image coordinate positions corresponding to each of the multiple text blocks constitute the text block sequence corresponding to the image;

[0012] The step involves inputting the image set, multiple text blocks corresponding to each image, and the image coordinates of each text block into a trained multimodal large language model for entity recognition, determining the entity words corresponding to each text block, including:

[0013] Using a trained multimodal large language model, feature extraction is performed on each image in the image set to determine the image feature vector corresponding to each image.

[0014] Feature extraction is performed on the text block sequence corresponding to each image to determine the text feature vector corresponding to each text block in the text block sequence.

[0015] The image feature vector and the text feature vector are concatenated, and entity classification prediction is performed using the self-attention mechanism of the multimodal large language model to determine the entity words corresponding to each text block.

[0016] Optionally, in the method described above, generating the structured data corresponding to the physical examination report based on the entity words corresponding to each text block includes:

[0017] By extracting entity relationships from the entity words corresponding to each text block, the entity relationships between the entity words are determined.

[0018] Named entity standardization is performed on each entity word, and structured data corresponding to the physical examination report is generated based on the entity relationships between each entity word.

[0019] Optionally, in the method described above, determining the entity relationships between entity words by extracting entity relationships from the entity words corresponding to each text block includes:

[0020] For each image, the entity words identified from the image are merged to obtain an entity word set, and the entity words in the entity word set are sorted according to the position information of each entity word in the entity word set; the position information of the entity word is the image coordinate position of the text block corresponding to the entity word;

[0021] Based on the row coordinates in the location information of each entity word, entity words with the same row coordinates or whose row coordinate differences are less than a preset threshold are grouped into a group of entity words, thus obtaining multiple entity word groups.

[0022] Based on the multiple entity phrases, the entity relationships between the entity phrases are determined.

[0023] Optionally, in the method described above, the step of performing named entity standardization on each entity word and generating structured data corresponding to the medical examination report based on the entity relationships between each entity word includes:

[0024] Based on a pre-defined standard medical terminology knowledge base, the standard terminology corresponding to each entity word is determined.

[0025] Structured data is generated according to a preset unified format based on the standard terminology corresponding to each entity word and the entity relationships between each entity word.

[0026] Optionally, as described above, the step of extracting multiple text blocks corresponding to each image in the image set and the image coordinates of each text block includes:

[0027] Using a pre-configured text detection model, text detection processing is performed on the image set to identify multiple text block images corresponding to each image in the image set and the image coordinate position corresponding to each text block image.

[0028] Using a pre-configured text recognition model, text recognition processing is performed on multiple text block images corresponding to each image to obtain the text block corresponding to each text block image; the image coordinate position of the text block is the same as the image coordinate position of the text block image.

[0029] Secondly, this application provides a structured analysis device for medical examination reports, comprising:

[0030] The report acquisition module is used to convert the acquired medical examination report into an image set; each page of the medical examination report corresponds to one image in the image set.

[0031] The text recognition module is used to extract multiple text blocks corresponding to each image in the image set, as well as the image coordinates of each text block.

[0032] The entity recognition module is used to input the image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block into the trained multimodal large language model for entity recognition, and to determine the entity word corresponding to each text block.

[0033] The report parsing module is used to generate structured data corresponding to the physical examination report based on the entity words corresponding to each text block.

[0034] Optionally, in the device described above, for each image, a sequence of text blocks corresponding to the image and the image coordinate positions corresponding to each of the multiple text blocks constitute the text block sequence corresponding to the image.

[0035] The entity recognition module includes an image processing unit, a text processing unit, and an entity classification unit;

[0036] The image processing unit is used to perform feature extraction processing on each image in the image set using a trained multimodal large language model, and determine the image feature vector corresponding to each image.

[0037] The text processing unit is used to perform feature extraction processing on the text block sequence corresponding to each image, and determine the text feature vector corresponding to each text block in the text block sequence.

[0038] The entity classification unit is used to concatenate the image feature vector and the text feature vector, and perform entity classification prediction through the self-attention mechanism of the multimodal large language model to determine the entity words corresponding to each text block.

[0039] Thirdly, this application provides an electronic device, the device including: a processor, and a memory communicatively connected to the processor;

[0040] The memory stores instructions that the computer executes;

[0041] The processor executes computer execution instructions stored in memory to implement the structured parsing method for physical examination reports described in any of the above embodiments.

[0042] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the structured parsing method for physical examination reports described in any of the above embodiments.

[0043] Compared with the prior art, this application has the following beneficial effects:

[0044] The method of this application converts the acquired medical examination report into an image set, avoiding the problem of inconsistent file formats. Then, it extracts multiple text blocks corresponding to each image in the image set, along with the image coordinates of each text block, to construct the spatial relationship between the text and the page layout. By inputting the image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block into a trained multimodal large language model for entity recognition, multimodal information fusion is achieved, constructing a complete semantic space of the medical examination report, thereby more accurately determining the entity words corresponding to each text block. Based on the entity words corresponding to each text block, structured data corresponding to the medical examination report is generated, which can accurately adapt to different report formats, improving the universality of structured parsing of medical examination reports and thus improving the accuracy of the structured parsing results. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 A flowchart illustrating a structured parsing method for a medical examination report provided in this application embodiment;

[0047] Figure 2 This is a schematic diagram of the structure of a multimodal large language model provided in an embodiment of this application;

[0048] Figure 3 A schematic diagram illustrating the result obtained after performing named entity recognition on the total examination section of a physical examination report, as provided in this embodiment of the application;

[0049] Figure 4 A schematic diagram illustrating the result obtained after performing named entity recognition on the examination items in a physical examination report, as provided in an embodiment of this application;

[0050] Figure 5 A schematic diagram illustrating the result obtained after performing named entity recognition on the imaging examination section of a physical examination report, as provided in an embodiment of this application;

[0051] Figure 6 A schematic diagram of a structured analysis device for a medical examination report provided in this application embodiment;

[0052] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and accompanying drawings. It should be particularly noted that the embodiments described in this application are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0054] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0055] With increasing health awareness, more and more people are undergoing regular health checkups. Currently, these reports are typically displayed separately through channels provided by the medical examination institutions; that is, the results can be viewed on the institution's app or website after the examination. However, since individuals may undergo checkups at different times and institutions, and the format of reports varies across institutions (paper reports, PDF files, photos, etc.), a key technical challenge is how to aggregate and analyze the results obtained from checkups at different institutions to support subsequent tasks such as health record creation and intelligent diagnosis.

[0056] On the one hand, with the development of mobile internet, users can easily store various forms of medical examination reports on their mobile devices. On the other hand, with the increasing maturity of artificial intelligence technologies such as deep learning, some methods have been developed to use related technologies to structure and parse medical examination reports from different institutions into a unified structure.

[0057] As described earlier, current methods or tools for structuring medical examination reports almost exclusively focus on text-level operations, neglecting crucial document layout and style information. This invention uses a multimodal large language model to jointly model the relationship between text and layout in medical examination reports, while also incorporating visual information from the reports into the model using image features. This method enhances the model's document understanding capabilities and improves the versatility of structured medical examination report parsing.

[0058] Through research, the inventors proposed a structured parsing method, device, equipment, and storage medium for medical examination reports. By using a trained multimodal large language model, the relationship between text and layout in the medical examination report is jointly constructed. Furthermore, image features are used to incorporate the style information of the medical examination report into the model, which can enhance the model's document understanding ability and improve the accuracy of structured parsing of medical examination reports.

[0059] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0060] See Figure 1 The figure is a flowchart illustrating a structured parsing method for a medical examination report provided in an embodiment of this application. Figure 1 As shown, the method includes:

[0061] S101: Convert the obtained medical examination report into an image set.

[0062] Each page of the medical examination report corresponds to one image in its respective image set.

[0063] In this embodiment, the obtained medical examination report is the medical examination report uploaded by the user, and the format of the medical examination report includes, but is not limited to, PDF, JPG image, PNG image, mobile phone screenshot and photo, and the source of the medical examination report is not limited to medical examination institution or hospital.

[0064] Understandably, if the user uploads a photograph of their medical examination report, factors such as the brightness and angle of the photograph may cause discrepancies between the photograph and the original report page, leading to inaccuracies in the recognized text and its position. Therefore, preprocessing adjustments such as rotation, cropping, brightness, or curvature of the photograph can be performed to obtain more accurate text and its corresponding position within the medical examination report in subsequent steps.

[0065] S102: Extract the multiple text blocks corresponding to each image in the image set, as well as the image coordinates of each text block.

[0066] In this embodiment, Optical Character Recognition (OCR) technology is used to convert each image in the image set into a sequence of text blocks with coordinate location information. The text block sequence contains multiple text blocks in the image and the coordinate location of each text block in the image.

[0067] As a feasible approach, the specific steps for extracting multiple text blocks corresponding to each image in the image set, as well as the image coordinates of each text block, may include:

[0068] Using a pre-configured text detection model, text detection processing is performed on the image set to identify multiple text block images corresponding to each image in the image set, as well as the image coordinate position of each text block image.

[0069] In this embodiment, a pre-configured text detection model is used to detect text in each image, identifying the regions where text blocks are located, resulting in multiple text block images. Each text block image carries the coordinates corresponding to its location within the image. Specifically, using the top-left corner of the current image as the origin, the coordinates of the top-left and bottom-right corners of the text block images in the current image can be obtained. These two coordinates can then be used to determine the unique region of the text block image.

[0070] The pre-configured text detection model can adopt the Progressive Scale Expansion Network (PSENet) model. The PSENet model is a scene text detection model that has good detection performance for text of various shapes and has good detection performance for naturally photographed report images that may have some curvature or deformation.

[0071] Using a pre-configured text recognition model, text recognition processing is performed on multiple text block images corresponding to each image to obtain the text block corresponding to each text block image.

[0072] The image coordinates of the text block are the same as the image coordinates of the corresponding text block image.

[0073] In this embodiment, a pre-configured text recognition model is used to perform text recognition processing on each text block image to identify the text content in each text block image, thereby obtaining the text block corresponding to each text block image. Furthermore, relative to each image, since each text block is the text content of the corresponding text block image, the image coordinate position corresponding to that text block image can be used as the image coordinate position of the text block, and multiple text blocks in an image and their corresponding image coordinate positions can be combined to form a sequence of text blocks corresponding to that image.

[0074] The pre-configured text recognition model can be a Convolutional Recurrent Neural Network (CRNN) model. CRNN is an end-to-end text recognition model that performs well with text of variable length. In this embodiment, the pre-configured CRNN model can be fine-tuned on a labeled dataset of medical examination reports to improve the text recognition performance of the reports.

[0075] In this embodiment, a pre-configured text detection model is used to perform text detection processing on the image set, identifying multiple text block images corresponding to each image in the image set and the image coordinate position corresponding to each text block image. This can accurately locate the text region and preserve spatial layout information. Then, a pre-configured text recognition model is used to perform text recognition processing on the multiple text block images corresponding to each image, obtaining the text block corresponding to each text block image. This can provide basic data for entity recognition of subsequent multimodal large language models, enabling the model to perceive the spatial distribution of text and improve the accuracy of medical examination report parsing.

[0076] S103: Input the image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block into the trained multimodal large language model for entity recognition, and determine the entity words corresponding to each text block.

[0077] In this embodiment, the trained multimodal large language model adopts a text-image multimodal Transformer structure. The model employs a 24-layer Transformer encoder with a 16-head self-attention mechanism. Text and images are respectively encoded using a text encoder and an image encoder to obtain text features and image features, which are then input into the multimodal large language model to predict entity words in the medical examination report. For example... Figure 3As shown, the Text Encoder is a text encoder, the Image Encoder is an image encoder, and the Transformer Encoder is an encoder. Specifically, for the input text portion, multiple text blocks corresponding to each image obtained using OCR technology, along with the image coordinates of each text block, are input into the text encoder of the multimodal large language model. The embedding layer includes word embedding, 1D position embedding, and 2D position embedding. For the input image portion, the image set is segmented and then input into the image encoder. The embedding layer includes patch embedding, 1D position embedding, and 2D position embedding. The text features output from the embedding layer are then concatenated with the image features and input into the Transformer encoder, thus obtaining the multimodal large language model in this embodiment.

[0078] As one feasible approach, for each image, the multiple text blocks corresponding to the image and the image coordinates of each text block constitute a sequence of text blocks corresponding to the image; the specific implementation steps of S103, "inputting the image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block into the trained multimodal large language model for entity recognition to determine the entity word corresponding to each text block," may include:

[0079] S1031: Using the trained multimodal large language model, feature extraction is performed on each image in the image set to determine the image feature vector corresponding to each image.

[0080] In this embodiment, for each image, the current image is divided into blocks, and then the trained multimodal large language model is used to obtain image features by linear mapping of the obtained image blocks. Specifically, the image is first scaled to a uniform size, such as 224×224, and then the image is divided into blocks of fixed size, such as 16×16. The image feature sequence and the one-dimensional position vector of the image block are obtained through linear mapping, thus obtaining the image feature vector.

[0081] S1032: Perform feature extraction processing on the text block sequence corresponding to each image to determine the text feature vector corresponding to each text block in the text block sequence.

[0082] In this embodiment, byte-pair encoding is used to segment the text blocks in the text block sequence to obtain word vectors. These word vectors are then combined with the one-dimensional and two-dimensional position vectors of the words determined according to the reading order of the text blocks to obtain the text feature vector corresponding to the text block. Specifically, based on the image coordinates of the text block obtained from the OCR results, the image coordinates are normalized, and the representations of the four embedding sub-layers (x, y, w, h) of the normalized coordinates are calculated. Here, x is the horizontal coordinate of the text block in the corresponding page image, y is the vertical coordinate of the text block in the corresponding page image, w is the width of the text block, and h is the height of the text block. Finally, the representations of the above four embedding sub-layers are added together to obtain the two-dimensional position vector.

[0083] S1033: The image feature vector and the text feature vector are concatenated, and entity classification prediction is performed through the self-attention mechanism of the multimodal large language model to determine the entity words corresponding to each text block.

[0084] In this embodiment, the text feature vector sequence and the image feature vector sequence are concatenated along the token dimension to form a multimodal input sequence containing text and image information. This sequence is then input into the Transformer encoder of the multimodal large language model, where a self-attention mechanism is used to enable the interaction between text block features and corresponding image region features, capturing the spatial association and semantic dependency between text and images. After the Transformer encoder outputs the context enhancement features corresponding to each text block, a fully connected layer is used to combine entity category labels for classification prediction, ultimately determining entity words for each text block.

[0085] For example, entity words are divided into four parts, totaling fifteen entity categories:

[0086] 1) User information section: Name, age, medical examination number, medical examination date;

[0087] 2) Final Inspection Section: Final Inspection;

[0088] 3) Project Inspection Section: Inspection items, inspection indicators, inspection results, units, reference range, and summary;

[0089] 4) Imaging examination section: examination items, examination indicators, examination results, and summary.

[0090] It is understood that the multimodal large language model in this embodiment is trained based on an annotated dataset of medical examination reports. During model training, the training data needs to be as rich as possible to improve the model's generalization and versatility. The training dataset in this embodiment includes approximately 130,000 medical examination reports from 15 medical examination institutions. Since the layout of medical examination reports varies between different institutions, this embodiment uses different rules for batch labeling of reports from different institutions, followed by manual sampling. This method efficiently completes the annotation while ensuring the quality of the data annotation.

[0091] In this embodiment, a trained multimodal large language model is used to extract features from each image in the image set, determining the image feature vector corresponding to each image. Feature extraction is then performed on the text block sequence corresponding to each image, determining the text feature vector corresponding to each text block in the text block sequence. The image feature vector and the text feature vector are concatenated, and entity classification prediction is performed using the self-attention mechanism of the multimodal large language model to determine the entity word corresponding to each text block. This method deeply integrates text, page layout, and image visual information to construct a complete semantic space, thereby improving the accuracy of entity recognition.

[0092] S104: Generate structured data corresponding to the physical examination report based on the entity words corresponding to each text block.

[0093] In this embodiment, based on the entity words corresponding to each text block, the entity words in each page of the report are merged, and the text blocks are sorted according to the image coordinate position in the reading order and grouped by line coordinate. Then, the entity relationships between the entity words are extracted to determine the structured data of the physical examination report.

[0094] In this embodiment, the acquired medical examination report is converted into an image set to avoid the problem of inconsistent file formats. Then, multiple text blocks corresponding to each image in the image set, along with the image coordinates of each text block, are extracted to construct the spatial relationship between the text and the page layout. By inputting the image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block into a trained multimodal large language model for entity recognition, multimodal information fusion is achieved, constructing a complete semantic space for the medical examination report. This allows for more accurate determination of the entity words corresponding to each text block. Based on the entity words corresponding to each text block, structured data corresponding to the medical examination report is generated, which can accurately adapt to different report formats, improving the universality of structured parsing of medical examination reports and thus enhancing the accuracy of the structured parsing results.

[0095] As an achievable approach, the specific implementation steps of "generating structured data corresponding to the physical examination report based on the entity words corresponding to each text block" in S104 include:

[0096] By extracting entity relationships from the entity words corresponding to each text block, the entity relationships between entity words are determined.

[0097] In this embodiment, the entity words identified from the image are merged to obtain an entity word set. The entity words in the entity word set are sorted according to the position information of each entity word. Then, based on the row coordinates in the position information of each entity word, entity words with the same row coordinates or a row coordinate difference less than a preset threshold are grouped into an entity word group, resulting in multiple entity word groups.

[0098] Named entity standardization is performed on each entity term, and structured data corresponding to the physical examination report is generated based on the entity relationships between each entity term.

[0099] In this embodiment, based on a preset standard medical terminology knowledge base, the standard term corresponding to each entity word is determined; then, based on the standard term corresponding to each entity word and the entity relationship between each entity word, structured data is generated in a preset unified form.

[0100] In this embodiment, by extracting entity relationships from the entity words corresponding to each text block, the entity relationships between entity words are determined, and the logical relationships between text blocks can be accurately identified. Then, named entity standardization is performed on each entity word to unify medical terminology, and structured data corresponding to the physical examination report is generated based on the entity relationships between each entity word. This achieves universal parsing of different physical examination reports and improves the consistency of physical examination report data.

[0101] As one possible approach, the specific steps for determining the entity relationships between entity words by extracting entity relationships from the entity words corresponding to each text block can include:

[0102] For each image, the entity words identified from the image are merged to obtain an entity word set. Then, the entity words in the entity word set are sorted according to the position information of each entity word in the entity word set. The position information of the entity word is the image coordinate position of the text block corresponding to the entity word.

[0103] In this embodiment, since the entity recognition process is performed page by page according to the image corresponding to each page of the medical examination report, while the entity word relationship extraction in this step needs to be based on the complete medical examination report, after obtaining the entity words in the image corresponding to each page of the report, it is necessary to merge the entity words from different pages. Specifically, when merging pages, the image corresponding to the first page is used as the reference, and the coordinates of the images on other pages are scaled according to their length and width, thereby scaling the coordinates of the text blocks. Finally, the entity words of each page are concatenated, keeping the x-axis coordinate of the text block on each page unchanged, while the y-axis coordinate needs to be recalculated with the top of the first page as the origin.

[0104] Based on the row coordinates in the location information of each entity word, entity words with the same row coordinates or whose row coordinate differences are less than a preset threshold are grouped into a single entity word group, resulting in multiple entity word groups.

[0105] In this embodiment, the row coordinate information of each entity word in the image is first extracted. The row position is usually represented by the y-value of the text block coordinates or the vertical center point coordinates. All entity words are arranged in ascending order of row coordinates to ensure the processing order. A preset row spacing threshold is used as the standard to determine whether they belong to the same row or adjacent rows. For example, it is set to 15 pixels based on the document font size and the empirical value of the row height. An entity word group list is initialized. The sorted entity words are traversed. For the current entity word, if the entity word group list is empty, a new group is created and the entity word is added. If the list is not empty, the difference between the row coordinate of the current entity word and the row coordinate of the first entity word in the last entity word group is calculated. If the difference is 0 or less than the preset threshold, it is added to the group. Otherwise, a new group is created. After the traversal is completed, all entity words that meet the row coordinate association conditions are divided into corresponding word groups to form an entity word group set based on spatial proximity.

[0106] Based on multiple entity phrases, the entity relationships between entity words are determined.

[0107] In this embodiment, common entity relationship types in the medical field are first defined, such as "test item-result value", "indicator-reference range", and "value-unit". Preset relationship determination rules are also defined based on the layout features of the physical examination report. For example, within the same entity phrase, the left-order entity is often of the "test item" category, and the right-order entity is often of the "result" category. For each entity phrase, the entity words are arranged in ascending order of their column coordinates in the image to restore the text reading order, and the named entity category of each entity word is extracted. For example, "test indicator", "value result", "reference range", and "unit". Based on the preset rules and entity category combination, pairwise relationship matching is performed on the entity words within the phrase. For example, if there are "test indicator" and "value result" entities in the phrase, and the former's x-coordinate is smaller than the latter's, then a "corresponding result" relationship is determined between them. Furthermore, domain knowledge can be combined to perform pairwise relationship matching on the entity words within the phrase. For example, when the right side of a "value result" entity is immediately adjacent to a "unit" entity, a "attribute association" relationship is determined between them, thus achieving structured extraction of entity relationships.

[0108] In this embodiment, for each image, the entity words identified from the image are merged to obtain an entity word set. Then, based on the position information of each entity word in the entity word set, the entity words in the entity word set are sorted. The position information of the entity word is the image coordinate position of the text block corresponding to the entity word. Then, based on the row coordinates in the position information of each entity word, entity words with the same row coordinates or a row coordinate difference less than a preset threshold are grouped into an entity word group, resulting in multiple entity word groups. Finally, based on the multiple entity word groups, the entity relationship between the entity words is determined, which improves the accuracy of entity relationship extraction.

[0109] For example, entity relation extraction in the final inspection section, such as... Figure 3 As shown, since the final inspection section contains only one entity, namely the "final inspection" entity, all text blocks corresponding to the blue boxes in the diagram are "final inspection" entities. Therefore, by concatenating all "final inspection" entities in the reading order, the structured data of the final inspection section can be obtained.

[0110] For entity relationship extraction in the project inspection section, such as Figure 4 As shown, the inspection item names are generally located in the table header, and the location of each table can be identified by the inspection item names. Tables typically contain multiple indicators and their corresponding inspection results, units, and ranges. The inspection indicators and their corresponding inspection results, units, and ranges are usually located in the same row. Whether the entities are in the same row can be used to determine if there is a correspondence between the inspection indicators, inspection results, units, and ranges. For example, in... Figure 7The inspection items are classified as "General Inspection," which includes seven indicators such as height, weight, and body mass index (BMI), seven results such as 176.5, 71.7, and 22.8, seven units such as cm and kg, and four reference ranges such as 18.5-22.3 and 90-139. This step requires matching the indicators, results, units, and reference ranges one by one. For the result 176.5, we can determine that its distance from the target height on the y-axis is the smallest; for the unit cm, we can determine that its distance from the target height on the y-axis is the smallest. For each result, unit, and reference range, we calculate the closest indicator, and finally obtain multiple relationships such as "height" - "176.5" - "cm" - "", "systolic blood pressure" - "123" - "mmHg" - "90-139". Finally, we merge the entity words "inspection items" and "summary" to obtain the structured data of the inspection items.

[0111] For entity relation extraction in the imaging examination section, such as Figure 5 As shown, the "Inspection Items" are generally located at the top, such as... Figure 5 As shown in the blue box; the "Summary" is usually located at the end, such as... Figure 5 The purple box in the middle shows the "Inspection Indicators" and "Inspection Results". The "Inspection Indicators" are as follows: Figure 5 As shown in the green box, the "Inspection Results" are as follows: Figure 5 As shown in the red boxes, the three red boxes represent "Inspection Result 1", "Inspection Result 2", and "Inspection Result 3" respectively. In this embodiment, the inspection indicators and results are first arranged in reading order, and then the inspection result is matched with its preceding inspection indicator. For example, after arranging them in reading order, we get: "Inspection Indicator 1", "Inspection Result 1", "Inspection Result 2", and "Inspection Result 3". Matching each inspection result with its preceding inspection indicator gives the relationship: "Inspection Indicator 1" - "Inspection Result 1 + Inspection Result 2 + Inspection Result 3". Finally, by merging the entity words "Inspection Item" and "Summary", the structured data of the image inspection section can be obtained.

[0112] As a possible approach, different institutions / hospitals may present different names for the same examination item or indicator in their medical examination reports. For example, "Cancer Antigen 50 (CA50)," "Cancer Tumor Antigen 50," "Glycan Antigen 50 Measurement," and "Carbohydrate Antigen 50 Measurement (CA50)" all refer to the same indicator. Different names for the same indicator can cause problems for downstream work in structuring medical examination reports. For instance, in health record creation, if the indicator name in the medical examination report does not match the indicator name in the database, it can lead to a significant amount of manual work and even prevent the health record creation process from proceeding. Therefore, to improve the universality of this method, the specific implementation steps for generating structured data corresponding to the medical examination report based on the entity relationships between each entity term, including:

[0113] Based on a pre-defined standard medical terminology knowledge base, the standard terminology corresponding to each entity word is determined.

[0114] Structured data is generated according to a pre-defined unified format based on the standard terminology corresponding to each entity word and the entity relationships between each entity word.

[0115] In this embodiment, based on the constructed standard medical terminology knowledge base, the indicator entity words in the physical examination report and the indicator names in the standard medical terminology knowledge base are input into the classification model to obtain the similarity score between the two indicator names; then, the indicator entity words are mapped to the standard terms with the highest similarity scores. The classification model in this embodiment can adopt a bidirectional encoder representation from transformers (BERT) network structure.

[0116] In this embodiment, based on a preset standard medical terminology knowledge base, the standard terminology corresponding to each entity word is determined; based on the standard terminology corresponding to each entity word and the entity relationships between each entity word, structured data is generated in a preset unified form, which can improve the standardization and uniformity of the physical examination report parsing results.

[0117] See Figure 6 The figure is a schematic diagram of a structured analysis device for a medical examination report provided in an embodiment of this application. Figure 6 As shown, the device 20 includes a report acquisition module 21, a text recognition module 22, an entity recognition module 23, and a report parsing module 24.

[0118] The report acquisition module 21 converts the acquired medical examination report into an image set; each page of the report corresponds to one image in the image set. The text recognition module 22 extracts multiple text blocks corresponding to each image in the image set, along with the image coordinates of each text block. The entity recognition module 23 inputs the image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block into a trained multimodal large language model for entity recognition, determining the entity words corresponding to each text block. The report parsing module 24 generates structured data corresponding to the medical examination report based on the entity words corresponding to each text block.

[0119] The structured analysis device for physical examination reports provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0120] Furthermore, based on the above embodiments, for each image, the multiple text blocks corresponding to the image and the image coordinate positions corresponding to each text block constitute the text block sequence corresponding to the image; the entity recognition module 23 includes an image processing unit, a text processing unit, and an entity classification unit.

[0121] Specifically, the image processing unit uses a trained multimodal large language model to extract features from each image in the image set, determining the image feature vector corresponding to each image; the text processing unit extracts features from the text block sequence corresponding to each image, determining the text feature vector corresponding to each text block in the text block sequence; and the entity classification unit concatenates the image feature vector and the text feature vector, and performs entity classification prediction through the self-attention mechanism of the multimodal large language model, determining the entity word corresponding to each text block.

[0122] The structured analysis device for physical examination reports provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0123] Furthermore, based on the above embodiments, when the entity recognition module 23 performs feature extraction processing on the text block sequence corresponding to each image and determines the text feature vector corresponding to each text block in the text block sequence, the entity recognition module 23 is specifically used to generate a one-dimensional position vector corresponding to each text block based on the recognition order of the text blocks in the text block sequence; determine a two-dimensional position vector corresponding to each text block based on the image coordinate position corresponding to each text block in the text block sequence; perform byte-by-byte encoding word segmentation processing on the text block sequence to generate a word vector corresponding to each text block; and for each text block, add the word vector, the one-dimensional position vector, and the two-dimensional position vector of the text block to obtain the text feature vector corresponding to the text block.

[0124] The structured analysis device for physical examination reports provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0125] Furthermore, based on the above embodiments, the report parsing module 24 is specifically used to extract entity relationships from the entity words corresponding to each text block to determine the entity relationships between entity words; to perform named entity standardization on each entity word; and to generate structured data corresponding to the physical examination report based on the entity relationships between each entity word.

[0126] The structured analysis device for physical examination reports provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0127] Furthermore, based on the above embodiments, when the report parsing module 24 is used to extract entity relationships from the entity words corresponding to each text block to determine the entity relationships between entity words, the report parsing module 24 is used to merge the entity words identified from the image for each image to obtain an entity word set, and sort the entity words in the entity word set according to the position information of each entity word in the entity word set; the position information of the entity word is the image coordinate position of the text block corresponding to the entity word; based on the row coordinates in the position information of each entity word, entity words with the same row coordinates or a row coordinate difference less than a preset threshold are grouped into an entity word group to obtain multiple entity word groups; and the entity relationships between entity words are determined based on the multiple entity word groups.

[0128] The structured analysis device for physical examination reports provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0129] Furthermore, based on the above embodiments, when the report parsing module 24 is used to perform named entity standardization on each entity word and generate structured data corresponding to the physical examination report according to the entity relationship between each entity word, the report parsing module 24 is specifically used to determine the standard terminology corresponding to each entity word based on the preset standard medical terminology knowledge base; and generate structured data in a preset unified form based on the standard terminology corresponding to each entity word and the entity relationship between each entity word.

[0130] The structured analysis device for physical examination reports provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0131] Furthermore, based on the above embodiments, the text recognition module 22 is specifically used to perform text detection processing on the image set using a pre-configured text detection model, to identify multiple text block images corresponding to each image in the image set and the image coordinate position corresponding to each text block image; using the pre-configured text recognition model, text recognition processing is performed on the multiple text block images corresponding to each image to obtain the text block corresponding to each text block image; the image coordinate position of the text block is the same as the image coordinate position corresponding to the text block image.

[0132] The structured analysis device for physical examination reports provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0133] See Figure 7 The figure is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, including:

[0134] Memory 11 is used to store computer programs;

[0135] The processor 12 is used to implement the steps of the structured parsing method for a physical examination report as described in any of the above method embodiments when executing the computer program.

[0136] In this embodiment, the device can be an in-vehicle computer, a PC (Personal Computer), or a terminal device such as a smartphone, tablet computer, handheld computer, or portable computer.

[0137] The device may include a memory 11, a processor 12, and a bus 13.

[0138] The memory 11 includes at least one type of readable storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the device, such as the hard disk of the device. In other embodiments, the memory 11 may be an external storage device of the device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, Flash Card, etc. Furthermore, the memory 11 may include both internal and external storage units of the device. The memory 11 can be used not only to store application software and various types of data installed on the device, such as program code for executing structured parsing methods for medical examination reports, but also to temporarily store data that has been output or will be output. In some embodiments, the processor 12 may be a central processing unit (CPU).

[0139] In some embodiments, processor 12 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip, used to run program code stored in memory 11 or process data, such as program code for executing a structured parsing method for a medical examination report.

[0140] This bus 13 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0141] Furthermore, the device may also include a network interface 14, which may optionally include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), typically used to establish communication connections between the device and other electronic devices.

[0142] Optionally, the device may further include a user interface 15, which may include a display, an input unit such as a keyboard, and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the device and to display a visual user interface.

[0143] Figure 7 Only devices with components 11-15 are shown; those skilled in the art will understand that... Figure 7 The structure shown does not constitute a limitation on the device and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0144] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a computer-readable storage medium storing computer instructions for causing the computer to execute the structured parsing method for physical examination reports as described in any of the above embodiments.

[0145] The computer-readable media in this application embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0146] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the structured parsing method of the physical examination report as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0147] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for methods, apparatuses, electronic devices, and media, since they are basically similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The methods, apparatuses, electronic devices, and media described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0148] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A structured parsing method for physical examination reports, characterized in that, include: The acquired medical examination report is converted into an image set; each page of the medical examination report corresponds to one image in the image set. Extract the multiple text blocks corresponding to each image in the image set, as well as the image coordinates of each text block; The image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block are input into the trained multimodal large language model for entity recognition to determine the entity words corresponding to each text block. Based on the entity words corresponding to each text block, the structured data corresponding to the physical examination report is generated.

2. The method according to claim 1, characterized in that, For each image, the multiple text blocks corresponding to the image and the image coordinate positions corresponding to each of the multiple text blocks constitute the text block sequence corresponding to the image; The step involves inputting the image set, multiple text blocks corresponding to each image, and the image coordinates of each text block into a trained multimodal large language model for entity recognition, determining the entity words corresponding to each text block, including: Using a trained multimodal large language model, feature extraction is performed on each image in the image set to determine the image feature vector corresponding to each image. Feature extraction is performed on the text block sequence corresponding to each image to determine the text feature vector corresponding to each text block in the text block sequence. The image feature vector and the text feature vector are concatenated, and entity classification prediction is performed using the self-attention mechanism of the multimodal large language model to determine the entity words corresponding to each text block.

3. The method according to any one of claims 1-2, characterized in that, The step of generating structured data corresponding to the physical examination report based on the entity words corresponding to each text block includes: By extracting entity relationships from the entity words corresponding to each text block, the entity relationships between the entity words are determined. Named entity standardization is performed on each entity word, and structured data corresponding to the physical examination report is generated based on the entity relationships between each entity word.

4. The method according to claim 3, characterized in that, The step of extracting entity relationships from the entity words corresponding to each text block to determine the entity relationships between them includes: For each image, the entity words identified from the image are merged to obtain an entity word set, and the entity words in the entity word set are sorted according to the position information of each entity word in the entity word set; the position information of the entity word is the image coordinate position of the text block corresponding to the entity word; Based on the row coordinates in the location information of each entity word, entity words with the same row coordinates or whose row coordinate differences are less than a preset threshold are grouped into a group of entity words, thus obtaining multiple entity word groups. Based on the multiple entity phrases, the entity relationships between the entity phrases are determined.

5. The method according to claim 3, characterized in that, The step of performing named entity standardization on each entity word and generating structured data corresponding to the medical examination report based on the entity relationships between each entity word includes: Based on a pre-defined standard medical terminology knowledge base, the standard terminology corresponding to each entity word is determined. Structured data is generated according to a preset unified format based on the standard terminology corresponding to each entity word and the entity relationships between each entity word.

6. The method according to claim 1, characterized in that, The step of extracting multiple text blocks corresponding to each image in the image set and the image coordinates of each text block includes: Using a pre-configured text detection model, text detection processing is performed on the image set to identify multiple text block images corresponding to each image in the image set and the image coordinate position corresponding to each text block image. Using a pre-configured text recognition model, text recognition processing is performed on multiple text block images corresponding to each image to obtain the text block corresponding to each text block image; the image coordinate position of the text block is the same as the image coordinate position of the text block image.

7. A structured analysis device for physical examination reports, characterized in that, include: The report acquisition module is used to convert the acquired medical examination report into an image set; each page of the medical examination report corresponds to one image in the image set. The text recognition module is used to extract multiple text blocks corresponding to each image in the image set, as well as the image coordinates of each text block. The entity recognition module is used to input the image set, the multiple text blocks corresponding to each image, and the image coordinates of each text block into the trained multimodal large language model for entity recognition, and to determine the entity word corresponding to each text block. The report parsing module is used to generate structured data corresponding to the physical examination report based on the entity words corresponding to each text block.

8. The apparatus according to claim 7, characterized in that, For each image, the multiple text blocks corresponding to the image and the image coordinate positions corresponding to each of the multiple text blocks constitute the text block sequence corresponding to the image; The entity recognition module includes an image processing unit, a text processing unit, and an entity classification unit; The image processing unit is used to perform feature extraction processing on each image in the image set using a trained multimodal large language model, and determine the image feature vector corresponding to each image. The text processing unit is used to perform feature extraction processing on the text block sequence corresponding to each image, and determine the text feature vector corresponding to each text block in the text block sequence. The entity classification unit is used to concatenate the image feature vector and the text feature vector, and perform entity classification prediction through the self-attention mechanism of the multimodal large language model to determine the entity words corresponding to each text block.

9. An electronic device, characterized in that, The device includes: a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 6.