Method and apparatus for structurally integrating pathological reports

The method structures pathology reports using multi-modal data processing and knowledge graphs to address the challenges of unstructured text, improving data integration and accuracy in medical research and clinical applications.

CN119920495BActive Publication Date: 2025-07-15SHENZHEN SHENGQIANG TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510401423.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-15
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

Traditional pathological reports are presented in text form, and information extraction is difficult and difficult to integrate and analyze, making it difficult to use medical data efficiently.

Method used

By automatically structuring the pathological data, including standardization, layered annotation and knowledge graph construction of multimodal medical data, pre-trained multi-dimensional relationship extraction network and dual-modal feature fusion recognition model, identifying handwritten fonts and seal covering areas, and constructing structured pathological reports.

Benefits of technology

It has achieved rapid extraction of key information from massive pathological reports, improved the accuracy and efficiency of medical research and clinical decision-making, and improved the utilization value of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920495B_ABST
    Figure CN119920495B_ABST
Patent Text Reader

Abstract

This application proposes a method and device for structurally integrating pathological reports, including the following steps: obtaining multimodal medical data, standardizing each medical term in the medical data to obtain standardized data; annotating the standardized data in a hierarchical annotation manner to obtain annotated data, and using a pre-trained multi-dimensional relation extraction network to extract entities and entity relations from the annotated data to obtain entity relation data; constructing a knowledge graph for the entity relation data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathological report. Through the automatic structural processing of pathological data, this solution can quickly extract key information from a large number of pathological reports, provide a high-quality data foundation for subsequent data analysis and applications, and improve the accuracy and efficiency of medical research and clinical decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a method and device for structurally integrating pathological reports. Background Art

[0002] In today's era, medical informatization is booming at an unprecedented speed, having a profound impact on all aspects of the medical field. As a crucial basis in the disease diagnosis process, the importance of pathological reports is self-evident. It carries the doctor's professional judgment on the patient's pathological condition and is related to many key matters such as the formulation of subsequent treatment plans and the prediction of the disease development trend.

[0003] However, traditionally, most pathological reports have been presented in a simple text form. Such text-based pathological reports have exposed many serious problems in practical applications. First of all, it is extremely difficult to extract information from pathological reports. Pathological reports contain a large amount of complex and key information, such as the patient's basic physical condition indicators, specific characteristics of the lesion site, detailed cell morphology, etc. However, this information is randomly distributed in the text. Extracting specific key information from it is like looking for a needle in a haystack, which requires a lot of time and effort and often fails to ensure the accuracy of extraction.

[0004] Secondly, it is also a major problem that data is difficult to integrate and analyze. Due to the lack of a standardized structure in the text form, there are significant differences in the content organization, terminology usage, etc. of different pathological reports. This leads to numerous obstacles when attempting to summarize and comprehensively analyze the data of many pathological reports. For example, in some reports, the description of the lesion degree is relatively casual, while in others, different reference standards are used, making it very difficult to conduct unified and effective analysis after data integration.

[0005] In the context of the big data era today, the value of data has become increasingly prominent. The medical field urgently needs to make full use of a large amount of medical data to provide strong support for important tasks such as disease research, clinical decision-making support, and medical quality assessment. However, these drawbacks of traditional pathological reports make it difficult to efficiently utilize these valuable medical data. Summary of the Invention

[0006] An embodiment of this application provides a method for structurally integrating pathological reports. Through automatic structural processing of pathological data, it can quickly extract key information from a large number of pathological reports, provide a high-quality data foundation for subsequent data analysis and applications, and improve the accuracy and efficiency of medical research and clinical decision-making.

[0007] In a first aspect, an embodiment of the present application provides a method for structurally integrating a pathological report, and the method includes:

[0008] Obtain multimodal medical data, standardize each medical term in the medical data to obtain standardized data, and disambiguate based on the context semantic relationship of each medical term during the standardization process;

[0009] Annotate the standardized data in a hierarchical annotation manner to obtain annotated data, and use a pre-trained multi-dimensional relationship extraction network to extract entities and entity relationships from the annotated data to obtain entity relationship data. Among them, during the hierarchical annotation process, the diagnostic conclusion layer, anatomical structure layer, pathological feature layer, and molecular marker layer of the standardized data are respectively annotated;

[0010] Construct a knowledge graph for the entity relationship data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathological report, where the entity dimension includes the medical entity information of each patient, the time dimension includes the pathological evolution process of each patient, the space dimension includes the spatial coordinate information of each patient's pathological section in the body, and the evidence dimension includes the credibility weight of the annotated data of each patient.

[0011] In a second aspect, an embodiment of the present application provides a device for structurally integrating a pathological report, including:

[0012] An acquisition module, configured to obtain multimodal medical data, standardize each medical term in the medical data to obtain standardized data, and disambiguate based on the context semantic relationship of each medical term during the standardization process;

[0013] An annotation module, configured to annotate the standardized data in a hierarchical annotation manner to obtain annotated data, and use a pre-trained multi-dimensional relationship extraction network to extract entities and entity relationships from the annotated data to obtain entity relationship data. Among them, during the hierarchical annotation process, the diagnostic conclusion layer, anatomical structure layer, pathological feature layer, and molecular marker layer of the standardized data are respectively annotated;

[0014] An integration module, configured to construct a knowledge graph for the entity relationship data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathological report, where the entity dimension includes the medical entity information of each patient, the time dimension includes the pathological evolution process of each patient, the space dimension includes the spatial coordinate information of each patient's pathological section in the body, and the evidence dimension includes the credibility weight of the annotated data of each patient.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute a method for structurally integrating a pathological report.

[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium. A computer program is stored in the readable storage medium. The computer program includes program codes for controlling a process to execute the process, and the process includes a method for structurally integrating a pathological report.

[0017] The main contributions and innovations of the present invention are as follows:

[0018] By obtaining multimodal medical data and performing standardized processing, using context semantic relationship disambiguation, and combining hierarchical annotation, multi-dimensional relationship extraction, and knowledge graph construction technologies, the embodiment of the present application can obtain a structured pathological report, providing comprehensive and accurate information for medical diagnosis and research. At the same time, the constructed bimodal feature fusion recognition model can effectively identify handwritten fonts and stamp-covered areas, improving the accuracy of data extraction. In addition, the standardized method based on different mapping tables and weight settings achieves more accurate term standardization for disease conditions and drug information. The division of labor and cooperation among the acquisition, annotation, and integration modules make the entire process efficient and orderly, ultimately improving the efficiency, accuracy, and clinical application value of medical data processing.

[0019] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0021] Figure 1 is a flowchart of a method for structurally integrating a pathological report according to an embodiment of the present application;

[0022] Figure 2 is a structural block diagram of a device for structurally integrating a pathological report according to an embodiment of the present application;

[0023] Figure 3 is a schematic hardware structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0025] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0026] Embodiment 1

[0027] An embodiment of the present application provides a method for structurally integrating a pathological report. Through the automatic structural processing of pathological data, key information can be quickly extracted from a large number of pathological reports, providing a high-quality data basis for subsequent data analysis and applications, and improving the accuracy and efficiency of medical research and clinical decision-making. Specifically, referring to Figure 1 , the method includes:

[0028] Obtain multimodal medical data, standardize each medical term in the medical data to obtain standardized data, and disambiguate based on the context semantic relationship of each medical term during the standardization process;

[0029] Annotate the standardized data in a hierarchical annotation manner to obtain annotated data, and use a pre-trained multi-dimensional relationship extraction network to extract entities and entity relationships from the annotated data to obtain entity relationship data. Among them, during the hierarchical annotation process, the diagnosis conclusion layer, anatomical structure layer, pathological feature layer, and molecular marker layer of the standardized data are respectively annotated;

[0030] Construct a knowledge graph for the entity relationship data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathological report, where the entity dimension includes the medical entity information of each patient, the time dimension includes the pathological evolution process of each patient, the space dimension includes the spatial coordinate information of each patient's pathological section in the body, and the evidence dimension includes the credibility weight of the annotated data of each patient.

[0031] In some embodiments, a medical data interface engine is established to obtain multimodal medical data, where the multimodal medical data includes DICOM images, digital pathology slides (WSI), LIS text reports, structured fields of electronic medical records, etc.

[0032] Furthermore, this solution also uses the HL7 FHIR standard for cross-system data mapping to solve the compatibility problem of heterogeneous data sources.

[0033] In some specific embodiments, the multimodal data may include some PDF-scanned versions of medical data, which are often handwritten by doctors and authenticated by stamping. The stamp may cover some fonts, so if only the scanned version of the PDF is used for text recognition, the accuracy will be relatively low. Therefore, this solution constructs a pre-trained dual-modal feature fusion recognition model to recognize the handwritten fonts and the areas covered by stamps in the multimodal medical data. The dual-modal feature fusion recognition model includes a parallel handwritten script recognition network and a semantic reasoning network. The handwritten script recognition network is trained with a large number of handwritten fonts for recognizing handwritten fonts. Prior knowledge of the pathology report structure is embedded in the semantic reasoning network, so as to recognize the areas covered by stamps through context reasoning. A dynamic fusion mechanism is set in the dual-modal feature fusion recognition model. When recognizing handwritten fonts, the weight of the handwritten script recognition network is set to the first weight and the weight of the semantic reasoning network is set to the second weight through the dynamic fusion mechanism; when recognizing the areas covered by stamps, the weight of the semantic reasoning network is set to the first weight and the weight of the handwritten script recognition network is set to the second weight through the dynamic fusion mechanism. The first weight is greater than the second weight.

[0034] Specifically, the handwritten script recognition network extracts stroke-level details through 12 layers of residual convolution, such as the ink connection points in cursive words, the topological structures of doctors' special shorthand symbols, etc., so as to accurately recognize handwritten fonts.

[0035] Specifically, the semantic reasoning network is of the Transformer structure. The prior knowledge introduced in the Transformer structure includes the fixed positions of the diagnosis conclusion area, the standard arrangement order of immunohistochemical indicators, etc.

[0036] Exemplarily, the first weight is 70% and the second weight is 30%. That is to say, when handwritten fonts are detected, the weight of the handwritten script recognition network is set to 70% and the weight of the semantic reasoning network is set to 30% through the dynamic fusion mechanism. When the areas covered by stamps are detected, the weight of the handwritten script recognition network is set to 30% and the weight of the semantic reasoning network is set to 70% through the dynamic fusion mechanism.

[0037] Specifically, the pre-trained bimodal feature fusion recognition model effectively improves the recognition accuracy of handwritten fonts and stamped areas in multimodal medical data, enhances the adaptability of the model to different recognition scenarios, and provides a strong guarantee for the accurate extraction and subsequent analysis and application of medical data.

[0038] In some embodiments, to avoid the low resolution of the PDF scanned version of the pathological report making it difficult to see clearly, this solution constructs a pre-trained super-resolution adversarial network to repair the blurred text.

[0039] Specifically, the generator of the super-resolution adversarial network is a U-Net structure. A medical text edge enhancement module is added to the decoder to repair the confusing features between numbers and English letters; document layout consistency detection is introduced into the discriminator to ensure that the repaired text maintains the paragraph structure and table alignment features of the original report.

[0040] Furthermore, the super-resolution adversarial network performs repair through a resolution adaptive strategy. When the DPI of the input image is lower than the repair threshold, the input image is textually repaired by the super-resolution adversarial network.

[0041] In some embodiments, a basic medical term mapping table and an extended term mapping table are constructed to standardize each medical term in the medical data. Among them, the medical term mapping table is constructed based on international medical term standards, and the extended term mapping table is constructed based on the local diagnosis expression habits of the hospital.

[0042] Furthermore, a dynamically updated term mapping table is additionally constructed to standardize each medical term in the medical data. Among them, an active learning mechanism is used to learn new medical terms in the medical field, and a dynamically updated term mapping table is constructed based on the new medical terms.

[0043] Specifically, the medical field is a field that is constantly developing and innovating. New diseases, new diagnostic techniques, and treatment methods are constantly emerging, followed by a large number of new terms. The dynamic update layer uses an active learning mechanism to capture these new terms in real time. The active learning mechanism uses advanced machine learning algorithms and data mining techniques to automatically identify potential new terms from massive medical literature, clinical medical records, and other data sources. Once new terms are discovered, these terms will be submitted to pathological experts for review. Pathological experts, relying on their professional knowledge and rich experience, strictly check the accuracy, rationality, and applicability of the new terms in the medical field. After being confirmed by pathological experts, the new medical terms can be added to the dynamically updated term mapping table.

[0044] Specifically, the process of standardizing medical terms is to convert the medical expressions representing the same disease in different pathological reports into the same expression in a mapping manner, so as to facilitate subsequent structured arrangement and provide assistance for disease diagnosis.

[0045] In some embodiments, when standardizing each medical term, the character similarity between the medical term and each element in the mapping table is calculated to obtain the similarity calculation result, and the context semantic relationship of each medical term is captured through a context encoder to obtain the context calculation result. The similarity calculation result and the context calculation result are integrated to obtain the mapping object of each medical term to complete the standardization.

[0046] Specifically, when calculating similarity, the special rules for medical abbreviations are optimized. For example, "Ca" is mapped to "cancer" instead of the chemical element calcium.

[0047] Specifically, the context encoder is a bidirectional LSTM structure. The context encoder is used to capture the context features that appear in medical terms. For example, "atypia" specifically refers to ductal epithelial dysplasia in breast pathology, while in digestive pathology it refers to intestinal metaplasia.

[0048] Furthermore, timeliness weighting matching is performed on each element in the extended term mapping table. That is to say, the probability that an element newly added to the extended term mapping table is matched as the mapped object is higher.

[0049] Furthermore, when standardizing the condition information, the weight of the similarity calculation result is set as the third weight, and the weight of the context calculation result is set as the fourth weight; when standardizing the drug information, the weight of the similarity calculation result is set as the fourth weight, and the weight of the context calculation result is set as the third weight, and the third weight is greater than the fourth weight.

[0050] Specifically, the third weight in this solution is 60% and the fourth weight is 40%. Since tumor information has a clear and standardized expression system, the globally used TNM staging system, each stage is represented by a fixed combination of letters and numbers, so character similarity is extremely crucial in matching, which directly determines whether the staging information can be accurately recognized; however, drug information is not only numerous, but may also have different meanings and usage methods in different medical scenarios. For the same drug, in different disease conditions and combined medication regimens, doctors' expressions may vary. For example, "aspirin" may have different combinations and descriptions in different scenarios of cardiovascular disease prevention and antipyretic analgesia. Therefore, different weights are set for the context calculation result and the similarity calculation result when standardizing the condition information and drug information.

[0051] In some specific embodiments, since different medical terms may have different meanings in the medical field, for example, "degree of differentiation" can be used to represent the grading of cancers in different organs, it is necessary to disambiguate it by combining the context information before and after to determine whether to adopt the WHO 2019 digestive system tumor grading standard or the ISUP prostate cancer grading system.

[0052] In some specific embodiments, the diagnosis conclusion layer includes the nature of the disease, the anatomical structure layer includes the specific body part where the lesion occurs, the pathological feature layer includes the microscopic features of the lesion tissue, and the molecular marker layer represents the biological features of the disease.

[0053] Exemplarily, the information marked in the diagnosis conclusion layer is malignant tumor, benign lesion, pre-cancerous lesion, etc., the information marked in the anatomical structure layer is the upper lobe of the left lung, the right anterior lobe of the liver, etc., the information marked in the pathological feature layer is glandular duct structure, mitotic count, etc., and the information marked in the molecular marker layer is HER2(3+), PD-L1(22C3), etc.

[0054] Furthermore, the multi-dimensional relationship extraction network is a hybrid architecture of a graph neural network and a conditional random field. In the multi-dimensional relationship extraction network, basic entity recognition is first performed through the conditional random field, and then the topological relationship between entities is modeled based on the graph neural network.

[0055] In some specific embodiments, the entity dimension includes core nodes such as patients, lesions, and treatments. The time dimension includes the pathological evolution process from biopsy to surgery and then to recurrence. The space dimension includes spatial coordinate information such as a certain pathological section being a section of the liver area. The evidence dimension includes the credibility weights of the labeled data for each patient. For example, in the diagnosis of lung cancer, the weight of the histopathological examination result is relatively high, while the weight of the suspicious shadow found by imaging examination is relatively low. Through the evidence dimension, the reliability of the diagnosis conclusion can be intuitively judged, helping doctors weigh the diagnosis results and make decisions.

[0056] In some embodiments, when the constructed structured pathological report detects "positive vascular cancer thrombus", the system automatically associates the adjuvant chemotherapy regimen recommended by the NCCN guidelines.

[0057] In some embodiments, the constructed structured pathological report of this solution automatically performs contradiction detection. For example, it can detect the contradictory combination of "ER positive" and "triple-negative breast cancer" diagnosis, thereby assisting doctors in making judgments.

[0058] Embodiment 2

[0059] Based on the same concept, referring to Figure 2 , this application also proposes a device for structurally integrating pathological reports, including:

[0060] An acquisition module, configured to acquire multimodal medical data, standardize each medical term in the medical data to obtain standardized data, and perform disambiguation based on the context semantic relationship of each medical term during the standardization process;

[0061] A labeling module, configured to label the standardized data in a hierarchical labeling manner to obtain labeled data, and use a pre-trained multi-dimensional relationship extraction network to extract entities and entity relationships from the labeled data to obtain entity relationship data, wherein, during the hierarchical labeling process, the diagnostic conclusion layer, anatomical structure layer, pathological feature layer, and molecular marker layer of the standardized data are respectively labeled;

[0062] An integration module, configured to construct a knowledge graph for the entity relationship data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathological report, wherein the entity dimension includes the medical entity information of each patient, the time dimension includes the pathological evolution process of each patient, the space dimension includes the spatial coordinate information of the pathological sections of each patient in the body, and the evidence dimension includes the credibility weight of the labeled data of each patient.

[0063] Embodiment III

[0064] This embodiment further provides an electronic device, referring to Figure 3 , including a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0065] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or a specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application.

[0066] Among them, the memory 404 may include a mass memory 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0067] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.

[0068] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the methods for structurally integrating pathological reports in the above embodiments.

[0069] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.

[0070] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0071] The input / output device 408 is used to input or output information. In this embodiment, the input information can be multimodal medical data, etc., and the output information can be a structured pathological report, etc.

[0072] Optionally, in this embodiment, the above processor 402 can be set to execute the following steps through a computer program:

[0073] Obtain multimodal medical data, standardize each medical term in the medical data to obtain standardized data, and disambiguate based on the context semantic relationship of each medical term during the standardization process;

[0074] Annotate the standardized data in a hierarchical annotation manner to obtain annotated data, and use a pre-trained multi-dimensional relationship extraction network to extract entities and entity relationships from the annotated data to obtain entity relationship data. Among them, during the hierarchical annotation process, the diagnostic conclusion layer, anatomical structure layer, pathological feature layer, and molecular marker layer of the standardized data are respectively annotated;

[0075] Construct a knowledge graph for the entity relationship data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathology report, where the entity dimension includes the medical entity information of each patient, the time dimension includes the pathological evolution process of each patient, the space dimension includes the spatial coordinate information of the pathological sections of each patient in the body, and the evidence dimension includes the credibility weight of the annotation data of each patient.

[0076] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and alternative embodiments, and will not be repeated here.

[0077] Generally, various embodiments can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. However, the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller, or other computing devices, or some combination thereof.

[0078] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or implemented by hardware, or implemented by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to execute the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, at this point, it should be noted that any block in the logical flow, as Figure 3 shown, can represent a program step, or an interconnected logical circuit, block, and function, or a combination of program steps and logical circuits, blocks, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.

[0079] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0080] The above embodiments only illustrate several implementation manners of the present application, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for structurally integrating a pathological report, characterized in that, Including the following steps: Obtain multimodal medical data, standardize each medical term in the medical data to obtain standardized data, and disambiguate based on the context semantic relationship of each medical term during the standardization process. Among them, construct a pre-trained bimodal feature fusion recognition model to recognize the handwritten fonts and the areas covered by seals in the multimodal medical data. The bimodal feature fusion recognition model includes a parallel handwritten handwriting recognition network and a semantic reasoning network. The handwritten handwriting recognition network is trained with a large number of handwritten fonts to recognize handwritten fonts. Prior knowledge of the pathological report structure is embedded in the semantic reasoning network, so as to identify the areas covered by seals through context reasoning. And a dynamic fusion mechanism is set in the bimodal feature fusion recognition model. When recognizing handwritten fonts, the weight of the handwritten handwriting recognition network is set to the first weight and the weight of the semantic reasoning network is set to the second weight through the dynamic fusion mechanism; when recognizing the areas covered by seals, the weight of the semantic reasoning network is set to the first weight and the weight of the handwritten handwriting recognition network is set to the second weight through the dynamic fusion mechanism, and the first weight is greater than the second weight; Label the standardized data in a hierarchical annotation manner to obtain annotated data, and use a pre-trained multi-dimensional relation extraction network to extract entities and entity relationships from the annotated data to obtain entity relationship data. Among them, during the hierarchical annotation process, the diagnosis conclusion layer, anatomical structure layer, pathological feature layer, and molecular marker layer of the standardized data are respectively annotated; Construct a knowledge graph for the entity relationship data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathological report. Among them, the entity dimension includes the medical entity information of each patient, the time dimension includes the pathological evolution process of each patient, the space dimension includes the spatial coordinate information of the pathological sections of each patient in the body, and the evidence dimension includes the credibility weight of the annotated data of each patient.

2. The method for structurally integrating a pathological report according to claim 1, wherein Construct a basic medical term mapping table and an extended term mapping table to standardize each medical term in the medical data. Among them, the medical term mapping table is constructed based on the international medical term standard, and the extended term mapping table is constructed based on the local diagnosis expression habits of the hospital.

3. The method for structurally integrating a pathological report according to claim 1, characterized in that When standardizing each medical term, calculate the character similarity between the medical term and the elements in each mapping table to obtain a similarity calculation result, capture the context semantic relationship of each medical term through a context encoder to obtain a context calculation result, and integrate the similarity calculation result and the context calculation result to obtain the mapping object of each medical term to complete the standardization.

4. A method for structurally integrating a pathological report according to claim 3, characterized in that, When standardizing the condition information, set the weight of the similarity calculation result to the third weight and the weight of the context calculation result to the fourth weight; when standardizing the drug information, set the weight of the similarity calculation result to the fourth weight and the weight of the context calculation result to the third weight, and the third weight is greater than the fourth weight.

5. The method for structurally integrating a pathological report according to claim 1, wherein The diagnosis conclusion layer includes the nature of the disease, the anatomical structure layer includes the specific body part where the lesion occurs, the pathological feature layer includes the microscopic features of the diseased tissue, and the molecular marker layer represents the biological features of the disease.

6. An apparatus for structurally integrating a pathological report, characterized in that, Comprising: An acquisition module, configured to acquire multimodal medical data, standardize each medical term in the medical data to obtain standardized data, and disambiguate based on the context semantic relationship of each medical term during the standardization process. Among them, a pre-trained bimodal feature fusion recognition model is constructed to recognize handwritten fonts and seal-covered areas in multimodal medical data. The bimodal feature fusion recognition model includes a parallel handwritten character recognition network and a semantic reasoning network. The handwritten character recognition network is trained with a large number of handwritten fonts for recognizing handwritten fonts. Prior knowledge of the pathological report structure is embedded in the semantic reasoning network, so as to identify the seal-covered area through context reasoning. And a dynamic fusion mechanism is set in the bimodal feature fusion recognition model. When recognizing handwritten fonts, the weight of the handwritten character recognition network is set to the first weight, and the weight of the semantic reasoning network is set to the second weight through the dynamic fusion mechanism; when recognizing the seal-covered area, the weight of the semantic reasoning network is set to the first weight, and the weight of the handwritten character recognition network is set to the second weight through the dynamic fusion mechanism, and the first weight is greater than the second weight; A labeling module, configured to label the standardized data in a hierarchical labeling manner to obtain labeled data, and use a pre-trained multi-dimensional relationship extraction network to extract entities and entity relationships from the labeled data to obtain entity relationship data. Among them, during the hierarchical labeling process, the diagnosis conclusion layer, anatomical structure layer, pathological feature layer, and molecular marker layer of the standardized data are respectively labeled; An integration module, configured to construct a knowledge graph for the entity relationship data in the entity dimension, time dimension, space dimension, and evidence dimension to obtain a structured pathological report. Among them, the entity dimension includes the medical entity information of each patient, the time dimension includes the pathological evolution process of each patient, the space dimension includes the spatial coordinate information of the pathological section of each patient in the body, and the evidence dimension includes the credibility weight of the labeled data of each patient.

7. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute a method for structurally integrating a pathological report according to any one of claims 1-5.

8. A readable storage medium, characterized in that, A computer program is stored in the readable storage medium, and the computer program includes program codes for controlling a process to execute the process, and the process includes a method for structurally integrating a pathological report according to any one of claims 1-5.

Citation Information

Patent Citations

  • Obstetrical midwifery decision-making model construction method based on mapping knowledge domain

    CN119153073A