Method, system, terminal and storage medium for extracting labels from structured radiology reports

By constructing a structured corpus and a natural language corpus of structured imaging reports, and combining the pathological characteristics of the examination site and free text, the imaging report labels are automatically extracted, which solves the problem of inaccurate label extraction in the existing technology and achieves higher label extraction accuracy and a wider range of applications.

CN115458110BActive Publication Date: 2025-09-26ZHONGSHAN HOSPITAL AFFILIATED TO FUDAN UNIV XIAMEN HOSPITAL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210972696.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2025-09-26
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

The accuracy of label extraction of diagnostic conclusion text in existing structured imaging reports is low, especially in diagnostic scenarios of accidental discoveries and lack of expert consensus, where reliance on manual experience leads to inaccurate label extraction.

Method used

By obtaining the examination sites and free text in the structured imaging reports, the pathological characteristics are determined, a structured corpus is constructed, and corpus analysis is performed in combination with the natural language corpus to automatically extract report tags.

Benefits of technology

It improves the accuracy of label extraction for structured imaging reports, meets complex reporting requirements, and expands its application scope in scientific research and teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115458110B_ABST
    Figure CN115458110B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, terminal, and storage medium for extracting labels from structured imaging reports. The method comprises: determining pathological characteristics based on the examination site in the structured imaging report; obtaining free text in the structured imaging report and determining a structured corpus based on the free text and the pathological characteristics; performing a corpus query based on the examination site to obtain a natural language corpus; performing corpus analysis on the diagnostic conclusion text, the structured corpus, and the natural language corpus in sequence, and extracting report labels from the structured imaging report based on the corpus analysis results. The present invention combines the content of the free text with the description of the imaging manifestations in the structured imaging report to obtain a structured corpus, and then extracts labels from the diagnostic conclusion text in combination with the natural language corpus, thereby improving the accuracy of label extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, system, terminal and storage medium for extracting labels from structured imaging reports. Background Art

[0002] Structured imaging reports feature labeling of imaging manifestation descriptions and can be used in a wide range of applications, including scientific research and teaching. In actual applications, not all imaging report content can be fully labeled. The following are two common scenarios where text is still needed. The first is for certain incidental findings or imaging manifestation types that are not covered by the current structured part. If the diagnostic physician wishes to describe them, they can only use text. The second is in diagnostic scenarios where there is a lack of expert consensus. The diagnosis cannot be automatically generated based on the imaging description, and the diagnostic physician still needs to manually enter the diagnostic conclusion text. In order to improve the accuracy of labeling of structured imaging reports, the problem of label extraction of the diagnostic conclusion text in structured imaging reports has received increasing attention.

[0003] In the existing process of extracting labels from diagnostic conclusion text in structured imaging reports, labels are generally extracted based on manual experience, resulting in low accuracy of label extraction. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method, system, terminal and storage medium for extracting labels from structured imaging reports, aiming to solve the problem of low accuracy in the existing extraction of labels from structured imaging reports.

[0005] The embodiment of the present invention is implemented as follows: a method for extracting labels from structured imaging reports, the method comprising:

[0006] Obtain the examination site in the structured imaging report and determine the pathological characteristics based on the examination site;

[0007] Obtaining free text in the structured imaging report, and determining a structured corpus based on the free text and the pathological characteristics, wherein the free text is a doctor's supplementary description of the imaging findings in the structured imaging report;

[0008] Performing a corpus query based on the inspection site to obtain a natural language corpus;

[0009] The diagnosis conclusion text in the structured imaging report is sequentially analyzed with the structured corpus and the natural language corpus, and the report label of the structured imaging report is extracted according to the corpus analysis results.

[0010] Furthermore, determining a structured corpus based on the free text and the pathological features includes:

[0011] Performing a structured unit query based on the pathological features to obtain a first structured unit, and querying a structured unit corresponding to the imaging structured report to obtain a second structured unit;

[0012] Acquire the corpora corresponding to the first structural unit and the second structural unit respectively to obtain a first corpus and a second corpus;

[0013] Obtaining a text position of the free text in the structured radiology report, and performing a corpus query based on the text position to obtain a third corpus;

[0014] The structured corpus is generated according to the first corpus, the second corpus, and the third corpus.

[0015] Furthermore, the corpus query is performed according to the text position to obtain a third corpus, including:

[0016] Obtaining a structural identifier of the second structural unit, and obtaining a paragraph tag and a title tag in the text position;

[0017] Matching the structured identifier, the paragraph tag, and the title tag with a pre-stored corpus query table to obtain a first sub-corpus;

[0018] Obtaining an associated structural unit of the first structural unit, and matching the associated structural unit with the corpus query table to obtain a second sub-corpus;

[0019] The third corpus is generated according to the first sub-corpus and the second sub-corpus.

[0020] Furthermore, after generating the third corpus according to the first sub-corpus and the second sub-corpus, the method further includes:

[0021] Obtaining the imaging description type of the diagnosis conclusion text, and matching the imaging description type with the corpus query table to obtain a third sub-corpus;

[0022] The third sub-corpus is added to the third corpus.

[0023] Furthermore, determining the pathological characteristics according to the examination site includes:

[0024] The site code of the inspection site is obtained, and the site code is matched with a pre-stored coding relationship tree to obtain the pathological feature, wherein the coding relationship tree stores the corresponding relationship between different site codes and corresponding pathological features.

[0025] Furthermore, after the diagnosis conclusion text in the structured imaging report is subjected to corpus analysis in sequence with the structured corpus and the natural language corpus, the method further includes:

[0026] A locally pre-stored general corpus is obtained, and corpus analysis is performed on the text diagnosis conclusion and the general corpus.

[0027] Another object of an embodiment of the present invention is to provide a system for extracting labels from structured imaging reports, the system comprising:

[0028] a feature determination module, configured to obtain the examination site in the structured imaging report and determine the pathological features based on the examination site;

[0029] a corpus determination module, configured to obtain free text from the structured imaging report, determine a structured corpus based on the free text and the pathological features, and perform a corpus query based on the examination site to obtain a natural language corpus, wherein the free text is a doctor's supplementary description of the imaging findings in the structured imaging report;

[0030] The label extraction module is used to perform corpus analysis on the diagnosis conclusion text in the structured imaging report with the structured corpus and the natural language corpus in sequence, and extract the report label of the structured imaging report according to the corpus analysis results.

[0031] Furthermore, the corpus determination module is further configured to:

[0032] Performing a structured unit query based on the pathological features to obtain a first structured unit, and querying a structured unit corresponding to the imaging structured report to obtain a second structured unit;

[0033] Acquire the corpora corresponding to the first structural unit and the second structural unit respectively to obtain a first corpus and a second corpus;

[0034] Obtaining a text position of the free text in the structured radiology report, and performing a corpus query based on the text position to obtain a third corpus;

[0035] The structured corpus is generated according to the first corpus, the second corpus, and the third corpus.

[0036] Another object of an embodiment of the present invention is to provide a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0037] Another object of an embodiment of the present invention is to provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0038] In an embodiment of the present invention, pathological characteristics are determined by examining the site, and a structured corpus can be automatically determined based on the free text and the pathological characteristics. By performing a corpus query on the examination site, the natural language corpus corresponding to the structured imaging report can be effectively determined. By sequentially performing corpus analysis on the diagnosis conclusion text with the structured corpus and the natural language corpus, the report label in the diagnosis conclusion text can be effectively extracted based on the corpus analysis results. In this embodiment, the content of the free text is combined with the description of the imaging manifestations of the structured imaging report to obtain a structured corpus, and then the label of the diagnosis conclusion text is extracted in combination with the natural language corpus, thereby improving the accuracy of the label extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of a method for extracting labels from structured imaging reports provided by the first embodiment of the present invention;

[0040] Figure 2 1 is a schematic diagram of the structure of the coding relationship tree provided by the first embodiment of the present invention;

[0041] Figure 3 Schematic diagram of the structure of a pancreatic focal peripheral invasion CDE provided by the first embodiment of the present invention;

[0042] Figure 4 Schematic diagram of the CDE structure of pancreatic focal lesions provided by the first embodiment of the present invention;

[0043] Figure 5 is a flow chart of a method for extracting labels from structured imaging reports provided by a second embodiment of the present invention;

[0044] Figure 6 1 is a schematic structural diagram of an imaging structured report label extraction system provided by a third embodiment of the present invention;

[0045] Figure 7 It is a schematic structural diagram of a terminal device provided in the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0047] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0048] Example 1

[0049] See also Figure 1 , is a flow chart of a method for extracting labels from an imaging structured report provided by a first embodiment of the present invention. The method for extracting labels from an imaging structured report can be applied to any terminal device or system. The method for extracting labels from an imaging structured report comprises the following steps:

[0050] Step S10, obtaining the examination site in the structured imaging report, and determining the pathological characteristics according to the examination site;

[0051] Among them, the underlying structure of the imaging manifestation description part of the imaging structured report can be divided into the following levels: first, the examination site (tissue / organ) level, and second, the description of specific physiological / pathological characteristics under the tissue / organ. For example, the liver is a description of an organ, and the description of its physiological / pathological characteristics includes diffuse descriptions: fatty liver, cirrhosis, polycystic liver, nonspecific diffuse liver disease, viral hepatitis, liver dysplasia, liver postoperative changes, etc.; focal lesions include: liver cancer, cysts, hemangiomas, echinococcosis, abscesses, liver metastases, etc. The above-mentioned diffuse and focal descriptions are made into independent structured units, referred to as CDE (common data element). All imaging description attributes of the above CDE are marked using RADLEX or SNOMED codes.

[0052] In this step, the basic information of the imaging structured report, the corresponding tissue / organ information and coding, and the pathological feature category information and coding to be described by the structured unit to which the imaging structured report belongs are obtained. The pathological feature can also be a physiological feature. For example, the imaging structured report is a general report for abdominal organs, using MR examination technology. The examination site obtained is the pancreas, and the structured unit to which the imaging structured report belongs is the pancreatic focal lesion.

[0053] Furthermore, before step S10, the method further includes:

[0054] The coding relationship tree is constructed using the coding, and the physiological / pathological characteristics described under each tissue site are expressed based on the coding relationship tree to determine the type of CDE that can appear under the tissue site. For example, see Figure 2The organ structure has subordinate contents such as position, quantity, distribution, and peripheral invasion. Peripheral invasion is a pathological feature of tissue parts. If there is a CDE of type peripheral invasion and the tissue organ is the pancreas, then the corpus corresponding to the CDE can be used in the pancreas part.

[0055] Construct a corpus of CDE, extract the corpus from each structured unit used to describe certain tissue parts or physiological / pathological characteristics of tissue parts, or directly input the corpus that can describe tissue parts and physiological / pathological characteristics to obtain a corpus query table.

[0056] For example, see Figure 3 The CDE in the figure is pancreatic focal invasion, which has two attribute codes: pancreas and peripheral invasion. According to the built-in logic of the structured report, the corpus content of lesion invasion, common bile duct, duodenum, stomach, spleen, etc. can be extracted. Figure 3 When extracting labels from the content, words such as lesion invasion and common bile duct have higher priority.

[0057] In this embodiment, attribute codes representing the medical meanings are also assigned to the corpus. For the corpus extracted from a specific structured unit, the coding of the CDE itself is naturally inherited. After the coding is obtained, the applicable corpus can be found through the coding relationship tree. Figure 3 The corpus generated by CDE has codes such as pancreas and peripheral invasion.

[0058] Optionally, in this step, determining the pathological characteristics according to the inspection site includes:

[0059] Obtaining the site code of the inspection site, and matching the site code with a pre-stored coding relationship tree to obtain the pathological feature,

[0060] Among them, the coding relationship tree stores the correspondence between different part codes and corresponding pathological characteristics. According to the part code of the examination part, the corresponding pathological characteristics are queried from the coding relationship tree. For example, the pathological characteristics include RADS classification describing the examination part, postoperative changes, developmental variations, and other items describing the properties of the examination part, such as size, morphology, substance, etc.

[0061] Step S20, obtaining free text in the structured imaging report, and determining a structured corpus based on the free text and the pathological features;

[0062] The free text is the doctor's supplementary description of the imaging findings in the structured imaging report. When designing the structured report, an edit box is added for the required tissues and organs, or the structured units describing physiological / pathological characteristics, for the doctor to enter the supplementary description of the imaging findings to obtain the free text.

[0063] For example, see Figure 4 ,The figure shows a CDE describing a focal lesion of the pancreas. This CDE has lesion characteristics, pancreas and other attribute codes. An “other edit box” is ,added to it to input supplementary image findings. The free text is ,obtained by obtaining the content in the “other edit box”.

[0064] In this embodiment, at the imaging description level, the free text addition position is set in layers, and the free text content of each layer is related to the description of the structured control under the layer. For example, taking the liver as an example, a row of free text boxes is added under the liver. The free text box can be used by diagnostic doctors to add imaging manifestation types that are not yet in the above-mentioned CDE list. The above-mentioned imaging manifestation types are often rare, and there is no need to make them into structured forms and place them in the interface. However, in a small number of cases, some doctors still believe that they are clinically significant and need to be described. Therefore, the content in the free text box is an independent type of imaging description of the liver. There may be a small number of descriptions of subordinate traits. The concepts / synonyms / grammatical structures contained in the content are limited to the CDE type corpus under the organ. Whether using natural language processing (NLP) for analysis or manual training, it is easy to disassemble it into CDE-like labels. For CDE types that frequently appear in free text and have not yet been made into CDE structured reports, this embodiment will prompt the structured report designer to add the CDE as a regular option.

[0065] In this embodiment, a free text box is also set under the subordinate CDE of tissue / organ for the diagnostic doctor to supplement the morphological description of CDE. For example, in the free text box under the liver / postoperative changes of the liver CDE, the diagnostic doctor may enter "the resection range is S7, S8", etc. The concepts / synonyms / grammatical structures contained in its content are limited to the supplementary description corpus under the CDE. Whether it is analyzed using NLP or manual training, it is easy to decompose it into CDE-like labels. For CDE subordinate attributes that appear frequently but are not yet included in the existing CDE, this embodiment will prompt the structured report designer to add the attributes of the CDE.

[0066] Step S30, performing a corpus query according to the inspection part to obtain a natural language corpus;

[0067] The natural language corpus obtained by the query is a personalized NLP library pre-set for the inspection site;

[0068] Step S40, performing corpus analysis on the diagnosis conclusion text in the structured imaging report with the structured corpus and the natural language corpus in sequence, and extracting the report tag of the structured imaging report according to the corpus analysis results;

[0069] Among them, by sequentially performing corpus analysis on the diagnostic conclusion text in the structured report of imaging with the structured corpus and the natural language corpus, the report tags in the diagnostic conclusion text can be effectively extracted based on the corpus analysis results. In this embodiment, a diagnostic conclusion text box is set under the imaging diagnosis column. The diagnostic conclusion text box is used to facilitate the diagnostic doctor to fill in the diagnostic conclusion text. Taking the CT scan diagnosis of liver cancer as an example, the diagnostic doctor may describe detailed contents such as fatty liver, cirrhosis, cysts, and multiple small liver cancers in the chapter of imaging findings, but the diagnosis is likely to only describe small liver cancers and cirrhosis, while ignoring other unimportant imaging manifestations. Therefore, when the imaging manifestations are described in a structured manner, the text of the diagnostic content is likely to be a subset of the imaging description content type, thereby limiting the corpus content of the diagnostic conclusion text. In this case, whether it is analyzed using NLP or manually trained, it is easy to decompose it into standard concepts and correspond to RADLEX / SNOMED diagnostic codes.

[0070] Optionally, in this step, after performing corpus analysis on the diagnostic conclusion text in the structured imaging report with the structured corpus and the natural language corpus in sequence, the step further includes:

[0071] A local pre-stored general corpus is obtained, and the text diagnosis conclusion is subjected to corpus analysis with the general corpus (general NLP library).

[0072] In this step, the NLP results are constrained according to the structured corpus (CDE corpus). Concepts / synonyms / grammatical structures contained in the corpus have higher result priority, for example:

[0073] 1. When the examination site is the pancreas, the most appropriate personalized NLP library is the pancreas NLP library;

[0074] 2. Obtain the CDE corpus determined in step S20;

[0075] 3. Perform NLP analysis, and prioritize matching each position in the diagnostic conclusion text with the content in the CDE corpus, followed by the content in the pancreatic NLP library, and finally with the content in the general NLP library, thereby improving the accuracy of diagnostic conclusion text label extraction.

[0076] Furthermore, in this embodiment, the basic information of the structured imaging report, the corresponding tissue / organ information, etc. can be used to match the personalized NLP dictionary library. The personalized NLP dictionary library is trained based on various scenarios, examination purposes, examination techniques, examination sites, tissues / organs, etc., and can improve the accuracy and quality of NLP for specific scenarios.

[0077] In this embodiment, labels are extracted from the diagnostic conclusion text based on structured image representations and labeled image representations (labels extracted from free text). The scope of analysis is limited by the description of the imaging representations, thereby improving the accuracy of label extraction. The corpus is used to constrain the NLP matching results under the current tissue organ, further improving the NLP recognition rate and recognition quality.

[0078] By designing free text boxes to describe complex additional information according to CDE modules in structured reports, it can not only meet the needs of reports that are more complex than the existing structured reports, but also limit the vocabulary and grammatical scope of NLP analysis based on the adjacent characteristics of CDE, and use NLP technology to convert the diagnostic conclusion text into code, while meeting the needs of scientific research and teaching. It greatly expands the scope of use of structured reports, and the structured design can be continuously improved based on NLP analysis to make it more sustainable.

[0079] In this embodiment, the pathological characteristics are determined by the examination site, and a structured corpus can be automatically determined based on the free text and the pathological characteristics. By performing a corpus query on the examination site, the natural language corpus corresponding to the structured imaging report can be effectively determined. By performing corpus analysis on the diagnosis conclusion text with the structured corpus and the natural language corpus in sequence, the report label in the diagnosis conclusion text can be effectively extracted based on the corpus analysis results. In this embodiment, the content of the free text is combined with the image manifestation description of the structured imaging report to obtain a structured corpus, and then the label of the diagnosis conclusion text is extracted in combination with the natural language corpus, thereby improving the accuracy of label extraction.

[0080] Example 2

[0081] See also Figure 5 , is a flow chart of a method for extracting labels from an imaging structured report provided by a second embodiment of the present invention. This embodiment is used to further refine step S20 in the first embodiment, including the steps of:

[0082] Step S21, performing a structured unit query based on the pathological feature to obtain a first structured unit, and querying a structured unit corresponding to the imaging structured report to obtain a second structured unit;

[0083] The structured unit query is performed based on the pathological characteristics to obtain the structured unit corresponding to the pathological characteristics related to the examination site, thereby obtaining the first structured unit. The structured unit corresponding to the imaging structured report is queried to obtain the second structured unit to which the current imaging structured report belongs.

[0084] Step S22, respectively obtaining the corpora corresponding to the first structural unit and the second structural unit to obtain a first corpus and a second corpus;

[0085] The first corpus and the second corpus are obtained by matching the structured identifiers of the first structured unit and the second structured unit with the corpus query table respectively to obtain the corpora corresponding to the first structured unit and the second structured unit;

[0086] Step S23, obtaining the text position of the free text in the structured radiology report, and performing a corpus query based on the text position to obtain a third corpus;

[0087] The third corpus is obtained by obtaining the text position of the free text in the structured radiology report and performing a corpus query based on the text position to query a corpus applicable to the free text. Optionally, in this step, performing a corpus query based on the text position to obtain the third corpus includes:

[0088] Obtaining a structural identifier of the second structural unit, and obtaining a paragraph tag and a title tag in the text position;

[0089] Matching the structured identifier, the paragraph tag, and the title tag with a pre-stored corpus query table to obtain a first sub-corpus;

[0090] Among them, by matching structured tags, paragraph tags, and title tags with the corpus query table, a corpus applicable to supplementary imaging findings at specific tissues and organs is queried. This free text may be used to describe independent types of imaging descriptions under the examination site, such as the pancreas as an "organ structure". It is more likely to contain a corpus generated by CDEs of a specific disease type, such as RADS classification, postoperative changes, developmental variations, and other CDEs;

[0091] Obtaining an associated structural unit of the first structural unit, and matching the associated structural unit with the corpus query table to obtain a second sub-corpus;

[0092] Among them, by matching the associated structured units with the corpus query table, a corpus suitable for supplementary imaging findings at the structured units of pathological characteristics is queried. For example, the edit box of free text in the CDE of pancreatic focal lesions has attribute encoding of lesion characteristics. Therefore, a corpus of CDE types such as size, morphology, substance, location, number, and surrounding invasion may be used.

[0093] The third corpus is generated according to the first sub-corpus and the second sub-corpus.

[0094] Furthermore, in this step, after generating the third corpus according to the first sub-corpus and the second sub-corpus, the step further includes:

[0095] Obtaining the imaging description type of the diagnosis conclusion text, and matching the imaging description type with the corpus query table to obtain a third sub-corpus;

[0096] adding the third sub-corpus to the third corpus;

[0097] Among them, by obtaining the imaging description type of the diagnosis conclusion text and matching the imaging description type with the corpus query table, the corpus applicable to the diagnosis conclusion text is queried. The diagnosis conclusion text is most likely a subset of the imaging description content type. Therefore, all CDE corpora related to the diagnosis conclusion text can be used as its constraints.

[0098] Step S24, generating the structured corpus according to the first corpus, the second corpus and the third corpus;

[0099] The structured corpus is obtained by combining the first corpus, the second corpus and the third corpus.

[0100] In this embodiment, a structured unit query is performed based on pathological characteristics to obtain structured units corresponding to pathological characteristics related to the examination site, thereby obtaining a first structured unit. The structured unit corresponding to the imaging structured report is queried to obtain a second structured unit to which the current imaging structured report belongs. The structured identifiers of the first structured unit and the second structured unit are matched with the corpus query table respectively to obtain the corpora corresponding to the first structured unit and the second structured unit, thereby obtaining a first corpus and a second corpus. The text position of the free text in the imaging structured report is obtained, and a corpus query is performed based on the text position to query the corpus applicable to the free text, thereby obtaining the third corpus. The structured corpus is obtained by combining the first corpus, the second corpus, and the third corpus.

[0101] Example 3

[0102] See also Figure 6 , is a schematic diagram of the structure of a radiology structured report label extraction system 100 provided in a third embodiment of the present invention, comprising: a feature determination module 10, a corpus determination module 11, and a label extraction module 12, wherein:

[0103] The feature determination module 10 is used to obtain the examination site in the structured imaging report and determine the pathological features according to the examination site.

[0104] The feature determination module 10 is further configured to obtain a site code of the inspection site and match the site code with a pre-stored coding relationship tree to obtain the pathological feature, wherein the coding relationship tree stores the correspondence between different site codes and corresponding pathological features.

[0105] The corpus determination module 11 is used to obtain the free text in the structured imaging report, determine the structured corpus based on the free text and the pathological characteristics, and perform a corpus query based on the examination site to obtain a natural language corpus. The free text is a doctor's supplementary description of the imaging findings in the structured imaging report.

[0106] The corpus determination module 11 is further configured to: perform a structured unit query based on the pathological characteristics to obtain a first structured unit, and query a structured unit corresponding to the imaging structured report to obtain a second structured unit;

[0107] Acquire the corpora corresponding to the first structural unit and the second structural unit respectively to obtain a first corpus and a second corpus;

[0108] Obtaining a text position of the free text in the structured radiology report, and performing a corpus query based on the text position to obtain a third corpus;

[0109] The structured corpus is generated according to the first corpus, the second corpus, and the third corpus.

[0110] Optionally, the corpus determination module 11 is further configured to: obtain a structural identifier of the second structural unit, and obtain a paragraph tag and a title tag in the text position;

[0111] Matching the structured identifier, the paragraph tag, and the title tag with a pre-stored corpus query table to obtain a first sub-corpus;

[0112] Obtaining an associated structural unit of the first structural unit, and matching the associated structural unit with the corpus query table to obtain a second sub-corpus;

[0113] The third corpus is generated according to the first sub-corpus and the second sub-corpus.

[0114] Furthermore, the corpus determination module 11 is further configured to: obtain the imaging description type of the diagnosis conclusion text, and match the imaging description type with the corpus query table to obtain a third sub-corpus;

[0115] The third sub-corpus is added to the third corpus.

[0116] The label extraction module 12 is used to perform corpus analysis on the diagnosis conclusion text in the structured imaging report with the structured corpus and the natural language corpus in sequence, and extract the report label of the structured imaging report according to the corpus analysis results.

[0117] The tag extraction module 12 is further configured to obtain a locally pre-stored general corpus and perform corpus analysis on the text diagnosis conclusion and the general corpus.

[0118] In this embodiment, the pathological characteristics are determined by the examination site, and a structured corpus can be automatically determined based on the free text and the pathological characteristics. By performing a corpus query on the examination site, the natural language corpus corresponding to the structured imaging report can be effectively determined. By performing corpus analysis on the diagnosis conclusion text in sequence with the structured corpus and the natural language corpus, the report label in the diagnosis conclusion text can be effectively extracted based on the corpus analysis results. In this embodiment, the content of the free text is combined with the image manifestation description of the structured imaging report to obtain a structured corpus, and then the label of the diagnosis conclusion text is extracted in combination with the natural language corpus, thereby improving the accuracy of label extraction.

[0119] Example 4

[0120] Figure 7 This is a structural block diagram of a terminal device 2 provided in the fourth embodiment of the present application. Figure 7 As shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for a method for extracting labels from structured radiographic reports. When the processor 20 executes the computer program 22, the steps of each embodiment of the method for extracting labels from structured radiographic reports are implemented.

[0121] Exemplarily, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to implement the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.

[0122] The processor 20 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0123] The memory 21 may be an internal storage unit of the terminal device 2, such as a hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 2. Furthermore, the memory 21 may include both an internal storage unit of the terminal device 2 and an external storage device. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is about to be output.

[0124] In addition, the functional modules in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0125] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable storage medium may include: any entity or device that can carry computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunications signals.

[0126] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for extracting labels from structured imaging reports, characterized in that: The method comprises: Obtain the examination site in the structured imaging report and determine the pathological characteristics based on the examination site; Obtaining free text in the structured imaging report, and determining a structured corpus based on the free text and the pathological characteristics, wherein the free text is a doctor's supplementary description of the imaging findings in the structured imaging report; Performing a corpus query based on the inspection site to obtain a natural language corpus; Performing corpus analysis on the diagnosis conclusion text in the structured imaging report in sequence with the structured corpus and the natural language corpus, and extracting the report label of the structured imaging report according to the corpus analysis results; The determining of a structured corpus according to the free text and the pathological features includes: Performing a structured unit query based on the pathological features to obtain a first structured unit, and querying a structured unit corresponding to the imaging structured report to obtain a second structured unit; Acquire the corpora corresponding to the first structural unit and the second structural unit respectively to obtain a first corpus and a second corpus; Obtaining a text position of the free text in the structured radiology report, and performing a corpus query based on the text position to obtain a third corpus; generating the structured corpus according to the first corpus, the second corpus, and the third corpus; The performing of a corpus query according to the text position to obtain a third corpus includes: Obtaining a structural identifier of the second structural unit, and obtaining a paragraph tag and a title tag in the text position; Matching the structured identifier, the paragraph tag, and the title tag with a pre-stored corpus query table to obtain a first sub-corpus; Obtaining an associated structural unit of the first structural unit, and matching the associated structural unit with the corpus query table to obtain a second sub-corpus; The third corpus is generated according to the first sub-corpus and the second sub-corpus.

2. The method for extracting labels from structured imaging reports according to claim 1, wherein: After generating the third corpus according to the first sub-corpus and the second sub-corpus, the method further includes: Obtaining the imaging description type of the diagnosis conclusion text, and matching the imaging description type with the corpus query table to obtain a third sub-corpus; The third sub-corpus is added to the third corpus.

3. The method for extracting labels from structured imaging reports according to claim 1, wherein: Determining the pathological characteristics according to the examination site includes: The site code of the inspection site is obtained, and the site code is matched with a pre-stored coding relationship tree to obtain the pathological feature, wherein the coding relationship tree stores the corresponding relationship between different site codes and corresponding pathological features.

4. The method for extracting labels from structured imaging reports according to any one of claims 1 to 3, wherein: After performing corpus analysis on the diagnosis conclusion text in the structured imaging report in sequence with the structured corpus and the natural language corpus, the method further includes: A locally pre-stored general corpus is obtained, and corpus analysis is performed on the diagnosis conclusion text and the general corpus.

5. A system for extracting labels from structured imaging reports, characterized in that: The system comprises: a feature determination module, configured to obtain the examination site in the structured imaging report and determine the pathological features based on the examination site; a corpus determination module, configured to obtain free text from the structured imaging report, determine a structured corpus based on the free text and the pathological features, and perform a corpus query based on the examination site to obtain a natural language corpus, wherein the free text is a doctor's supplementary description of the imaging findings in the structured imaging report; a label extraction module, configured to perform corpus analysis on the diagnostic conclusion text in the structured imaging report in sequence with the structured corpus and the natural language corpus, and extract the report label of the structured imaging report according to the corpus analysis results; The corpus determination module is further configured to: perform a structured unit query based on the pathological features to obtain a first structured unit, and query a structured unit corresponding to the imaging structured report to obtain a second structured unit; Acquire the corpora corresponding to the first structural unit and the second structural unit respectively to obtain a first corpus and a second corpus; Obtaining a text position of the free text in the structured radiology report, and performing a corpus query based on the text position to obtain a third corpus; generating the structured corpus according to the first corpus, the second corpus, and the third corpus; The corpus determination module is further configured to: obtain a structural identifier of the second structural unit, and obtain a paragraph tag and a title tag in the text position; Matching the structured identifier, the paragraph tag, and the title tag with a pre-stored corpus query table to obtain a first sub-corpus; Obtaining an associated structural unit of the first structural unit, and matching the associated structural unit with the corpus query table to obtain a second sub-corpus; The third corpus is generated according to the first sub-corpus and the second sub-corpus.

6. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Method and system for extracting information from electronic documents

    CN103294764A

  • Realization method and system for electronic medical record post-structuring and auxiliary diagnosis

    CN106383853A