Electronic medical record retrieval method and apparatus

By extracting image feature descriptions from imaging description text and predicting lesion types, generating search prompt text, and filtering out important medical record fragments, the problem of existing systems being unable to understand the unique language structure of imaging is solved, thus improving the accuracy and efficiency of medical record retrieval.

CN121561129BActive Publication Date: 2026-05-12BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV
Filing Date
2025-08-29
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing electronic medical record retrieval systems cannot accurately understand the unique linguistic structure and imaging features of radiology, resulting in insufficient accuracy and relevance of retrieval results and an inability to select medical record fragments in accordance with the unique standards of the radiology field.

Method used

By extracting image feature descriptions from imaging description text, predicting lesion types, generating search prompt text, filtering out target medical record fragments whose importance values ​​meet preset values, and performing structured processing to display them in the medical record search interface.

Benefits of technology

实现了精准理解影像学描述文本中的影像征象描述,提高了病历检索的效率和精准度,减少了医生的工作负担,为影像分析提供更全面、更智能化的支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121561129B_ABST
    Figure CN121561129B_ABST
Patent Text Reader

Abstract

The application discloses an electronic medical record retrieval method and device. The method comprises the following steps: analyzing an imaging description text, determining a lesion type described by the imaging description text, and extracting an imaging sign description related to the lesion type from the imaging description text; extracting an imaging sign description related to the lesion type from the imaging description text; predicting a lesion type corresponding to a target object based on the imaging sign description; generating a retrieval prompt text based on the lesion type; retrieving a plurality of candidate medical record segments related to the lesion type based on the retrieval prompt text; and accurately retrieving medical record segments based on imaging characteristics to avoid omissions or false detections caused by ambiguous natural language expressions or incomplete retrieval conditions. Based on the importance values of the plurality of candidate medical record segments, the plurality of target medical record segments with the highest matching degree with the retrieval prompt text are selected from the plurality of candidate medical record segments, thereby improving the efficiency and accuracy of medical record retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and in particular relates to electronic medical record retrieval methods and devices. Background Technology

[0002] In the field of medical imaging, radiologists need to transform complex visual information from CT, MRI, and X-rays into image description text, and combine this with patient clinical medical record information to generate reliable imaging reports. These image description texts typically contain a large amount of technical terminology and specific imaging features. Therefore, the ability to quickly extract key information and perform targeted electronic medical record searches is a crucial need for doctors and researchers.

[0003] However, traditional electronic medical record retrieval relies heavily on keyword matching, which often fails to understand the unique language structure of radiology and specific imaging features in radiological descriptions, and cannot identify factors that lead to insufficient accuracy and relevance in the search results. Summary of the Invention

[0004] In view of this, this application provides an electronic medical record retrieval method, device, and storage medium that solves or partially solves the above-mentioned technical problems.

[0005] In a first aspect, embodiments of this application provide an electronic medical record retrieval method, the method comprising:

[0006] In response to the imaging description text entered by the user for the target object in the medical record retrieval interface, the imaging sign description is extracted from the imaging description text, and the imaging sign description is used to describe the imaging manifestations of the lesion tissue.

[0007] The image features are analyzed to predict the lesion type corresponding to the target object;

[0008] Based on the lesion type, a search prompt text is generated, and multiple candidate medical record fragments that match the search prompt text are retrieved from the electronic medical record text corresponding to the target object.

[0009] Determine the importance value corresponding to each candidate medical record segment, the importance value being used to characterize the degree of clinical impact of the candidate medical record segment on the imaging problems described in the imaging description text;

[0010] Based on the importance values ​​corresponding to each of the multiple candidate medical record segments, multiple target medical record segments whose importance values ​​meet preset values ​​are selected from the multiple candidate medical record segments;

[0011] The multiple target medical record fragments are processed into structured medical record information to obtain multiple structured medical record information, which is then displayed in the medical record retrieval interface.

[0012] Secondly, embodiments of this application provide an electronic medical record retrieval device, the device comprising:

[0013] The extraction module is used to extract image feature descriptions from the image feature description text input by the user for the target object in the medical record retrieval interface. The image feature descriptions are used to describe the image manifestations corresponding to the lesion tissue.

[0014] The prediction module is used to analyze the image feature description and predict the lesion type corresponding to the target object;

[0015] The generation module is used to generate search prompt text based on the lesion type, and retrieve multiple candidate medical record fragments that match the search prompt text from the electronic medical record text corresponding to the target object;

[0016] A determination module is used to determine the importance value corresponding to each candidate medical record segment, wherein the importance value is used to characterize the degree of clinical impact of the candidate medical record segment on the imaging problems described in the imaging description text;

[0017] The filtering module is used to filter out multiple target medical record segments whose importance values ​​meet preset values ​​from the multiple candidate medical record segments based on the importance values ​​corresponding to each of the multiple candidate medical record segments.

[0018] The processing module is used to perform structured processing on the multiple target medical record fragments to obtain multiple structured medical record information, and to display the multiple structured medical record information in the medical record retrieval interface.

[0019] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor performs the steps of the electronic medical record retrieval method as described in the first aspect.

[0020] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the steps of the electronic medical record retrieval method as described in the first aspect.

[0021] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, cause the processor to implement the steps in the electronic medical record retrieval method provided in embodiments of this application.

[0022] The electronic medical record retrieval scheme provided in this application embodiment can provide users with accurate electronic medical record retrieval services. In specific implementation, users can input imaging description text for a target object in the medical record retrieval interface to trigger a medical record retrieval request. In response to the imaging description text input by the user for the target object in the medical record retrieval interface, imaging sign descriptions are extracted from the imaging description text. These imaging sign descriptions describe the imaging manifestations corresponding to the lesion tissue. Next, the imaging sign descriptions are analyzed to predict the lesion type corresponding to the target object. Then, based on the lesion type, retrieval prompt text is generated, and multiple candidate medical record fragments matching the retrieval prompt text are retrieved from the electronic medical record text corresponding to the target object. Then, the importance value corresponding to each candidate medical record fragment is determined. The importance value is used to characterize the clinical impact of the candidate medical record fragment on the imaging problem described in the imaging description text. Based on the importance values ​​corresponding to each of the multiple candidate medical record fragments, multiple target medical record fragments whose importance values ​​meet preset values ​​are selected from the multiple candidate fragments. Finally, the multiple target medical record fragments are processed into structured medical record information to obtain multiple structured medical record information, which is then displayed in the medical record retrieval interface.

[0023] In the above technical solution, by analyzing the imaging description text, the type of lesion described in the imaging description text is determined. Imaging feature descriptions related to the lesion type are extracted from the imaging description text. Then, based on the imaging feature descriptions, the lesion type corresponding to the target object is predicted. Based on the lesion type, a search suggestion text is generated. Based on the search suggestion text, candidate medical record fragments related to the lesion type are retrieved. This achieves accurate understanding of the imaging feature descriptions in the imaging description text and accurate retrieval of candidate medical record fragments related to the lesion type described in the imaging description text based on unique imaging features. This avoids omissions or false detections caused by ambiguous natural language expressions or incomplete search conditions. Furthermore, during the retrieval process, multiple candidate medical record segments that match the semantics are first selected from multiple medical record segments corresponding to the target object. Then, based on the importance values ​​of each of the multiple candidate medical record segments, multiple target medical record segments with higher importance are selected from the multiple candidate medical record segments. This can improve the efficiency and accuracy of medical record retrieval, reduce the workload of doctors, and provide more comprehensive and intelligent support for image analysis. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0025] Figure 1 A flowchart illustrating an electronic medical record retrieval method provided for an exemplary embodiment of this application;

[0026] Figure 2 A flowchart illustrating the process of rewriting imaging description text in a medical record retrieval request, provided as an exemplary embodiment of this application;

[0027] Figure 3 A flowchart illustrating a process for generating search suggestion text based on lesion type, provided as an exemplary embodiment of this application;

[0028] Figure 4 An application diagram of an electronic medical record retrieval method provided as an exemplary embodiment of this application;

[0029] Figure 5 A schematic diagram of the structure of an electronic medical record retrieval device provided as another exemplary embodiment of this application;

[0030] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “said,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one.

[0033] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0034] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.

[0035] In the field of medical imaging, the core task of radiologists is to transform complex visual information (such as CT, MRI, X-ray, etc.) into image description text, and to combine these image description texts with the patient's clinical history to form valuable imaging reports.

[0036] However, current information retrieval systems lack the ability to understand the unique text in radiological reports, specifically in the following two aspects:

[0037] Firstly, image description texts have unique linguistic structures and image sign expression systems. Current information retrieval systems cannot accurately understand image description texts and image sign characteristics, thus failing to accurately retrieve medical record fragments related to these image sign characteristics from the patient's corresponding electronic medical record text.

[0038] Specifically, imaging descriptions use a language system completely different from that of other medical departments, with strict norms for expressing signs and multidimensional spatial information encoding characteristics. For example, an imaging description such as "a 2.3×1.8cm round nodule in the posterior segment of the right upper lobe, with a CT value of approximately 42 HU, rough edges, lobulation and spiculation, and pleural traction" simultaneously includes precise three-dimensional spatial location ("posterior segment of the right upper lobe"), quantitative measurement data ("2.3×1.8cm", "CT value approximately 42 HU"), morphological features ("round", "rough edges"), and specific imaging signs ("lobulation", "spiculation", "pleural traction"). These imaging-specific terms and expressions are almost non-existent in medical records from other medical departments, and existing general information retrieval systems cannot understand the complexity of this professional language. Moreover, combinations of specific imaging features have special diagnostic significance. For example, the combination of "spiculated signs + lobulation signs + pleural traction" strongly suggests malignancy in the assessment of pulmonary nodules. The semantic meaning of such imaging feature combinations far exceeds the simple superposition of individual terms. However, existing information retrieval systems lack the ability to understand and analyze information related to combinations of imaging features.

[0039] Secondly, the imaging field has unique standards for assessing the importance of clinical information. When searching for medical record fragments, it is impossible to combine these unique standards to screen the fragments, thus making it impossible to accurately retrieve medical record fragments related to the imaging features from the patient's corresponding electronic medical record text.

[0040] In the process of imaging evaluation, the importance of specific clinical information differs significantly from that in other departments. For example, in the evaluation of pulmonary nodules, the patient's smoking history, occupational exposure history (such as asbestos, uranium mining, etc.), and history of cancer are far more important than other routine clinical information; while in the evaluation of intracranial hemorrhage, recent anticoagulant use, platelet count, and coagulation function test results are decisive. Unlike internal medicine, the importance assessment of clinical information in radiology is highly dependent on the specific location and findings of the current examination. The same piece of clinical information (such as "long-term aspirin use") may have completely different importance depending on different imaging findings. Currently, general information retrieval systems cannot dynamically adjust the importance weight of clinical information according to different types of imaging examinations and findings, forcing radiologists to spend a lot of time sifting through irrelevant information to find key clues.

[0041] To address the aforementioned technical issues, this application provides an electronic medical record retrieval scheme. In this scheme, imaging features related to the lesion type described in the imaging description text are extracted. Based on these imaging features, the lesion type corresponding to the target object is predicted. Then, based on the lesion type, a search suggestion text is generated. This text is used to retrieve relevant medical record fragments from the patient's electronic medical record. This achieves accurate understanding of the imaging features described in the imaging description text and accurately retrieves candidate medical record fragments related to the lesion type described in the text based on unique imaging features. This avoids omissions or false positives caused by ambiguous natural language expressions or incomplete search conditions. Furthermore, after obtaining multiple candidate medical record fragments, more important target medical record fragments can be selected from them according to their importance value, allowing users to quickly view important medical history information. This not only achieves automatic retrieval of medical record information, reducing the workload of doctors, but also improves the accuracy of the retrieved medical record fragments through two rounds of filtering, while providing more comprehensive and intelligent support for image analysis.

[0042] The specific implementation of the electronic medical record retrieval scheme provided in this application is described below with reference to specific embodiments.

[0043] Figure 1 This is a flowchart illustrating an electronic medical record retrieval method provided as an exemplary embodiment of this application. Figure 1 As shown, specifically, the method may include the following steps:

[0044] 101. In response to the imaging description text entered by the user for the target object in the medical record retrieval interface, extract the image sign description from the imaging description text.

[0045] 102. Analyze the description of imaging features to predict the lesion type corresponding to the target object.

[0046] 103. Based on the lesion type, generate search suggestion text and retrieve multiple candidate medical record fragments that match the search suggestion text from the electronic medical record text corresponding to the target object.

[0047] 104. Determine the importance value corresponding to each candidate medical record segment. The importance value is used to characterize the degree of influence of the candidate medical record segment on the imaging problems described in the imaging description text.

[0048] 105. Based on the importance values ​​of each of the multiple candidate medical record segments, select multiple target medical record segments from the multiple candidate medical record segments whose importance values ​​meet the preset values.

[0049] 106. Perform structured processing on multiple target medical record fragments to obtain structured medical record information, and display the structured medical record information in the medical record retrieval interface.

[0050] In practical applications, when users perform imaging analysis on a target object, in order to accurately obtain medical record information relevant to the needs of this imaging analysis, users can trigger a medical record retrieval request for the target object. Specifically, users can enter imaging description text for the target object in the medical record retrieval interface to trigger a medical record retrieval request to the electronic medical record retrieval device. After receiving the imaging description text entered by the user, the electronic medical record retrieval device retrieves the imaging features described in the imaging description text. The imaging feature description describes the imaging manifestations corresponding to the lesion tissue.

[0051] In an optional embodiment, in order to extract the image sign description from the image description text more accurately, the image description text can be analyzed first to determine the lesion type corresponding to the target object, and the image sign description related to the lesion type can be extracted from the image description text.

[0052] Typically, an image description text contains a large number of technical terms and specific imaging features. The purpose of retrieving medical record fragments is to further retrieve medical history information related to the imaging problem, so as to analyze the imaging problem in combination with the medical history information and determine the lesion information corresponding to the target object. Therefore, in order to retrieve multiple target medical record fragments related to the image description text more accurately and efficiently from the electronic medical records corresponding to the target object, during the retrieval, the lesion information corresponding to the lesion tissue can be extracted from the image description text first, and the lesion type described by the image description text can be determined based on the lesion information corresponding to the lesion tissue.

[0053] For example, suppose the user inputs the imaging description text as "A round or oval nodule is seen in the upper lobe of the right lung, approximately 2.3cm × 1.8cm in size, with rough edges, lobulation, and spiculation, and pleural traction." After analyzing this imaging description text, the lesion information extracted is "A round or oval nodule is seen in the upper lobe of the right lung." Based on this lesion information, the corresponding lesion type is determined to be a pulmonary nodule.

[0054] After determining the lesion type corresponding to the lesion tissue described in the imaging description text, the imaging feature description related to that lesion type is extracted from the imaging description text. The imaging feature description describes the imaging manifestations corresponding to the lesion tissue and includes at least one imaging feature.

[0055] In extracting descriptions of imaging features, the process begins by identifying specific imaging features from the user-inputted text. Then, descriptions of lesion types are extracted from these identified features. Continuing with the example above, the identified features include spatial location information ("right upper lobe"), measurement data ("2.3cm × 1.8cm"), morphological characteristics ("round or oval", "rough edges"), and specific imaging features ("lobulation sign", "spiculated sign", "pleural traction"). The extracted descriptions of these features are then "lobulation sign", "spiculated sign", and "pleural traction".

[0056] Next, the descriptions of imaging features are analyzed to predict the lesion type corresponding to the target object. Here, lesion type refers to the specific classification of a medically relevant disease. In practice, this can be achieved by combining an imaging feature database, which stores various imaging feature characteristics and their corresponding lesion types. Alternatively, a pre-trained lesion type prediction model can be used to predict the lesion type of the target object. Or, a medical knowledge graph and the imaging feature characteristics included in the imaging feature description can be combined to predict the lesion type of the target object.

[0057] When conducting medical record searches, the essence is still to obtain medical history information that may lead to the imaging problems described in the imaging description text or related medical history information. In order to retrieve relevant medical history information from the electronic medical records of the target object more accurately, we can first predict the possible lesion types of the target object based on the imaging sign description, and then generate search prompt text based on the lesion type.

[0058] Medical record retrieval differs from traditional information retrieval. When conducting a medical record retrieval, it's not only necessary to find medical record fragments that semantically match the extracted keywords, but also to retrieve other medical history information associated with those keywords. Therefore, during a medical record retrieval, explicit search suggestions can be generated directly. These suggestions can include possible corresponding lesion types, key information requiring attention, and risk factors leading to the lesion type. For example, the generated search suggestion might be: "The lesion is suspected to be a malignant lung tumor. Key areas of focus include: ① smoking history ② occupational exposure history ③ previous cancer history ④ lung infection history ⑤ previous imaging records of the same site."

[0059] After the search suggestion text is generated, the electronic medical record text corresponding to the target object is obtained based on the search suggestion text, and multiple candidate medical record fragments that match the search suggestion text are retrieved from the electronic medical record text corresponding to the target object.

[0060] However, in practical applications, the electronic medical record text corresponding to a target object contains a large amount of medical history information about that target object. Therefore, in order to quickly and accurately retrieve medical record information related to the imaging description text entered by the user from the electronic medical record, the electronic medical record text can be segmented after obtaining the target object's electronic medical record text. In specific implementation, the electronic medical record text can be segmented according to themes, text structures, language features, etc.

[0061] After obtaining multiple medical record fragments, a preliminary screening process is performed to select candidate medical record fragments that semantically match the search prompt text. This screening can be achieved by analyzing the semantic relevance between the multiple medical record fragments and the search prompt text.

[0062] In practical applications, when the electronic medical record text corresponding to the target object contains a large amount of medical record information related to the search prompt text, if these multiple candidate medical record fragments are sent to the user simultaneously, the user will need to select the required relevant information from the multiple candidate medical record fragments again, which will affect the user's work efficiency. Therefore, in this embodiment, after selecting multiple candidate medical record fragments that match the semantics of the search prompt text, multiple target medical record fragments corresponding to the search prompt text can be selected again from the multiple candidate medical record fragments based on the degree of influence of the multiple candidate medical record fragments on the imaging problems described in the imaging description text.

[0063] Specifically, a medical knowledge graph can be used to determine the importance value for each candidate medical record segment. The importance value characterizes the degree of influence of the candidate medical record segment on the imaging problems described in the imaging descriptive text.

[0064] This process involves pre-creating a medical knowledge graph, which can be directly used for medical record retrieval. In practice, data such as medical literature, clinical guidelines, and electronic medical records can be collected, and noisy data is removed. The collected data is then rewritten to use standardized terminology. The rewritten data is then labeled with medical entities, relationships, and attributes. This can be done by defining entity classes (e.g., diseases, drugs) and relationship classes (e.g., "interactions"). Medical entity, relationship, and attribute information is extracted from each rewritten data set based on medical rules. Based on this information, a graph structure is generated, and a graph database is selected to store it. Nodes in the graph structure represent medical entities, and edges represent relationships between these entities.

[0065] The medical knowledge graph stores multiple medical entities and their relationships, allowing for the determination of the importance value for each candidate medical record segment. For example, a medical knowledge graph can be represented as a set of triples: KG = (h, r, t) | h, t ∈ Entities, r ∈ Relations, where h represents the head entity, r represents the relationship between entities, and t represents the tail entity. Entities and Relations represent the set of entities and relationships in the knowledge graph, respectively. Then, based on the importance values ​​of each candidate medical record segment, multiple target medical record segments whose importance values ​​meet preset values ​​are selected.

[0066] In a specific implementation, in an optional embodiment, multiple candidate medical record segments can be sorted sequentially according to their respective importance values ​​from largest to smallest, and the top N medical record segments can be selected as the target medical record segments. The value of N can be set according to actual needs.

[0067] In another optional embodiment, more refined features can be extracted from multiple candidate medical record fragments and search prompt text. Combining the importance values ​​corresponding to each of the multiple candidate medical record fragments, the matching degree between the multiple candidate medical record fragments and the search prompt text is determined based on the features corresponding to each of the multiple candidate medical record fragments and the features corresponding to the search prompt text. The multiple candidate medical record fragments are then sorted in descending order of matching degree between the multiple candidate medical record fragments and the search prompt text, and the top N medical record fragments are selected as target medical record fragments.

[0068] Finally, multiple target medical record fragments are processed into structured medical record information, which is then displayed in the medical record retrieval interface. This allows users to obtain the information they need more quickly and clearly.

[0069] Furthermore, when displaying multiple structured medical records, the system can sort them according to a preset imaging assessment priority rule, resulting in a list of sorted structured medical records. These sorted records are then displayed sequentially in the medical record retrieval interface. The imaging assessment priority rule determines the priority of each structured medical record based on the level of focus in the image analysis.

[0070] This application embodiment analyzes the imaging description text to determine the lesion type described in the text, extracts imaging features related to the lesion type from the text, and then predicts the lesion type corresponding to the target object based on the imaging feature description. Based on the lesion type, it generates search suggestion text, and retrieves candidate medical record fragments related to the lesion type based on the search suggestion text. This achieves accurate understanding of the imaging feature descriptions in the imaging description text and accurately retrieves candidate medical record fragments related to the lesion type described in the text based on unique imaging features, thus avoiding omissions or false detections caused by ambiguous natural language expressions or incomplete search conditions. Furthermore, during the search, multiple semantically matching candidate medical record fragments are first selected from multiple medical record fragments corresponding to the target object. Then, based on the importance values ​​of each candidate medical record fragment, multiple target medical record fragments with higher importance are selected from the candidate medical record fragments. This improves the efficiency and accuracy of medical record retrieval, reduces the workload of doctors, and provides more comprehensive and intelligent support for image analysis.

[0071] In practical applications, each user has their own preferred way of describing things, so users may use different terms, abbreviations, or dialects to describe the same medical phenomenon. Rewriting can transform these non-standardized descriptions into standardized language commonly used in the medical field, making them easier to retrieve.

[0072] Therefore, in this embodiment, after obtaining the image description text input by the user, the image description text can be rewritten to convert it into a standardized description text, and then image feature descriptions can be extracted from the image description text. This rewriting can be based on preset text conversion rules, or a machine learning model or a large language model can be used to rewrite the image description text.

[0073] Furthermore, to facilitate quick understanding and extraction of key information from user input during subsequent searches, the standardized description text can be converted into structured search suggestion text after it is obtained. This can be achieved by using predefined rule templates to extract key information from the imaging description text and populate it into structured fields.

[0074] As described above, in this embodiment, rewriting the imaging description text mainly includes converting the imaging description text into a standardized description text and converting the standardized description text into structured information. For example, the imaging description text input by the user is: "A high-density nodular shadow is visible in the right lung field, with rough edges, lobulation and spiculation around it, and pleural traction." Rewriting this text yields the standardized description text "Right lung nodule, irregular edges, with lobulation, spiculation and pleural traction." Converting this into structured search prompt text "Symptom: Lung nodule, Nodule location: Right lung, Nodule shape: Irregular edges, with lobulation, spiculation and pleural traction."

[0075] By rewriting the imaging description text entered by the user, a structured search suggestion text is obtained. This allows for a clearer extraction of key information from the user's input, thereby avoiding omissions or false positives caused by vague natural language expressions or incomplete search conditions.

[0076] To facilitate understanding of the specific implementation process of rewriting the imaging description text in the medical record retrieval request in the above embodiments, the following provides an exemplary description of how to rewrite the imaging description text in the medical record retrieval request using a machine learning model to obtain structured retrieval prompt text.

[0077] Figure 2 This is a schematic diagram illustrating a process for rewriting imaging description text in a medical record retrieval request, provided as an exemplary embodiment of this application. For example... Figure 2 As shown in the embodiments of this application, a method is provided to rewrite the imaging description text in a medical record retrieval request to obtain structured retrieval prompt text. Specifically, the method may include the following steps:

[0078] 201. Input the image description text into the pre-trained text rewriting model to obtain the rewritten text corresponding to each word in the image description text.

[0079] 202. Generate structured search suggestion text based on the rewritten text corresponding to each word.

[0080] In practice, the acquired image description text is input into a pre-trained text rewriting model to obtain standardized entity words corresponding to each entity word in the image description text. Based on the rewritten text corresponding to each entity word, a structured image description text is generated.

[0081] In the text rewriting model, each entity word in the imaging description text is encoded to obtain a word vector corresponding to each entity word. The word vectors, medical entity embedding vectors, and relation embedding vectors between medical entities corresponding to each entity word are then fused to obtain a fused vector representation for each entity word. This fused vector representation is then decoded to obtain the rewritten text corresponding to each entity word. The medical entity embedding vector represents the semantic information of each entity word within the medical knowledge graph. The relation embedding vector represents the relationships between the medical entities corresponding to each entity word.

[0082] In practical applications, a text rewriting model can be pre-trained to rewrite the image description text input by the user. This text rewriting model can include an encoding layer, a fusion layer, and a decoding layer. The encoding layer consists of multiple encoders, and the decoding layer consists of multiple decoders. The encoding layer is mainly used to encode the image description text to obtain vector representations of each word in the text. These vector representations include not only the semantic information of the word but also the contextual information (the semantic information of the preceding words).

[0083] Optionally, the encoding layer can use a multi-layer Transformer encoder. Each layer includes a self-attention mechanism and a feedforward neural network, which can model the semantic dependencies of the image description text and obtain the vector representation of the image description text. That is, Encoder(Q) = [h1, h2, ..., hn]. Where Q represents the image description text, and hi represents the encoded representation of the i-th word.

[0084] The fusion layer is mainly used to introduce medical entity embedding and relation embedding on the output of the encoding layer, and to integrate the information of the medical knowledge graph into the vector representation corresponding to the imaging description text, so as to improve the model's ability to understand and process medical problems.

[0085] In an optional embodiment, the fusion layer can fuse the vector representations corresponding to each word, the medical entity embeddings corresponding to each word, and the entity relationship embeddings together using the following formula: ei = hi + Emb(entityi) + Emb(relationi). Here, ei represents the medical entity corresponding to the i-th word, and relationi represents the relationship type between entities. Emb(entityi) is the medical entity embedding vector corresponding to the i-th word. The medical entity embedding vector can be obtained through pre-training or training on a specific task, reflecting the semantic information of the medical entity in the medical knowledge graph. Emb(relationi) is the relationship embedding corresponding to the i-th word. If two words have a relationship in the medical knowledge graph, the embedding vector of that relationship, Emb(relationi), is added to the word representation. Relation embedding can capture the semantic associations between entities, enhancing the model's understanding of medical logic.

[0086] The decoding layer can use a multi-layer Transformer decoder, which is mainly used to generate rewritten text corresponding to each word in the imaging description text. Specifically, the encoder attends to the output of the encoding layer and the medical knowledge representation through an attention mechanism, while simultaneously generating the text word by word using an autoregressive approach.

[0087] The training process of the text rewriting model is roughly the same as the usage process of the text rewriting model described above, and can be referred to the above description, so it will not be repeated here.

[0088] After obtaining the standardized entity words corresponding to each entity word in the image description text using a pre-trained text rewriting model, a structured image description text can be generated based on the standardized entity words corresponding to each entity word.

[0089] By rewriting the imaging description text using the above method, not only is the semantic and contextual information of each word in the imaging description text combined, but also the knowledge of the medical entities corresponding to each entity word and the relationships between entities are incorporated. The standardized entity words obtained in this way can express the medical observation results more accurately, professionally, and clearly, while improving the efficiency and effectiveness of subsequent automated processing.

[0090] After obtaining the structured radiographic description text, the imaging features related to the lesion type are extracted from it. Specifically, based on the structured radiographic description text, a second cue word is generated; this cue word is input into a large language model, which analyzes the structured radiographic description text to determine the lesion type corresponding to the lesion tissue described in the text; and imaging features related to the lesion type are extracted from the structured radiographic description text.

[0091] In practical applications, combinations of specific imaging features have special diagnostic significance. For example, the combination of "spiculated signs + lobulation signs + pleural traction" strongly suggests malignancy in the assessment of pulmonary nodules. The semantic meaning of such a combination of signs far exceeds the simple summation of individual terms. Therefore, when multiple imaging features are extracted, the combination of these features can be analyzed to predict the lesion type corresponding to the target object.

[0092] The following describes a possible specific implementation method for analyzing the image sign description and predicting the lesion type corresponding to the target object when the extracted image sign description includes multiple image sign features.

[0093] The specific implementation of analyzing imaging features to predict the lesion type corresponding to the target object can be as follows: if an entry corresponding to a combination of multiple imaging feature characteristics is retrieved from the imaging feature database, the lesion type in the entry is determined as the lesion type corresponding to the target object. The imaging feature database stores the correspondence between various imaging feature combinations and their corresponding lesion types.

[0094] In practice, an image sign database can be pre-built, which stores the correspondence between various image sign features and their corresponding lesion types, as well as the correspondence between various image sign combinations and their corresponding lesion types.

[0095] As described above, by utilizing a pre-built database of imaging features and analyzing the lesion types corresponding to combinations of imaging features, we can better understand the specific clinical significance of the imaging features in the imaging description text entered by the user when they appear in combination.

[0096] In another scenario, if no entry corresponding to a combination of multiple imaging features is retrieved from the imaging feature database, the imaging description text is analyzed in conjunction with the medical knowledge graph and the lesion type described in the imaging description text to predict the lesion type corresponding to the target object.

[0097] One possible approach to analyzing imaging description text and predicting the lesion type of a target object is as follows: retrieve knowledge fragments related to lesion type from a medical knowledge graph; generate a first prompt word based on the knowledge fragment, imaging description text, and lesion type; input the first prompt word into a large language model to obtain the lesion type output by the large language model.

[0098] The specific implementation process of generating search suggestion text will be described in detail below with reference to the following embodiments.

[0099] Figure 3This is a flowchart illustrating a process for generating search suggestion text based on lesion type, provided as an exemplary embodiment of this application. For example... Figure 3 As shown in the embodiment of this application, a method for generating search prompt text based on lesion type is provided. Specifically, the method may include the following steps:

[0100] 301. Access the medical knowledge base to identify the risk factors that lead to the type of lesion.

[0101] 302. Fill in the lesion type and risk factors into the preset search prompt template to generate search prompt text.

[0102] When generating search suggestion text, risk factors that lead to the type of lesion can be identified so that the electronic medical record text can be directly searched to see if these risk factors are present, in order to further determine the lesion information corresponding to the target object.

[0103] One approach is to utilize a medical knowledge base, including a medical knowledge graph, to identify risk factors leading to different disease types. When using the medical knowledge graph to determine risk factors for disease types, multiple medical entity terms (named entities containing medical information) can be identified within the imaging description text. Named Entity Recognition (NER) technology can then be used to extract key medical entity terms. For example, if the imaging description text is: "Patient's chest CT scan shows a peripheral mass in the right upper lobe of the lung," the extracted medical entity terms would be {lung, mass, right upper lobe, peripheral}.

[0104] Specifically, in an optional embodiment, the process of recognizing multiple medical entity words included in the imaging description text may include: inputting the imaging description text into a pre-trained medical entity recognition model to obtain multiple medical entity words, and the medical entity recognition model is used to identify the medical entity words included in the imaging description text.

[0105] Then, the target entity nodes corresponding to each of the multiple medical entity terms are searched from the medical knowledge graph. To facilitate the rapid and accurate retrieval of target medical record fragments that match the imaging description text, multiple medical entity terms included in the imaging description text can be linked with medical entities in the medical knowledge graph, so as to find the target entity nodes corresponding to each of the multiple medical entity terms from the medical knowledge graph.

[0106] Next, the structural information of the medical knowledge graph can be used to calculate the importance score corresponding to each target entity node. This can be achieved by calculating the degree centrality and PageRank value of each target entity node. Specifically, the degree centrality of each target entity node in the medical knowledge graph is determined based on the number of directly connected edges in the graph.

[0107] In an optional embodiment, the degree centrality of each target entity node in the medical knowledge graph can be determined using the following formula: Centrality(e_i) = (deg(e_i)) / (|Entities| - 1). Here, deg(e_i) represents the degree of the target entity node e_i in the knowledge graph (i.e., the number of edges directly connected to it).

[0108] After determining the degree centrality of each target entity node, the key focus fields corresponding to the image description text are then determined based on the degree centrality of each target entity node. That is, the key focus fields here may be the key points that the user is interested in, and this method can be used to determine the key content that the user is interested in.

[0109] Next, based on the key focus fields, relevant factor information corresponding to the key focus fields is retrieved from the medical knowledge graph. Then, based on the lesion type and relevant factor information, risk factors leading to the lesion type are identified. In practical applications, the natural language input by the user may only contain partial information. Therefore, to more comprehensively retrieve the medical record information required by the user, key clinical information related to the key focus fields can be expanded upon.

[0110] In one optional embodiment, the specific implementation of searching for relevant factor information corresponding to a key focus field from a medical knowledge graph can be as follows: A query statement is generated based on the key focus field. This query statement is then used to search for relevant factor information corresponding to the key focus field from the medical knowledge graph. Based on multiple medical entity terms and relevant factor information, structured search suggestion text is generated.

[0111] For example, for the key focus field "lung mass", you can use a SPARQL query to find related information such as symptoms, signs, examinations, and treatments.

[0112] SELECT ?rel ?factor WHERE {:lung mass?rel ?factor .?rel rdfs:subPropertyOf* :related factors}

[0113] In this query, "lung mass" represents the primary medical entity of interest, and "related factors" represents a high-level abstract relationship type. The query results will return all clinical factors related to "lung mass" and their corresponding specific relationships. Finally, the lesion type and risk factors are filled into a preset search suggestion template to generate search suggestion text.

[0114] In this embodiment, multiple medical entity terms included in the imaging description text are identified, and target entity nodes corresponding to each medical entity term are found from a medical knowledge graph. Based on the number of directly connected edges of each target entity node in the medical knowledge graph, the degree centrality of each target entity node in the medical knowledge graph is determined. Based on the degree centrality of each target entity node, the key focus fields corresponding to the imaging description text are determined. Based on the key focus fields, relevant factor information corresponding to the key focus fields is searched from the medical knowledge graph. Based on multiple medical entity terms and relevant factor information, risk factors leading to the lesion type are determined. Then, structured search suggestion text is generated based on the risk factors of the lesion type. This enriches the relevant information in the search suggestion text, improving the retrieval of information earlier and enabling a more comprehensive retrieval of relevant medical record information from the target object's electronic medical record text for further analysis of the imaging report by the user.

[0115] The following provides a detailed description of one possible method for filtering multiple candidate medical record fragments in the above embodiments.

[0116] In practical applications, the electronic medical record text corresponding to the target object is retrieved from the medical record database, and the electronic medical record text is segmented to obtain multiple medical record fragments. The multiple medical record fragments and the search prompt text are input into a pre-trained text retrieval model to obtain multiple candidate medical record fragments that match the search prompt text.

[0117] This involves pre-training a text retrieval model for semantic relevance filtering. This pre-trained model can then be used to filter multiple medical record fragments to identify candidate fragments that semantically match the search prompt text. Specifically, after obtaining multiple medical record fragments, these fragments and the search prompt text are input into the pre-trained text retrieval model to obtain multiple candidate medical record fragments that semantically match the search prompt text.

[0118] In the text retrieval model, the retrieval prompt text is semantically encoded to obtain the first semantic feature vector corresponding to the retrieval prompt text. Multiple medical record fragments are semantically encoded to obtain the second semantic feature vectors corresponding to each medical record fragment. The similarity between the second semantic feature vector corresponding to each medical record fragment and the first semantic feature vector corresponding to the retrieval prompt text is calculated. Based on a preset similarity threshold, multiple candidate medical record fragments are selected from the multiple medical record fragments.

[0119] In this embodiment of the application, a text retrieval model is used to initially screen multiple medical record fragments. This allows for more accurate selection of multiple candidate medical record fragments that match the semantics of the retrieval prompt text from the multiple medical record fragments corresponding to the target object, thereby improving the accuracy of subsequent retrieval results.

[0120] In practical applications, when the electronic medical record text corresponding to the target object contains multiple medical record information related to the same problem, or when there is a large amount of related medical record information, if these multiple candidate medical record fragments are sent to the user simultaneously, the user will need to select the required relevant information from the multiple candidate medical record fragments again, which will affect the user's work efficiency. Therefore, in this embodiment, after filtering out multiple candidate medical record fragments that match the semantics of the search prompt text, the importance value corresponding to each candidate medical record fragment can be further determined. Then, based on the importance value corresponding to each candidate medical record fragment, the multiple target medical record fragments with the highest matching degree with the search prompt text can be further filtered from the multiple candidate medical record fragments.

[0121] In one optional embodiment, the specific implementation process of determining the importance value corresponding to each candidate medical record segment using a medical knowledge graph may include: encoding each candidate medical record segment to obtain a second encoding vector corresponding to each candidate medical record segment; determining the weight value corresponding to each medical entity term in the target candidate medical record segment, where the target candidate medical record segment is any one of multiple candidate medical record segments; determining the association influence value (PageRank value) corresponding to each medical entity term in the target candidate medical record segment. The association influence value is used to characterize each medical entity term in the medical knowledge graph. Based on the weight value and PageRank value of each medical entity term in the target candidate medical record segment, the degree centrality score of the target candidate medical record segment is determined. Based on the number of preset high-risk medical entity terms contained in the target candidate medical record segment, the risk level score of the target candidate medical record segment is determined. Finally, based on the degree centrality score and the risk level score of the target candidate medical record segment, the importance value of the target candidate medical record segment is determined.

[0122] For example, in practical implementation, the degree centrality score can be calculated using the following formula:

[0123]

[0124] in, The degree centrality score corresponding to the candidate medical record segment. Candidate medical record fragments, It can be used to characterize candidate medical record fragments. In medical knowledge graphs Centrality measure in for The Middle A medical entity term, for The number of TCM entities For medical entities The weight (which changes with factors such as concept frequency and location), For medical entities PageRank value.

[0125] The risk level score can be calculated using the following formula:

[0126]

[0127] in, Used to represent candidate medical record fragments It can be used to characterize candidate medical record fragments. Does it contain the concept of high risk? For indicator functions, when The value is 1 if true, and 0 otherwise. This is a predefined set of high-risk concepts. The risk concept set is a collection of predefined emergency situations in radiology, including: aneurysm, hematoma, bleeding, hernia, dissection, corpus luteum rupture, etc. Once these occur, this information has the highest level of importance and should be seen first by the current diagnosing physician.

[0128] Combining the above two parts, candidate medical record fragments Importance level value:

[0129]

[0130] Among them, the two parameters wc and wk can be customized by the user. For example, in some hospitals that are emergency hospitals with many critically ill patients, the wk parameter can be appropriately reduced and the wc parameter can be increased to highlight the special information in the medical records and enable doctors to pay more attention to the patient's history. If the hospital is a general hospital with fewer critically ill patients and doctors are less vigilant about such patients, the wk parameter can be increased to make doctors pay more attention to critical patients.

[0131] The above calculations yielded a numerical score for the importance of each candidate medical record segment, reflecting its importance within the overall medical knowledge system, its relevance to imaging analysis conclusions, and the risk information it contains.

[0132] The process for determining the importance value of each candidate medical record segment is roughly the same. You can refer to the specific implementation method for determining the importance value of the target candidate medical record segment mentioned above, and there are no restrictions on this.

[0133] In response to the unique need for assessing the importance of clinical information in image analysis, in another optional embodiment, a mechanism specifically designed for assessing the importance of medical record segments discovered in images is designed. This mechanism determines the symptom correlation score corresponding to the target candidate medical record segment, and combines the symptom correlation score, the risk level score, and the symptom correlation score to determine the importance value of the target candidate medical record segment.

[0134] Specifically, firstly, the correlation between each entity term in the candidate medical record segment and the description of the imaging features is determined. Then, based on the weight values ​​corresponding to each medical entity term and the correlation between each entity term and the description of the imaging features, the feature correlation score corresponding to the target candidate medical record segment is determined. Based on the degree centrality score, risk level score, and feature correlation score corresponding to the target candidate medical record segment, the importance value corresponding to the target candidate medical record segment is determined. The weight values ​​are used to characterize the importance of the corresponding medical entity terms.

[0135] For example, in practical implementation, the symptom correlation score corresponding to the target candidate medical record segment can be calculated using the following formula:

[0136]

[0137] in, for The Middle A medical entity term, For medical entities The weight value, Represents medical entities Compared with current image features The correlation between these indicators is an assessment metric specifically designed for imaging diagnosis. For example, when the imaging description includes "spiculated signs," medical record segments related to "smoking history" will receive a higher correlation score.

[0138] Then, combining the above three components, candidate medical record fragments The specific importance value for imaging diagnosis is:

[0139]

[0140] in, , and These are weighted parameters that are dynamically adjusted based on different types of imaging examinations and clinical purposes. For example, for follow-up examinations of lung tumors, The weights will be set higher to emphasize clinical information relevant to the tumor. For trauma and emergency imaging assessments, The weights will be increased to highlight high-risk situations. This allows for weight adjustments based on specific scenes within the image.

[0141] Next, based on the importance values ​​of each of the multiple candidate medical record segments, multiple target medical record segments whose importance values ​​meet the preset values ​​are selected from the multiple candidate medical record segments.

[0142] In one optional embodiment, the specific implementation process of selecting multiple target medical record segments whose importance values ​​meet preset values ​​from multiple candidate medical record segments based on the importance values ​​corresponding to each of the multiple candidate medical record segments may include: inputting the importance values ​​corresponding to each of the multiple candidate medical record segments, the multiple candidate medical record segments, and the retrieval prompt text into a pre-trained segment ranking model to obtain multiple target medical record segments corresponding to the retrieval prompt text.

[0143] One approach is to pre-train a fragment ranking model. When ranking multiple candidate medical record fragments, this model can combine the importance values ​​and features of the multiple candidate medical record fragments to obtain the multiple target medical record fragments that have the highest matching degree with the search prompt text.

[0144] To facilitate understanding of the specific implementation process of the above electronic retrieval, an example based on a specific application scenario will be provided. For example... Figure 4As shown, in specific implementation, when a radiologist enters the imaging description text "A round nodule is seen in the upper lobe of the right lung, approximately 2.3cm × 1.8cm in size, with rough edges, lobulation and spiculation, and pleural traction" into the medical record retrieval interface, the electronic medical record retrieval device, after receiving the imaging text entered by the user, executes the following medical record retrieval processing flow specifically tailored to the imaging characteristics:

[0145] Step 401: Identify the image signature expressions in the image description text and extract the image signature descriptions from the image description text.

[0146] The identified signs included spatial location information ("right upper lobe"), measurement data ("2.3cm × 1.8cm"), morphological characteristics ("round or oval", "rough edges"), and imaging signs ("lobulation sign", "spiculated sign", "pleural traction"). The extracted imaging signs were described as "lobulation sign", "spiculated sign", and "pleural traction".

[0147] Step 402: Perform semantic analysis on the image feature descriptions to identify combinations of multiple image feature characteristics.

[0148] Specifically, the combination of "lobulation sign + spiculation sign + pleural traction" was identified.

[0149] Step 403: Based on the combination of multiple identified imaging features, predict the lesion type corresponding to the target object. Then, generate targeted search prompt text based on the lesion type.

[0150] Specifically, the system accesses a medical knowledge base to identify risk factors that lead to the specific type of lesion. The lesion type and risk factors are then entered into a pre-defined search suggestion template to generate search suggestion text. For example, the generated search suggestion text might be: "Suspicious malignant lung tumor, with particular attention to: ① smoking history ② occupational exposure history ③ previous cancer history ④ lung infection history ⑤ previous imaging records of the same site."

[0151] Step 404: Retrieve multiple candidate medical record fragments that match the search prompt text from the electronic medical record text corresponding to the target object.

[0152] Step 405: Determine the importance value corresponding to each candidate medical record segment. Based on the importance values ​​of multiple candidate medical record segments, select multiple target medical record segments whose importance values ​​meet preset values. The importance value characterizes the clinical impact of the candidate medical record segment on the imaging problems described in the imaging description text.

[0153] In the candidate medical record segment importance assessment phase, specific importance assessment parameters can be set for the imaging findings of pulmonary nodules, as follows: =0.3, =0.3, =0.4, where Clinical information particularly relevant to the current imaging findings was given special weighting. Based on the above parameters, a medical record segment containing the statement "The patient has a 30-year smoking history, one pack per day, and was hospitalized 5 years ago for a lung infection" was assessed for importance. Because this segment contained information about a "smoking history" highly correlated with the "spiculated sign + lobulation sign," its sign relevance score was [highly weighted]. The score was as high as 0.85, and the final importance value was calculated to be 0.76, which is significantly higher than the 0.42 score obtained by the general assessment method.

[0154] Step 406: Perform structuring processing on multiple target medical record fragments to obtain multiple structured medical record information, and display the multiple structured medical record information in the medical record retrieval interface.

[0155] The selected target medical record segments undergo structured image analysis, organizing the information into categories of specific interest to the image analysis, such as "risk factors," "previous relevant examinations," and "relevant treatment history," and highlighting information highly relevant to the current image features in the interface.

[0156] Figure 5 This is a schematic diagram of the structure of an electronic medical record retrieval device, provided as another exemplary embodiment of this application. (See diagram below.) Figure 5 As shown, the device includes: an extraction module 11, a prediction module 12, a generation module 13, a determination module 14, a filtering module 15, and a processing module 16.

[0157] Extraction module 11 is used to extract image feature descriptions from the image feature description text input by the user for the target object in the medical record retrieval interface, the image feature descriptions being used to describe the image manifestations corresponding to the lesion tissue.

[0158] The prediction module 12 is used to analyze the image feature description and predict the lesion type corresponding to the target object.

[0159] The generation module 13 is used to generate a search prompt text based on the lesion type, and to retrieve multiple candidate medical record fragments that match the search prompt text from multiple electronic medical record texts corresponding to the target object.

[0160] The determination module 14 is used to determine the importance value corresponding to each candidate medical record segment, wherein the importance value is used to characterize the degree of clinical impact of the candidate medical record segment on the imaging problem described in the imaging description text.

[0161] The filtering module 15 is used to filter out multiple target medical record segments whose importance values ​​meet preset values ​​from the multiple candidate medical record segments based on the importance values ​​corresponding to each of the multiple candidate medical record segments.

[0162] The processing module 16 is used to perform structured processing on the multiple target medical record fragments to obtain multiple structured medical record information, and to display the multiple structured medical record information in the medical record retrieval interface.

[0163] Optionally, the image sign description includes multiple image sign features; wherein, the prediction module 12 is specifically used to: if an entry corresponding to the combination of the multiple image sign features is retrieved from the image sign database, then the lesion type in the entry is determined as the lesion type corresponding to the target object; the image sign database stores the correspondence between various image sign combinations and their corresponding lesion types; if no entry corresponding to the combination of the multiple image sign features is retrieved from the image sign database, then the image description text is analyzed in conjunction with the medical knowledge graph and the lesion type described in the image description text to predict the lesion type corresponding to the target object.

[0164] Optionally, the prediction module 12 is specifically used to: retrieve knowledge fragments related to the lesion type from the medical knowledge graph; generate a first prompt word based on the knowledge fragment, the imaging description text, and the lesion type; input the first prompt word into a large language model to obtain the lesion type output by the large language model.

[0165] Optionally, the generation module 13 is specifically used to: call a medical knowledge base to determine the risk factors that lead to the lesion type; fill the lesion type and the risk factors into a preset search prompt template to generate search prompt text.

[0166] Optionally, the generation module 13 is specifically used for: identifying multiple medical entity words included in the imaging description text, wherein the medical entity words are named entity words with medical information; searching for target entity nodes corresponding to each of the multiple medical entity words from the medical knowledge graph; determining the degree centrality of each target entity node in the medical knowledge graph based on the number of directly connected edges of each target entity node in the medical knowledge graph; determining the key focus field corresponding to the imaging description text based on the degree centrality of each target entity node; searching for relevant factor information corresponding to the key focus field from the medical knowledge graph based on the key focus field; and determining the risk factors leading to the lesion type based on the lesion type and the relevant factor information.

[0167] Optionally, the generation module 13 is specifically used to: generate a query statement based on the key focus field; and use the query statement to search for relevant factor information corresponding to the key focus field from the medical knowledge graph.

[0168] Optionally, the determining module 14 is specifically configured to: encode each candidate medical record segment to obtain an encoding vector corresponding to each candidate medical record segment; for a target candidate medical record segment, identify each medical entity term contained in the target candidate medical record segment, wherein the target candidate medical record segment is any one of the plurality of candidate medical record segments; determine the weight value corresponding to each medical entity term, wherein the weight value is used to characterize the importance of the corresponding medical entity term; determine the association influence value corresponding to each medical entity term; wherein the association influence value is used to characterize the importance and association influence of each medical entity term in the medical knowledge graph; determine the degree centrality score corresponding to the target candidate medical record segment based on the weight value and the association influence value corresponding to each medical entity term; determine the risk level score corresponding to the target candidate medical record segment based on the number of preset high-risk medical entity terms contained in the target candidate medical record segment; and determine the importance value corresponding to the target candidate medical record segment based on the degree centrality score and the risk level score corresponding to the target candidate medical record segment.

[0169] Optionally, the determining module 14 is specifically used to: determine the correlation degree between each entity word and the image sign description, wherein the correlation degree is used to characterize the degree of association between the entity information corresponding to the entity word and the image sign description; determine the sign correlation score corresponding to the target candidate medical record segment based on the weight value corresponding to each medical entity word and the correlation degree between each entity word and the image sign description; and determine the importance value corresponding to the target candidate medical record segment based on the degree centrality score, the risk level score, and the sign correlation score corresponding to the target candidate medical record segment.

[0170] Optionally, the generation module 13 is specifically configured to: retrieve electronic medical record text corresponding to the target object from the medical record database, and segment the electronic medical record text to obtain multiple medical record fragments; input the multiple medical record fragments and the search prompt text into a pre-trained text retrieval model to obtain multiple candidate medical record fragments that match the search prompt text; wherein, in the text retrieval model, the search prompt text is semantically encoded to obtain a first semantic feature vector corresponding to the search prompt text, the multiple medical record fragments are semantically encoded to obtain a second semantic feature vector corresponding to each of the multiple medical record fragments, the similarity between the second semantic feature vector corresponding to each medical record fragment and the first semantic feature vector corresponding to the search prompt text is calculated, and multiple candidate medical record fragments are selected from the multiple medical record fragments according to a preset similarity threshold.

[0171] Optionally, before extracting the image sign description from the image description text, the extraction module 11 is further configured to: rewrite the image description text to obtain a structured image description text; wherein, extracting the image sign description from the image description text includes: generating a second prompt word based on the structured image description text; inputting the second prompt word into a large language model so that the large language model can determine the lesion type corresponding to the lesion tissue described in the structured image description text by analyzing the structured image description text; and extracting the image sign description related to the lesion type from the structured image description text.

[0172] Optionally, the extraction module 11 is specifically used for: inputting the imaging description text into a pre-trained text rewriting model to obtain standardized entity words corresponding to each entity word in the imaging description text; generating structured imaging description text based on the rewritten text corresponding to each entity word; wherein, in the text rewriting model, each entity word in the imaging description text is encoded to obtain word vectors corresponding to each entity word in the imaging description text, the word vectors corresponding to each entity word, the medical entity embedding vectors corresponding to each entity word, and the relation embedding vectors between medical entities corresponding to each entity word are fused to obtain the fused vector representations corresponding to each entity word, and the fused vector representations corresponding to each entity word are decoded to obtain the rewritten text corresponding to each entity word; the medical entity embedding vectors are used to represent the semantic information of each entity word in the medical knowledge graph; the relation embedding vectors are used to represent the association relationships between medical entities corresponding to each entity word.

[0173] Optionally, the extraction module 11 is specifically used to: input the target candidate medical record fragment into a pre-trained medical entity recognition model to obtain multiple medical entity words, wherein the medical entity recognition model is used to identify the medical entity words included in the descriptive text.

[0174] Optionally, the filtering module 15 is specifically used to: input the importance values ​​corresponding to the plurality of candidate medical record segments, the plurality of candidate medical record segments, and the search prompt text into a pre-trained segment ranking model to obtain a plurality of target medical record segments that have the highest matching degree with the search prompt text.

[0175] Optionally, the processing module 16 is specifically used to: sort the multiple structured medical record information according to a preset imaging assessment priority rule to obtain multiple sorted structured medical record information; and sequentially display the multiple sorted structured medical record information in the medical record retrieval interface.

[0176] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device includes: a memory 21 and a processor 22; wherein,

[0177] The memory 21 is used to store programs;

[0178] The processor 22, coupled to the memory, is configured to execute the program stored in the memory for:

[0179] In response to the imaging description text entered by the user for the target object in the medical record retrieval interface, the imaging sign description is extracted from the imaging description text, and the imaging sign description is used to describe the imaging manifestations of the lesion tissue.

[0180] The image features are analyzed to predict the lesion type corresponding to the target object;

[0181] Based on the lesion type, a search prompt text is generated, and multiple candidate medical record fragments that match the search prompt text are retrieved from multiple electronic medical record texts corresponding to the target object.

[0182] Determine the importance value corresponding to each candidate medical record segment, the importance value being used to characterize the degree of clinical impact of the candidate medical record segment on the imaging problems described in the imaging description text;

[0183] Based on the importance values ​​corresponding to each of the multiple candidate medical record segments, multiple target medical record segments whose importance values ​​meet preset values ​​are selected from the multiple candidate medical record segments;

[0184] The multiple target medical record fragments are processed into structured medical record information to obtain multiple structured medical record information, which is then displayed in the medical record retrieval interface.

[0185] The aforementioned memory 21 can be configured to store various other data to support operation on a computing device. Examples of such data include instructions for any application or method used to operate on a computing device. Memory 21 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0186] When the processor 22 executes the program in the memory 21, in addition to the functions described above, it can also perform other functions, as detailed in the descriptions of the preceding embodiments.

[0187] Furthermore, such as Figure 6 As shown, the electronic device also includes: display 23, power supply component 24, communication component 25, and other components. Figure 6 The diagram only shows some components and does not imply that the electronic device includes only these components. Figure 6 The components shown.

[0188] Accordingly, embodiments of this application also provide a readable storage medium storing a computer program, which, when executed by a computer, can implement the steps or functions of the electronic medical record retrieval method provided in the above embodiments.

[0189] Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, enable the processor to implement the steps in the above method embodiments.

[0190] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0191] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An electronic medical record retrieval method, characterized in that, include: In response to the imaging description text entered by the user for the target object in the medical record retrieval interface, the imaging sign description is extracted from the imaging description text. The imaging sign description is used to describe the imaging manifestations of the lesion tissue and includes multiple imaging sign features. If an entry corresponding to the combination of the multiple image features is retrieved from the image feature database, the lesion type in the entry is determined as the lesion type corresponding to the target object; the image feature database stores the correspondence between various image feature combinations and their corresponding lesion types. If no table entry corresponding to the combination of the multiple image features is retrieved from the image feature database, the image description text is analyzed in conjunction with the medical knowledge graph and the lesion type described in the image description text to predict the lesion type corresponding to the target object. Identify multiple medical entity words included in the imaging description text, wherein the medical entity words are named entity words with medical information; From the medical knowledge graph, find the target entity node corresponding to each of the multiple medical entity terms; The degree centrality of each target entity node in the medical knowledge graph is determined based on the number of directly connected edges in the medical knowledge graph. Based on the degree centrality of each target entity node, the key focus fields corresponding to the image description text are determined; Based on the key focus fields, relevant factor information corresponding to the key focus fields is retrieved from the medical knowledge graph; Based on the lesion type and the related factor information, the risk factors leading to the lesion type are identified; The disease type and risk factors are filled into a preset search prompt template to generate search prompt text, and multiple candidate medical record fragments that match the search prompt text are retrieved from the electronic medical record text corresponding to the target object. Determine the importance value corresponding to each candidate medical record segment, the importance value being used to characterize the degree of clinical impact of the candidate medical record segment on the imaging problems described in the imaging description text; Based on the importance values ​​corresponding to each of the multiple candidate medical record segments, multiple target medical record segments whose importance values ​​meet preset values ​​are selected from the multiple candidate medical record segments; The multiple target medical record fragments are processed into structured medical record information to obtain multiple structured medical record information, which is then displayed in the medical record retrieval interface.

2. The method according to claim 1, characterized in that, The analysis of the imaging description text, combining medical knowledge graphs and lesion types, to predict the lesion type corresponding to the target object includes: Retrieve knowledge fragments related to the lesion type from the medical knowledge graph; A first prompt word is generated based on the knowledge fragment, the imaging description text, and the lesion type; The first prompt word is input into the large language model to obtain the lesion type output by the large language model.

3. The method according to claim 1, characterized in that, The step of searching for relevant factor information corresponding to the key focus field from the medical knowledge graph includes: Based on the key fields of interest, generate a query statement; Use the query statement to find relevant factor information corresponding to the key focus field from the medical knowledge graph.

4. The method according to claim 1, characterized in that, Determining the importance value corresponding to each candidate medical record segment includes: Each candidate medical record segment is encoded to obtain an encoding vector corresponding to each candidate medical record segment; For a target candidate medical record segment, identify each medical entity word contained in the target candidate medical record segment, wherein the target candidate medical record segment is any one of the plurality of candidate medical record segments; Determine the weight value corresponding to each medical entity term, whereby the weight value is used to characterize the importance of the corresponding medical entity term; Determine the association influence value corresponding to each medical entity term; the association influence value is used to characterize the importance and association influence of each medical entity term in the medical knowledge graph. Based on the weight value and the association influence value of each medical entity term, the degree centrality score of the target candidate medical record segment is determined. Based on the number of preset high-risk medical entity words contained in the target candidate medical record segment, the risk level score corresponding to the target candidate medical record segment is determined; Based on the degree centrality score and the risk level score of the target candidate medical record segment, the importance value of the target candidate medical record segment is determined.

5. The method according to claim 4, characterized in that, The step of determining the importance value of the target candidate medical record segment based on the degree centrality score and the risk level score of the target candidate medical record segment includes: The correlation degree between each entity word and the image feature description is determined, and the correlation degree is used to characterize the degree of association between the entity information corresponding to the entity word and the image feature description; Based on the weight value corresponding to each medical entity term and the correlation between each entity term and the image sign description, the sign correlation score corresponding to the target candidate medical record segment is determined; The importance value of the target candidate medical record segment is determined based on the degree centrality score, the risk level score, and the symptom correlation score corresponding to the target candidate medical record segment.

6. The method according to claim 1, characterized in that, The step of retrieving multiple candidate medical record fragments from the electronic medical record text corresponding to the target object that match the search prompt text includes: Retrieve the electronic medical record text corresponding to the target object from the medical record database, and segment the electronic medical record text to obtain multiple medical record fragments; The multiple medical record fragments and the search suggestion text are input into a pre-trained text retrieval model to obtain multiple candidate medical record fragments that match the search suggestion text; In the text retrieval model, the retrieval prompt text is semantically encoded to obtain a first semantic feature vector corresponding to the retrieval prompt text. The multiple medical record fragments are semantically encoded to obtain a second semantic feature vector corresponding to each of the multiple medical record fragments. The similarity between the second semantic feature vector corresponding to each medical record fragment and the first semantic feature vector corresponding to the retrieval prompt text is calculated. Based on a preset similarity threshold, multiple candidate medical record fragments are selected from the multiple medical record fragments.

7. The method according to claim 1, characterized in that, Before extracting image feature descriptions from the image description text, the method further includes: The image description text is rewritten to obtain a structured image description text; The step of extracting image feature descriptions from the image description text includes: Based on the structured image description text, a second prompt word is generated; The second prompt word is input into the large language model so that the large language model can analyze the structured imaging description text to determine the lesion type corresponding to the lesion tissue described in the structured imaging description text. Extract the image feature descriptions related to the lesion type from the structured imaging description text.

8. The method according to claim 7, characterized in that, The process of rewriting the image description text to obtain a structured image description text includes: The image description text is input into a pre-trained text rewriting model to obtain standardized entity words corresponding to each entity word in the image description text; Based on the rewritten text corresponding to each entity word, a structured image description text is generated; In the text rewriting model, each entity word in the imaging description text is encoded to obtain the word vector corresponding to each entity word. The word vectors corresponding to each entity word, the medical entity embedding vectors corresponding to each entity word, and the relation embedding vectors between medical entities corresponding to each entity word are fused to obtain the fused vector representation corresponding to each entity word. The fused vector representation corresponding to each entity word is then decoded to obtain the rewritten text corresponding to each entity word. The medical entity embedding vector is used to represent the semantic information of each entity word in the medical knowledge graph; the relation embedding vector is used to represent the association relationship between the medical entities corresponding to each entity word.

9. The method according to claim 4, characterized in that, The identification of each medical entity term contained in the target candidate medical record fragment includes: The target candidate medical record fragment is input into a pre-trained medical entity recognition model to obtain multiple medical entity words. The medical entity recognition model is used to identify medical entity words included in the descriptive text.

10. The method according to claim 1, characterized in that, The step of selecting multiple target medical record segments whose importance values ​​meet preset values ​​from the multiple candidate medical record segments based on their respective importance values ​​includes: The importance values ​​of each of the multiple candidate medical record segments, the multiple candidate medical record segments, and the search prompt text are input into a pre-trained segment ranking model to obtain multiple target medical record segments that have the highest matching degree with the search prompt text.

11. The method according to claim 1, characterized in that, Displaying the multiple structured medical record information entries in the medical record retrieval interface includes: According to the preset imaging assessment priority rules, the multiple structured medical record information is sorted to obtain multiple sorted structured medical record information. The sorted structured medical records are displayed sequentially in the medical record retrieval interface.

12. An electronic device, characterized in that, include: Memory and processor; The memory is used to store a computer program; the processor, coupled to the memory, is used to execute the computer program to implement the steps of the electronic medical record retrieval method according to any one of claims 1-11.