Big model-based report interpretation method and device, equipment and medium
By using a large model-based approach to extract and interpret textual information from medical test reports, and combining a pre-set knowledge base and a report interpretation model, the problems of accuracy and scalability in report interpretation are solved, and efficient identification and interpretation of indicator information are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2024-11-11
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to guarantee accuracy and scalability in report interpretation, especially when processing medical test reports, as they cannot efficiently identify and interpret indicator information within the text.
A large model-based approach is adopted to extract text information from the target report image, identify indicator names and their indication information, retrieve knowledge of unidentified indicators using a pre-set knowledge base, and input the text information, indicator knowledge, and pre-set prompt text into the report interpretation model for interpretation.
While ensuring the accuracy of report interpretation, it improves the scalability and efficiency of report interpretation, especially in the automatic interpretation of medical test reports, where it can accurately identify and interpret indicator information.
Smart Images

Figure CN119559656B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of deep learning and natural language processing technology, and specifically to a report interpretation method, apparatus, electronic device, computer-readable storage medium, and computer program product based on a large model. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] With the rise of artificial intelligence technology, it has demonstrated remarkable capabilities in image recognition and text analysis. Related artificial intelligence technologies have shown potential to outperform traditional methods in multiple studies and clinical trials.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for interpreting reports based on a large model.
[0006] According to one aspect of this disclosure, a report interpretation method based on a large model is provided, comprising: extracting text information from a target report image; identifying indicator names and corresponding indication information in the text information, wherein the indication information is used to indicate whether the indicator corresponding to the indicator name is within a normal range; responding to a first indicator name in the text information containing an indicator name for which no indication information has been identified, retrieving the first indicator name from a preset knowledge base to obtain indicator knowledge corresponding to the first indicator name, wherein the indicator knowledge includes an indicator reference range; and inputting the text information, indicator knowledge, and preset prompt text into a report interpretation model, so that, guided by the preset prompt text, the report interpretation model generates report interpretation information of the target report image based on the text information and indicator knowledge.
[0007] According to another aspect of this disclosure, a report interpretation device based on a large model is provided, comprising: a first extraction unit configured to extract text information from a target report image; a first recognition unit configured to recognize an indicator name and corresponding indication information in the text information, the indication information indicating whether the indicator corresponding to the indicator name is within a normal range; a first retrieval unit configured to, in response to a first indicator name in the text information containing an indicator name for which no indication information was recognized, retrieve the first indicator name in a preset knowledge base to obtain indicator knowledge corresponding to the first indicator name, the indicator knowledge including an indicator reference range; and a first generation unit configured to input the text information, indicator knowledge, and preset prompt text into a report interpretation model, so that, guided by the preset prompt text, the report interpretation model generates report interpretation information of the target report image based on the text information and indicator knowledge.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned large-model-based report interpretation method.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above-described large-model-based report interpretation method.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the aforementioned report interpretation method based on a large model.
[0011] According to one or more embodiments of this disclosure, the scalability of report interpretation can be improved while ensuring the accuracy of report interpretation.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;
[0015] Figure 2 A flowchart illustrating a large-model-based report interpretation method according to an embodiment of this disclosure is shown;
[0016] Figure 3 A flowchart illustrating the identification of indicator names and indicator information according to embodiments of the present disclosure is shown;
[0017] Figure 4 A flowchart of a large-model-based report interpretation method according to an exemplary embodiment of the present disclosure is shown;
[0018] Figure 5 A structural block diagram of a large-model-based report interpretation apparatus according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0022] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0023] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0024] Figure 1A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0025] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the large model-based report interpretation method of this disclosure.
[0026] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0027] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0028] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to upload report images and send report interpretation requests. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to users through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0029] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0030] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0031] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0032] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0033] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.
[0034] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0035] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0036] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0037] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0038] According to embodiments of this disclosure, such as Figure 2 As shown, a report interpretation method based on a large model is provided, including: step S201, extracting text information from the target report image; step S202, identifying the indicator name and the corresponding indication information in the text information, the indication information being used to indicate whether the indicator corresponding to the indicator name is within the normal range; step S203, in response to a first indicator name in the text information containing an unidentified indication information, retrieving the first indicator name from a preset knowledge base to obtain the indicator knowledge corresponding to the first indicator name, the indicator knowledge including the indicator reference range; and step S204, inputting the text information, indicator knowledge, and preset prompt text into the report interpretation model, so that, guided by the preset prompt text, the report interpretation model generates report interpretation information for the target report image based on the text information and indicator knowledge.
[0039] Therefore, by recognizing the text information in the target report image, the indicator names and corresponding instruction information are obtained. For the first indicator name where no instruction information is recognized, indicator knowledge is retrieved. The text information, indicator knowledge, and preset prompt text are then input into the report interpretation model. Guided by the preset prompt text, the model analyzes the text information and indicator knowledge to generate report interpretation information. This approach ensures the accuracy of report interpretation while improving its scalability.
[0040] In some embodiments, the target report image may be an image of a medical test report uploaded by the user through a client.
[0041] In some embodiments, a user can send a report interpretation request to the server at the same time as uploading the image. After receiving the report interpretation request and the report image, the server can interpret the report using the large model-based report interpretation method described above.
[0042] In some embodiments, text information in a target report image can be extracted using optical character recognition (OCR) technology.
[0043] In some embodiments, the target report image may be preprocessed (e.g., image enhancement) before text information extraction to improve the accuracy of text information extraction.
[0044] In some embodiments, text information can be extracted by combining operations such as layout analysis, heading level recognition, and table recognition.
[0045] Specifically, through layout analysis, this method analyzes and understands the structure of documents such as medical test reports. It breaks down the content layout and identifies text fragments in different areas through document segmentation, region classification, and line and character extraction. By recognizing headings and other hierarchical structures in medical test reports, it accurately identifies key content such as departmental information, contributing to improved accuracy in subsequent medical examination reports. Furthermore, it identifies tables through methods such as table region detection, row and column segmentation, cell segmentation and extraction, and table structure analysis, obtaining text fragments corresponding to each table region, further optimizing the accuracy of indicator extraction results in subsequent medical test reports.
[0046] Subsequently, based on this, the indicator names and their corresponding indication information in the text information can be identified. In some embodiments, the indication information is used to indicate whether the indicator corresponding to the indicator name is within the normal range. The indication information can be, for example, a reference range of the normal value of an indicator, or an abnormal indicator marker used to identify whether the indicator is abnormal, such as "+", "-", "↑", "↓", etc.
[0047] In some embodiments, the indicator name and indication information in the text information can be identified by natural language processing technology (such as named entity recognition technology), and the correspondence between the indicator name and indication information can be determined by judging whether the indicator name and indication information belong to the same text segment.
[0048] If the text information includes one or more first indicator names for which no corresponding indication information has been identified, relevant indicator knowledge can be retrieved from the preset knowledge base for the first indicator name.
[0049] In some embodiments, the preset knowledge base can be a knowledge graph or a database, which pre-stores information such as the indicator name, synonyms of the indicator name, reference range, unit, clinical significance, and knowledge reference source for multiple indicators.
[0050] In some embodiments, retrieving relevant indicator knowledge for a first indicator name in a preset knowledge base can be done by performing text matching on a pre-stored indicator name and its synonyms in the preset knowledge base based on the first indicator name, thereby obtaining the indicator knowledge corresponding to the first indicator name in the preset knowledge base.
[0051] In some embodiments, the retrieval of relevant indicator knowledge for the first indicator name in the preset knowledge base can be based on semantic similarity. The preset indicator name with a semantic similarity greater than the preset similarity to the first indicator name is matched in the preset knowledge base, thereby obtaining the corresponding indicator knowledge.
[0052] In some embodiments, retrieving relevant indicator knowledge for a first indicator name in a preset knowledge base can first involve performing precise text matching on pre-stored indicator names and their synonyms in the preset knowledge base based on the first indicator name. If no corresponding indicator is found through text matching, then, based on semantic similarity, preset indicator names in the preset knowledge base with a semantic similarity greater than a preset similarity are matched to obtain the corresponding indicator knowledge. This makes the obtained indicator knowledge more accurate.
[0053] In some embodiments, while identifying the indicator name and indication information, natural language processing techniques (such as named entity recognition) can be used to further identify information such as report names and sample sources in the text information, and this information can also be used as query terms when performing indicator retrieval. This makes the acquired indicator knowledge more accurate.
[0054] In some embodiments, indicator knowledge may include the indicator reference range. In some embodiments, indicator knowledge may also include information such as the report to which the indicator belongs, its clinical significance, and the source of the knowledge reference.
[0055] In some embodiments, after obtaining the indicator knowledge corresponding to the first indicator name for all unidentified indication information, the text information, indicator knowledge, and preset prompt text are input into the report interpretation model. Guided by the preset prompt text, the report interpretation model extracts the indicator name, indicator value, and reference range from the text information and indicator knowledge, and analyzes and interprets the extracted information of each indicator in conjunction with the relevant indicator knowledge to generate report interpretation information.
[0056] In some embodiments, the above-mentioned report interpretation model can be a general large language model.
[0057] In some embodiments, the aforementioned report interpretation model can be a large language model obtained through supervised fine-tuning training based on a general large language model. Each sample data point may include textual information from the sample report, relevant indicator knowledge, and sample report interpretation information.
[0058] In some embodiments, preset prompt text can be used to prompt the report interpretation model to extract information such as indicator names, indicator values, indication information, and reference ranges from text information and indicator knowledge, and to analyze and interpret the extracted information of each indicator in combination with relevant indicator knowledge to generate report interpretation information.
[0059] In some embodiments, if all the extracted indicator names in the text information contain corresponding indication information, the text information and the preset prompt text can be directly input into the report interpretation model to obtain the report interpretation information generated by the report interpretation model.
[0060] In some embodiments, such as Figure 3 As shown, identifying the indicator name and the corresponding indication information in the text information may include: step S301, performing named entity recognition on the text information to obtain the indicator name; step S302, matching the text fragments containing the indicator name in the text information based on a preset regular expression to identify the reference interval information corresponding to the indicator name; and step S303, in response to identifying the reference interval information, determining the reference interval information as the indication information.
[0061] In some embodiments, identifying the indicator names and corresponding indication information in the text information can be achieved by first performing named entity recognition on the text information to obtain all indicator names in the text information, and then determining the text segment where each indicator name is located based on each indicator name (e.g., a line of text in the detection report corresponding to the indicator name), and matching the text segment based on a preset regular expression to obtain the reference interval information corresponding to the corresponding indicator name.
[0062] In some embodiments, the regular expression may include text formats of indicator names and reference interval information, used to match text formats with the same format in each text segment, and to determine the text information corresponding to the reference interval information portion as the reference interval information corresponding to the corresponding indicator name.
[0063] Therefore, the names of various indicators in the text information are first obtained through named entity recognition; then, based on the preset regular expression, the text segments containing each indicator name are matched, so as to accurately and efficiently identify the reference interval information corresponding to each indicator name.
[0064] In some embodiments, if a certain indicator name cannot be matched with reference interval information through regular expressions, the corresponding indicator knowledge can be obtained by searching the preset knowledge base for that indicator name.
[0065] In some embodiments, identifying the indicator name and the corresponding indication information in the text information may further include: in response to the absence of reference interval information, performing semantic understanding on the text fragment to identify the abnormal indicator marker corresponding to the indicator name; and in response to the identification of the abnormal indicator marker, determining the abnormal indicator marker as indication information.
[0066] In some embodiments, if a certain indicator name is not matched with reference interval information by regular expression, the text segment containing the indicator name can be further semantically understood and analyzed based on semantic understanding to determine whether the text segment contains abnormal indicator markers, such as "+", "-", "↑", "↓", etc.
[0067] Therefore, when the reference interval information is not identified, the text fragment can be further semantically understood to determine whether it contains an abnormal indicator marker corresponding to the indicator name. This allows for accurate and comprehensive acquisition of the original indication information of each indicator in the report, improving the accuracy of subsequent report interpretation.
[0068] In some embodiments, when a certain indicator name has neither its corresponding reference range information nor its corresponding abnormal indicator marker identified, a search can be performed on the indicator name in a preset knowledge base to obtain its corresponding indicator knowledge.
[0069] In some embodiments, acquiring the target report image may include: in response to receiving a first image, classifying the first image to determine the image category of the first image; and in response to determining that the image category of the first image is a report category, identifying the first image as the target report image.
[0070] The image classification described above can be achieved based on a trained image classification model. This model can categorize input images into report categories and non-report categories (e.g., landscape images, handwritten note images, etc.), and can be trained on multiple sample images that have been labeled with image categories.
[0071] Therefore, in response to receiving the first report image uploaded by the user, it first determines whether the image is a report type image, thereby quickly filtering out interfering images and avoiding the waste of resources caused by subsequent processing of interfering images.
[0072] In some embodiments, determining the first image as the target report image in response to determining that the image category of the first image is a report category may include: extracting text from the first image to obtain first text in the first image in response to determining that the image category of the first image is a report category; and determining the first image as the target report image in response to determining that the first text does not contain preset keywords.
[0073] One approach is to preset some keywords and then filter out report images containing those keywords by performing keyword matching.
[0074] In some embodiments, the preset keywords can be keywords such as "discharge", "admission", and "medical record" used to indicate the type of report. These keywords can be used to quickly filter out redundant reports that do not need to be interpreted.
[0075] In some embodiments, the preset keywords can be specific keywords from reports such as electrocardiogram reports, hearing tests, and psychological tests that are not suitable for machine interpretation. These keywords can be used to quickly filter out reports that are not suitable for machine interpretation.
[0076] Therefore, by further recognizing the text in the report image and further filtering the report image through keyword matching, interfering document images or reports that are not suitable for machine interpretation are removed. This avoids the waste of resources caused by subsequent processing of the above images and further improves the accuracy of report interpretation.
[0077] In some embodiments, in response to the identification of an image that is not a report type or a report image containing the keywords mentioned above, subsequent information extraction and report interpretation operations will no longer be performed on the report image, and the user will be given corresponding prompts, such as, "Image X and Image X do not appear to be detection reports. You can try changing the images or skipping to continue interpretation."
[0078] In some embodiments, a user can upload multiple first images at once through a client. If one or more target report images are obtained after recognizing the multiple first images, subsequent report interpretation operations can be performed based solely on the recognized one or more target report images.
[0079] In some embodiments, the preset prompt text may include a first prompt text and a second prompt text. The report interpretation information generated by the report interpretation model based on the text information and indicator knowledge, under the guidance of the preset prompt text, may include: obtaining abnormal indicator information based on the text information and indicator knowledge under the guidance of the first prompt text; and generating report interpretation information based on at least the abnormal indicator information under the guidance of the second prompt text.
[0080] In some embodiments, the preset prompt text may include prompt text corresponding to multiple steps, such as a first prompt text corresponding to the abnormal indicator information extraction step and a second prompt text corresponding to the report interpretation information generation step. For example, the first prompt text may be "Please extract the abnormal indicators in the <Report Information> and the corresponding values and reference ranges of the indicators based on the <Report Information> and <Medical Reference Knowledge>"; the second prompt text may be "Provide a professional medical interpretation of the extracted abnormal items / locations." Here, the aforementioned report information corresponds to the text information in the target report image, and the aforementioned medical reference knowledge corresponds to indicator knowledge.
[0081] Therefore, by setting preset prompt text, the report interpretation model can extract abnormal indicator information step by step, and generate report interpretation information based on the abnormal indicator information, thereby further improving the accuracy of report interpretation, reducing the amount of data processing in the generation of report interpretation information, and improving the efficiency of report interpretation information generation.
[0082] In some embodiments, the first prompt text can be further configured to guide the report interpretation model to prioritize the use of textual information in the target report image for abnormal indicator identification, and then to identify abnormal indicators based on reference interval information in the indicator knowledge. For example, the first prompt text could be: "When the <report information> contains a normal reference interval, and the detection result of a certain examination item or site is not within this interval, this examination item or site can be marked as 'abnormal item / site'; if the <report information> does not provide a normal reference interval, it is matched with the reference range in <medical reference knowledge>, and examination items or sites not within this reference interval are marked as 'abnormal item / site'." Thus, by prioritizing the use of the original information in the target report image, the accuracy of report interpretation can be improved.
[0083] In some embodiments, the preset prompt text may further include a third prompt text. The report interpretation information, generated by the report interpretation model based on the text information and indicator knowledge under the guidance of the preset prompt text, may further include: obtaining user age information from the text information under the guidance of the third prompt text; and wherein, obtaining abnormal indicator information based on the text information and indicator knowledge under the guidance of the report interpretation model and the first prompt text may include: obtaining abnormal indicator information based on the text information, indicator knowledge, and user age information under the guidance of the first prompt text and the report interpretation model.
[0084] In some embodiments, the preset prompt text may sequentially include a third prompt text, a first prompt text, and a second prompt text. The third prompt text can be used to guide the report interpretation model to extract user age information from the text information; for example, the third prompt text could be "Extract user age from <report information>".
[0085] In some embodiments, the first prompt text can be further configured to guide the report interpretation model to extract abnormal indicator information based on different information according to the user's age information.
[0086] In some exemplary embodiments, the first prompt text can be further configured as follows: "When the <Report Information> contains a normal reference range, and the test result of a certain examination item or site is not within this range, the examination item or site can be marked as 'abnormal item / site'; if the <Report Information> does not provide a normal reference range, determine whether the examinee's age is less than 18 years old. If the examinee is less than 18 years old, end the current step and directly execute the next step; if the examinee's age is not less than 18 years old, match it with the reference range in <Medical Reference Knowledge>, and mark the examination item or site that is not within this reference range as 'abnormal item / site'."
[0087] Therefore, by setting the preset prompt text, the report interpretation model can further extract user age information and extract abnormal indicator information based on the user age information, thereby further improving the accuracy of abnormal indicator information extraction and thus improving the accuracy of report interpretation.
[0088] In some embodiments, the preset prompt text may further include prompt text to guide the report interpretation model in extracting the report name and user gender information. Furthermore, the first prompt text may be further configured to, when the text information does not provide a normal reference range, obtain the reference range information of the relevant indicator from the indicator knowledge based on information such as the report name, user age, and user gender, and determine whether the corresponding indicator is an abnormal indicator.
[0089] In some exemplary embodiments, the first prompt text may be: "When the <Report Information> contains a normal reference range, and the test result of a certain examination item or site is not within this range, this examination item or site can be marked as 'abnormal item / site'; if the <Report Information> does not provide a normal reference range, determine whether the examinee's age is less than 18 years old. If the examinee is less than 18 years old, end the current step and directly proceed to the next step; if the examinee's age is not less than 18 years old, use the examination item name (report name), age, gender, etc. in the <Report Information> to match the reference range in <Medical Reference Knowledge>, and mark the examination item or site that is not within this reference range as 'abnormal item / site'."
[0090] In some embodiments, the prompt text used to guide the report interpretation model in extracting the report name may further include prompt text to guide the report interpretation model to refine the report name by referring to indicator knowledge. For example, it could be "Describe the report name in as much detail as possible, such as specifying it as blood routine, urine routine, etc., and you can refer to <Medical Reference Knowledge> for further refinement."
[0091] In some embodiments, the preset prompt text may further include a fourth prompt text. The report interpretation information generated by the report interpretation model based on the text information and indicator knowledge, guided by the preset prompt text, may further include: determining, guided by the fourth prompt text, whether the text information contains interpretable report information; and wherein, obtaining abnormal indicator information based on the text information and indicator knowledge, guided by the first prompt text, through the report interpretation model, may include: in response to determining that the text information contains report information, obtaining abnormal indicator information based on the report information and indicator knowledge, guided by the first prompt text, through the report interpretation model.
[0092] In some embodiments, the preset prompt text may sequentially include a fourth prompt text, a third prompt text, a first prompt text, and a second prompt text. The fourth prompt text can be used to guide the report interpretation model to determine whether the text information contains interpretable report information.
[0093] In some embodiments, interpretable report information may include, for example, pathology reports, radiological examination results, MRI results, and diagnostic conclusions of preset examination items. The preset examination items are used to determine the range of reports suitable for interpretation by the report interpretation model; for example, electrocardiogram reports, hearing tests, and psychological tests are not within this range. In some exemplary embodiments, the fourth prompt text may include "Determine whether the <report information> contains medical test data and results of the <preset examination items>. If so, continue to the next step."
[0094] Therefore, by setting preset prompt text, the report interpretation model can determine whether the current target report is interpretable before information extraction. Only when interpretable report information is included can subsequent information extraction and interpretation be performed, thereby further improving the accuracy of report interpretation.
[0095] In some embodiments, the fourth prompt text may also include, "If there is no specific medical test data and results in the <Report Information>, please respond with: 'Sorry, I can only recognize medical test reports. You can try uploading a medical report again.' and then do not continue with the following steps." This guides the report interpretation model to send the corresponding prompt to the user when it recognizes that the text information does not contain interpretable report information, and to stop proceeding with subsequent report interpretation steps. This further improves the user experience while avoiding resource waste.
[0096] In some embodiments, the preset prompt text may also include prompt text to guide the report interpretation model in generating health advice and recommended departments. It is understood that those skilled in the art can customize the prompt content in the preset prompt text according to actual needs, and no restrictions are imposed here.
[0097] Figure 4 A flowchart of a large-model-based report interpretation method according to an exemplary embodiment of this disclosure is shown.
[0098] In some exemplary embodiments, such as Figure 4 As shown, the report interpretation method based on the large model may include: Step S401, determining whether the input first image is the target report image; Step S402, in response to the first image not being the target report image, ending the report interpretation and providing corresponding prompts to the user; Step S403, in response to the first image being the target report image, extracting report information from the text information of the target report image, wherein the report information includes the indicator name, indicator indication information, report name, user age, etc.; Step S404, for the first indicator name that does not contain corresponding indication information in the text information of the target report image, retrieving indicator knowledge corresponding to the first indicator name from the preset knowledge base; Step S405, inputting the text information of the target report image, the retrieved indicator knowledge, and the preset prompt text into the report interpretation model, so that the report interpretation model first determines whether the text information contains interpretable report information; Step S406, in response to the text information not containing interpretable report information, ending the report interpretation and providing corresponding prompts to the user; Step S407, in response to the text information containing interpretable report information, generating report interpretation information based on the text information and indicator knowledge.
[0099] In some embodiments, such as Figure 5As shown, a report interpretation device 500 based on a large model is also provided, including: a first extraction unit 510 configured to extract text information from a target report image; a first recognition unit 520 configured to recognize an indicator name and corresponding indication information in the text information, the indication information being used to indicate whether the indicator corresponding to the indicator name is within a normal range; a first retrieval unit 530 configured to, in response to a first indicator name in the text information containing an indicator name for which no indication information was recognized, retrieve the first indicator name in a preset knowledge base to obtain indicator knowledge corresponding to the first indicator name, the indicator knowledge including an indicator reference range; and a first generation unit 540 configured to input the text information, indicator knowledge, and preset prompt text into a report interpretation model, so that, guided by the preset prompt text, the report interpretation model generates report interpretation information of the target report image based on the text information and indicator knowledge.
[0100] in, Figure 5 Each unit of the device 500 shown can be connected to a reference. Figure 2 The steps described correspond to those in the large-model-based report interpretation method. Therefore, the operations, features, and advantages described above for the large-model-based report interpretation method also apply to apparatus 500 and its constituent units. For the sake of brevity, some operations, features, and advantages will not be repeated here.
[0101] In some embodiments, the first identification unit may include: a first acquisition subunit configured to perform named entity recognition on the text information to obtain the indicator name; a first identification subunit configured to match text segments containing the indicator name in the text information based on a preset regular expression to identify the reference interval information corresponding to the indicator name; and a first determination subunit configured to determine the reference interval information as indication information in response to the identification of the reference interval information.
[0102] In some embodiments, the first identification unit may further include: a second identification subunit configured to perform semantic understanding on the text fragment in response to the absence of reference interval information to identify an abnormal indicator marker corresponding to the indicator name; and a second determination subunit configured to determine the abnormal indicator marker as indication information in response to the identification of the abnormal indicator marker.
[0103] In some embodiments, acquiring the target report image may include: in response to receiving a first image, classifying the first image to determine the image category of the first image; and in response to determining that the image category of the first image is a report category, identifying the first image as the target report image.
[0104] In some embodiments, determining the first image as the target report image in response to determining that the image category of the first image is a report category may include: extracting text from the first image to obtain first text in the first image in response to determining that the image category of the first image is a report category; and determining the first image as the target report image in response to determining that the first text does not contain preset keywords.
[0105] In some embodiments, the preset prompt text may include a first prompt text and a second prompt text, and the first generation unit may include: a second acquisition subunit configured to acquire abnormal indicator information based on text information and indicator knowledge under the guidance of the first prompt text through a report interpretation model; and a first generation subunit configured to generate report interpretation information based on at least the abnormal indicator information under the guidance of the second prompt text through a report interpretation model.
[0106] In some embodiments, the preset prompt text may further include a third prompt text, and the first generation unit may further include: a third acquisition subunit, configured to acquire user age information in the text information under the guidance of the third prompt text through a report interpretation model; and wherein the second acquisition subunit may be further configured to: acquire abnormal indicator information based on the text information, indicator knowledge and user age information under the guidance of the first prompt text through a report interpretation model.
[0107] In some embodiments, the preset prompt text may further include a fourth prompt text, and the first generation unit may further include: a third determining subunit, configured to determine, under the guidance of the fourth prompt text, whether the text information contains interpretable report information through a report interpretation model; and wherein the second obtaining subunit may be further configured to: in response to determining that the text information contains report information, obtain abnormal indicator information based on the report information and indicator knowledge under the guidance of the first prompt text through a report interpretation model.
[0108] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0109] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0110] refer to Figure 6The present invention describes a structural block diagram of an electronic device 600 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0111] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0112] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, output unit 607, storage unit 608, and communication unit 609. Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 607 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 608 can include, but is not limited to, disk and optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0113] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the large model-based report interpretation method described above. For example, in some embodiments, the large model-based report interpretation method described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the large model-based report interpretation method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the above-described large model-based report interpretation method by any other suitable means (e.g., by means of firmware).
[0114] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0115] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0116] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0118] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0119] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0120] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0121] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A report interpretation method based on a large model, comprising: Extract text information from the target report image; Named entity recognition is performed on the text information to obtain the indicator names in the text information; Based on a preset regular expression, the text fragments containing the indicator name in the text information are matched to identify the reference interval information corresponding to the indicator name; In response to the identification of the reference interval information, the reference interval information is determined as the indication information corresponding to the indicator name, and the indication information is used to indicate whether the indicator corresponding to the indicator name is within the normal range; In response to the failure to identify the reference interval information, semantic understanding is performed on the text fragment to identify the abnormal indicator marker corresponding to the indicator name; In response to the identification of the abnormal indicator marker, the abnormal indicator marker is determined as the indication information; In response to the text information containing a first indicator name for which the indication information was not identified, the first indicator name is retrieved from a preset knowledge base to obtain indicator knowledge corresponding to the first indicator name, the indicator knowledge including an indicator reference range. as well as The text information, the indicator knowledge, and the preset prompt text are input into the report interpretation model. Guided by the preset prompt text, the report interpretation model generates report interpretation information for the target report image based on the text information and the indicator knowledge. The preset prompt text includes multiple prompt texts corresponding to multiple steps, enabling the report interpretation model to determine whether the text information contains interpretable report information according to the multiple prompt texts, extract anomaly indicators, and determine the report interpretation information of the target report image based on the anomaly indicators. The interpretable report information includes preset check items used to determine whether the report interpretation model is suitable for interpretation. Furthermore, the report interpretation model is guided to prioritize the use of the text information in the target report image to identify abnormal indicators in response to determining that the preset check item exists in the text information, and then to identify abnormal indicators based on the indicator reference range.
2. The method according to claim 1, wherein, The acquisition of the target report image includes: In response to receiving a first image, performing image classification on the first image to determine the image category of the first image; and In response to determining that the image category of the first image is a report category, the first image is identified as the target report image.
3. The method according to claim 2, wherein, The step of determining the first image as the target report image in response to determining that the image category of the first image is a report category includes: In response to determining that the image category of the first image is a report category, text extraction is performed on the first image to obtain the first text in the first image; and In response to determining that the first text does not contain preset keywords, the first image is identified as the target report image.
4. The method according to any one of claims 1 to 3, wherein, The preset prompt text includes a first prompt text and a second prompt text. The report interpretation information, generated by the report interpretation model based on the text information and the indicator knowledge under the guidance of the preset prompt text, includes: Through the report interpretation model, guided by the first prompt text, abnormal indicator information is obtained based on the text information and the indicator knowledge; and Guided by the second prompt text, the report interpretation model generates report interpretation information based at least on the abnormal indicator information.
5. The method according to claim 4, wherein, The preset prompt text also includes a third prompt text. The report interpretation information, generated by the report interpretation model based on the text information and the indicator knowledge under the guidance of the preset prompt text, further includes: Using the report interpretation model, guided by the third prompt text, the user's age information is obtained from the text information; and wherein, The step of obtaining abnormal indicator information through the report interpretation model, guided by the first prompt text, based on the text information and the indicator knowledge, includes: Guided by the first prompt text, the abnormal indicator information is obtained through the report interpretation model based on the text information, the indicator knowledge, and the user's age information.
6. The method according to claim 4, wherein, The preset prompt text also includes a fourth prompt text. The report interpretation information, generated by the report interpretation model based on the text information and the indicator knowledge under the guidance of the preset prompt text, further includes: Using the report interpretation model, guided by the fourth prompt text, it is determined whether the text information contains interpretable report information; and wherein, The step of obtaining abnormal indicator information through the report interpretation model, guided by the first prompt text, based on the text information and the indicator knowledge, includes: In response to determining that the text information contains the report information, the abnormal indicator information is obtained through the report interpretation model, guided by the first prompt text, based on the report information and the indicator knowledge.
7. A report interpretation device based on a large model, comprising: The first extraction unit is configured to extract text information from the target report image; The first acquisition subunit is configured to perform named entity recognition on the text information to obtain the indicator names in the text information. The first identification subunit is configured to match text segments containing the indicator name in the text information based on a preset regular expression, so as to identify the reference interval information corresponding to the indicator name; The first determining subunit is configured to determine the reference interval information as indication information corresponding to the indicator name in response to the recognition of the reference interval information, wherein the indication information is used to indicate whether the indicator corresponding to the indicator name is within the normal range; The second identification subunit is configured to perform semantic understanding on the text segment in response to the failure to identify the reference interval information, so as to identify the abnormal indicator marker corresponding to the indicator name; The second determining subunit is configured to determine the abnormal indicator marker as the indication information in response to the identification of the abnormal indicator marker; A first retrieval unit is configured to, in response to the text information containing a first indicator name for which the indication information was not identified, retrieve the first indicator name from a preset knowledge base to obtain indicator knowledge corresponding to the first indicator name, the indicator knowledge including an indicator reference range; and The first generation unit is configured to input the text information, the indicator knowledge, and the preset prompt text into a report interpretation model, so that, guided by the preset prompt text, the report interpretation model generates report interpretation information of the target report image based on the text information and the indicator knowledge. The preset prompt text includes multiple prompt texts corresponding to multiple steps, enabling the report interpretation model to determine whether the text information contains interpretable report information, extract abnormal indicators, and determine the report interpretation information of the target report image based on the abnormal indicators. The interpretable report information includes preset check items for determining suitable interpretation by the report interpretation model. Furthermore, the report interpretation model is guided to, in response to determining the presence of the preset check items in the text information, prioritize the use of text information in the target report image for abnormal indicator identification, and then perform abnormal indicator identification based on the indicator reference range.
8. The apparatus according to claim 7, wherein, The acquisition of the target report image includes: In response to receiving a first image, performing image classification on the first image to determine the image category of the first image; and In response to determining that the image category of the first image is a report category, the first image is identified as the target report image.
9. The apparatus according to claim 8, wherein, The step of determining the first image as the target report image in response to determining that the image category of the first image is a report category includes: In response to determining that the image category of the first image is a report category, text extraction is performed on the first image to obtain the first text in the first image; and In response to determining that the first text does not contain preset keywords, the first image is identified as the target report image.
10. The apparatus according to any one of claims 7 to 9, wherein, The preset prompt text includes a first prompt text and a second prompt text, and the first generation unit includes: The second acquisition subunit is configured to, through the report interpretation model and guided by the first prompt text, acquire abnormal indicator information based on the text information and the indicator knowledge; and The first generation subunit is configured to generate the report interpretation information based on at least the abnormal indicator information, guided by the second prompt text and through the report interpretation model.
11. The apparatus according to claim 10, wherein, The preset prompt text further includes a third prompt text, and the first generation unit further includes: The third acquisition subunit is configured to, through the report interpretation model and guided by the third prompt text, acquire the user's age information from the text information; and wherein, The second acquisition subunit is further configured as follows: Guided by the first prompt text, the abnormal indicator information is obtained through the report interpretation model based on the text information, the indicator knowledge, and the user's age information.
12. The apparatus according to claim 10, wherein, The preset prompt text further includes a fourth prompt text, and the first generation unit further includes: The third determining subunit is configured to, under the guidance of the fourth prompt text and through the report interpretation model, determine whether the text information contains interpretable report information; and wherein, The second acquisition subunit is further configured as follows: In response to determining that the text information contains the report information, the abnormal indicator information is obtained through the report interpretation model, guided by the first prompt text, based on the report information and the indicator knowledge.
13. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.