Medical report generation model obtaining method and medical report generation method and device
By training a medical report generation model and utilizing the image features of sample images and reports, the problem of long diagnostic and analysis times for doctors has been solved, enabling fast and accurate medical report generation.
Patent Information
- Application Number
- CN202410575156.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-11-11
AI Technical Summary
Doctors spend a lot of time diagnosing and analyzing medical images and writing medical reports, resulting in low efficiency in generating medical reports.
By acquiring sample medical images and reports, the first processing network is used to extract image features of tissue structures, and the second processing network is trained to generate a medical report generation model for rapid medical report generation.
It improves the efficiency of medical report generation, ensures the accuracy and structure of reports, and reduces the workload of doctors.
Smart Images

Figure CN120932798A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method for obtaining a medical report generation model, a medical report generation method, and an apparatus. Background Technology
[0002] Medical reports are used to describe the diagnostic results of tissue structures in medical images. For example, a diagnostic report may include the pathological condition of the tissue structure, the location of the lesion, and the affected tissue structures. Generally, medical images of the internal tissue structures of an organism or a part of an organism are first obtained, and then diagnostic analysis of these images is performed to generate a medical report. However, doctors need to spend a significant amount of time diagnosing and analyzing medical images and writing medical reports, resulting in low efficiency in medical report generation. Summary of the Invention
[0003] This application provides a method for obtaining a medical report generation model, a method for generating a medical report, and an apparatus for generating a medical report. The medical report generation model can be obtained, and a medical report can be generated through the medical report generation model. The technical solution includes the following contents.
[0004] Firstly, a method for obtaining a medical report generation model is provided, the method comprising:
[0005] Acquire sample medical images and sample medical reports, wherein the sample medical reports are used to describe the diagnostic results of at least one sample tissue structure in the sample medical images;
[0006] The sample medical images and the first processing network include multiple first query features to determine the first image features of multiple tissue structures, the multiple tissue structures including the at least one sample tissue structure, each first query feature represents the visual information of a tissue structure, and each first image feature of a tissue structure represents the image content in the sample medical images related to the tissue structure.
[0007] Based on the sample medical report and the first image features of each tissue structure, the first processing network is trained to obtain the second processing network;
[0008] Based on the second processing network, a medical report generation model is obtained, which is used to generate the target medical report.
[0009] Secondly, a method for generating medical reports is provided, the method comprising:
[0010] Acquire target medical images;
[0011] The medical report generation model is invoked, and based on the target medical image and the multiple target query features included in the medical report generation model, multiple target features of tissue structures are determined. The medical report generation model is trained according to the method shown in the first aspect. Any target query feature is used to characterize the content characteristics of an image of a tissue structure, and any target feature of a tissue structure is used to characterize the image content in the target medical image related to the tissue structure.
[0012] The medical report generation model is invoked to generate a target medical report based on the target features of each tissue structure. The target medical report is used to describe the diagnostic results of at least one target tissue structure in the target medical image, and the multiple tissue structures include the at least one target tissue structure.
[0013] Thirdly, a device for acquiring a medical report generation model is provided, the device comprising:
[0014] An acquisition module is used to acquire sample medical images and sample medical reports, wherein the sample medical report is used to describe the diagnostic results of at least one sample tissue structure in the sample medical images;
[0015] The determination module is used to determine first image features of multiple tissue structures through the sample medical image and multiple first query features included in the first processing network. The multiple tissue structures include the at least one sample tissue structure. Each first query feature represents the visual information of a tissue structure, and each first image feature of a tissue structure represents the image content in the sample medical image related to the tissue structure.
[0016] The training module is used to train the first processing network based on the first image features of the sample medical report and the various tissue structures to obtain the second processing network;
[0017] The acquisition module is further configured to acquire a medical report generation model based on the second processing network, the medical report generation model being used to generate a target medical report.
[0018] In one possible implementation, the determining module is used to segment the sample medical image to obtain multiple sample image blocks; extract image features of each sample image block; and for any tissue structure, determine a first image feature of the tissue structure by using the image features of each sample image block and a first query feature of the tissue structure.
[0019] In one possible implementation, the determining module is configured to determine a first similarity between the image features of each sample image block and the first query feature of the arbitrary organizational structure; and based on the first similarity, extract the first image feature of the arbitrary organizational structure from the image features of each sample image block.
[0020] In one possible implementation, the training module is used to determine the first text features of each organizational structure through the sample medical report and multiple second query features included in the third processing network. Each second query feature represents the descriptive information of an organizational structure, and the first text features of any organizational structure represent the text content related to the organizational structure in the sample medical report. Based on the first text features and first image features of each organizational structure, the first processing network is trained to obtain the second processing network.
[0021] In one possible implementation, the training module is used to segment the sample medical report to obtain multiple sample text segments, each of which is used to describe the diagnostic result of a sample tissue structure; extract text features from each sample text segment; and determine the first text features of each tissue structure by using the text features of each sample text segment and the second query features of multiple tissue structures.
[0022] In one possible implementation, the training module is configured to, for any given organizational structure, determine a second similarity between the text features of each sample text segment and a second query feature of the given organizational structure; and, based on each second similarity, extract a first text feature of the given organizational structure from the text features of each sample text segment.
[0023] In one possible implementation, the training module is used to determine second image features of each organizational structure through a first momentum network, wherein the first momentum network has the same structure as the first processing network, and the network parameters of the first momentum network are obtained by updating the network parameters of the first processing network; to determine second text features of each organizational structure through a second momentum network, wherein the second momentum network has the same structure as the third processing network, and the network parameters of the second momentum network are obtained by updating the network parameters of the third processing network; and to train the first processing network based on the first text features, second text features, first image features, and second image features of each organizational structure to obtain the second processing network.
[0024] In one possible implementation, the training module is configured to: determine a third similarity between any first text feature and any second image feature; obtain first annotation information characterizing whether the first text feature and the second image feature correspond to the same organizational structure; determine a fourth similarity between any second text feature and any first image feature; obtain second annotation information characterizing whether the second text feature and the first image feature correspond to the same organizational structure; and train the first processing network based on multiple third similarities, multiple fourth similarities, multiple first annotation information, and multiple second annotation information to obtain a second processing network.
[0025] In one possible implementation, the training module is used to determine a masked medical report based on the sample medical report, the masked medical report being obtained by masking multiple first characters in the sample medical report; a fourth processing network determines the generation probability of each first character based on the masked medical report and the first image features of each tissue structure; and the first processing network is trained based on the generation probability of each first character to obtain a second processing network.
[0026] In one possible implementation, the masked medical report includes multiple masked text segments, which are obtained by masking the first character in a sample text segment, the sample text segment being used to describe the diagnostic result of a sample tissue structure;
[0027] The training module is used to extract text features of any masked text segment through a fourth processing network; determine the first image features of the sample organization structure corresponding to the masked text segment from the first image features of each organization structure through the fourth processing network; fuse the text features of the masked text segment and the first image features of the corresponding sample organization structure to obtain a first fused feature; and determine the generation probability of the first character in the masked text segment based on the first fused feature.
[0028] In one possible implementation, the acquisition module is configured to determine sample features of each tissue structure using the sample medical image and multiple third query features included in the second processing network, wherein any third query feature represents the visual information of a tissue structure, and any sample feature of a tissue structure represents the image content in the sample medical image related to the tissue structure; and, based on the sample features of each tissue structure, determine the generation probability of each second character in the sample medical report using a fifth processing network; and, based on the generation probability of each second character, train a neural network model to obtain a medical report generation model, wherein the neural network model includes the second processing network and the fifth processing network.
[0029] In one possible implementation, the acquisition module is configured to, for the first second character, determine the generation probability of the first second character based on the sample features of each organizational structure through a fifth processing network; and for non-first second characters, determine the generation probability of the non-first second character based on the sample features of each organizational structure and each second character preceding the non-first second character through the fifth processing network.
[0030] In one possible implementation, the acquisition module is configured to determine semantic features based on each second character preceding the non-first second character through the fifth processing network, the semantic features representing the semantic content represented by each second character preceding the non-first second character; fuse the semantic features and the sample features of each organizational structure through the fifth processing network to obtain a second fused feature; and determine the generation probability of the non-first second character based on the second fused feature through the fifth processing network.
[0031] Fourthly, a medical report generation device is provided, the device comprising:
[0032] The acquisition module is used to acquire the target medical image;
[0033] The determination module is used to call the medical report generation model, and determine the target features of multiple tissue structures based on the target medical image and the multiple target query features included in the medical report generation model. The medical report generation model is trained according to the method shown in the first aspect. Any target query feature is used to characterize the content characteristics of an image of a tissue structure, and any target feature of a tissue structure is used to characterize the image content in the target medical image related to the any tissue structure.
[0034] The generation module is used to call the medical report generation model to generate a target medical report based on the target features of each tissue structure. The target medical report is used to describe the diagnostic results of at least one target tissue structure in the target medical image, and the plurality of tissue structures includes the at least one target tissue structure.
[0035] In one possible implementation, the determining module is used to call a medical report generation model to segment the target medical image to obtain multiple target image blocks; extract image features of each target image block; and for any given tissue structure, determine the target features of the tissue structure by using the image features of each target image block and the target query features of the tissue structure.
[0036] Fifthly, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, which is loaded and executed by the processor to enable the electronic device to acquire the medical report generation model described in the first aspect or to implement the medical report generation method described in the second aspect.
[0037] In a sixth aspect, a computer-readable storage medium is also provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to enable an electronic device to acquire the medical report generation model shown in the first aspect above or to implement the medical report generation method shown in the second aspect above.
[0038] In a seventh aspect, a computer program is also provided, wherein the computer program is at least one, and the at least one computer program is loaded and executed by a processor to enable an electronic device to acquire the medical report generation model shown in the first aspect or to implement the medical report generation method shown in the second aspect.
[0039] Eighthly, a computer program product is also provided, wherein at least one computer program is stored in the computer program product, the at least one computer program being loaded and executed by a processor, so as to enable an electronic device to acquire the medical report generation model shown in the first aspect above or to implement the medical report generation method shown in the second aspect above.
[0040] The technical solution provided in this application brings at least the following beneficial effects:
[0041] In the technical solution provided in this application, a sample medical report is used to describe the diagnostic results of at least one sample tissue structure in a sample medical image. That is, the sample medical report analyzes the sample medical image based on tissue structure, exhibiting a highly structured characteristic. Based on this, by using the sample medical image and the first query features of each tissue structure included in the first processing network, the first image features of each tissue structure are determined. This achieves the goal of guiding the analysis of the sample medical image based on the image content characteristics of the tissue structure, enabling the first image features of the tissue structure to characterize the image content of that tissue structure in the sample medical image. This not only improves the accuracy of the first image features but also makes the first image features more consistent with the structured characteristics of the sample medical report. By training the first processing network with the first image features of the sample medical report and each tissue structure, the alignment of text content and image content at the tissue structure level is achieved. This allows the trained second processing network to extract image features conducive to generating medical reports, thereby enabling the medical report generation model determined based on the second processing network to accurately generate medical reports and improve the efficiency of medical report generation. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a schematic diagram of a computer system for obtaining a medical report generation model or generating a medical report, as provided in an embodiment of this application.
[0044] Figure 2 This is a flowchart of a method for obtaining a medical report generation model provided in an embodiment of this application;
[0045] Figure 3 This is a schematic diagram of the structure of an encoder provided in an embodiment of this application;
[0046] Figure 4 This is a schematic diagram of the structure of an encoding block provided in an embodiment of this application;
[0047] Figure 5 This is a flowchart of a medical report generation method provided in an embodiment of this application;
[0048] Figure 6 This is a schematic diagram illustrating the generation of a medical report according to an embodiment of this application;
[0049] Figure 7This is a flowchart illustrating the training process of a medical report generation model provided in an embodiment of this application.
[0050] Figure 8 This is a schematic diagram of a pre-training stage provided in an embodiment of this application;
[0051] Figure 9 This is a schematic diagram of an optimized training phase provided in an embodiment of this application;
[0052] Figure 10 This is a schematic diagram illustrating the generation of a CT report according to an embodiment of this application;
[0053] Figure 11 This is a schematic diagram illustrating another CT report generation method provided in this application embodiment;
[0054] Figure 12 This is a schematic diagram of a CT image and CT report provided in an embodiment of this application;
[0055] Figure 13 This is a schematic diagram of the structure of a medical report generation model acquisition device provided in an embodiment of this application;
[0056] Figure 14 This is a schematic diagram of the structure of a medical report generation device provided in an embodiment of this application;
[0057] Figure 15 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;
[0058] Figure 16 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0060] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0061] First, the abbreviations and key terms involved in the embodiments of this application will be explained and described.
[0062] CT (Computed Tomography) imaging: CT imaging is a technique that combines X-ray scan projection data with reconstruction mathematics and computer technology to obtain medical images based on slice information. CT images record basic information about the object, the model of the CT equipment, and specific parameters of the CT imaging (such as pixel resolution).
[0063] Medical Report Generation (RG) is a technology that automatically generates corresponding text reports using medical images. This technology combines computer vision and natural language processing methods to help doctors and medical professionals generate accurate medical reports more efficiently. By analyzing medical images, such as those acquired through MRI (Nuclear Magnetic Resonance Imaging), CT imaging, or X-rays, it automatically extracts key information from the images and converts it into easily understandable text descriptions. Typically, medical reports can include basic information such as the location, size, and shape of lesions, as well as pathological features and possible diagnostic recommendations.
[0064] Masked Language Modeling (MLM): MLM is a technique that masks certain words in text and predicts the masked words using contextual information. In some cases, it can also be combined with images related to the text to predict the masked words.
[0065] Contrastive Learning (CL): CL can be applied in image-text pre-training to establish connections between images and text. By pairing images and text and training the model using a contrastive loss function, the model learns the similarities and differences between images and text. Generally, the training data can be divided into positive and negative sample pairs. Positive samples refer to images and text that are related, while negative sample pairs refer to images and text that are not related. Training the model using a contrastive loss function allows the model to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs.
[0066] In the medical field, medical images of the internal structures of an organism or a part of an organism can be obtained first. These images are then analyzed to generate a medical report, which describes the diagnostic results of the tissue structures observed in the images. However, doctors need to spend a significant amount of time analyzing medical images and writing medical reports, resulting in low efficiency in obtaining these reports.
[0067] Based on this, embodiments of this application provide a method for obtaining a medical report generation model and a method for generating medical reports, which can quickly and accurately generate medical reports through the medical report generation model, thereby improving the efficiency of obtaining medical reports.
[0068] like Figure 1 As shown, Figure 1 This is a schematic diagram of a computer system for obtaining a medical report generation model or generating a medical report, as provided in an embodiment of this application. The computer system includes a terminal device 101 and a server 102. The method for obtaining a medical report generation model or generating a medical report provided in this embodiment can be executed by the terminal device 101, by the server 102, or by both the terminal device 101 and the server 102. This embodiment does not limit the execution of this method.
[0069] Terminal device 101 is equipped with a client for displaying sample medical images, sample medical reports, etc., and server 102 provides background services for this client. In one possible implementation, server 102 undertakes the main computational work, and terminal device 101 undertakes the secondary computational work. Alternatively, server 102 undertakes the secondary computational work, and terminal device 101 undertakes the main computational work. Or, terminal device 101 and server 102 can collaborate on computation using a distributed computing architecture.
[0070] Optionally, the terminal device 101 can be any electronic device product capable of human-computer interaction with the user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. For example, the terminal device 101 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, PC (Personal Computer), mobile phone, PDA (Personal Digital Assistant), wearable device, PPC (Pocket PC), smart car system, smart TV, etc.
[0071] Terminal device 101 can refer to one of a plurality of terminal devices. This embodiment uses terminal device 101 as an example. Those skilled in the art will know that the number of terminal devices 101 can be more or less. For example, there may be only one terminal device 101, or there may be dozens or hundreds of terminal devices 101, or more. This application embodiment does not limit the number or type of terminal devices 101.
[0072] Server 102 can be a single server, a server cluster consisting of multiple servers, or any of the following: a cloud computing platform or a virtualization center. This embodiment of the application does not limit this. Server 102 communicates directly or indirectly with terminal device 101 via a wired or wireless network. Server 102 has data receiving, data processing, and data sending functions. Of course, server 102 may also have other functions, which are not limited in this embodiment of the application.
[0073] In this embodiment, the target object 103 can interact with the terminal device 101 via an interactive interface. In one possible implementation, the target object 103 inputs the target medical image 104 into the terminal device 101 through the interactive interface, and the terminal device 101 sends the target medical image 104 to the server 102 via a wireless network or a wired network. The server 102 has pre-configured the interface according to the specified parameters. Figure 2 An embodiment of the method for obtaining a relevant medical report generation model is provided, in which the medical report generation model is trained based on sample medical images and sample medical reports. After the server 102 receives the target medical image 104, it inputs the target medical image 104 into the medical report generation model, and then processes it according to the... Figure 5 An embodiment of a related medical report generation method generates and outputs a target medical report through a medical report generation model. Server 102 sends the target medical report to terminal device 101 via a wireless or wired network. Terminal device 101 displays the target medical image 104 and the target medical report 105 through an interactive interface.
[0074] Those skilled in the art should understand that the terminal device 101 and server 102 described above are merely illustrative examples. Other existing or future terminal devices or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0075] The various optional embodiments of this application can be implemented based on artificial intelligence (AI) technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making functions.
[0076] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0077] like Figure 2 As shown, Figure 2 This is a flowchart illustrating a method for obtaining a medical report generation model according to an embodiment of this application. This method can be applied to the aforementioned computer system. For ease of description, the terminal device 101 or server 102 executing the method for obtaining the medical report generation model in this embodiment is referred to as an electronic device, and this method can be executed by an electronic device. Figure 2 As shown, the method includes the following steps.
[0078] Step 201: Obtain the sample medical image and sample medical report, wherein the sample medical report is used to describe the diagnostic results of at least one sample tissue structure in the sample medical image.
[0079] Sample medical images are images obtained by non-invasively acquiring the internal tissue structure of an organism or a part of an organism. Generally, these images can be acquired using techniques such as Computed Tomography (CT) or Magnetic Resonance Imaging (MRI). Sample medical images include X-ray films (images acquired via X-ray), CT scans (images acquired via CT imaging), MRI scans (images acquired via MRI), and ultrasound images (images acquired via ultrasound). The method by which electronic devices acquire sample medical images is not limited here. For example, an electronic device may be connected to an acquisition device used to acquire sample medical images, and the acquisition device may transmit the acquired sample medical images to the electronic device in real time. Alternatively, the electronic device may acquire sample medical images input by a user, or it may acquire sample medical images via the internet.
[0080] A medical image of a sample includes an image region of at least one sample tissue structure. Any sample tissue structure is a tissue structure, which is a part of an organism divided according to its structure, function, etc. For example, the lungs, heart, and liver have different structures and functions and can be regarded as three tissue structures.
[0081] Each medical image sample corresponds to a medical image sample report, which includes the diagnostic results for each tissue structure. The diagnostic results for any given tissue structure describe its pathological condition, location of disease, etc. The method by which electronic devices acquire medical image samples is not limited here; for example, electronic devices can acquire user-inputted medical image samples, or they can acquire medical image samples via the internet.
[0082] Step 202: Using the sample medical images and the multiple first query features included in the first processing network, determine the first image features of multiple tissue structures, including at least one sample tissue structure. Each first query feature represents the visual information of a tissue structure, and each first image feature of a tissue structure represents the image content in the sample medical images related to any tissue structure.
[0083] This application does not limit the structure or number of parameters of the first processing network. The first processing network may include at least one network layer such as a convolutional layer, attention layer, activation layer, pooling layer, and linear layer. The first processing network includes multiple first query features of organizational structures, which can characterize the content characteristics of the image of the organizational structure. Since the content of the image of the organizational structure is the visual representation of the organizational structure, the first query features of the organizational structure characterize the visual features of the organizational structure and can reflect visual information such as the location, texture, color, and size of the organizational structure.
[0084] The number of tissue structures is not limited here. For example, there are nine tissue structures: thoracic cage, ribs, lungs, heart, pleura, liver, kidneys, thyroid gland, and others (remaining tissue structures other than the aforementioned eight, including the spleen, spine, etc.). These tissue structures include the tissue structures of each sample; for example, the sample tissue structures include the lungs, heart, and liver.
[0085] Sample medical images can be input into a first processing network. For any given tissue structure, the first processing network extracts features from the sample medical image based on a first query feature of that tissue structure, obtaining the first image feature of that tissue structure. It is understood that different structures of the first processing network will result in different feature extraction methods, which are not limited here. Since the first query feature of a tissue structure characterizes its visual features, guiding feature extraction from the sample medical image using this first query feature enables the extraction of the image content of that tissue structure from the sample medical image, resulting in the first image feature of the tissue structure. Based on this, the first image feature of the tissue structure can characterize the image content of that tissue structure in the sample medical image, including visual information such as the tissue structure's location, texture, color, and size.
[0086] In one possible implementation, step 202 includes steps A1 to A3 (not shown in the figure).
[0087] Step A1: Segment the sample medical images to obtain multiple sample image blocks.
[0088] This application does not limit the method of segmenting sample medical images. For example, the sample medical image can be segmented into a set number of sample image blocks, or it can be segmented into several sample image blocks of a set size. The set number or set size can be preset based on human experience, or it can be data input by the target object or randomly determined data.
[0089] Step A2: Extract image features from each sample image patch.
[0090] In this example, the first processing networks with different structures and sizes extract image features from sample image patches in different ways. The following section describes the first processing network as including... Figure 3 Using the encoder shown as an example, the implementation principle of step A2 is explained.
[0091] like Figure 3 As shown, the sample medical image is input into the encoder, which includes a segmentation layer to divide the sample medical image into multiple sample image blocks. The encoder also includes a multi-stage feature extraction network. The first stage of the feature extraction network includes a linear layer and M encoding blocks. The l-th stage of the second to L+1 stages includes an integration layer and N encoding blocks. L, M, and N are all positive integers, and l takes any value from 2 to L+1. For example, L equals 3, M equals 2, and N equals 3. The linear layer performs linear mapping operations, the encoding blocks perform feature extraction operations, and the integration layer performs dimensionality reduction operations.
[0092] For the first stage feature extraction network, the input includes various sample image patches. For any sample image patch, a linear layer performs a linear mapping to obtain the embedding feature (i.e., embedding) of the sample image patch. The embedding feature is a high-dimensional vector used to represent the content of the sample image patch. Feature extraction is performed on the embedding feature of the sample image patch through M encoding blocks to obtain the first feature of the sample image patch. The structure and number of parameters of the encoding blocks are not limited here; different structures and parameter numbers lead to different feature extraction methods. The following section uses the encoding block as an example... Figure 4 Taking the structure shown as an example, the feature extraction process is illustrated. If the structure of the coding block is... Figure 4 The structure shown is then Figure 3The encoder shown is also known as a Swing Transformer-based encoder.
[0093] like Figure 4 As shown, the encoding block consists of two parts, each including a normalization layer, a multi-head self-attention module (MSA), a normalization layer, and a multi-layer perceptron (MLP). Specifically, the MSA in the first part is a window-based multi-head self-attention module (W-MSA), and the MSA in the second part is a shifted window-based multi-head self-attention module (SW-MSA).
[0094] After inputting the embedding features of the sample image patch into the first coding block of M coding blocks, feature extraction is first performed on the embedding features of the sample image patch through the first part. First, the embedding features of the sample image patch are standardized through the first normalization layer to obtain the first processing result. Then, W-MSA is used to divide the first processing result into multiple small blocks of the same size. Attention processing is performed on each small block based on a self-attention mechanism to obtain the processing result of each small block. The processing results of each small block are then concatenated to obtain the second processing result. Since W-MSA obtains small blocks of the same size, the size of each small block can be regarded as the size of a regular window; therefore, W-MSA is a multi-head self-attention processing based on a regular window. By performing attention processing on the small blocks, features of local regions are extracted, which helps improve the accuracy of feature representation. Since the first normalization layer is skip-connected to W-MSA, the embedding features of the sample image patch are fused with the second processing result to obtain the first fusion result, reducing information loss caused by feature extraction. Next, the first fusion result is standardized through the second normalization layer to obtain the third processing result. The third processing result is then processed by a multilayer perceptron to extract features, resulting in a fourth processing result. Since the second normalization layer is skipped from the multilayer perceptron, the first fusion result is fused with the fourth processing result to obtain a second fusion result, thereby reducing information loss caused by feature extraction.
[0095] Next, feature extraction is performed on the second fusion result in the second part. First, the second fusion result is standardized through the first normalization layer to obtain the fifth processing result. Then, SW-MSA is used to divide the fifth processing result into multiple small blocks of different sizes. Attention processing is applied to each small block based on a self-attention mechanism to obtain the processing result for each small block. The processing results of each small block are then concatenated to obtain the sixth processing result. Since SW-MSA obtains small blocks of different sizes, each small block can be regarded as information within the window of the fifth processing result obtained through a sliding window. Therefore, SW-MSA is a multi-head self-attention processing based on a sliding window. Building upon the attention processing of small blocks of the same size using W-MSA, attention processing is applied to small blocks of different sizes using SW-MSA, enabling the exchange of information in local regions and improving the accuracy of feature representation. Since the first normalization layer is skipped with SW-MSA, the second fusion result is fused with the sixth processing result to obtain the third fusion result, reducing information loss caused by feature extraction. Then, the third fusion result is standardized through the second normalization layer to obtain the seventh processing result. The seventh processing result is then processed by a multilayer perceptron to extract features, resulting in the eighth processing result. Since the second normalization layer is skip-connected to the multilayer perceptron, the third fusion result is fused with the eighth processing result to obtain the fourth fusion result, thus reducing information loss caused by feature extraction. The output of the first coding block in the M coding blocks includes the fourth fusion result.
[0096] For a non-first coding block among the M coding blocks, the input of the non-first coding block includes the output of the previous coding block. Feature extraction can be performed on the embedded features of the sample image block in the same way as the first coding block, and by extracting features from the output of the previous coding block through the non-first coding blocks, the output of the first-stage feature extraction network is obtained, which is the first feature of the sample image block.
[0097] For the feature extraction network in the l-th stage, the input includes the first features of each sample image patch. For any sample image patch, the first features of the sample image patch are reduced in dimensionality through an integration layer to obtain the dimensionality-reduced features, which will have an increased number of channels. The second features of the sample image patch can be obtained by sequentially extracting features from the embedded features of the sample image patch using the same method as extracting features from the first encoding block, and then performing feature extraction multiple times on the dimensionality-reduced features through N encoding blocks.
[0098] The encoder also includes an integration layer concatenated after the feature extraction network in the last stage. This integration layer reduces the dimensionality of the second features of the sample image patch output by the feature extraction network in the last stage to obtain the image features of the sample image patch.
[0099] In the above feature extraction process, multiple stages of feature extraction are performed on the embedded features of the sample image patch to continuously reduce the feature dimensionality and complexity, thereby improving the feature representation ability. The image features of the sample image patch obtained after multiple stages of feature extraction are high-dimensional features that can represent information such as the content, texture, color, and size of the sample image patch.
[0100] Step A3: For any given organizational structure, determine the first image feature of the organizational structure by using the image features of each sample image block and the first query feature of the organizational structure.
[0101] For any given sample image patch, based on the first query feature of any tissue structure, image sub-features of the tissue structure are extracted from the image features of that sample image patch. These image sub-features are used to characterize the visual information of the tissue structure in the sample image patch. The various image sub-features of the tissue structure are then concatenated to obtain the first image feature of the tissue structure. Since the image features of a sample image patch characterize information such as its content, texture, color, and size, and a sample image patch may contain more than just tissue structure information, extracting image sub-features of the tissue structure separately from the image features of each sample image patch enables fine-grained extraction of visual information of the tissue structure from local regions of the sample medical image. This results in higher accuracy of the image sub-feature representation, thereby improving the accuracy of the first image feature of the tissue structure and ultimately enhancing the accuracy of the medical report generation model.
[0102] Optionally, the first processing network further includes a selection network, which includes at least one of attention layers, linear layers, and nonlinear layers. The image features of each sample image patch and the first query features of each tissue structure are input into the selection network, and the selection network outputs the first image features of each tissue structure. In one possible implementation, the image features of any sample image patch and its positional features are concatenated to obtain the concatenated features of the sample image patch. The positional features of the sample image patch are used to characterize the position of the sample image patch in the sample medical image. For example, if the sample image patch is the i-th (i is 0 or a positive integer) image patch in the sample medical image, then the positional features of the sample image patch are used to characterize i. The concatenated features of each sample image patch and the first query features of each tissue structure are input into the selection network, and the selection network outputs the first image features of each tissue structure.
[0103] In an exemplary embodiment, step A3 includes: determining a first similarity between the image features of each sample image block and a first query feature of any organization; and extracting the first image feature of any organization from the image features of each sample image block based on the first similarity.
[0104] For any given sample image patch, a first similarity score can be calculated between the image features of that sample image patch and the first query feature of any organizational structure using a similarity function, network layers, etc. The first similarity score is then multiplied by the image features of the sample image patch, and the result, or a weighted sum of the multiplication results, is used as a sub-feature of the organizational structure, thus extracting sub-features from the image features of the sample image patch that are visually similar to the organizational structure. Alternatively, if the first similarity score is not less than a similarity threshold, the image features of that sample image patch are used as sub-features of the organizational structure, thus selecting image features from sample image patches that are visually similar to the organizational structure. The similarity threshold is a pre-set value based on human experience, or it can be one of the first similarity scores among the first similarity scores between the image features of each sample image patch and the first query feature of any organizational structure.
[0105] Using the above method, at least one image sub-feature of any tissue structure can be determined. These sub-features are then concatenated to obtain the first image feature of the tissue structure. By extracting features similar to the visual information of the tissue structure from the image features of sample image blocks, or by selecting features similar to the visual information of the tissue structure, the features used to characterize the tissue structure are accurately determined, improving the accuracy of the first image feature of the tissue structure and thus enhancing the accuracy of the medical report generation model.
[0106] If we sort the image features of each sample image patch in descending order and determine the first similarity between them and the first query feature of any organizational structure, and use the Kth (K is a positive integer) first similarity as the similarity threshold, then for each organizational structure, we can select the image features of K sample image patches as the image sub-features of that organizational structure. Assume there are N... s If there are multiple organizational structures, then a total of K×N can be selected. s Each image sub-feature is used to stitch together the image sub-features to obtain the first image feature of each tissue structure.
[0107] Optionally, use To characterize N s (N s The first query feature of (positive integer) organizational structures. Represents the i-th (i takes values from 1 to N) s The first query feature of an organizational structure (any one of the items in the list). Using... Characterizing N v (N v Image features of image patches (where v is a positive integer) of sample image blocks. j Characterize the j-th (j takes values from 1 to N) vThe image features of each sample image block are used as an example. The cross-attention network is used to calculate the cross-attention result between the image features of each sample image block and the first query feature of each tissue structure. The calculation process is shown in the following formulas (1) and (2).
[0108]
[0109]
[0110] in, Representing three linear projection matrices, T is the symbol for the transpose matrix, and softmax is a normalized exponential function. That is, A v It includes N s ×N v A matrix of n elements, where each element is a real number. v The first similarity between the image features representing each sample image patch and the first query features representing each tissue structure. Characterizing N s The first image feature of an organizational structure, s i v The first image feature characterizing the i-th tissue structure.
[0111] Step 203: Based on the sample medical report and the first image features of each tissue structure, train the first processing network to obtain the second processing network.
[0112] In this embodiment, a first loss can be determined based on the first image features of the sample medical report and various tissue structures. The first processing network is then trained once using this first loss to obtain a trained first processing network. If the trained first processing network meets a first termination condition, it is used as a second processing network. If the trained first processing network does not meet the first termination condition, it is used as the first processing network for the next training iteration. The training continues according to steps 201 to 203 until the trained first processing network meets the first termination condition, thus obtaining the second processing network.
[0113] This application does not limit the content of the first termination condition being met by the first processing network after training. For example, the first termination condition being met by the first processing network after training includes: the number of training iterations of the first processing network after training reaches the first number, or the performance index of the first processing network after training is not less than a first index, etc. Here, the first number or the first index can be a value preset based on human experience, or it can be a value input by the target object.
[0114] In one possible implementation, step 203 includes steps B1 to B2 (not shown in the figure).
[0115] Step B1: Using the sample medical report and multiple second query features included in the third processing network, determine the first text features of each organizational structure. Any second query feature is used to characterize the descriptive information of any organizational structure, and the first text feature of any organizational structure is used to characterize the text content related to any organizational structure in the sample medical report.
[0116] This application does not limit the structure or number of parameters of the third processing network. The third processing network may include at least one network layer such as a convolutional layer, attention layer, activation layer, pooling layer, and linear layer. The third processing network includes multiple second query features of organizational structures, which can characterize the content features of the text of the organizational structure. Since the content of the text of the organizational structure is information describing the organizational structure, the second query features of the organizational structure can reflect descriptive information such as the location, texture, color, and size of the organizational structure.
[0117] A sample medical report can be input into a third processing network. For any given organizational structure, the third processing network extracts features from the sample medical report based on the second query features of that organizational structure, obtaining the first textual features of that organizational structure. It is understood that different structures of the third processing network will result in different feature extraction methods, which are not limited here. By guiding feature extraction from the sample medical report through the second query features of the organizational structure, the content of that organizational structure can be extracted from the sample medical report, yielding the first textual features of the organizational structure. Based on this, the first textual features of the organizational structure can characterize information such as the location, texture, color, and size of the organizational structure in the sample medical report.
[0118] In an exemplary embodiment, step B1 includes steps B11 to B13 (not shown in the figures).
[0119] Step B11: Segment the sample medical report to obtain multiple sample text segments. Any sample text segment can be used to describe the diagnostic results of a sample tissue structure.
[0120] Generally, medical reports are highly structured, meaning they typically describe the pathological condition and location of disease in each tissue structure separately. Based on this, sample medical reports can be segmented according to their tissue structure, resulting in multiple sample text segments. Each sample text segment corresponds to a specific tissue structure and describes at least one diagnostic result, such as the condition or location of disease in that tissue structure. Any two sample text segments may correspond to the same or different tissue structures. Each sample text segment must contain at least one sentence, or the number of characters in any sample text segment must not exceed a predetermined character count, which is a value pre-set based on human experience.
[0121] In one possible implementation, keywords for each organizational structure are pre-set based on human experience, or keywords for each organizational structure are obtained through the internet. For any sentence in the sample medical report, if the sentence contains a keyword for a certain organizational structure, then that organizational structure is identified as the sample organizational structure, and the sentence is the sample text segment corresponding to that organizational structure. Optionally, if the organizational structure includes the aforementioned "other," and no keyword for "other" has been set or obtained, then if a sentence does not contain any of the keywords, then that sentence is identified as the sample text segment corresponding to "other." By segmenting the sample medical report into sample text segments through keyword matching, the prior knowledge of each organizational structure is used to quickly and accurately segment sentences for each organizational structure, thus accelerating the segmentation efficiency.
[0122] By segmenting the sample medical reports according to their tissue structure, the segmentation method is more in line with the characteristics of medical reports. This facilitates the subsequent extraction of the first text features of each tissue structure, which can improve the accuracy of the medical report generation model during the subsequent training process based on the first text features of each tissue structure.
[0123] Step B12: Extract the text features of each sample text segment.
[0124] In this example, the methods for extracting text features from sample text segments differ depending on the structure and size of the third processing network. The implementation principle of step B12 is illustrated below, taking an example where the third processing network includes at least one cascaded first extraction module (e.g., the third processing network includes six cascaded first extraction modules), and any one of the first extraction modules includes a bidirectional self-attention (Bi Self-Att) layer and a feed-forward layer.
[0125] In this embodiment, each character in any sample text segment can be converted into a character vector to obtain the embedding feature (i.e., embedding) of the sample text segment. The embedding feature includes the character vector of each character, which is used to represent the content of the sample text segment and is a high-dimensional vector.
[0126] For the first extraction module, the embedded features of the sample text segment are input. For any character vector, a bidirectional self-attention layer performs attention processing on the character vector of that character and the character vectors of the remaining characters, obtaining the attention processing result for that character. Wherein, if the same sample organization structure corresponds to at least one sample text segment, the remaining characters include not only all characters in the sample text segment containing that character, but also characters in all sample text segments corresponding to the same sample organization structure, excluding the sample text segment containing that character. The attention processing result for any character is used to describe the semantics of that character in the context of the remaining characters. Through attention processing, information interaction between a character and the remaining characters is achieved to explore the relationship between the remaining characters and the character, which is beneficial for extracting the semantics of the sample text segment. Through a feedforward layer, the attention processing results of each character are fused to integrate the semantics of each character, obtaining the output result of the first extraction module.
[0127] For non-first extraction modules, the output of the previous first extraction module is input into the non-first extraction module. Following the feature processing method used by the first extraction module for embedding features of the sample text segment, the output of the previous first extraction module is processed by the non-first extraction module to obtain the output of the non-first extraction module.
[0128] The final output of the first extraction module is the text features of the sample text segment, used to characterize its semantic content. By continuously extracting features from each first extraction module, the semantics of each character in the sample text segment are continuously extracted and fused, resulting in increasingly higher semantic accuracy of the sample text segment.
[0129] Step B13: Determine the first text features of each organizational structure by using the text features of each sample text segment and the second query features of multiple organizational structures.
[0130] For any given sample text segment, based on the second query feature of any organizational structure, textual sub-features of the organizational structure are extracted from the text features of that sample text segment. These textual sub-features are used to characterize the descriptive information of the organizational structure in the sample text segment. The various textual sub-features of the organizational structure are concatenated to obtain the first textual feature of the organizational structure. Since the textual features of a sample text segment characterize the descriptive information of that segment, and a sample text segment may contain more than just organizational structure information, by extracting textual sub-features of the organizational structure separately from the text features of each sample text segment, fine-grained extraction of descriptive information of the organizational structure from local sentences in the sample medical report is achieved. This results in higher accuracy of the textual sub-features, thereby improving the accuracy of the first textual feature of the organizational structure and ultimately improving the accuracy of the medical report generation model.
[0131] Optionally, step B13 includes steps B131 to B132 (not shown in the figure).
[0132] Step B131: For any organizational structure, determine the second similarity between the text features of each sample text segment and the second query feature of any organizational structure.
[0133] For any given sample text segment, the second similarity between the text features of that sample text segment and the second query feature of any organizational structure can be calculated using similarity functions, network layers, etc.
[0134] Step B132: Based on each second similarity, extract the first text feature of any organizational structure from the text features of each sample text segment.
[0135] In this embodiment, each sample text segment corresponds to a second similarity. This second similarity can be multiplied by the text features of the sample text segment, and the result of the multiplication, or a weighted sum of the multiplication results, can be used as a text sub-feature of the organizational structure. This allows for the extraction of text sub-features similar to the descriptive information of the organizational structure from the text features of the sample text segment. Alternatively, if the second similarity is not less than a similarity threshold, the text features of the sample text segment are used as text sub-features of the organizational structure, thus selecting text features of sample text segments similar to the descriptive information of the organizational structure. The similarity threshold is a pre-set value based on human experience, or it can be one of the second similarities between the text features of each sample text segment and the second query feature of any organizational structure.
[0136] Using the above method, at least one textual sub-feature of any organizational structure can be determined. These sub-features are then concatenated to obtain the first textual feature of the organizational structure. By extracting features similar to the descriptive information of the organizational structure from the textual features of sample text segments, or by selecting features similar to the descriptive information of the organizational structure, the features used to characterize the organizational structure are accurately determined. This improves the accuracy of the first textual feature of the organizational structure, thereby enhancing the accuracy of the medical report generation model.
[0137] Optionally, use To characterize N s (N s The second query feature is (a positive integer) organizational structures. Represents the i-th (i takes values from 1 to N) s The second query feature of an organizational structure (any one of the items in the list). Using... Characterizing N t (N t The text features of (a positive integer) sample text segments, t k Represents the k-th (k takes values from 1 to N) t The text features of any one of the sample text segments (M) can be obtained through a masked cross-attention network, based on the mask M. strct The first text features of each organizational structure are calculated according to the following formulas (3) and (4).
[0138]
[0139]
[0140] in, Representing three linear projection matrices, T is the symbol for the transpose matrix, and softmax is a normalized exponential function. That is, A t It includes N s ×N t A matrix of n elements, where each element is a real number. t The second similarity is used to characterize the text features of each sample text segment and the second query features of each organizational structure. Characterizing N s The first textual feature of an organizational structure The first text feature representing the i-th organizational structure.
[0141] Mask M strct It is a matrix used to represent whether each text feature and each second query feature corresponds to the same organizational structure. In other words, the mask M... strctIt includes multiple elements. Each element represents whether a text feature and a second query feature correspond to the same organizational structure. If they correspond to the same organizational structure, the element is 1 or close to 1; if they correspond to different organizational structures, the element is 0 or close to 0. This is achieved through a mask M. strct This allows text features with the same organizational structure and second query features to exchange information, resulting in a higher second similarity for features with the same organizational structure and a lower second similarity for features with different organizational structures, which helps improve the accuracy of the first text features.
[0142] Step B2: Based on the first text features and first image features of each organizational structure, train the first processing network to obtain the second processing network.
[0143] In this embodiment, a first loss can be determined based on the first text features and first image features of each organizational structure. The first processing network is then trained at least once using the first loss to obtain a second processing network. The training process is not described in detail here. The method for determining the first loss is not limited here. For example, since the first text features and first image features of the same organizational structure should represent the same content, a sub-loss for that organizational structure can be determined using the first text features and first image features of the same organizational structure. The sub-losses of each organizational structure are then added, averaged, or weighted to obtain the first loss. Alternatively, the first loss can be determined and the first processing network trained based on the first loss according to steps B21 to B23 below. That is, step B2 includes steps B21 to B23 (not shown in the figure).
[0144] Step B21: Determine the second image features of each tissue structure through the first momentum network. The first momentum network has the same structure as the first processing network, and the network parameters of the first momentum network are obtained by updating the network parameters of the first processing network.
[0145] Momentum update is a commonly used technique for updating network parameters, referring to updating network parameters at a relatively slow rate. In this embodiment, the network parameters of the first processing network are assumed to be θ. q Before the first momentum network obtained from the momentum update, the network parameters of the original momentum network are θ. k After the momentum update yields the first momentum network, the network parameters θ of the first momentum network... k+1 Satisfy: θ k+1 ←mθ k +(1-m)θ q Where m is a hyperparameter, which can be preset based on human experience or adaptively adjusted as momentum is updated. Optionally, m = 0.99.
[0146] As can be seen from the momentum update process, since the original momentum network has a larger proportion of network parameters, the update speed of the network parameters is relatively slow. This makes the network parameters of the first momentum network close to those of the original momentum network, resulting in features output by the first momentum network being similar to those output by the original momentum network, with a low degree of feature variation. Subsequently, when training the first processing network using the features output by the first momentum network, the low degree of feature variation reduces the possibility of the first processing network collapsing, thereby improving the performance of the first processing network.
[0147] Since the first momentum network and the first processing network have the same structure, after inputting the sample medical image into the first momentum network, the second image features of each tissue structure can be determined through the first momentum network based on the fifth query features of the sample medical image and the multiple tissue structures included in the first momentum network, according to the implementation principle of step 202. The determination process will not be elaborated here. The fifth query features of the tissue structure are used to characterize the content characteristics of the tissue structure image, and the second image features of the tissue structure are used to characterize the image content related to the tissue structure in the sample medical image.
[0148] Step B22: Determine the second text features of each organizational structure through the second momentum network. The second momentum network has the same structure as the third processing network, and the network parameters of the second momentum network are obtained by updating the network parameters of the third processing network.
[0149] In this embodiment, following the implementation principle of step B21, the network parameters of the second momentum network are obtained by weighted summation of the network parameters of the original momentum network before the third processing network and the second momentum network obtained from the momentum update. By reducing the update speed of the network parameters, the degree of change in the network output features is reduced, thereby reducing the possibility of the first processing network collapsing and improving the performance of the first processing network.
[0150] Since the second momentum network and the third processing network have the same structure, after inputting the sample medical report into the second momentum network, following the implementation principle of step B1, the second text features of each organizational structure can be determined through the second momentum network based on the sixth query features of the multiple organizational structures included in the sample medical report and the second momentum network. The determination process will not be elaborated here. The sixth query features of the organizational structure are used to characterize the content characteristics of the text of the organizational structure, and the second text features of the organizational structure are used to characterize the text content related to the organizational structure in the sample medical report.
[0151] Step B23: Based on the first text features, second text features, first image features, and second image features of each organizational structure, train the first processing network to obtain the second processing network.
[0152] In this embodiment, a first loss can be determined based on the first text features, second text features, first image features, and second image features of each organizational structure. The first processing network is then trained at least once using the first loss to obtain a second processing network. The training process is not described in detail here. Optionally, since the first text features, second text features, first image features, and second image features of the same organizational structure should represent the same content, a sub-loss for that organizational structure can be determined using the first text features, second text features, first image features, and second image features of the same organizational structure. The first loss is obtained by adding, averaging, or weighting the sub-losses of each organizational structure.
[0153] Alternatively, step B23 may include steps B231 through B233 (not shown in the figure).
[0154] Step B231: For any first text feature and any second image feature, determine the third similarity between the first text feature and the second image feature, and obtain first annotation information to characterize whether the first text feature and the second image feature correspond to the same organizational structure.
[0155] In this embodiment, any first text feature and any second image feature may correspond to the same organizational structure or different organizational structures. Based on this, a third similarity between the first text feature and the second image feature can be calculated using a similarity algorithm. This embodiment does not limit the similarity algorithm; for example, the similarity algorithm can be any similarity algorithm such as cosine similarity or Euclidean distance.
[0156] Alternatively, the similarity algorithm is as follows: Characterize the a-th image feature, G represents the a-th image feature. v (·) represents a function that linearly maps image features, g t (·) represents the function that performs a linear mapping on text features, and T represents the sign of the transpose matrix. Similarity representation. Through linear mapping, high-dimensional image features and high-dimensional text features can be mapped to low-dimensional representations, for example, to 128 dimensions, thus accelerating the computational efficiency of similarity.
[0157] Based on the above similarity algorithm, the i-th first text feature can be calculated. and the m-th second image feature Third similarity between Furthermore, the electronic device can acquire first annotation information regarding whether the i-th first text feature and the m-th second image feature correspond to the same organizational structure. Optionally, the first annotation information is data in one-hot encoded form. That is, if the first text feature and the second image feature correspond to the same organizational structure, the first annotation information is a first annotation value; if the first text feature and the second image feature correspond to different organizational structures, the first annotation information is a second annotation value. The first annotation value and the second annotation value are different numerical values; for example, the first annotation value is 1, and the second annotation value is 0.
[0158] In some cases, a medical report generation model is trained using sample medical images and sample medical reports from multiple sample objects. In this case, the first annotation information is used to characterize whether any first text feature and any second image feature correspond to the same sample object and the same tissue structure. Optionally, if the first text feature and the second image feature correspond to the same sample object and the same tissue structure, then the first annotation information is a first annotation value; if the first text feature and the second image feature correspond to at least one of different sample objects or different tissue structures, then the first annotation information is a second annotation value.
[0159] Step B232: For any second text feature and any first image feature, determine the fourth similarity between the second text feature and the first image feature, and obtain second annotation information to characterize whether the second text feature and the first image feature correspond to the same organizational structure.
[0160] Any second text feature and any first image feature may correspond to the same organizational structure or different organizational structures. Based on this, a fourth similarity score can be calculated between the second text feature and the first image feature using a similarity algorithm. Based on the similarity algorithm mentioned above, the i-th first image feature can be calculated. and the m-th second text feature The fourth similarity between
[0161] Furthermore, the electronic device can obtain second annotation information regarding whether the i-th first image feature and the m-th second text feature correspond to the same organizational structure. If the first image feature and the second text feature correspond to the same organizational structure, then the second annotation information is the first annotation value; if the first image feature and the second text feature correspond to different organizational structures, then the second annotation information is the second annotation value.
[0162] In some cases, a medical report generation model is trained using sample medical images and sample medical reports from multiple sample objects. In this case, the second annotation information is used to characterize whether any first image feature and any second text feature correspond to the same sample object and the same tissue structure. Optionally, if the first image feature and the second text feature correspond to the same sample object and the same tissue structure, then the second annotation information is the first annotation value; if the first image feature and the second text feature correspond to at least one of different sample objects or different tissue structures, then the second annotation information is the second annotation value.
[0163] Step B233: Based on multiple third similarities, multiple fourth similarities, multiple first annotation information and multiple second annotation information, train the first processing network to obtain the second processing network.
[0164] The first loss can be determined based on multiple third similarities, multiple fourth similarities, multiple first annotation information, and multiple second annotation information.
[0165] Optionally, if the first annotation information is a first annotation value, then the first text feature and the second image feature corresponding to the first annotation information are a positive sample pair; if the first annotation information is a second annotation value, then the first text feature and the second image feature corresponding to the first annotation information are a negative sample pair. Similarly, if the second annotation information is a first annotation value, then the first image feature and the second text feature corresponding to the second annotation information are a positive sample pair; if the second annotation information is a second annotation value, then the first image feature and the second text feature corresponding to the second annotation information are a negative sample pair.
[0166] If the first label value is 1 and the second label value is 0, then each third similarity and each fourth similarity need to be normalized so that the normalized third similarity and the normalized fourth similarity are between 0 and 1.
[0167] Optionally, for the i-th first text feature and the m-th second image feature Third similarity between The third similarity is normalized according to formula (5) as shown below. For the i-th first image feature... and the m-th second text feature The fourth similarity between The fourth similarity is normalized according to the formula (6) shown below.
[0168]
[0169]
[0170] in, The third similarity is represented by the normalized value. The fourth similarity is represented after normalization. Assume p m include and Then p m The number is N s ×N q One. That is to say, N s ×N q p m The set p formed satisfies: exp represents an exponential function with base e. τ is a learnable parameter, which in this embodiment can be referred to as the temperature parameter. N s N represents the number of tissue structures. q This represents the sum of the number of positive and negative sample pairs corresponding to any given organizational structure. It should be noted that the calculation processes shown in formulas (5) and (6) above can be obtained using a normalization function (such as the softmax function).
[0171] Next, based on the normalized third similarity, the normalized fourth similarity, the first annotation information, and the second annotation information, the first loss is calculated. Optionally, the first loss is shown in the following formula (7).
[0172]
[0173] in, Characterizing the first loss, The symbol representing the average. Characterizing the i-th first text feature and the m-th second image feature The first annotation information between them. Characterizing the i-th first image feature and the m-th second text feature The second annotation information is between the parameters. H represents the sign of the cross-entropy loss function. The meanings of the remaining parameters are described above and will not be repeated here.
[0174] Subsequently, a second processing network is obtained by training a first processing network based on the first loss. The training method has been described above and will not be repeated here. By calculating the third similarity between the first text feature and the second image feature, and the fourth similarity between the second text feature and the first image feature, fine-grained semantic alignment of the first text feature, second image feature, second text feature, and first image feature with the same organizational structure is achieved. This allows for fine-grained differentiation of the first text feature, second image feature, second text feature, and first image feature with different organizational structures. This enables the second processing network, trained using each third and fourth similarity, to distinguish content with different organizational structures, thus improving the accuracy of the second processing network.
[0175] Furthermore, if positive sample pairs correspond to the same sample object and the same organizational structure, and negative sample pairs correspond to at least one of different sample objects or different organizational structures, then the trained second processing network can also distinguish information from different objects, resulting in higher accuracy.
[0176] The preceding text described the process of training the first processing network based on the third processing network to obtain the second processing network. In one possible implementation, the second processing network can be obtained by training the first processing network based on the fourth processing network, as shown in steps C1 to C3 (not shown in the figure). That is, step 203 includes steps C1 to C3.
[0177] Step C1: Determine the masked medical report based on the sample medical report. The masked medical report is obtained by masking multiple first characters in the sample medical report.
[0178] In natural language processing, masking is a technique that replaces characters or words with special markers. These special markers are not limited here; for example, they can be zero vectors, unit vectors, or the [MASK] symbol. In this embodiment, the sample medical report includes multiple characters, including multiple first characters, where any two first characters may be consecutive or non-consecutive in the sample medical report. Optionally, each first character in the sample medical report can be replaced with a special marker to obtain a masked medical report. Alternatively, if at least two consecutive first characters among the multiple first characters can form a word, then each first character forming the word is replaced with a special marker. Furthermore, each first character that cannot form a word is replaced with a respective feature marker. In this way, a masked medical report is obtained.
[0179] Step C2: The generation probability of each first character is determined by the fourth processing network based on the first image features of the masked medical report and each tissue structure.
[0180] This application does not limit the structure or number of parameters of the fourth processing network. The fourth processing network may include at least one network layer such as a convolutional layer, attention layer, activation layer, pooling layer, linear layer, and feedforward layer. The first image features of the masked medical report and various tissue structures can be input into the fourth processing network, which then determines the generation probability of each first character. The generation probability of the first character character represents the likelihood of determining the first character through the fourth processing network; the higher the generation probability of the first character, the higher the likelihood of determining the first character through the fourth processing network. Optionally, the generation probability of the first character is greater than or equal to 0 and less than or equal to 1. It is understood that different structures of the fourth processing network will result in different methods for determining the generation probability of the first character. One possible implementation is shown below.
[0181] In this example, the masked medical report includes multiple masked text segments, which are obtained by masking the first character in a sample text segment. The sample text segment is used to describe the diagnostic result of a sample tissue structure. Step C2 includes: for any masked text segment, extracting the text features of any masked text segment through a fourth processing network; determining the first image features of the sample tissue structure corresponding to any masked text segment from the first image features of each tissue structure through the fourth processing network; fusing the text features of any masked text segment and the first image features of the corresponding sample tissue structure to obtain a first fused feature; and determining the generation probability of the first character in any masked text segment based on the first fused feature.
[0182] As mentioned above, the medical report is segmented according to the sample tissue structure to obtain multiple sample text segments. For each sample text segment, the first character in the sample text segment can be replaced with a special character as in step C1 to obtain a masked text segment. Next, each character in the masked text segment is converted into a character vector to obtain the embedding feature of the masked text segment. This embedding feature includes the character vector of each character and is used to represent the content of the masked text segment; it is a high-dimensional vector. Then, the embedding feature of the masked text segment and the first image features of each tissue structure are input into the fourth processing network. In this example, the fourth processing network includes at least one cascaded second extraction module (e.g., the fourth processing network includes six cascaded second extraction modules), and any second extraction module includes a bidirectional self-attention layer, a cross-attention layer, and a feedforward layer.
[0183] For the first second extraction module, the embedding features of the masked text segment and the first image features of each organizational structure are input into the first second extraction module. The input of the bidirectional self-attention layer includes the embedding features of the masked text segment. For the character vector of any character in these embedding features, the bidirectional self-attention layer performs attention processing on the character vector of that character and the character vectors of the remaining characters, obtaining the attention processing result for that character. If the same sample organizational structure corresponds to at least one masked text segment, the remaining characters include not only all characters in the masked text segment containing any given character, but also all characters in the masked text segments corresponding to the same sample organizational structure, excluding the masked text segment containing any given character. The attention processing result for any character is used to describe the semantics of that character in the context of the remaining characters. Through attention processing, information interaction between a character and the remaining characters is achieved to explore the relationship between the remaining characters and the given character, which is beneficial for extracting the semantics of the masked text segment. Based on this, the output of the bidirectional self-attention layer is the text features of the masked text segment, used to characterize the semantics of the masked text segment. The input to the cross-attention layer includes the output of the bidirectional self-attention layer and the first image features of each organizational structure. Since the masked text segment is obtained by masking the sample text segment, and the sample text segment contains content related to the sample organizational structure, the first image features of the sample organizational structure can be determined from the first image features of each organizational structure through the cross-attention layer. Cross-attention processing is then performed on the text features of the masked text segment and the first image features of the sample organizational structure to obtain the first fusion feature. This first fusion feature represents the fusion information after fusing the semantic content of the masked text segment and the visual information of the sample organizational structure. The cross-attention layer aligns the semantic content and visual information of the same organizational structure, enabling the visual information to compensate for the semantic content, which helps improve the generation probability of the first character and thus improves the training effect of the first processing network. The first fusion feature is input to the feedforward layer, which sequentially performs linear transformations and activations on the first fusion feature to further extract the features of the fusion information, resulting in the output of the feedforward layer. Extracting the features of the fusion information through the feedforward layer reduces feature complexity and improves the representational ability of the features. The output of the feedforward layer is also the output of the first second extraction module.
[0184] For non-first second extraction modules, the output of the previous second extraction module and the first image features of each tissue structure are input into the non-first second extraction module. Following the feature processing method used by the first second extraction module for the embedding features of the masked text segment and the first image features of each tissue structure, the non-first second extraction module performs feature processing on the output of the previous second extraction module and the first image features of each tissue structure to obtain the output of the non-first second extraction module.
[0185] In this example, the output of the last second extraction module is referred to as the first target feature, which is used to characterize the semantics of the masked text segment and the visual information of the corresponding image. The fourth processing network also includes a prediction network, which determines the generation probability of each character in the masked text segment based on the first target feature. Since these characters include the first character, the generation probability of each first character can be determined by the prediction network.
[0186] By continuously extracting features from the embedded features of the masked text segment and the first image features of each organizational structure, the semantic content and visual information of each organizational structure are fully integrated. This allows the visual information to compensate for the semantic content, reducing the complexity of the first target features while improving the ability of the first target features to represent the semantic content and visual information of the sample organizational structure. This increases the probability of generating the first character and thus improves the training effect of the first processing network.
[0187] Step C3: Based on the generation probability of each first character, train the first processing network to obtain the second processing network.
[0188] Specifically, based on the generation probability of each first character, the first loss is calculated according to the function formula of the first loss. This first loss can be denoted as: The formula for this function can be any of the following: cross-entropy function, relative entropy function, mean squared error function, etc., which will not be elaborated here. Next, the first processing network is trained based on the first loss to obtain the second processing network. The training method has been described above and will not be repeated here.
[0189] In this embodiment, the first image features of each tissue structure can characterize the visual content of the tissue structure involved in the sample medical image, and the sample medical report and the sample medical image should correspond to the same content. Therefore, by using the first image features of each tissue structure and the masked medical report, the model is guided to learn the generation probability of the masked first character, thereby reconstructing the sample medical report. Training the first processing network with the generation probability of the first character helps to optimize the first processing network in a direction that aligns the output image features of each tissue structure with the sample medical report, thus optimizing the first processing network in a direction that makes the output features conducive to generating the sample medical report. This facilitates the subsequent determination of the medical report generation model based on the trained second processing network, improving the performance of the medical report generation model.
[0190] In practical applications, the training methods shown in steps B1 to B2 can be combined with the training methods shown in steps C1 to C3. That is, after training the first processing network based on the third processing network according to steps B1 to B2, the trained first processing network can be further trained based on the fourth processing network according to steps C1 to C3 to obtain the second processing network. In this case, the aforementioned methods can be obtained respectively. (which can be called the second sub-loss) and (This can be called the first sub-loss), according to Determine the first loss The first processing network is trained based on the first loss to obtain the second processing network. Optionally, during the training of the first processing network, at least one of the third or fourth processing networks can be trained simultaneously.
[0191] Step 204: Based on the second processing network, obtain the medical report generation model, which is used to generate the target medical report.
[0192] Since the second processing network is trained based on the first processing network, they have the same structure and function, differing only in their network parameters. Optionally, the second processing network can serve as an image extraction network, followed by a text generation network to obtain a medical report generation model. The structure and parameters of the text generation network are not limited here. The second processing network extracts features from various tissue structures in medical images, with each tissue structure's features representing its visual information in the medical image. A text generation network can be appended after the second processing network to generate text content based on the features of any tissue structure, and the concatenation of these text contents yields the medical report.
[0193] In one possible implementation, step 204 includes steps D1 to D3 (not shown in the figure).
[0194] Step D1: Through the sample medical images and the multiple third query features included in the second processing network, determine the sample features of each tissue structure. The third query feature of any tissue structure is used to characterize the content characteristics of the image of any tissue structure, and the sample feature of any tissue structure is used to characterize the image content related to any tissue structure in the sample medical images.
[0195] The sample medical image can be input into the second processing network. For any tissue structure, the second processing network extracts features from the sample medical image based on the third query feature of the tissue structure, thus obtaining the sample features of the tissue structure. Since the second processing network and the first processing network differ only in their network parameters, and their structures and functions are the same, the implementation method of step D1 can be found in the description of step 202, and will not be repeated here. Because the third query feature of the tissue structure represents the semantic characteristics of the tissue structure, the feature extraction of the sample medical image is guided by the third query feature of the tissue structure, realizing the extraction of the visual information of the tissue structure from the sample medical image, thus obtaining the sample features of the tissue structure. Based on this, the sample features of the tissue structure can characterize the image content of the tissue structure in the sample medical image, including visual information such as the location, texture, color, and size of the tissue structure.
[0196] Step D2: Using the fifth processing network, the generation probability of each second character in the sample medical report is determined based on the sample characteristics of each tissue structure.
[0197] This application does not limit the structure or number of parameters of the fifth processing network. The fifth processing network may include at least one network layer such as convolutional layers, attention layers, activation layers, pooling layers, linear layers, and feedforward layers. Sample features of various organizational structures can be input into the fifth processing network to determine the generation probability of each second character. The generation probability of the second character is used to characterize the likelihood of determining the second character through the fifth processing network; the higher the generation probability of the second character, the higher the likelihood of determining the second character through the fifth processing network. Optionally, the generation probability of the second character is greater than or equal to 0 and less than or equal to 1. It is understood that different structures of the fifth processing network will result in different methods for determining the generation probability of the second character. One possible implementation is shown below.
[0198] In this example, the fifth processing network and the fourth processing network have the same structure. Optionally, the network parameters of the fifth processing network are initialized based on the network parameters of the fourth processing network to improve the training efficiency of the network. For ease of description, the implementation process of step D2 is illustrated below, taking an example where the fifth processing network includes at least one cascaded third extraction module (e.g., the fifth processing network includes six cascaded third extraction modules), and any third extraction module includes a self-attention layer, a cross-attention layer, and a feedforward layer. Step D2 includes steps D21 to D22 (not shown in the figure).
[0199] Step D21: For the first and second characters, the generation probability of the first and second characters is determined by the fifth processing network based on the sample features of each organizational structure.
[0200] In this embodiment, sample features and predefined features of various organizational structures can be input into a fifth processing network to determine the generation probability of the first second character. The predefined features are pre-set vectors, which can be zero vectors, unit vectors, etc. The implementation principle of step D21 is similar to that of step D22. Since the generation probability of the second character is described in detail in step D2, it will not be repeated here. Furthermore, the processing method for the predefined features is similar to the processing method for "all second characters before the first second character" in step D22.
[0201] Step D22: For non-first second characters, the fifth processing network determines the generation probability of non-first second characters based on the sample features of each organizational structure and each second character preceding the non-first second character.
[0202] In this embodiment, sample features of various organizational structures and all second characters preceding the first second character can be input into a fifth processing network to determine second target features. The second target features characterize the semantics of the text segment composed of each character and the visual information of the corresponding image. The fifth processing network also includes a prediction network, which determines the generation probability of the non-first second character based on the second target features.
[0203] Optionally, step D22 includes: determining semantic features based on each second character preceding the non-first second character through the fifth processing network, the semantic features being used to characterize the semantic content represented by each second character preceding the non-first second character; fusing the semantic features and sample features of each organizational structure through the fifth processing network to obtain a second fused feature; and determining the generation probability of the non-first second character based on the second fused feature through the fifth processing network.
[0204] For the first third extraction module of the fifth processing network, the character vectors of all second characters preceding the first second character and the sample features of each organizational structure are input into the first third extraction module. The input of the self-attention layer includes the character vectors of each second character. For any second character's character vector, the self-attention layer performs attention processing on the character vector of that second character and the character vectors of the remaining characters preceding it, obtaining the attention processing result for that second character. This attention processing result is used to describe the semantics of the second character in the context of the remaining characters. Through attention processing, information interaction between a second character and the remaining second characters is achieved to explore the relationship between the remaining second characters and the first second character, which is beneficial for extracting the semantics of the text segment composed of the various second characters. Based on this, the output of the self-attention layer can be denoted as semantic features, used to characterize the semantics of the text segment composed of the various second characters. The input of the cross-attention layer includes semantic features and sample features of each organizational structure. The cross-attention layer performs cross-attention processing on the semantic features and sample features of each organizational structure to obtain the second fusion feature. The second fusion feature represents the fused information after fusing the semantic content of the text segment composed of the various second characters and the visual information of each organizational structure. Learning the relationship between semantic content and visual information through a cross-attention layer facilitates the generation of semantic content guided by visual information, thereby increasing the probability of generating the second character. The second fused feature is input to the feedforward layer, which sequentially performs linear transformations and activations on the second fused feature to further extract features from the fused information, yielding the output of the feedforward layer. Extracting features from the fused information through the feedforward layer reduces feature complexity and improves feature representation capabilities. The output of the feedforward layer is also the output of the first third extraction module.
[0205] For the non-first third extraction module, the output of the previous third extraction module and the sample features of each organizational structure are input into the non-first third extraction module. Following the feature processing method used by the first third extraction module for the character vectors of each second character before the non-first second character and the sample features of each organizational structure, the non-first third extraction module performs feature processing on the output of the previous third extraction module and the sample features of each organizational structure to obtain the output of the non-first third extraction module.
[0206] The output of the final third extraction module is the second target feature mentioned above. This second target feature can be used to determine the generation probability of non-first second characters. By continuously extracting features from samples of each second character preceding the first one and each organizational structure, the semantic content of each second character and the visual information of each organizational structure are fully integrated. This allows the generation of semantic content to be guided by visual information, thus improving the generation probability of the second character. Subsequently, the generation probabilities of each second character are used to train a medical report generation model, which helps improve the accuracy of the medical report generation model.
[0207] Step D3: Based on the generation probability of each second character, train the neural network model to obtain the medical report generation model. The neural network model includes a second processing network and a fifth processing network.
[0208] In this embodiment, the second loss can be calculated based on the generation probability of each second character according to the function formula of the second loss. The second loss can be denoted as... The second loss is satisfied: N t Indicates the number of the second character. t k Represents the k-th second character, where k takes values from 1 to N. t A positive number in t. 1:k-1 Represents the first to k-1th second characters. P(t) k |t 1:k-1 The expression represents the probability of generating the k-th second character based on the first to k-1 second characters. This function can be any of the following: cross-entropy function, relative entropy function, mean square error function, etc., which will not be elaborated further here.
[0209] Next, the neural network model is trained once using the second loss to obtain the trained neural network model. If the trained neural network model meets the second termination condition, it is used as the medical report generation model. If the trained neural network model does not meet the second termination condition, it is used as the neural network model for the next training iteration, and the neural network model is trained again according to steps D1 to D3, until the trained neural network model meets the second termination condition and the medical report generation model is obtained.
[0210] This application does not limit the content of the trained neural network model satisfying the second termination condition. For example, the trained neural network model satisfying the second termination condition includes: the number of training iterations of the trained neural network model reaches a second number, or the performance index of the trained neural network model is not less than a second index, etc. Here, the second number or the second index can be a value preset based on human experience, or it can be a value input by the target object.
[0211] The medical report generation model is used to generate target medical reports based on target medical images. The generation method can be found in [link to relevant documentation]. Figure 5 The description of the relevant embodiments will not be repeated here.
[0212] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions. For example, the sample medical images and sample medical reports involved in this application were obtained with full authorization.
[0213] In the above method, the sample medical report describes the diagnostic results of at least one sample tissue structure in the sample medical image. That is, the sample medical report analyzes the sample medical image based on tissue structure, exhibiting a highly structured characteristic. Based on this, by using the sample medical image and the first query features of each tissue structure included in the first processing network, the first image features of each tissue structure are determined. This achieves the goal of guiding the analysis of the sample medical image based on the image content characteristics of the tissue structure, enabling the first image features of the tissue structure to characterize the image content of that tissue structure in the sample medical image. This not only improves the accuracy of the first image features but also makes the first image features more consistent with the structured characteristics of the sample medical report. By training the first processing network with the first image features of the sample medical report and each tissue structure, the text content and image content are aligned at the tissue structure level. This allows the trained second processing network to extract image features conducive to generating medical reports, thereby enabling the medical report generation model determined based on the second processing network to accurately generate medical reports and improve the efficiency of medical report generation.
[0214] like Figure 5 As shown, Figure 5 This is a flowchart illustrating a medical report generation method provided in an embodiment of this application. This method can be applied to the aforementioned computer system. For ease of description, the terminal device 101 or server 102 executing the medical report generation method in this embodiment is referred to as an electronic device, and this method can be executed by an electronic device. Figure 5 As shown, the method includes the following steps.
[0215] Step 501: Acquire the target medical image.
[0216] This application does not limit the method by which the electronic device acquires the target medical image. For example, the electronic device is connected to an acquisition device for acquiring the target medical image, and the acquisition device transmits the acquired target medical image to the electronic device in real time. Alternatively, the electronic device can acquire the target medical image input by the user, or the electronic device can acquire the target medical image via the Internet. The content of the target medical image is similar to the content of the sample medical image, as described in step 201, and will not be repeated here.
[0217] Step 502: Invoke the medical report generation model. Based on the target medical image and multiple target query features included in the medical report generation model, determine the target features of multiple tissue structures. Any target query feature is used to characterize the content characteristics of an image of a tissue structure, and any target feature of a tissue structure is used to characterize the image content in the target medical image related to any tissue structure.
[0218] Among them, the medical report generation model is based on and Figure 2 The relevant medical report generation model is obtained through training. The medical report generation model includes an image extraction network, which is trained based on either a first processing network or a second processing network. Therefore, the implementation principle of step 502 is similar to that of step 202, as described in the description of step 202, and will not be repeated here.
[0219] Optionally, step 502 includes: segmenting the target medical image to obtain multiple target image blocks; extracting image features of each target image block using a medical report generation model; and for any given tissue structure, determining the target features of that tissue structure using the image features of each target image block and the target query features of that tissue structure. The implementation principles of the aforementioned steps are similar to those of steps A1 to A3, and will not be repeated here.
[0220] Step 503: Invoke the medical report generation model and generate a target medical report based on the target features of each tissue structure. The target medical report is used to describe the diagnostic results of at least one target tissue structure in the target medical image. Multiple tissue structures include at least one target tissue structure.
[0221] In this embodiment, the medical report generation model further includes a text generation network, which can be trained based on the fifth processing network. Based on this, the text generation network, following the implementation principle shown in step D21, determines the first predicted feature based on the target features of each tissue structure. The first predicted feature is used to characterize the content after fusing the visual information of each tissue structure in the target medical image. The text generation network determines the generation probability of each character in the dictionary based on the first predicted feature, and the character with the highest generation probability is taken as the first character in the target medical report. Next, for non-first characters in the target medical report, the text generation network, following the implementation principle shown in step D22, determines non-first predicted features based on the target features of each tissue structure and the character vectors of the characters preceding the non-first character. The non-first predicted features are used to characterize the content after fusing the semantic content of the characters preceding the non-first character and the visual information of each tissue structure in the target medical image. The text generation network determines the generation probability of each character in the dictionary based on the non-first predicted features, and the character with the highest generation probability is taken as the non-first character in the target medical report.
[0222] Following the above method, the various characters in the target medical report are continuously generated, thus obtaining the target medical report. The target medical report describes the diagnostic result of at least one target tissue structure in the target medical image; the multiple tissue structures mentioned above include each target tissue structure. For example... Figure 6 As shown, medical images are input into the medical report generation model. The medical report generation model generates a medical report according to the implementation principles shown in steps 501 to 503. The medical report analyzes the medical images from the perspective of multiple tissue structures such as the chest, lungs, and heart, and has a highly structured feature.
[0223] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions. For example, the target medical images involved in this application were obtained with full authorization.
[0224] The above method determines the target features of each tissue structure by utilizing the target query features of the target medical image and the medical report generation model. This enables the image content characteristics of the tissue structure to guide the parsing of the target medical image, allowing the target features of the tissue structure to characterize the image content of that tissue structure in the target medical image, thus improving the accuracy of the target features. Subsequently, the medical report generation model can accurately and efficiently generate target medical reports based on the target features of each tissue structure. Furthermore, the structured nature of the target medical reports facilitates the understanding of each tissue structure.
[0225] The above describes the method for obtaining the medical report generation model and the method for generating medical reports according to the embodiments of this application from the perspective of method steps. The following describes the method in detail with reference to specific scenarios. The method of the embodiments of this application is applicable to the medical field. In practical applications, a medical report generation model can be trained based on sample medical images such as X-ray films, CT scans, MRI scans, and ultrasound scans, and corresponding sample medical reports, according to the method of the embodiments of this application. The corresponding target medical report can then be generated based on the target medical images such as X-ray films, CT scans, MRI scans, and ultrasound scans using the medical report generation model. For ease of description, the method of the embodiments of this application will be illustrated below using a scenario of CT scans (i.e., CT images) and CT reports as an example.
[0226] like Figure 7 As shown, in this example, the first processing network can be pre-trained based on sample CT images and sample CT reports to obtain the second processing network. Then, the second and fifth processing networks are optimized and trained to obtain the CT report generation model. The pre-training stage and the optimization training stage are described below. The pre-training stage includes steps 701 to 711.
[0227] Step 701: Uniformly segment the sample CT image to obtain multiple sample image blocks. The segmentation method can be found in the description of step A1, and will not be repeated here.
[0228] Step 702 involves processing each sample image block using a first processing network to obtain the first image features of each tissue structure. The implementation of step 702 can be found in the descriptions of steps A2 to A3, and will not be repeated here.
[0229] Step 703 involves updating the network parameters of the first processing network to obtain a first momentum network. This first momentum network is then used to process each sample image block to obtain the second image features of each tissue structure. The implementation of step 703 is described in step B21 and will not be repeated here.
[0230] Step 704: Segment the sample CT report according to the tissue structure to obtain multiple sample text segments. The segmentation method can be found in the description of step B11, and will not be repeated here.
[0231] Step 705 involves processing each sample text segment using a third processing network to obtain the first text features of each organizational structure. The implementation of step 705 is described in steps B12 and B13, and will not be repeated here.
[0232] Step 706 involves updating the network parameters of the third processing network to obtain a second momentum network. This second momentum network is then used to process each sample text segment to obtain the second text features of each organizational structure. The implementation of step 706 can be found in the description of step B22, and will not be repeated here.
[0233] Step 707: Based on the first image features, second image features, first text features, and second text features of each organizational structure, a first sub-loss is obtained. Wherein, the first sub-loss... The calculation method can be found in the description of step B23, and will not be repeated here.
[0234] Step 708: Obtain the masked CT report by masking the sample CT report. The masking method can be found in the description of step C1, and will not be repeated here.
[0235] Step 709: Using the fourth processing network, the generation probability of each masked character is determined based on the first image features of each tissue structure and the masked CT report. The masked character is the first character mentioned above. The implementation principle of step 709 can be found in the description of step C2, and will not be repeated here.
[0236] Step 710: Determine the second sub-loss based on the generation probability of each masked character. The second sub-loss... The calculation method can be found in the description of step C3, and will not be repeated here.
[0237] Step 711: Train the first processing network based on the first sub-loss and the second sub-loss to obtain the second processing network. Specifically, the first sub-loss... Second loss Perform a weighted summation to obtain the first loss. for example, The second processing network is obtained by training the first processing network using the first loss.
[0238] The pre-training process corresponding to steps 701 to 711 above can be as follows: Figure 8As shown. On one hand, sample CT images are input into a first processing network, which includes an encoder and a selection network. The encoder is used to segment the sample CT images to obtain multiple sample image blocks, and to extract features from each sample image block to obtain image features 801. The image features 801 of any sample image block are concatenated with the positional features of that sample image block to obtain the concatenated features of that sample image block. The first processing network includes first query features Q of multiple tissue structures. v The stitching features of each sample image patch and the first query feature Q of multiple tissue structures are combined. v The input is selected by a network, which determines the stitching features of each sample image patch and the first query feature Q of multiple tissue structures. v The first similarity between them. For each organizational structure, the selection network selects K first similarities not less than the similarity threshold from each first similarity. Based on the splicing features of the sample image blocks corresponding to the K first similarities, the first image features of the organizational structure are determined. In this way, the first image features 802 of each organizational structure are obtained. In addition, the selection network determines the first image features 803 of each organizational structure based on each first similarity and the splicing features of each sample image block, according to the implementation principle of formulas (1) and (2) mentioned above.
[0239] On the other hand, the sample CT report is input into a third processing network, which includes six cascaded first extraction modules. Each first extraction module includes a bidirectional self-attention layer and a feedforward layer. The third processing network is used to segment the sample CT report according to the tissue structure to obtain multiple sample text segments. For any sample text segment, the bidirectional self-attention layer of the first first extraction module performs self-attention processing on the segment, and the feedforward layer fuses the self-attention processing results to obtain the output of the first first extraction module. For subsequent first extraction modules, the output of the previous first extraction module is processed according to the processing method of the first first extraction module to obtain the output of the subsequent first extraction module. The output of the last first extraction module is the text feature of the sample text segment. In this way, the text features 804 of each sample text segment can be determined. In this example, the same tissue structure corresponds to at least two sample text segments. Based on this, cross-attention processing is performed on the text features 804 of each sample text segment corresponding to the same tissue structure to obtain the cross-attention processing result 805, which is used to characterize the feature similarity between every two text features 804. Furthermore, the third processing network includes second query features Q for multiple tissue structures. t Based on the implementation principles of formulas (3) and (4) mentioned above, the second query feature Q based on multiple organizational structures... t805. Cross-attention processing results, 804. Text features of each sample text segment, and 806. Determine the first text features of each organizational structure.
[0240] On the other hand, based on the Exponential Moving Average (EMA) method, a first momentum network is obtained by updating the momentum of the network parameters of the first processing network, and a second momentum network is obtained by updating the momentum of the network parameters of the third processing network based on the EMA method. Figure 8 The momentum network shown includes a first momentum network and a second momentum network. The momentum network can determine queue 1 based on sample CT images, following the processing method of the first processing network. Queue 1 includes the second image features of each tissue structure. The momentum network can also determine queue 2 based on sample CT reports, following the processing method of the third processing network. Queue 2 includes the second text features of each tissue structure. Then, based on the first image features 803, the second image features, the first text features 806, and the second text features of each tissue structure, a first sub-loss is obtained.
[0241] Furthermore, a masked CT report is obtained by masking the sample CT report. The masked CT report and the first image features 802 of each tissue structure are input into the fourth processing network. The fourth processing network includes six cascaded second extraction modules, each of which includes a bidirectional self-attention layer, a cross-attention layer, and a feedforward layer. The fourth processing network is used to segment the masked CT report according to the tissue structure to obtain multiple masked text segments. In this example, the same tissue structure corresponds to at least two masked text segments. Based on this, the bidirectional self-attention layer of the first second extraction module performs self-attention processing on each masked text segment corresponding to the same tissue structure. The cross-attention layer performs cross-attention processing on the self-attention processing result and the first image features 802 corresponding to the same tissue structure. The feedforward layer fuses the cross-attention processing results corresponding to each tissue structure to obtain the output of the first second extraction module. For non-first second extraction modules, the output of the previous second extraction module is processed according to the processing method of the first second extraction module to obtain the output of the non-first second extraction module. The generation probability of each masked character is determined based on the output of the last second extraction module, and the second sub-loss is determined based on these character generation probabilities.
[0242] Next, according to The first loss was calculated. The first processing network is trained by using the first loss to obtain the trained first processing network, and based on... Figure 8The pre-training process shown involves training the first processing network again after initial training. Through multiple training iterations, a second processing network is obtained. A fifth processing network is then concatenated after the second processing network to obtain a neural network model. This neural network model is then optimized and trained to obtain the medical report generation model. The optimization training phase includes steps 712 to 714.
[0243] Step 712: Process each sample image block using the second processing network to obtain sample features of each tissue structure. The processing method is described in step D1 and will not be repeated here.
[0244] Step 713: Based on the sample features of each tissue structure, the fifth processing network determines the generation probability of each character in the sample CT report. The implementation principle of step 713 can be found in the description of step D2, and will not be repeated here.
[0245] Step 714: Using the generation probabilities of each character in the sample CT report, train the second and fifth processing networks to obtain the CT report generation model. The implementation principle of step 714 can be found in the description of step D3, and will not be repeated here.
[0246] The optimization training process corresponding to steps 712 to 714 above can be as follows: Figure 9 As shown. In this embodiment, sample CT images are input into a second processing network. Since the second processing network is trained based on the first processing network, the functions of the second and first processing networks are similar. That is, the second processing network includes an encoder and a selection network. The encoder is used to extract features from each sample image block of the sample CT image to obtain image features 901 for each sample image block. The image features 901 of each sample image block are concatenated with the positional features to obtain the concatenated features of each sample image block. The second processing network includes third query features Q for multiple tissue structures. v The stitching features of each sample image patch and the third query feature Q of multiple tissue structures are combined. v Input the network selection and determine the sample features of each organizational structure 902 by selecting the network.
[0247] The fifth processing network consists of six cascaded third extraction modules, each of which includes a self-attention layer, a cross-attention layer, and a feedforward layer. Since the fifth processing network has a similar structure to the fourth processing network, its network parameters can be initialized using the parameters of the fourth processing network.
[0248] For the first character in the sample CT report, the character vector of the initial character (i.e., the defined features mentioned above) and the sample features 902 of each tissue structure are input into the fifth processing network. For the first third extraction module, the character vector of the initial character is processed by a self-attention layer, and the result of the self-attention processing and the sample features 902 of each tissue structure are processed by a cross-attention layer. The result of the cross-attention processing is then fused by a feedforward layer to obtain the output of the first third extraction module. For subsequent third extraction modules, the output of the previous third extraction module is processed according to the processing method of the first third extraction module to obtain the output of the subsequent third extraction module. The generation probability of the first character in the sample CT report is determined based on the output of the last third extraction module.
[0249] For characters not starting from the first character in the sample CT report, the character vectors of all characters preceding the first character and the sample features 902 of each tissue structure are input into the fifth processing network. Following the processing method of the fifth processing network for the initial character's character vector and the sample features 902 of each tissue structure, the fifth processing network processes the character vectors of each character and the sample features 902 of each tissue structure to obtain the generation probability of characters not starting from the first character in the sample CT report.
[0250] Following the above method, the generation probability of each character in the sample CT report can be determined. Based on the generation probability of each character, the second loss is determined. The neural network model includes a second processing network and a fifth processing network. The trained neural network model is obtained by training the neural network model using a second loss function, and based on... Figure 9 The optimized training process shown involves training the trained neural network model again. Through multiple training iterations, a CT report generation model is obtained.
[0251] In practical applications, such as Figure 10 As shown, the CT report generation model can be deployed on the backend. After frontend A acquires CT images, it can send the CT images to the backend. The backend inputs the CT images into the CT report generation model, generates a CT report, and sends the CT report to frontend B. Frontend A and frontend B can be the same or different frontends.
[0252] Since the CT report generation model is trained based on a neural network model, its processing of CT images (i.e., target CT images) is similar to how the neural network model processes sample CT images. Specifically, as follows... Figure 11 As shown.
[0253] In this example, the CT report generation model includes an image extraction network and a text generation network. The image extraction network is trained based on a second processing network, and the text generation network is trained based on a fifth processing network. Based on this, the target CT image is input into the image extraction network, which includes an encoder and a selection network. The encoder extracts features from each target image patch of the target CT image, obtaining image features 1001 for each target image patch. The image features 1001 of each target image patch are concatenated with location features to obtain the concatenated features of each target image patch. The image extraction network includes target query features Q for multiple tissue structures. v The stitching features of each target image patch and the target query features Q of multiple tissue structures are combined. v Input the network selection and determine the target features 1002 for each organizational structure by selecting the network.
[0254] The text generation network comprises six cascaded modules. Each extraction module includes a self-attention layer, a cross-attention layer, and a feedforward layer. The character vectors of the initial characters and the target features 1002 of each organizational structure are input into the text generation network. For the first module, the self-attention layer performs self-attention processing on the character vectors of the initial characters. The cross-attention layer performs cross-attention processing on the result of the self-attention processing and the target features 1002 of each organizational structure. The feedforward layer fuses the cross-attention processing result to obtain the output of the first module. For subsequent modules, the processing method of the first module is followed, and the output of the previous module is processed to obtain the output of the subsequent modules. Based on the output of the last module, the generation probability of each character in the dictionary is determined, and the character with the highest generation probability is selected as the first character in the target CT report.
[0255] For characters not starting with the first character in the target CT report, the character vectors of all characters preceding the first character and the target features 1002 of each tissue structure are input into the text generation network. Following the method used by the text generation network to generate the first character, the text generation network generates the non-first characters in the target CT report based on the character vectors of each character and the target features 1002 of each tissue structure.
[0256] By following the above method, each character in the target CT report is generated step by step, thus obtaining the target CT report.
[0257] In the pre-training phase, on the one hand, by selecting the first query features based on each tissue structure, the first image features of each tissue structure are determined, enabling the extraction of the most informative image content related to the tissue structure from sample CT images. Similarly, based on the second query features of each tissue structure, the first text features of each tissue structure are determined, enabling the extraction of the most informative text content related to the tissue structure from sample CT reports. Using the first text features and first image features of each tissue structure, a first sub-loss is determined, and the first processing network is trained based on this first sub-loss, aligning the image content and text content of the tissue structure, allowing the trained first processing network to extract features more accurately. On the other hand, using the first image features of each tissue structure, the generation probability of each masked character in the masked CT report is determined. Based on these generation probabilities, a second sub-loss is determined, and the first processing network is trained based on this second sub-loss. This helps optimize the first processing network in a direction that makes the extracted features more conducive to CT report generation, resulting in the trained first processing network extracting features that are more beneficial for CT report generation. Subsequently, by fine-tuning the first and fifth processing networks after training during the optimization training phase, the trained CT report generation model can accurately generate target CT reports, and the target CT reports have highly structured characteristics.
[0258] Furthermore, in the training environment of the CT report generation model in this application embodiment, several related models were also trained, which can also generate CT reports based on CT images. Optionally, the training environment is as follows: training is performed using four GPUs (Graphics Processing Units) with 32GB (gigabytes) of video memory. During training, the AdamW algorithm is used as the optimizer, with 600 training iterations in the pre-training phase and 150 training iterations in the optimization training phase. The initial learning rate is 1e-5 (i.e., 10 to the power of negative 5), and the number of training iterations in the warm-up phase is 10% of the total training iterations. After the warm-up phase, a linear learning rate scheduler is used to linearly adjust the learning rate. There are nine tissue structures: thoracic cage, ribs, lungs, heart, pleura, liver, kidneys, thyroid gland, and others (including tissue structures other than those mentioned above, such as spleen, spine, etc.). When calculating the first sub-loss, the parameter N mentioned above is used. q=64. A publicly available dataset was obtained via the internet, comprising 1804 chest CT images and corresponding CT reports. Each CT image contains approximately 300 images of 512×512 pixels. The dataset was randomly divided into training, validation, and test sets in a 3:1:1 ratio. The training set was used to adjust network parameters during the pre-training and optimization training phases, while the validation set was used to verify whether the model, after adjusting network parameters, exhibited overfitting or underfitting issues. Optionally, during the pre-training phase, the CT images were cropped to filter out background areas, resulting in foreground regions. The foreground regions were then downsampled by a multiple of 2 to further reduce memory consumption. The downsampled foreground regions were divided into 36 sample image blocks, and each sample image block underwent further processing. The test set was used to test the CT report generation model (i.e., Model 7) and related models (i.e., Models 1 to 6) of this application embodiment. The test results are shown in Table 1 below.
[0259] Table 1
[0260]
[0261]
[0262] Among them, BLEU-1 to BLEU-4 are four common evaluation metrics determined based on BLEU (Bilingual Evaluation Understudy). In this example, the higher the value of BLEU-1 to BLEU-4, the higher the model quality. METEOR (Metric for Evaluation of Translation with Explicit Ordering) is another evaluation metric. In this example, the higher the value of METEOR, the higher the model quality. ROUGE-L (Recall-Oriented Understudy for Gisting Evaluation - Longest Common Subsequence) is also an evaluation metric. In this example, the higher the value of ROUGE-L, the higher the model quality.
[0263] As can be seen from Table 1, the CT report generation model (i.e., Model 7) of this application embodiment achieved high values in all evaluation indicators, thus demonstrating high quality and the ability to generate high-quality CT reports. Figure 12As shown, CT Report 1 is the CT report corresponding to the CT images in the dataset, CT Report 2 is the CT report generated by the relevant model based on the CT images, and CT Report 3 is the CT report generated by the CT report generation model of this application based on the CT images. By comparing CT Reports 1 to 3, it can be seen that CT Report 3 is closer to CT Report 1 in terms of the division and description of tissue structure compared to CT Report 2. This indicates that the CT report generation model of this application can generate more accurate, more detailed, and more structured CT reports.
[0264] In addition, this application embodiment also verifies the impact of each loss on the CT report generation model through ablation experiments, and the verification results are shown in Table 2 below.
[0265] Table 2
[0266]
[0267] In Table 2, "1" corresponds to the first sub-loss. The "2" in Table 2 corresponds to the second sub-loss. In other words, model a is a CT report generation model trained using only the second loss, model b is a CT report generation model trained using the first and second sub-losses, model c is a CT report generation model trained using both the second and second sub-losses, and model d is a CT report generation model trained using the first, second, and third sub-losses. As shown in Table 2, model a has the worst quality. Models b and c are both of higher quality than model a, with model b having higher quality than model c, indicating that the first sub-loss improves model performance more effectively than the second sub-loss. Model d has higher quality than models a through c, indicating that the first, second, and third sub-losses used in the training process all improve model performance to varying degrees, thereby improving the quality of the generated CT reports.
[0268] Figure 13 The diagram shown is a structural schematic of a medical report generation model acquisition device provided in an embodiment of this application. Figure 13 As shown, the device includes:
[0269] The acquisition module 1301 is used to acquire sample medical images and sample medical reports, wherein the sample medical report is used to describe the diagnostic results of at least one sample tissue structure in the sample medical images;
[0270] The determination module 1302 is used to determine the first image features of multiple tissue structures through the sample medical image and the multiple first query features included in the first processing network. The multiple tissue structures include at least one sample tissue structure. Each first query feature represents the visual information of a tissue structure, and each first image feature of a tissue structure represents the image content in the sample medical image related to any tissue structure.
[0271] Training module 1303 is used to train a first processing network based on the first image features of sample medical reports and various tissue structures to obtain a second processing network;
[0272] The acquisition module 1304 is also used to acquire a medical report generation model based on the second processing network, and the medical report generation model is used to generate the target medical report.
[0273] In one possible implementation, the determining module 1302 is used to segment the sample medical image to obtain multiple sample image blocks; extract the image features of each sample image block; and for any tissue structure, determine the first image feature of any tissue structure by using the image features of each sample image block and the first query feature of any tissue structure.
[0274] In one possible implementation, the determining module 1302 is used to determine the first similarity between the image features of each sample image block and the first query feature of any organization structure; based on the first similarity, the first image feature of any organization structure is extracted from the image features of each sample image block.
[0275] In one possible implementation, the training module 1303 is used to determine the first text features of each organization structure through the sample medical report and multiple second query features included in the third processing network. Each second query feature represents the descriptive information of an organization structure, and each first text feature of an organization structure represents the text content related to any organization structure in the sample medical report. Based on the first text features and first image features of each organization structure, the first processing network is trained to obtain the second processing network.
[0276] In one possible implementation, the training module 1303 is used to segment the sample medical report to obtain multiple sample text segments, each of which is used to describe the diagnostic result of a sample tissue structure; extract the text features of each sample text segment; and determine the first text features of each tissue structure through the text features of each sample text segment and the second query features of multiple tissue structures.
[0277] In one possible implementation, the training module 1303 is used to determine, for any given organizational structure, a second similarity between the text features of each sample text segment and the second query features of any given organizational structure; and based on each second similarity, to extract the first text features of any given organizational structure from the text features of each sample text segment.
[0278] In one possible implementation, the training module 1303 is used to determine the second image features of each organizational structure through a first momentum network, the first momentum network having the same structure as the first processing network, and the network parameters of the first momentum network being updated through the network parameters of the first processing network; to determine the second text features of each organizational structure through a second momentum network, the second momentum network having the same structure as the third processing network, and the network parameters of the second momentum network being updated through the network parameters of the third processing network; and to train the first processing network based on the first text features, second text features, first image features, and second image features of each organizational structure to obtain the second processing network.
[0279] In one possible implementation, the training module 1303 is used to determine a third similarity between any first text feature and any second image feature, and obtain first annotation information representing whether any first text feature and any second image feature correspond to the same organizational structure; for any second text feature and any first image feature, determine a fourth similarity between any second text feature and any first image feature, and obtain second annotation information representing whether any second text feature and any first image feature correspond to the same organizational structure; and train a first processing network based on multiple third similarities, multiple fourth similarities, multiple first annotation information and multiple second annotation information to obtain a second processing network.
[0280] In one possible implementation, training module 1303 is used to determine a masked medical report based on a sample medical report, the masked medical report being obtained by masking multiple first characters in the sample medical report; a fourth processing network determines the generation probability of each first character based on the masked medical report and the first image features of each tissue structure; and a first processing network is trained based on the generation probability of each first character to obtain a second processing network.
[0281] In one possible implementation, the masked medical report includes multiple masked text segments, which are obtained by masking the first character in a sample text segment. The sample text segment is used to describe the diagnostic results of a sample tissue structure.
[0282] The training module 1303 is used to extract the text features of any masked text segment through the fourth processing network; determine the first image features of the sample organization structure corresponding to any masked text segment from the first image features of each organization structure through the fourth processing network; fuse the text features of any masked text segment and the first image features of the corresponding sample organization structure to obtain the first fused feature; and determine the generation probability of the first character in any masked text segment based on the first fused feature.
[0283] In one possible implementation, the acquisition module 1301 is used to determine the sample features of each tissue structure through the sample medical image and multiple third query features included in the second processing network. Each third query feature represents the visual information of a tissue structure, and each sample feature of a tissue structure represents the image content related to any tissue structure in the sample medical image. Based on the sample features of each tissue structure, the generation probability of each second character in the sample medical report is determined through the fifth processing network. Based on the generation probability of each second character, the neural network model is trained to obtain a medical report generation model, which includes the second processing network and the fifth processing network.
[0284] In one possible implementation, the acquisition module 1301 is used to determine the generation probability of the first second character based on the sample features of each organizational structure through the fifth processing network for the first second character; and to determine the generation probability of the non-first second character based on the sample features of each organizational structure and each second character preceding the non-first second character through the fifth processing network.
[0285] In one possible implementation, the acquisition module 1301 is used to determine semantic features based on each second character before the non-first second character through a fifth processing network, wherein the semantic features represent the semantic content represented by each second character before the non-first second character; to fuse the semantic features and sample features of each organizational structure through the fifth processing network to obtain a second fused feature; and to determine the generation probability of the non-first second character based on the second fused feature through the fifth processing network.
[0286] In the aforementioned apparatus, the sample medical report describes the diagnostic results of at least one sample tissue structure in the sample medical image. In other words, the sample medical report analyzes the sample medical image based on tissue structure, exhibiting a highly structured characteristic. Based on this, by using the sample medical image and the first query features of each tissue structure included in the first processing network, the first image features of each tissue structure are determined. This enables the analysis of the sample medical image to be guided by the image content characteristics of the tissue structure, allowing the first image features of the tissue structure to characterize the image content of that tissue structure in the sample medical image. This not only improves the accuracy of the first image features but also makes them more consistent with the structured characteristics of the sample medical report. By training the first processing network using the sample medical report and the first image features of each tissue structure, the text content and image content are aligned at the tissue structure level. This allows the trained second processing network to extract image features conducive to generating medical reports, thereby enabling the medical report generation model determined based on the second processing network to accurately generate medical reports and improving the efficiency of medical report generation.
[0287] It should be understood that the above Figure 13 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.
[0288] Figure 14 The diagram shown is a structural schematic of a medical report generation device provided in an embodiment of this application. Figure 14 As shown, the device includes:
[0289] Acquisition module 1401 is used to acquire target medical images;
[0290] Module 1402 is used to invoke the medical report generation model. Based on the target medical image and multiple target query features included in the medical report generation model, it determines the target features of multiple tissue structures. The medical report generation model then uses these features in conjunction with... Figure 2 The relevant method implementation examples are trained to obtain that any target query feature is used to characterize the content characteristics of an image of an tissue structure, and any target feature of an tissue structure is used to characterize the image content of a target medical image related to any tissue structure.
[0291] The generation module 1403 is used to call the medical report generation model to generate a target medical report based on the target features of each tissue structure. The target medical report is used to describe the diagnostic results of at least one target tissue structure in the target medical image, and the multiple tissue structures include at least one target tissue structure.
[0292] In one possible implementation, the determination module 1402 is used to call the medical report generation model to segment the target medical image and obtain multiple target image blocks; extract the image features of each target image block; and for any organization structure, determine the target features of any organization structure by using the image features of each target image block and the target query features of any organization structure.
[0293] In the aforementioned device, by utilizing the target query features of various tissue structures included in the target medical image and medical report generation model, the target features of each tissue structure are determined. This enables the analysis of the target medical image to be guided by the image content characteristics of the tissue structure, allowing the target features of the tissue structure to characterize the image content of that tissue structure in the target medical image, thus improving the accuracy of the target features. Subsequently, the target medical report can be accurately and efficiently generated based on the target features of each tissue structure using the medical report generation model. Furthermore, the target medical report has a structured characteristic, which is beneficial for understanding the situation of each tissue structure.
[0294] It should be understood that the above Figure 14 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.
[0295] Figure 15 A structural block diagram of a terminal device 1500 provided in an exemplary embodiment of this application is shown. The terminal device 1500 includes a processor 1501 and a memory 1502.
[0296] Processor 1501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1501 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1501 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0297] The memory 1502 may include one or more computer-readable storage media, which may be non-transitory. The memory 1502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1502 are used to store at least one computer program, which is executed by the processor 1501 to implement the method for obtaining the medical report generation model and the method for generating medical reports provided in the method embodiments of this application.
[0298] In some embodiments, the terminal device 1500 may also optionally include: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, memory 1502, and peripheral device interface 1503 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1503 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1504, a display screen 1505, a camera assembly 1506, an audio circuit 1507, and a power supply 1508.
[0299] Peripheral interface 1503 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1501 and memory 1502. In some embodiments, processor 1501, memory 1502 and peripheral interface 1503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1501, memory 1502 and peripheral interface 1503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0300] The radio frequency (RF) circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1504 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1504 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1504 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0301] Display screen 1505 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1505 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1501 for processing. In this case, display screen 1505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1505, disposed on the front panel of terminal device 1500; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal device 1500 or in a folded design; in still other embodiments, display screen 1505 may be a flexible display screen, disposed on a curved or folded surface of terminal device 1500. Furthermore, display screen 1505 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1505 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0302] The camera assembly 1506 is used to acquire images or videos. Optionally, the camera assembly 1506 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0303] The audio circuit 1507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1501 for processing, or input to the radio frequency circuit 1504 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal device 1500. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1501 or the radio frequency circuit 1504 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1507 may also include a headphone jack.
[0304] Power supply 1508 is used to power the various components in terminal device 1500. Power supply 1508 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1508 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0305] In some embodiments, the terminal device 1500 further includes one or more sensors 1509. The one or more sensors 1509 include, but are not limited to: an acceleration sensor 1511, a gyroscope sensor 1512, a pressure sensor 1513, an optical sensor 1514, and a proximity sensor 1515.
[0306] Accelerometer 1511 can detect the magnitude of acceleration along the three axes of a coordinate system established by terminal device 1500. For example, accelerometer 1511 can be used to detect the components of gravitational acceleration along the three axes. Processor 1501 can control display screen 1505 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1511. Accelerometer 1511 can also be used for games or for acquiring user motion data.
[0307] The gyroscope sensor 1512 can detect the orientation and rotation angle of the terminal device 1500. The gyroscope sensor 1512, in conjunction with the accelerometer sensor 1511, can collect 3D motion data from the user on the terminal device 1500. Based on the data collected by the gyroscope sensor 1512, the processor 1501 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0308] The pressure sensor 1513 can be disposed on the side bezel of the terminal device 1500 and / or on the lower layer of the display screen 1505. When the pressure sensor 1513 is disposed on the side bezel of the terminal device 1500, it can detect the user's grip signal on the terminal device 1500, and the processor 1501 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1513. When the pressure sensor 1513 is disposed on the lower layer of the display screen 1505, the processor 1501 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1505. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0309] Optical sensor 1514 is used to collect ambient light intensity. In one embodiment, processor 1501 can control the display brightness of display screen 1505 based on the ambient light intensity collected by optical sensor 1514. Specifically, when the ambient light intensity is high, the display brightness of display screen 1505 is increased; when the ambient light intensity is low, the display brightness of display screen 1505 is decreased. In another embodiment, processor 1501 can also dynamically adjust the shooting parameters of camera assembly 1506 based on the ambient light intensity collected by optical sensor 1514.
[0310] The proximity sensor 1515, also known as a distance sensor, is typically located on the front panel of the terminal device 1500. The proximity sensor 1515 is used to detect the distance between the user and the front of the terminal device 1500. In one embodiment, when the proximity sensor 1515 detects that the distance between the user and the front of the terminal device 1500 is gradually decreasing, the processor 1501 controls the display screen 1505 to switch from a screen-on state to a screen-off state; when the proximity sensor 1515 detects that the distance between the user and the front of the terminal device 1500 is gradually increasing, the processor 1501 controls the display screen 1505 to switch from a screen-off state to a screen-on state.
[0311] Those skilled in the art will understand that Figure 15 The structure shown does not constitute a limitation on the terminal device 1500, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0312] Figure 16This is a schematic diagram of the server structure provided in the embodiments of this application. The server 1600 can vary considerably due to different configurations or performance. It may include one or more processors 1601 and one or more memories 1602. The one or more memories 1602 store at least one computer program, which is loaded and executed by the one or more processors 1601 to implement the medical report generation model acquisition method and medical report generation method provided in the above-described method embodiments. For example, the processor 1601 is a CPU. Of course, the server 1600 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1600 may also include other components for implementing device functions, which will not be elaborated here.
[0313] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program, which is loaded and executed by a processor to enable an electronic device to implement the above-described method for obtaining a medical report generation model and the method for generating a medical report.
[0314] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0315] In an exemplary embodiment, a computer program is also provided, which is at least one such computer program, loaded and executed by a processor, to enable an electronic device to implement the above-described method for obtaining a medical report generation model and the method for generating a medical report.
[0316] In an exemplary embodiment, a computer program product is also provided, which stores at least one computer program that is loaded and executed by a processor to enable an electronic device to implement the above-described method for obtaining a medical report generation model and the method for generating a medical report.
[0317] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0318] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0319] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for obtaining a medical report generation model, characterized in that, The method includes: Acquire sample medical images and sample medical reports, wherein the sample medical reports are used to describe the diagnostic results of at least one sample tissue structure in the sample medical images; The sample medical images and the first processing network include multiple first query features to determine the first image features of multiple tissue structures, the multiple tissue structures including the at least one sample tissue structure, each first query feature characterizes the visual information of a tissue structure, and each first image feature of a tissue structure characterizes the image content in the sample medical images related to the tissue structure. Based on the sample medical report and the first image features of each tissue structure, the first processing network is trained to obtain the second processing network; Based on the second processing network, a medical report generation model is obtained, which is used to generate the target medical report.
2. The method according to claim 1, characterized in that, The process of determining first image features of multiple tissue structures using the sample medical images and multiple first query features included in the first processing network includes: The medical image sample is segmented to obtain multiple sample image blocks; Extract image features from each sample image patch; For any given organizational structure, the first image feature of the given organizational structure is determined by combining the image features of each sample image block with the first query feature of the given organizational structure.
3. The method according to claim 2, characterized in that, The step of determining the first image feature of any organizational structure by using the image features of each sample image block and the first query feature of any organizational structure includes: Determine the first similarity between the image features of each sample image block and the first query feature of any organizational structure; Based on each first similarity, the first image feature of any one tissue structure is extracted from the image features of each sample image block.
4. The method according to claim 1, characterized in that, The process of training the first processing network based on the first image features of the sample medical report and the various tissue structures to obtain the second processing network includes: The first text features of each organizational structure are determined by the sample medical report and the third processing network including multiple second query features. Each second query feature represents the descriptive information of an organizational structure, and the first text feature of any organizational structure represents the text content related to the organizational structure in the sample medical report. Based on the first text features and first image features of each organizational structure, the first processing network is trained to obtain the second processing network.
5. The method according to claim 4, characterized in that, The method of determining the first textual features of each tissue structure through multiple second query features included in the sample medical report and the third processing network includes: The sample medical report is segmented to obtain multiple sample text segments, and any one of the sample text segments is used to describe the diagnostic result of a sample tissue structure. Extract text features from each sample text segment; The first text features of each organizational structure are determined by using the text features of each sample text segment and the second query features of multiple organizational structures.
6. The method according to claim 5, characterized in that, The step of determining the first text features of each organizational structure by using the text features of each sample text segment and the second query features of multiple organizational structures includes: For any given organizational structure, determine the second similarity between the text features of each sample text segment and the second query feature of the given organizational structure; Based on each second similarity, the first text feature of any organizational structure is extracted from the text features of each sample text segment.
7. The method according to claim 4, characterized in that, The process of training the first processing network based on the first text features and first image features of each organizational structure to obtain the second processing network includes: The second image features of each tissue structure are determined by a first momentum network. The first momentum network has the same structure as the first processing network. The network parameters of the first momentum network are updated by the network parameters of the first processing network. The second text features of each organizational structure are determined by a second momentum network, which has the same structure as the third processing network. The network parameters of the second momentum network are updated by the network parameters of the third processing network. Based on the first text features, second text features, first image features, and second image features of each organizational structure, the first processing network is trained to obtain the second processing network.
8. The method according to claim 7, characterized in that, The process of training the first processing network based on the first text features, second text features, first image features, and second image features of the various organizational structures to obtain the second processing network includes: For any first text feature and any second image feature, determine the third similarity between the first text feature and the second image feature, and obtain first annotation information characterizing whether the first text feature and the second image feature correspond to the same organizational structure; For any second text feature and any first image feature, determine the fourth similarity between the second text feature and the first image feature, and obtain second annotation information characterizing whether the second text feature and the first image feature correspond to the same organizational structure; The first processing network is trained based on multiple third similarity values, multiple fourth similarity values, multiple first annotation information, and multiple second annotation information to obtain the second processing network.
9. The method according to claim 1, characterized in that, The process of training the first processing network based on the first image features of the sample medical report and the various tissue structures to obtain the second processing network includes: A masked medical report is determined based on the sample medical report, which is obtained by masking multiple first characters in the sample medical report; The fourth processing network determines the generation probability of each first character based on the first image features of the masked medical report and each tissue structure. Based on the generation probability of each first character, the first processing network is trained to obtain the second processing network.
10. The method according to claim 9, characterized in that, The masked medical report includes multiple masked text segments, which are obtained by masking the first character in a sample text segment. The sample text segment is used to describe the diagnostic result of a sample tissue structure. The step of determining the generation probability of each first character based on the masked medical report and the first image features of each tissue structure using the fourth processing network includes: For any masked text segment, the text features of the masked text segment are extracted through the fourth processing network; The fourth processing network determines the first image features of the sample organization structure corresponding to any masked text segment from the first image features of each organization structure, and fuses the text features of any masked text segment and the first image features of the corresponding sample organization structure to obtain the first fused feature. Based on the first fusion feature, the generation probability of the first character in any masked text segment is determined.
11. The method according to any one of claims 1 to 10, characterized in that, The step of obtaining a medical report generation model based on the second processing network includes: The sample features of each tissue structure are determined by the sample medical image and the multiple third query features included in the second processing network. Each third query feature represents the visual information of a tissue structure, and the sample features of any tissue structure represent the image content in the sample medical image related to the tissue structure. The fifth processing network determines the generation probability of each second character in the sample medical report based on the sample features of each tissue structure. Based on the generation probability of each of the second characters, a neural network model is trained to obtain a medical report generation model, wherein the neural network model includes the second processing network and the fifth processing network.
12. The method according to claim 11, characterized in that, The step of determining the generation probability of each second character in the sample medical report based on the sample features of each tissue structure through the fifth processing network includes: For the first and second characters, the generation probability of the first and second characters is determined by the fifth processing network based on the sample features of each organizational structure; For a non-first second character, the fifth processing network determines the generation probability of the non-first second character based on the sample features of each organizational structure and each second character preceding the non-first second character.
13. The method according to claim 12, characterized in that, The step of determining the generation probability of the non-first second character through the fifth processing network, based on the sample features of each organizational structure and each second character preceding the non-first second character, includes: The fifth processing network determines semantic features based on each second character preceding the non-first second character, and the semantic features represent the semantic content represented by each second character preceding the non-first second character. The second fused feature is obtained by fusing the semantic features and the sample features of each organizational structure through the fifth processing network. The fifth processing network determines the generation probability of the non-first second character based on the second fusion feature.
14. A method for generating medical reports, characterized in that, The method includes: Acquire target medical images; The medical report generation model is invoked, and based on the target medical image and multiple target query features included in the medical report generation model, target features of multiple tissue structures are determined. The medical report generation model is trained according to the method described in any one of claims 1 to 13. Each target query feature represents the visual information of a tissue structure, and each target feature of a tissue structure represents the image content in the target medical image related to the tissue structure. The medical report generation model is invoked to generate a target medical report based on the target features of each tissue structure. The target medical report is used to describe the diagnostic results of at least one target tissue structure in the target medical image, and the multiple tissue structures include the at least one target tissue structure.
15. The method according to claim 14, characterized in that, The medical report generation model, based on the target medical image and multiple target query features included in the medical report generation model, determines target features of multiple tissue structures, including: The target medical image is segmented using a medical report generation model to obtain multiple target image blocks; Extract image features from each target image patch; For any given organizational structure, the target features of that organizational structure are determined by combining the image features of each target image block with the target query features of that organizational structure.
16. A device for acquiring a medical report generation model, characterized in that, The device includes: An acquisition module is used to acquire sample medical images and sample medical reports, wherein the sample medical report is used to describe the diagnostic results of at least one sample tissue structure in the sample medical images; The determination module is used to determine first image features of multiple tissue structures through the sample medical image and multiple first query features included in the first processing network. The multiple tissue structures include the at least one sample tissue structure. Each first query feature represents the visual information of a tissue structure, and each first image feature of a tissue structure represents the image content in the sample medical image related to the tissue structure. The training module is used to train the first processing network based on the first image features of the sample medical report and the various tissue structures to obtain the second processing network; The acquisition module is further configured to acquire a medical report generation model based on the second processing network, the medical report generation model being used to generate a target medical report.
17. A medical report generation device, characterized in that, The device includes: The acquisition module is used to acquire the target medical image; The determination module is used to call the medical report generation model, and determine the target features of multiple tissue structures based on the target medical image and multiple target query features included in the medical report generation model. The medical report generation model is trained according to the method of any one of claims 1 to 13. Each target query feature represents the visual information of a tissue structure, and each target feature of a tissue structure represents the image content in the target medical image related to the tissue structure. The generation module is used to call the medical report generation model to generate a target medical report based on the target features of each tissue structure. The target medical report is used to describe the diagnostic results of at least one target tissue structure in the target medical image, and the plurality of tissue structures includes the at least one target tissue structure.
18. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to perform the method as described in any one of claims 1 to 15.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the electronic device to perform the method as described in any one of claims 1 to 15.
20. A computer program product, characterized in that, The computer program product stores at least one computer program, which is loaded and executed by a processor to enable the electronic device to perform the method as described in any one of claims 1 to 15.