Medical imaging report generation method, device, storage medium and computer equipment

By combining pre-trained image and text feature extraction models with a cross-attention mechanism for feature fusion, the problem of insufficient professionalism and accuracy of medical imaging reports in existing technologies is solved, and more professional and accurate medical imaging reports are generated.

CN118262855BActive Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410460966.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-17
Publication Date
2025-09-16
Estimated Expiration
2044-04-17

AI Technical Summary

Technical Problem

Existing automatic medical imaging report generation algorithms lack professionalism and accuracy, especially when processing medical images with single colors and subtle image differences, making it difficult to generate high-quality reports.

Method used

Pre-trained image feature extraction models and text feature extraction models are used to extract features from medical images and text data, and the cross-attention mechanism is combined for feature fusion. Medical imaging reports are generated through pre-trained disease type classification modules and report generation models.

Benefits of technology

It improves the professionalism and accuracy of medical imaging reports, can better describe abnormal disease sites, and generate professional and accurate reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118262855B_ABST
    Figure CN118262855B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical imaging report generation method, device, storage medium and computer equipment, and relates to the field of medical data processing technology. The method includes: obtaining medical imaging data and medical text data to be identified, performing feature extraction on the medical imaging data based on a pre-trained image feature extraction model to obtain image features, and performing feature extraction on the medical text data based on a pre-trained text feature extraction model to obtain text features. Then, the image features and text features are fused based on a pre-trained feature fusion model to obtain image-text fusion features. The image-text fusion features are classified and processed by a pre-trained disease type classification module to obtain disease features. Finally, a medical imaging report is generated using a pre-trained report generation model based on the image-text fusion features and disease features. The above method can generate accurate and professional medical imaging reports, providing a reference for doctors to read films.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical data processing, and in particular to a method, device, storage medium and computer equipment for generating a medical imaging report. Background Art

[0002] To assist radiologists and alleviate their burden, a growing number of studies are using artificial intelligence (AI) to automatically generate medical imaging reports. Existing algorithms for automatically generating medical imaging reports are primarily based on techniques used in image captioning for natural images. These algorithms extract image features using convolutional neural networks (CNNs), then feed them into recurrent neural networks (RNNs) or neural networks designed to process natural language. These features are then converted into corresponding text descriptions to produce the medical imaging report.

[0003] However, medical images are different from natural images. Medical images have a single color and the differences between images are relatively subtle. At the same time, medical imaging reports usually consist of a large number of long sentences including professional terms. Due to the particularity of the medical field, the professionalism and accuracy of imaging reports generated by existing solutions are insufficient. Summary of the Invention

[0004] In view of this, the present application provides a medical imaging report generation method, apparatus, storage medium and computer equipment, the main purpose of which is to solve the technical problem that medical imaging reports generated by traditional methods have relatively poor professionalism and accuracy.

[0005] According to a first aspect of the present invention, a method for generating a medical imaging report is provided, the method comprising:

[0006] Acquire medical image data and medical text data to be identified, wherein the medical text data includes medical history text data corresponding to the user to whom the medical image data belongs;

[0007] Perform feature extraction on medical image data based on a pre-trained image feature extraction model to obtain image features, and perform feature extraction on medical text data based on a pre-trained text feature extraction model to obtain text features;

[0008] Based on the pre-trained feature fusion model, image features and text features are fused to obtain image-text fusion features;

[0009] Classify the image-text fusion features through a pre-trained disease type classification module to obtain disease features, where the disease features include at least one disease type;

[0010] Based on the image-text fusion features and disease characteristics, a pre-trained report generation model is used to generate medical imaging reports.

[0011] Optionally, the medical image data to be identified is medical image data taken by the target user at multiple shooting angles, and the image features are used to characterize abnormal parts in the medical image data. The training method of the image feature extraction model includes:

[0012] Acquire medical imaging data taken by multiple sample users at multiple angles as medical imaging data samples, and construct a first pre-training model for image feature extraction based on the medical imaging data samples;

[0013] Taking medical imaging data samples as input and feature-level comparison as the goal, the teacher-student architecture is trained using the first pre-training model to obtain an image feature extraction model.

[0014] Optional training methods for text feature extraction models include:

[0015] Using the encoder module, a second pre-trained model for text feature extraction is constructed;

[0016] Acquire medical text data of multiple sample users as medical text data samples;

[0017] The second pre-training model is optimized and trained based on the medical text data samples to obtain a text feature extraction model.

[0018] Optional training methods for the feature fusion model and disease type classification module include:

[0019] Acquire medical image data taken by multiple sample users at multiple angles as medical image data samples, and acquire medical text data of multiple sample users as medical text data samples;

[0020] Perform feature extraction on medical image data samples based on an image feature extraction model to obtain image feature samples, and perform feature extraction on medical text data samples based on a text feature extraction model to obtain text feature samples;

[0021] Using the encoding and decoding model based on the cross-attention mechanism, a third pre-training model is constructed to fuse image features and text features;

[0022] Optimize and train the third pre-trained model based on the image feature samples and the text feature samples to obtain a feature fusion model;

[0023] Based on the feature fusion model, image feature samples and text feature samples are fused to obtain image-text fusion feature samples;

[0024] A multi-classification neural network is connected to the feature fusion model, and the multi-classification neural network is trained based on the image-text fusion feature samples to obtain a disease type classification module, where the output of the feature fusion model is the input of the multi-classification neural network.

[0025] Optionally, image features and text features are fused based on a pre-trained feature fusion model to obtain image-text fusion features, including:

[0026] The text feature is used as the query vector, and the image feature is used as the key vector and value vector. The attention distribution of the text feature to the image feature is calculated. The image feature vector is weighted and summed according to the attention distribution of the text feature to the image feature to obtain the fusion feature vector of the text feature and the image feature.

[0027] The image feature is used as the query vector, and the text feature is used as the key vector and value vector. The attention distribution of the image feature to the text feature is calculated. The text feature vector is weighted and summed according to the attention distribution of the image feature to the text feature to obtain the fusion feature vector of the image feature and the text feature.

[0028] The fusion feature vector of text features and image features and the fusion feature vector of image features and text features are weighted summed to obtain the image-text fusion feature.

[0029] Optionally, the image-text fusion features are classified using a pre-trained disease type classification module to obtain disease features, including:

[0030] Input the image-text fusion features into the disease type classification module to obtain the probability distribution of multiple disease types, and determine the classification results of the disease type based on the probability distribution and the preset classification conditions;

[0031] The disease status corresponding to each disease type is obtained based on the classification results, and the disease characteristics are obtained based on the disease type and disease status, wherein the disease status is used to characterize whether the disease type meets the classification conditions.

[0032] Optionally, the report generation model is an encoding / decoding model based on a superposition of multiple layers of masked multi-head self-attention layers and feedforward neural network layers. The pre-trained report generation model is used to generate a medical imaging report based on the image-text fusion features and disease characteristics, including:

[0033] The fusion features and disease features are input into the report generation model. The report generation model predicts the next word by calculating the hidden state of each word position, and then generates a complete medical imaging report.

[0034] According to a second aspect of the present invention, there is provided a medical imaging report generating device, the device comprising:

[0035] A data acquisition module is used to acquire medical image data and medical text data to be identified, wherein the medical text data includes medical history text data corresponding to the user to whom the medical image data belongs;

[0036] A feature extraction module is used to extract features from medical image data based on a pre-trained image feature extraction model to obtain image features, and to extract features from medical text data based on a pre-trained text feature extraction model to obtain text features;

[0037] The feature fusion module is used to fuse image features and text features based on the pre-trained feature fusion model to obtain image-text fusion features;

[0038] A classification module is used to classify the image-text fusion features using a pre-trained disease type classification module to obtain disease features, wherein the disease features include at least one disease type;

[0039] The report generation module is used to generate medical imaging reports based on image and text fusion features and disease characteristics using a pre-trained report generation model.

[0040] According to a third aspect of the present invention, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned medical imaging report generation method is implemented.

[0041] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned medical imaging report generating method when executing the program.

[0042] The present invention provides a medical imaging report generation method, apparatus, storage medium, and computer device. The method first obtains medical image data and medical text data to be identified, wherein the medical text data includes medical history text data corresponding to the user to whom the medical image data belongs. Feature extraction is performed on the medical image data based on a pretrained image feature extraction model to obtain image features, and feature extraction is performed on the medical text data based on a pretrained text feature extraction model to obtain text features. The medical feature extraction model can process single-color medical images to obtain image features that characterize abnormal areas in the medical images. The text features obtained from the medical text data contain the target user's past medical information. Comprehensive consideration of image and text features can improve the accuracy of report diagnosis and presentation. The method then fuses the image and text features based on a pretrained feature fusion model to obtain image-text fusion features. Feature fusion based on a cross-attention mechanism can complement and enhance the medical and text features, resulting in a richer and more comprehensive feature representation. The image-text fusion features are classified and processed using a pretrained disease type classification module to obtain disease features, wherein the disease features include at least one disease type. Finally, a medical imaging report is generated using the pretrained report generation model based on the image-text fusion features and disease features. The combination of fusion features and disease features can enable the model to focus on abnormal disease sites, enabling the model to better describe disease information and thus generate accurate medical imaging reports.

[0043] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0045] Figure 1 A schematic diagram showing a flow chart of a method for generating a medical imaging report provided by an embodiment of the present invention;

[0046] Figure 2 A schematic diagram of a flow chart of a method for fusing image and text features provided by an embodiment of the present invention is shown;

[0047] Figure 3 A schematic structural diagram of a medical imaging report generating device provided by an embodiment of the present invention is shown;

[0048] Figure 4A schematic structural diagram of a storage device and a computer device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0049] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0050] The terms used in the various embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "first", "second", "third", etc. in the description and claims of the present application are used to distinguish different objects and are not limited to describing a specific order and the scope of the embodiments of the present application.

[0051] In each embodiment of the present application, the user data obtained, such as medical imaging data, medical text data, etc., are all data obtained after user authorization. The present application obtains user data through a server or computer program for data processing and will not disclose the user's privacy data to a third party.

[0052] In one embodiment, Figure 1 As shown, a method for generating a medical imaging report is provided, which is described by taking the method applied to a computer device such as a server as an example, and includes the following steps:

[0053] 101. Obtain medical image data and medical text data to be identified.

[0054] Among them, the medical image data to be identified is the medical image data taken by the target user at multiple shooting angles, such as the data corresponding to multiple chest X-ray medical images taken from multiple directions such as the front, left, and right of the target user, or the data corresponding to multiple chest tomography medical images of the target user obtained based on electronic computed tomography. Obtaining medical image data of the target user from multiple angles can obtain rich medical image data features, and comprehensive analysis of medical image features from multiple angles can more comprehensively and accurately extract abnormal parts from the medical image data. Correspondingly, the medical text data includes the medical history text data corresponding to the user to whom the medical image data belongs, that is, the medical history text data of the target user. For example, after taking a medical image of user A, the medical image data of user A needs to be processed. At this time, the medical image data of user A and the medical history text data of user A are obtained, and then the medical image data of user A is analyzed in combination with the medical history text data.

[0055] 102. Perform feature extraction on medical image data based on a pre-trained image feature extraction model to obtain image features, and perform feature extraction on medical text data based on a pre-trained text feature extraction model to obtain text features.

[0056] In this embodiment, the medical feature extraction model can process single-color medical images by performing comparative learning on medical image data to obtain image features that characterize abnormal portions of the medical image data. For example, the medical feature extraction model can compare the target user's medical image data with normal medical image data to extract abnormal portions of the medical image data and obtain image features. Normal medical image data can be understood as medical image data corresponding to healthy body parts / organs. The extracted image features can be represented by 0 or 1. For example, data in the target user's medical image data that is identical to the normal medical image data can be marked as 0, and data that is different from the normal medical image data can be marked as 1, thereby obtaining the target user's image features. The "0" and "1" are merely examples to illustrate the embodiments of this application and do not constitute a limitation of the embodiments of this application. Accordingly, the text feature extraction model can extract features from the medical text data by learning from the medical text data to obtain text features containing the target user's past medical information. The text features can be represented as the segmented words obtained by the text feature extraction model after encoding the medical text data and the weights corresponding to each segmented word.

[0057] 103. Based on the pre-trained feature fusion model, image features and text features are fused to obtain image-text fusion features.

[0058] In this embodiment, a feature fusion model based on a cross-attention mechanism is used to fuse image features and text features to obtain image and text fusion features that reinforce each other. The image and text fusion features can be used to strengthen the features of abnormal parts, so that the model can focus on abnormal parts in subsequent processing. Figure 2 As shown in the figure, the image features and text features are respectively strengthened through the self-attention mechanism, and then the text features are strengthened with the image features and vice versa through the cross-attention mechanism. The mutually strengthened features are superimposed and normalized to obtain the image-text fusion features.

[0059] 104. Classify the image-text fusion features using a pre-trained disease type classification module to obtain disease features, wherein the disease features include at least one disease type.

[0060] In this embodiment, after obtaining the image-text fusion features, the image-text fusion features are input into a pre-trained disease type classification module, and the image-text fusion features are classified and processed by the disease type classification module to obtain disease features. The disease features are used to characterize the types of diseases that may exist in medical images. Classification is performed based on the fusion features obtained after complementing and reinforcing the image features and text features, so that the obtained disease types can be more accurate. Among them, the classified disease types can be represented by a vector consisting of at least one preset disease type and the disease state corresponding to each preset disease type. The disease state is used to indicate whether the corresponding preset disease type exists in the medical image of the target user. The disease state can be represented by 0 or 1. 0 indicates that the corresponding disease type does not meet the classification results, that is, the disease type does not exist in the medical image, and 1 indicates that the corresponding disease type meets the classification results, that is, the disease type exists in the medical image. For example, there are fourteen preset disease types, from disease type 1 to disease type 14. If only the preset disease type 1 meets the classification results after classifying the fusion features, the disease status of disease type 1 is set to 1, and the disease statuses of the remaining preset disease types 2 to disease type 14 are all set to 0. The disease features include each preset disease type and its corresponding disease status.

[0061] 105. Generate medical imaging reports using the pre-trained report generation model based on the image-text fusion features and disease characteristics.

[0062] In this embodiment, after obtaining the image-text fusion features and disease features, the image-text fusion features and disease features are input into the report generation model. The report generation model divides the fusion features and disease features into multiple position vectors, interacts at different positions through the self-attention mechanism, splices the fusion features and disease features, and then generates a medical imaging report based on the dependency relationship between each word and the context. The combination of fusion features and disease features can enable the model to focus on the abnormal parts of the disease, allowing the model to better describe disease information. Because the fusion features contain the content of text features extracted from the medical history text, the fusion features contain professional terms and professional text descriptions extracted from the medical history text, and the disease type in the disease feature is a preset disease type, which also uses professional terms. Therefore, accurate and professional medical imaging reports can be generated based on the fusion features and disease features.

[0063] The medical imaging report generation method provided in this embodiment first obtains medical image data and medical text data to be identified, wherein the medical text data includes medical history text data corresponding to the user to whom the medical image data belongs. Feature extraction is performed on the medical image data based on a pretrained image feature extraction model to obtain image features, and feature extraction is performed on the medical text data based on a pretrained text feature extraction model to obtain text features. The medical feature extraction model can process single-color medical images to obtain image features that characterize abnormal areas in the medical images. The text features obtained from the medical text data contain the target user's past medical information. Comprehensive consideration of image and text features can improve the accuracy of report diagnosis and presentation. The image and text features are then fused based on a pretrained feature fusion model to obtain image-text fusion features. Feature fusion based on a cross-attention mechanism can complement and enhance the medical and text features, resulting in a richer and more comprehensive feature representation. The image-text fusion features are classified and processed using a pretrained disease type classification module to obtain disease features, wherein the disease features include at least one disease type. Finally, a medical imaging report is generated using the pretrained report generation model based on the image-text fusion features and disease features. The combination of fusion features and disease features can enable the model to focus on abnormal disease sites, enabling the model to better describe disease information, thereby generating accurate medical imaging reports and providing reference for doctors to read films.

[0064] Furthermore, in order to fully illustrate the implementation process of this embodiment, the specific implementation methods of the above-mentioned embodiment are refined and expanded. In an optional embodiment, the training method of the image feature extraction model in step 102 can be specifically implemented by the following steps: obtaining medical imaging data taken by multiple sample users at multiple angles as medical imaging data samples, and constructing a first pre-training model for image feature extraction based on the medical imaging data samples; using the medical imaging data samples as input and feature-level comparison as the goal, using the first pre-training model to train the teacher-student architecture to obtain an image feature extraction model.

[0065] In the above embodiment, the feature extraction model can select a teacher-student architecture model for training. The student model in the teacher-student architecture is trained to imitate the behavior and output of the teacher model. The teacher model has strong performance and expression capabilities and can extract and capture more subtle features in the input data. The student model can obtain rich feature representations by imitating the output of the teacher model, thereby improving the performance and generalization ability of the model. Therefore, the feature extraction model based on the teacher-student architecture can make the feature level division finer, thereby obtaining more subtle feature representations. Regarding the training of the teacher-student architecture, first, medical imaging reports of multiple sample users taken at multiple angles are obtained as medical imaging data samples. The model that can be used for image feature extraction is trained using the medical imaging data samples to obtain a first pre-trained model, wherein the model for image feature extraction can select a neural network model. Then, with the medical imaging data samples as input and feature level comparison as the goal, the first pre-trained model is used as the teacher model in the teacher-student architecture to train the student model in the teacher-student architecture to obtain an image feature extraction model. The image feature extraction model extracts abnormal parts from the target user's medical imaging data by comparing the feature differences between the target user's medical imaging data and normal medical imaging data, and converts them into image features for output.

[0066] In an optional embodiment, the training method of the text feature extraction model in step 102 can be specifically implemented by the following steps: using the encoder module to construct a second pre-trained model for text feature extraction; obtaining medical text data of multiple sample users as medical text data samples; and optimizing and training the second pre-trained model based on the medical text data samples to obtain a text feature extraction model.

[0067] In the above embodiment, medical text data of multiple sample users is first obtained as medical text data samples, where each sample user is the same as the sample user from which the medical imaging data was obtained in the aforementioned embodiment. A second pre-trained model for text feature extraction is then determined. This second pre-trained model can utilize the encoding module of a transformer model. The second pre-trained model is trained using the medical text data samples, and the parameters of the second pre-trained model are optimized to obtain a text feature extraction model.

[0068] In an optional embodiment, the training method of the feature fusion model and disease type classification module in step 103 and step 104 can be specifically implemented by the following steps: obtaining medical image data taken by multiple sample users at multiple angles as medical image data samples, and obtaining medical text data of multiple sample users as medical text data samples; performing feature extraction on the medical image data samples based on the image feature extraction model to obtain image feature samples, and performing feature extraction on the medical text data samples based on the text feature extraction model to obtain text feature samples; using a coding and decoding model based on a cross-attention mechanism to construct a third pre-training model for fusing image features and text features; optimizing and training the third pre-training model based on the image feature samples and the text feature samples to obtain a feature fusion model; fusing the image feature samples and the text feature samples based on the feature fusion model to obtain image-text fusion feature samples; connecting a multi-classification neural network to the feature fusion model, and training the multi-classification neural network based on the image-text fusion feature samples to obtain a disease type classification module, wherein the output of the feature fusion model is the input of the multi-classification neural network.

[0069] In the above embodiment, medical image data and medical text data captured at multiple angles for multiple sample users are first obtained as medical image data samples and medical text data samples, respectively. Feature extraction is then performed on each medical image data sample and medical text data sample using a trained image feature extraction model and text feature extraction model, resulting in multiple image feature samples and multiple text feature samples. The image feature samples and text feature samples for each sample user correspond one-to-one. A coding / decoding model based on a cross-attention mechanism is then determined as a third pre-trained model for fusing image features and text features. The third pre-trained model is then trained using the image feature samples and text feature samples of each sample user, and the parameters of the third pre-trained model are adjusted to obtain a graphic-text feature fusion model.

[0070] After obtaining the image-text feature fusion model, based on a method similar to the above embodiment, the image feature samples and text feature samples of each sample user are fused through the trained image-text feature fusion model to obtain multiple image-text fusion feature samples. A multi-classification neural network is connected to the output position of the image-text feature fusion model as a classification module, and the classification module is trained with the image-text feature fusion samples to obtain a disease type classification module. Taking the transformer model as an example of an encoding and decoding model based on a cross-attention mechanism for feature fusion, a softmax layer can be connected after the output position of the transformer model as a classification module for classifying the image-text feature fusion model.

[0071] In an optional embodiment, in step 103, the image features and text features are fused based on a pre-trained feature fusion model to obtain a method for obtaining a text-image fusion feature. Specifically, the method can be implemented by the following steps: taking the text feature as the query vector, the image feature as the key vector and the value vector, calculating the attention distribution of the text feature to the image feature, and weighted summing the image feature vector according to the attention distribution of the text feature to the image feature to obtain a fusion feature vector of the text feature to the image feature; taking the image feature as the query vector, the text feature as the key vector and the value vector, calculating the attention distribution of the image feature to the text feature, and weighted summing the text feature vector according to the attention distribution of the image feature to the text feature to obtain a fusion feature vector of the image feature to the text feature; and performing weighted summing on the fusion feature vector of the text feature to the image feature and the fusion feature vector of the image feature to the text feature to obtain a text-image fusion feature.

[0072] In the above embodiment, the fusion of image features and text features based on the cross-attention mechanism can be implemented based on the transformer model, using the text feature as the query vector, the image feature as the key vector and the value vector, and calculating the attention distribution of the text feature to the image feature. The image feature vector is weighted and summed according to the attention distribution of the text feature to the image feature to obtain a fusion feature vector of the text feature to the image feature, wherein the fusion feature vector of the text feature to the image feature represents the fusion feature vector obtained by strengthening the text feature to the image feature. At the same time, using the image feature as the query vector, the text feature as the key vector and the value vector, and calculating the attention distribution of the image feature to the text feature, the text feature vector is weighted and summed according to the attention distribution of the image feature to the text feature to obtain a fusion feature vector of the image feature to the text feature, wherein the fusion feature vector of the image feature to the text feature represents the fusion feature vector obtained by strengthening the image feature to the text feature. Then, the fusion feature vector obtained by strengthening the text feature to the image feature and the fusion feature vector obtained by strengthening the image feature to the text feature are weighted and summed to obtain the image-text fusion feature.

[0073] In an optional embodiment, in step 104, the image-text fusion features are classified and processed by a pre-trained disease type classification module, and the method for obtaining disease characteristics can be specifically implemented by the following steps: inputting the image-text fusion features into the disease type classification module to obtain the probability distribution of multiple disease types, and determining the classification results of the disease types based on the probability distribution and preset classification conditions; obtaining the disease state corresponding to each disease type based on the classification results, and obtaining disease characteristics based on the disease type and disease state, wherein the disease state is used to characterize whether the disease type meets the classification conditions.

[0074] In the above embodiment, the image-text fusion features are input into the disease type classification module, and the correlation between the image-text fusion features and various disease types is calculated to obtain a probability distribution of multiple disease types. The disease type classification result is determined based on the probability distribution of each disease type and a preset classification condition, where the preset condition can be a preset probability threshold. Disease types with a probability greater than or equal to the preset threshold are considered to be disease types present in the medical image, i.e., target disease types, and disease types with a probability less than the preset threshold are considered to be disease types not present in the medical image, i.e., non-target disease types. The preset condition can also be set to select disease types based on a ranking of probability values. For example, the probability values ​​of each disease type are sorted from large to small, and the disease types corresponding to the top three probability values ​​are selected as target disease types, while the remaining disease types are considered non-target disease types. Based on the classification results, the disease status corresponding to each disease type is obtained. The disease status of the target disease type can be represented by 1, indicating that the disease type meets the preset classification condition, and the disease status of the non-target disease type can be represented by 0, indicating that the disease type does not meet the preset classification condition. Disease characteristics are obtained based on each disease type and the disease status corresponding to each disease type.

[0075] In an optional embodiment, the report generation model in step 105 is an encoding and decoding model based on the superposition of multiple layers of masked multi-head self-attention layers and feedforward neural network layers; the method of generating a medical imaging report using a pre-trained report generation model based on image-text fusion features and disease features can be specifically implemented by the following steps: inputting the fusion features and disease features into the report generation model, and the report generation model predicts the next word by calculating the hidden state of each word position, thereby generating a complete medical imaging report.

[0076] In the above embodiment, the transformer model can be used as the encoding and decoding model based on the superposition of multi-layer masked multi-head self-attention layers and feedforward neural network layers. The report generation model learns the dependencies between different positions in the fusion features and disease feature sequences, and encodes the information of each position of the input sequence on other positions by calculating the attention score of each position with other positions. During the decoding process, the report generation model generates a representation vector for each position in the sequence through the self-attention mechanism. The representation vector of the next position is then generated by combining the output of the encoder. Finally, the representation vector output by the decoder is mapped to the corresponding text through linear transformation and softmax function. This mechanism enables the model to consider the position information of the entire input sequence at the same time, rather than being limited to the local context, and can flexibly generate text with contextual consistency.

[0077] Further, as Figure 1 The specific implementation of the method shown in this embodiment provides a medical imaging report generating device, such as Figure 3 As shown, the device includes: a data acquisition module 31, a feature extraction module 32, a feature fusion module 33, a classification module 34, and a report generation module 35.

[0078] The data acquisition module 31 is configured to acquire medical image data and medical text data to be identified, wherein the medical text data includes medical history text data corresponding to the user to whom the medical image data belongs; the medical image data to be identified is medical image data taken by the target user at multiple shooting angles, and the image features are used to characterize abnormal parts in the medical image data;

[0079] A feature extraction module 32 may be used to extract features from medical image data based on a pre-trained image feature extraction model to obtain image features, and to extract features from medical text data based on a pre-trained text feature extraction model to obtain text features;

[0080] The feature fusion module 33 can be used to fuse image features and text features based on a pre-trained feature fusion model to obtain image-text fusion features;

[0081] The classification module 34 may be used to classify the image-text fusion features using a pre-trained disease type classification module to obtain disease features, wherein the disease features include at least one disease type;

[0082] The report generation module 35 can be used to generate a medical imaging report based on the image-text fusion features and disease characteristics using a pre-trained report generation model, wherein the report generation model is an encoding and decoding model based on the superposition of a multi-layer masked multi-head self-attention layer and a feedforward neural network layer.

[0083] In a specific application scenario, the feature extraction module 32 can be used to obtain medical imaging data taken by multiple sample users at multiple angles as medical imaging data samples, and construct a first pre-training model for image feature extraction based on the medical imaging data samples; using the medical imaging data samples as input and feature-level comparison as the goal, the teacher-student architecture is trained using the first pre-training model to obtain an image feature extraction model.

[0084] In a specific application scenario, the feature extraction module 32 can also be used to use the encoder module to construct a second pre-trained model for text feature extraction; obtain medical text data of multiple sample users as medical text data samples; and optimize and train the second pre-trained model based on the medical text data samples to obtain a text feature extraction model.

[0085] In a specific application scenario, the feature fusion module 33 can be specifically used to obtain medical image data taken by multiple sample users at multiple angles as medical image data samples, and obtain medical text data of multiple sample users as medical text data samples; perform feature extraction on the medical image data samples based on the image feature extraction model to obtain image feature samples, and perform feature extraction on the medical text data samples based on the text feature extraction model to obtain text feature samples; utilize the encoding and decoding model based on the cross-attention mechanism to construct a third pre-training model for fusing image features and text features; optimize and train the third pre-training model based on the image feature samples and the text feature samples to obtain a feature fusion model; fuse the image feature samples and the text feature samples based on the feature fusion model to obtain a graphic-text fusion feature sample.

[0086] In a specific application scenario, the classification module 34 can be used to connect a multi-classification neural network to the feature fusion model, train the multi-classification neural network based on the image-text fusion feature samples, and obtain a disease type classification module, wherein the output of the feature fusion model is the input of the multi-classification neural network.

[0087] In a specific application scenario, the feature fusion module 33 can also be used to use text features as query vectors, image features as key vectors and value vectors, calculate the attention distribution of text features to image features, and perform weighted summation of image feature vectors based on the attention distribution of text features to image features to obtain a fused feature vector of text features to image features; use image features as query vectors, text features as key vectors and value vectors, calculate the attention distribution of image features to text features, and perform weighted summation of text feature vectors based on the attention distribution of image features to text features to obtain a fused feature vector of image features to text features; perform weighted summation of the fused feature vector of text features to image features and the fused feature vector of image features to text features to obtain a text-image fusion feature.

[0088] In a specific application scenario, the classification module 34 can also be used to input the image and text fusion features into the disease type classification module to obtain the probability distribution of multiple disease types, and determine the classification results of the disease types based on the probability distribution and preset classification conditions; obtain the disease state corresponding to each disease type based on the classification results, and obtain the disease characteristics based on the disease type and disease state, wherein the disease state is used to characterize whether the disease type meets the classification conditions.

[0089] In a specific application scenario, the report generation module 35 can be used to input the fusion features and disease features into the report generation model. The report generation model predicts the next word by calculating the hidden state of each word position, thereby generating a complete medical imaging report.

[0090] It should be noted that for other corresponding descriptions of the functional units involved in the medical imaging report generation device provided in this embodiment, please refer to Figure 1 and Figure 2 The corresponding description in will not be repeated here.

[0091] Based on the above Figure 1 and Figure 2 The method shown in FIG. 1 is a method for performing the above-mentioned operation. Accordingly, this embodiment further provides a storage medium on which a computer program is stored. When the program is executed by a processor, the above-mentioned Figure 1 and Figure 2 The medical imaging report generation method shown.

[0092] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product. The software product to be identified can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each implementation scenario of the present application.

[0093] Based on the above Figure 1 and Figure 2 The method shown, and Figure 3 In order to achieve the above-mentioned purpose, the medical image report generating device embodiment shown in FIG. Figure 4 As shown, this embodiment also provides a computer device for generating medical imaging reports, which can be a personal computer, a server, a smart phone, a tablet computer, a smart watch, or other network devices, etc. The computer device includes a storage medium and a processor; the storage medium is used to store computer programs and an operating system; the processor is used to execute the computer program to achieve the above-mentioned Figure 1 The method shown.

[0094] Optionally, the computer device may further include an internal memory, a communication interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, a display, an input device such as a keyboard, etc. Optionally, the communication interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Wi-Fi interface), etc.

[0095] Those skilled in the art will understand that the computer device structure for identifying an operation action provided in this embodiment does not constitute a limitation on the computer device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0096] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the computer device hardware and the software resources to be identified, supporting the execution of the information processing program and other software and / or programs to be identified. The network communication module is used to enable communication between components within the storage medium and with other hardware and software in the information processing computer device.

[0097] From the above description of the embodiments, those skilled in the art will clearly understand that this application can be implemented using software plus the necessary general-purpose hardware platform, or it can be implemented using hardware. By applying the technical solution of this application, medical image data and medical text data to be identified are first obtained, wherein the medical text data includes the medical history text data corresponding to the user to whom the medical image data belongs. Feature extraction is performed on the medical image data based on a pre-trained image feature extraction model to obtain image features, and feature extraction is performed on the medical text data based on a pre-trained text feature extraction model to obtain text features. The image features and text features are then fused based on a pre-trained feature fusion model to obtain image-text fusion features. The image-text fusion features are classified and processed using a pre-trained disease type classification module to obtain disease features, wherein the disease features include at least one disease type. Finally, a medical image report is generated using a pre-trained report generation model based on the image-text fusion features and disease features. Compared to existing technologies, the medical feature extraction model can process single-color medical images to obtain image features that characterize abnormal areas in the medical images. The text features obtained from the medical text data include the target user's previous medical information. The combined consideration of image and text features can improve the accuracy of the report diagnosis and presentation. The feature fusion method based on the cross-attention mechanism can complement and strengthen medical features and text features to obtain a richer and more comprehensive feature representation. The combination of fused features and disease features can enable the model to focus on abnormal disease sites. The fused features and disease features contain professional terms and professional text expressions, which enable the model to better describe disease information, thereby generating accurate medical imaging reports and providing reference for doctors to read films.

[0098] Those skilled in the art will understand that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required to implement the present application. Those skilled in the art will understand that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the implementation scenario description, or can be changed accordingly and located in one or more devices different from the implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0099] The serial numbers of the above application are for descriptive purposes only and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure only discloses several specific implementation scenarios of the present application, but the present application is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present application.

Claims

1. A method for generating a medical imaging report, characterized in that: The method comprises: Acquire medical image data and medical text data to be identified, wherein the medical text data includes medical history text data corresponding to the user to whom the medical image data belongs; Performing feature extraction on the medical image data based on a pre-trained image feature extraction model to obtain image features, and performing feature extraction on the medical text data based on a pre-trained text feature extraction model to obtain text features; The image features and text features are fused based on a pre-trained feature fusion model to obtain image-text fusion features; Inputting the image-text fusion features into a pre-trained disease type classification module to obtain a probability distribution of multiple disease types, and determining a classification result of the disease type based on the probability distribution and preset classification conditions; obtaining a disease state corresponding to each disease type based on the classification results, and obtaining a disease feature based on the disease type and disease state, wherein the disease state is used to indicate whether the disease type meets the classification conditions, and the disease feature includes at least one disease type; Inputting the image-text fusion features and the disease features into a pre-trained report generation model, the report generation model predicts the next word by calculating the hidden state of each word position, and then generates a complete medical imaging report, wherein the report generation model is an encoding and decoding model based on the superposition of multi-layer masked multi-head self-attention layers and feedforward neural network layers; The training method of the feature fusion model and disease type classification module includes: Obtain medical image data taken by multiple sample users at multiple angles as medical image data samples, and obtain medical text data of multiple sample users as medical text data samples; perform feature extraction on the medical image data samples based on the image feature extraction model to obtain image feature samples, and perform feature extraction on the medical text data samples based on the text feature extraction model to obtain text feature samples; use an encoding and decoding model based on a cross-attention mechanism to construct a third pre-training model for fusing image features and text features; optimize and train the third pre-training model based on the image feature samples and text feature samples to obtain the feature fusion model; fuse the image feature samples and text feature samples based on the feature fusion model to obtain image-text fusion feature samples; connect a multi-classification neural network to the feature fusion model, train the multi-classification neural network based on the image-text fusion feature samples, and obtain the disease type classification module, wherein the output of the feature fusion model is the input of the multi-classification neural network.

2. The method according to claim 1, characterized in that The medical image data to be identified is medical image data taken by the target user at multiple shooting angles, and the image features are used to characterize abnormal parts in the medical image data; The training method of the image feature extraction model includes: Acquire medical image data taken by multiple sample users at multiple angles as medical image data samples, and construct a first pre-training model for image feature extraction based on the medical image data samples; Taking the medical image data sample as input and feature-level comparison as the goal, the teacher-student architecture is trained using the first pre-training model to obtain the image feature extraction model.

3. The method according to claim 1, characterized in that The training method of the text feature extraction model includes: Using the encoder module, a second pre-trained model for text feature extraction is constructed; Acquire medical text data of multiple sample users as medical text data samples; The second pre-training model is optimized and trained based on the medical text data sample to obtain the text feature extraction model.

4. The method according to claim 1 or 3, characterized in that The pre-trained feature fusion model is used to fuse the image features and text features to obtain image-text fusion features, including: Using the text feature as a query vector and the image feature as a key vector and a value vector, calculating the attention distribution of the text feature to the image feature, and performing weighted summation of the image feature vectors according to the attention distribution of the text feature to the image feature to obtain a fusion feature vector of the text feature to the image feature; Using the image feature as a query vector and the text feature as a key vector and a value vector, calculating the attention distribution of the image feature to the text feature, and performing weighted summation of the text feature vectors according to the attention distribution of the image feature to the text feature to obtain a fusion feature vector of the image feature to the text feature; The image-text fusion feature is obtained by performing weighted summation on the fusion feature vector of the text feature and the image feature and the fusion feature vector of the image feature and the text feature.

5. A medical imaging report generating device, characterized in that: The device comprises: A data acquisition module, configured to acquire medical image data to be identified and medical text data, wherein the medical text data includes medical history text data corresponding to the user to whom the medical image data belongs; a feature extraction module, configured to extract features from the medical image data based on a pre-trained image feature extraction model to obtain image features, and to extract features from the medical text data based on a pre-trained text feature extraction model to obtain text features; A feature fusion module, configured to fuse the image features and text features based on a pre-trained feature fusion model to obtain image-text fusion features; a classification module, configured to input the image-text fusion features into a pre-trained disease type classification module to obtain a probability distribution of multiple disease types, and determine a classification result of the disease type based on the probability distribution and preset classification conditions; obtain a disease state corresponding to each disease type based on the classification result, and obtain a disease feature based on the disease type and disease state, wherein the disease state is used to indicate whether the disease type meets the classification conditions, and the disease feature includes at least one disease type; A report generation module, configured to input the image-text fusion features and the disease features into a pre-trained report generation model. The report generation model predicts the next word by calculating the hidden state of each word position, thereby generating a complete medical imaging report. The report generation model is an encoding-decoding model based on a superposition of multi-layer masked multi-head self-attention layers and feedforward neural network layers. The feature fusion module is specifically used to obtain medical image data taken by multiple sample users at multiple angles as medical image data samples, and obtain medical text data of multiple sample users as medical text data samples; perform feature extraction on the medical image data samples based on the image feature extraction model to obtain image feature samples, and perform feature extraction on the medical text data samples based on the text feature extraction model to obtain text feature samples; use the encoding and decoding model based on the cross-attention mechanism to construct a third pre-training model for fusing image features and text features; optimize and train the third pre-training model based on the image feature samples and text feature samples to obtain the feature fusion model; fuse the image feature samples and text feature samples based on the feature fusion model to obtain image-text fusion feature samples; The classification module is specifically used to connect a multi-classification neural network to the feature fusion model, train the multi-classification neural network based on the image-text fusion feature samples, and obtain the disease type classification module, wherein the output of the feature fusion model is the input of the multi-classification neural network.

6. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Convolutional localization networks for intelligent captioning of medical images

    US20210241884A1

  • Medical report generation method and apparatus, model training method and apparatus, and device

    WO2023029817A1