A medical image processing method and device, electronic equipment and storage medium
By combining medical images and clinical text reports to generate image processing strategies, and using deep neural networks for precise registration and fusion, the problems of misalignment and blurred fusion of different modal medical images in traditional methods are solved, thereby improving the credibility and interpretability of the images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI MEDICAL IMAGE INSIGHTS INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-03
AI Technical Summary
In existing technologies, different modal medical images differ in terms of tissue contrast, resolution, and noise levels. Traditional registration methods lead to misalignment and blurred fusion, resulting in fused images that lack credibility and clinical interpretability.
By acquiring medical images and clinical text reports of the target area, an image processing strategy is generated using a language model-based strategy generation model. Combining visual and semantic features, image registration and fusion are performed, and a deep neural network model is used for accurate registration and information fusion.
It significantly improves the accuracy of anatomical structure registration between images of different modalities, enhances the credibility and interpretability of fused images, and provides more reliable clinical diagnostic evidence.
Smart Images

Figure CN122335708A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a medical image processing method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] Multimodal imaging technology can integrate the advantages of different imaging modalities in tissue contrast, metabolic function, and spatial structure, thereby improving the reliability of disease diagnosis. However, existing technologies still have the following shortcomings: 1. Images from different modalities vary greatly in terms of tissue contrast, resolution, and noise levels. Traditional registration methods (based on feature points or deep networks) often suffer from misalignment and blurred fusion. 2. Current medical image fusion typically relies solely on visual features and focuses on the fusion effect, resulting in a lack of reliability and clinical interpretability of the fused images. Summary of the Invention
[0003] This invention provides a medical image processing method, apparatus, electronic device, storage medium, and program product, which can significantly improve the registration accuracy of anatomical structures between images of different modalities and effectively enhance the credibility and interpretability of fused images.
[0004] According to one aspect of the present invention, a medical image processing method is provided, the method comprising: A first medical image of the target area and a corresponding clinical text report are acquired. Based on a pre-built strategy generation model, an image processing strategy is generated according to the first medical image and the clinical text report. The first medical image includes at least two medical images of different modalities. The strategy generation model is built based on a language model. The image processing strategy includes an image registration strategy and an image fusion strategy. The first medical image is registered according to the image registration strategy and the pre-built image registration model to obtain a second medical image and a deformation field; wherein, the image registration model is constructed based on a deep neural network model; the deformation field is used to describe the degree of displacement of each voxel in the first medical image during the registration process; The second medical image is fused according to the image fusion strategy, the deformation field, and the pre-built image fusion model to obtain a third medical image; wherein the image fusion model is constructed based on a deep neural network model.
[0005] According to another aspect of the present invention, a medical image processing apparatus is provided, the apparatus comprising: A strategy generation module is used to acquire a first medical image of the target area and a corresponding clinical text report of the first medical image, and generate an image processing strategy based on the first medical image and the clinical text report according to a pre-built strategy generation model; wherein, the first medical image includes at least two medical images of different modalities; the strategy generation model is built based on a language model; the image processing strategy includes an image registration strategy and an image fusion strategy; An image registration module is used to register the first medical image according to the image registration strategy and a pre-built image registration model to obtain a second medical image and a deformation field; wherein, the image registration model is constructed based on a deep neural network model; the deformation field is used to describe the displacement degree of each voxel in the first medical image during the registration process; An image fusion module is used to fuse the second medical image according to the image fusion strategy, the deformation field, and a pre-built image fusion model to obtain a third medical image; wherein the image fusion model is constructed based on a deep neural network model.
[0006] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the medical image processing method according to any embodiment of the present invention.
[0007] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the medical image processing method according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the medical image processing method as described in any of the embodiments of the present disclosure.
[0009] The technical solution of this invention involves acquiring a first medical image of a target area and a corresponding clinical text report, and generating an image processing strategy based on a pre-built strategy generation model, according to the first medical image and the clinical text report. The first medical image includes at least two medical images of different modalities. The strategy generation model is built based on a language model. The image processing strategy includes an image registration strategy and an image fusion strategy. The first medical image is registered according to the image registration strategy and the pre-built image registration model to obtain a second medical image and a deformation field. The image registration model is built based on a deep neural network model. The deformation field describes the displacement of each voxel in the first medical image during the registration process. The second medical image is fused according to the image fusion strategy, the deformation field, and the pre-built image fusion model to obtain a third medical image. The image fusion model is built based on a deep neural network model. The technical solution of this invention utilizes a strategy generation model based on a language model, combined with visual features in the first medical image and semantic features in the clinical text report, to generate an image processing strategy for the first medical image. This strategy guides subsequent image registration and image fusion processes, significantly improving the registration accuracy of anatomical structures between different modalities in the first medical image and effectively enhancing the credibility and interpretability of the fused image.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a medical image processing method provided according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of a medical image processing method provided according to Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the structure of a medical image processing device according to Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the medical image processing method of the present invention. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0015] Example 1 Figure 1 This is a flowchart of a medical image processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to processing multiple medical images of different modalities of a target area. The method can be executed by a medical image processing device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S110. Obtain a first medical image of the target area and a corresponding clinical text report, and generate an image processing strategy based on the first medical image and the clinical text report according to a pre-built strategy generation model; wherein, the first medical image includes at least two medical images of different modalities; the strategy generation model is built based on a language model; the image processing strategy includes an image registration strategy and an image fusion strategy.
[0016] The first medical image includes at least two medical images of different modalities, such as computed tomography (CT) images, magnetic resonance imaging (MRI) images, positron emission tomography (PET) images, ultrasound images, and digital radiography (DR) images. The corresponding clinical text report for the first medical image includes an explanatory description of the first medical image, such as a radiological diagnosis description and lesion location markings.
[0017] The strategy generation model, built upon a language model, outputs structured processing suggestions, i.e., image processing strategies, for the first medical image. These strategies include image registration and image fusion strategies, guiding the registration and fusion processes of the first medical image, respectively. Optionally, the image processing strategies can be at least one of image enhancement, semantic correction, and registration prior strategies. For example, image enhancement strategies might include histogram equalization of low-contrast MRI images, applying non-rigid motion correction to PET images containing motion artifacts, or suppressing metal artifacts in CT images; semantic correction strategies might involve local resampling or marking regions that do not match high uptake areas in PET images with CT anatomical structures as low-confidence regions; and registration prior strategies might apply higher deformation constraints to the "ventricular region" during registration.
[0018] In this embodiment of the invention, a first medical image of the target area and a corresponding clinical text report can be obtained first. Then, based on a pre-built strategy generation model, an image processing strategy, including an image registration strategy and an image fusion strategy, is generated according to the first medical image and the clinical text report. By utilizing a strategy generation model built based on a language model, combined with visual features in the first medical image and semantic features in the clinical text report, an image processing strategy for the first medical image is generated to guide the subsequent image registration and image fusion processes. This can significantly improve the registration accuracy of anatomical structures between different modalities in the first medical image and effectively enhance the credibility and interpretability of the fused image. Optionally, while obtaining the first medical image of the target area and the corresponding clinical text report, metadata corresponding to the first medical image can also be obtained, such as the model of the acquisition device, scanning parameters, reconstruction algorithm, etc., as input to the strategy generation model to participate in the generation process of the image processing strategy.
[0019] Optionally, the step of generating an image processing strategy based on a pre-built strategy generation model, according to the first medical image and the clinical text report, includes: performing image quality analysis on the first medical image to obtain image quality analysis results; wherein the image quality analysis includes at least one of noise detection and artifact detection; performing semantic consistency analysis on the first medical image and the clinical text report to obtain semantic consistency analysis results; and inputting the image quality analysis results and the semantic consistency analysis results into the pre-built strategy generation model to generate a structured image processing strategy.
[0020] Image quality analysis includes at least one of noise detection and artifact detection. Artifact detection includes motion artifact detection, truncation artifact detection, etc. Optionally, image quality analysis also includes detecting low-contrast regions in the image and annotating abnormal structures.
[0021] In this embodiment of the invention, an image processing strategy is generated based on a pre-built strategy generation model, according to a first medical image and a clinical text report. The specific process is as follows: First, a preset image processing method or a lightweight convolutional neural network can be used to perform image quality analysis on the first medical image to obtain the image quality analysis results. Second, semantic consistency analysis can be performed on the first medical image and the clinical text report. Specifically, a large language model can be used to parse the clinical text report, such as "a 2.3cm nodule is seen in the upper lobe of the right lung," and spatial-semantic alignment can be performed with the candidate lesion region detected in the first medical image to determine whether there are missing annotations, positional contradictions, or inconsistencies between modalities, thus obtaining the semantic consistency analysis results. Finally, the image quality analysis results and the semantic consistency analysis results can be input into the pre-built strategy generation model to guide it to output a standardized image processing strategy with structured prompts. For example, the structured prompts can be "Based on the following image quality issues and clinical semantic information, please generate an image processing strategy for CT and MRI images, in the format: {modality: [operation 1, operation 2], reason: '...'}". S120. Register the first medical image according to the image registration strategy and the pre-built image registration model to obtain the second medical image and the deformation field; wherein, the image registration model is built based on a deep neural network model; the deformation field is used to describe the displacement of each voxel in the first medical image during the registration process.
[0022] The image registration model is built upon a deep neural network model to achieve high-precision non-rigid registration between multimodal medical images. In this embodiment of the invention, the image registration model can not only utilize low-level pixel / gradient features, but also integrate high-level semantic priors provided by a policy generation model, thereby improving the accuracy and clinical reliability of anatomical structure alignment through joint optimization.
[0023] The deformation field is a three-dimensional vector field used to describe the degree of displacement of each voxel in the first medical image during the registration process. Specifically, it is used to describe the degree of displacement of each voxel in the first medical image, which is a floating image, during the registration process.
[0024] In this embodiment of the invention, when registering a first medical image according to an image registration strategy and a pre-built image registration model, it is necessary to first determine the fixed image and the floating image of the first medical image during registration. The fixed image is usually selected as the first medical image with clear anatomical structure and high spatial resolution, while the floating image refers to at least one image of another modality to be registered. Then, the fixed image and the floating image can be registered according to the image registration strategy and the pre-built image registration model to obtain a second medical image and a deformation field. Optionally, the image registration strategy can be used as a conditional input to the image registration model, such as encoding the image registration strategy as a vector and inputting it into the conditional layer of the image registration model to dynamically adjust the model's region of interest. For example, if the image registration strategy is "enhancing lesion edge alignment", then the weight of the edge gradient consistency term is added to the loss function of the image registration model.
[0025] Optionally, registering the first medical image according to the image registration strategy and a pre-built image registration model to obtain a second medical image and a deformation field includes: converting the image registration strategy into a pixel-level registration guide map through keyword extraction and spatial mapping; inputting the first medical image and the registration guide map into the pre-built image registration model to obtain the second medical image and deformation field output by the image registration model.
[0026] The registration guidance map, also known as the semantic guidance map, takes the form of a spatial attention mask, such as marking lesion areas with high weight (value = 1.0) and non-critical areas with low weight (value = 0.2), or a category probability map, such as "tumor = 0.9, ventricle = 0.7, normal tissue = 0.3".
[0027] In this embodiment of the invention, the image registration strategy can be transformed into a pixel-level registration guide map through keyword extraction and spatial mapping. Then, the first medical image and the registration guide map are input into a pre-constructed image registration model to obtain the second medical image and deformation field output by the image registration model, thereby achieving semantic-driven accurate registration of the first medical image.
[0028] S130. The second medical image is fused according to the image fusion strategy, deformation field and pre-built image fusion model to obtain the third medical image; wherein, the image fusion model is built based on a deep neural network model.
[0029] The image fusion model is built upon a deep neural network model and can perform image fusion based on a weighted attention mechanism or a cross-modal Transformer. For example, the structural design of the cross-modal Transformer is explained here: First, each second medical image can be divided into 3D patches, linearly projected into a token sequence, transforming the second medical image into structured sequence data that the Transformer can process. Second, a second medical image of one modality (e.g., MRI) can be used as the query, and second medical images of other modalities as key / value pairs, enabling deep feature interaction and information complementarity between images of different modalities. Finally, the fused tokens can be reconstructed into images through 3D deconvolution or a ViT decoder, remapping the high-level abstract features generated by cross-modal attention back into the image space to generate the final task output.
[0030] In this embodiment of the invention, after registering the first medical image according to the image registration strategy and the pre-built image registration model to obtain the second medical image and the deformation field, the second medical image can be fused according to the image fusion strategy, the deformation field, and the pre-built image fusion model to adaptively fuse the multimodal information in the second medical image, highlight the complementary features of key clinical regions (such as lesions), suppress low-quality or unreliable regions, and generate a structurally clear, semantically rich, and diagnostically reliable fused image as the third medical image. It should be noted that the third medical image needs to balance anatomical structure fidelity, functional / metabolic information preservation, and clinical semantic consistency.
[0031] Optionally, fusing the second medical image according to the image fusion strategy, the deformation field, and the pre-built image fusion model to obtain a third medical image includes: adjusting the attention weights of different regions in the second medical image during image fusion according to the image fusion strategy; inputting the second medical image and the deformation field into the pre-built image fusion model, and fusing the second medical image based on the adjusted attention weights through the image fusion model to obtain a third medical image.
[0032] In this embodiment of the invention, a second medical image is fused according to an image fusion strategy, a deformation field, and a pre-built image fusion model. Specifically, firstly, the attention weights of different regions in the second medical image are adjusted during image fusion based on the image fusion strategy. Specifically, the attention score of the corresponding pixel / patch can be increased based on the "high confidence region" label in the image fusion strategy. For example, if "PET images have high confidence in the tumor core area," then the feature weight of the PET image in that region is increased during image fusion. Then, the second medical image and the deformation field are input into the pre-built image fusion model, and the image fusion model fuses the second medical image based on the adjusted attention weights to obtain a third medical image.
[0033] The technical solution of this invention involves acquiring a first medical image of a target area and a corresponding clinical text report, and generating an image processing strategy based on a pre-built strategy generation model, according to the first medical image and the clinical text report. The first medical image includes at least two medical images of different modalities; the strategy generation model is built based on a language model; the image processing strategy includes an image registration strategy and an image fusion strategy; the first medical image is registered according to the image registration strategy and the pre-built image registration model to obtain a second medical image and a deformation field; the image registration model is built based on a deep neural network model; the deformation field describes the displacement degree of each voxel in the first medical image during the registration process; the second medical image is fused according to the image fusion strategy, the deformation field, and the pre-built image fusion model to obtain a third medical image; the image fusion model is built based on a deep neural network model. The technical solution of this invention utilizes a strategy generation model based on a language model, combined with visual features in the first medical image and semantic features in the clinical text report, to generate an image processing strategy for the first medical image. This strategy guides subsequent image registration and image fusion processes, significantly improving the registration accuracy of anatomical structures between different modalities in the first medical image and effectively enhancing the credibility and interpretability of the fused image.
[0034] Example 2 Figure 2 This is a flowchart of a medical image processing method provided in Embodiment 2 of the present invention. The embodiments of the present invention are optimized based on the above embodiments. Solutions not described in detail in the embodiments of the present invention are described in the above embodiments. Figure 2 As shown, the method includes: S210. Obtain the first medical image of the target area and the corresponding clinical text report, and generate an image processing strategy based on the first medical image and the clinical text report according to the pre-built strategy generation model.
[0035] S220. Perform preprocessing operations on at least some modalities of the first medical image according to the image preprocessing strategy.
[0036] Optionally, the image processing strategy also includes an image preprocessing strategy for preprocessing at least a portion of the first medical image.
[0037] In this embodiment of the invention, the strategy generation model generates an image processing strategy based on the first medical image and clinical text report, which also includes an image preprocessing strategy. The system can call the corresponding image processing module according to the image preprocessing strategy to perform preprocessing operations on at least some modalities of the first medical image to optimize the first medical image, which is equivalent to constructing a semantically aware preprocessing channel.
[0038] S230. Register the first medical image according to the image registration strategy and the pre-built image registration model to obtain the second medical image and the deformation field.
[0039] S240. The second medical image is fused according to the image fusion strategy, deformation field and pre-built image fusion model to obtain the third medical image.
[0040] S250. Determine the image quality index value of the third medical image and the semantic consistency detection result between the third medical image and the clinical text report; wherein, the image quality index includes at least one of structural similarity, information entropy, and contrast.
[0041] In this embodiment of the invention, after fusing the second medical image according to the image fusion strategy, deformation field, and pre-built image fusion model to obtain the third medical image, the credibility of the third medical image can be evaluated using image quality indicators and a semantic consistency diagnostic model. Specifically, image quality indicators such as structural similarity, information entropy, and contrast of the third medical image can be calculated, and semantic consistency detection between the third medical image and the clinical text report can be performed using the semantic consistency diagnostic model to obtain the semantic consistency detection results.
[0042] S260. Determine the image credibility score of the third medical image based on the image quality index value and semantic consistency detection results. If the image credibility score is less than the preset threshold, adjust the parameters of the image fusion model and re-fuse the second medical image.
[0043] In this embodiment of the invention, the image credibility score of the third medical image can be determined based on the image quality index value and semantic consistency detection result obtained in step S250. Optionally, corresponding scoring intervals can be set for the image quality index value and semantic consistency detection result respectively, and after obtaining the scores corresponding to both, the scores corresponding to both are weighted and summed to obtain the final image credibility score. It is understood that if the image credibility score is less than a preset threshold, the parameters of the image fusion model need to be adjusted, and the second medical image needs to be fused again to ensure the credibility of the output third medical image. The preset threshold can be set by those skilled in the art according to the actual situation, and this embodiment of the invention does not limit it.
[0044] S270. If the image credibility score is greater than a preset threshold, output the third medical image and generate an interpretation report of the third medical image; wherein, the interpretation report includes at least one of the image processing strategy, image credibility score, and target region in the third medical image; the target region is determined based on the semantic consistency detection result.
[0045] In this embodiment of the invention, if the image credibility score is greater than a preset threshold, a third medical image is output and an interpretation report of the third medical image is automatically generated. The interpretation report includes at least one of the following: image processing strategy, image credibility score, and target region in the third medical image. The target region is determined based on the semantic consistency detection result and may be a key area of interest in the third medical image, such as a lesion region.
[0046] The technical solution of this invention involves acquiring a first medical image of a target area and a corresponding clinical text report, and generating an image processing strategy based on a pre-built strategy generation model, according to the first medical image and the clinical text report; performing preprocessing operations on at least some modalities of the first medical image according to the image preprocessing strategy; registering the first medical image according to an image registration strategy and a pre-built image registration model to obtain a second medical image and a deformation field; fusing the second medical image according to an image fusion strategy, the deformation field, and a pre-built image fusion model to obtain a third medical image; and determining the image quality index value of the third medical image, as well as the third medical... The invention utilizes a language model-based strategy generation model, combined with visual features in the first medical image and semantic features in the clinical text report, to generate an image processing strategy for the first medical image. This strategy guides subsequent image registration and fusion processes, significantly improving the registration accuracy of anatomical structures between different modalities in the first medical image and effectively enhancing the credibility and interpretability of the fused image. The image credibility score of the third medical image is determined based on the image quality index value and the semantic consistency detection results. If the image credibility score is less than a preset threshold, the parameters of the image fusion model are adjusted, and the second medical image is fused again. If the image credibility score is greater than the preset threshold, the third medical image is output, and an interpretation report for the third medical image is generated. The interpretation report includes at least one of the following: image processing strategy, image credibility score, and target region in the third medical image. The target region is determined based on the semantic consistency detection results. Meanwhile, by determining the image credibility score of the third medical image based on the image quality index value and semantic consistency detection results, the credibility of the output third medical image can be ensured, providing a more reliable image basis for clinical diagnosis.
[0047] Example 3 Figure 3 This is a schematic diagram of the structure of a medical image processing device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: The strategy generation module 310 is used to acquire a first medical image of the target area and a clinical text report corresponding to the first medical image, and generate an image processing strategy based on the first medical image and the clinical text report according to a pre-built strategy generation model; wherein, the first medical image includes at least two medical images of different modalities; the strategy generation model is built based on a language model; the image processing strategy includes an image registration strategy and an image fusion strategy; Image registration module 320 is used to register the first medical image according to the image registration strategy and a pre-built image registration model to obtain a second medical image and a deformation field; wherein, the image registration model is constructed based on a deep neural network model; the deformation field is used to describe the degree of displacement of each voxel in the first medical image during the registration process; The image fusion module 330 is used to fuse the second medical image according to the image fusion strategy, the deformation field and the pre-built image fusion model to obtain a third medical image; wherein the image fusion model is constructed based on a deep neural network model.
[0048] Optionally, the image processing strategy may include at least one of image enhancement, semantic correction, and registration prior.
[0049] Optionally, the policy generation module 310 is specifically used for: The first medical image is subjected to image quality analysis to obtain image quality analysis results; wherein, the image quality analysis includes at least one of noise detection and artifact detection; Semantic consistency analysis is performed on the first medical image and the clinical text report to obtain the semantic consistency analysis results; The image quality analysis results and the semantic consistency analysis results are input into a pre-built strategy generation model to generate a structured image processing strategy.
[0050] Optionally, the image processing strategy further includes an image preprocessing strategy, which is used to preprocess at least a portion of the image in the first medical image; The device further includes: A preprocessing module is configured to perform preprocessing operations on at least a portion of the modalities of the first medical image according to the image preprocessing strategy.
[0051] Optional, the image registration module 320 is specifically used for: By extracting keywords and spatial mapping, the image registration strategy is transformed into a pixel-level registration guidance map; The first medical image and the registration guide map are input into a pre-constructed image registration model to obtain the second medical image and deformation field output by the image registration model.
[0052] Optional, the image fusion module 330 is specifically used for: According to the image fusion strategy, the attention weights of different regions in the second medical image are adjusted during image fusion. The second medical image and the deformation field are input into a pre-constructed image fusion model. The image fusion model fuses the second medical image based on the adjusted attention weights to obtain a third medical image.
[0053] Optionally, the device further includes: An image detection module is used to determine the image quality index value of the third medical image and the semantic consistency detection result between the third medical image and the clinical text report; wherein, the image quality index includes at least one of structural similarity, information entropy, and contrast; The parameter adjustment module is used to determine the image credibility score of the third medical image based on the image quality index value and the semantic consistency detection result. If the image credibility score is less than a preset threshold, the parameters of the image fusion model are adjusted, and the second medical image is fused again. Optionally, the device further includes: An image output module is used to output a third medical image and generate an interpretation report of the third medical image if the image credibility score is greater than a preset threshold; wherein the interpretation report includes at least one of the following: image processing strategy, image credibility score, and target region in the third medical image; the target region is determined based on the semantic consistency detection result.
[0054] The medical image processing apparatus provided in the embodiments of the present invention can execute the medical image processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0055] Example 4 Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0056] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0057] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0058] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as medical image processing methods.
[0059] In some embodiments, the medical image processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the medical image processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the medical image processing method by any other suitable means (e.g., by means of firmware).
[0060] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0061] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0062] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0063] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0064] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0065] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0066] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.
[0067] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0068] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A medical image processing method, characterized in that, The method includes: A first medical image of the target area and a corresponding clinical text report are acquired. Based on a pre-built strategy generation model, an image processing strategy is generated according to the first medical image and the clinical text report. The first medical image includes at least two medical images of different modalities. The strategy generation model is built based on a language model. The image processing strategy includes an image registration strategy and an image fusion strategy. The first medical image is registered according to the image registration strategy and the pre-built image registration model to obtain a second medical image and a deformation field; wherein, the image registration model is constructed based on a deep neural network model; the deformation field is used to describe the degree of displacement of each voxel in the first medical image during the registration process; The second medical image is fused according to the image fusion strategy, the deformation field, and the pre-built image fusion model to obtain a third medical image; wherein the image fusion model is constructed based on a deep neural network model.
2. The method according to claim 1, characterized in that, The image processing strategy includes at least one of the following: image enhancement, semantic correction, and registration prior.
3. The method according to claim 1, characterized in that, The image processing strategy generated based on the pre-built strategy generation model, according to the first medical image and the clinical text report, includes: The first medical image is subjected to image quality analysis to obtain image quality analysis results; wherein, the image quality analysis includes at least one of noise detection and artifact detection; Semantic consistency analysis is performed on the first medical image and the clinical text report to obtain the semantic consistency analysis results; The image quality analysis results and the semantic consistency analysis results are input into a pre-built strategy generation model to generate a structured image processing strategy.
4. The method according to claim 1, characterized in that, The image processing strategy further includes an image preprocessing strategy, which is used to preprocess at least a portion of the image in the first medical image; Before registering the first medical image according to the image registration strategy and the pre-built image registration model, the method further includes: Preprocessing operations are performed on at least some modalities of the first medical image according to the image preprocessing strategy.
5. The method according to claim 1, characterized in that, The step of registering the first medical image according to the image registration strategy and the pre-built image registration model to obtain the second medical image and the deformation field includes: By extracting keywords and spatial mapping, the image registration strategy is transformed into a pixel-level registration guidance map; The first medical image and the registration guide map are input into a pre-constructed image registration model to obtain the second medical image and deformation field output by the image registration model.
6. The method according to claim 1, characterized in that, The process of fusing the second medical image according to the image fusion strategy, the deformation field, and the pre-built image fusion model to obtain a third medical image includes: According to the image fusion strategy, the attention weights of different regions in the second medical image are adjusted during image fusion. The second medical image and the deformation field are input into a pre-constructed image fusion model. The image fusion model fuses the second medical image based on the adjusted attention weights to obtain a third medical image.
7. The method according to claim 1, characterized in that, After fusing the second medical image according to the image fusion strategy, the deformation field, and the pre-built image fusion model to obtain the third medical image, the method further includes: The image quality index value of the third medical image and the semantic consistency detection result between the third medical image and the clinical text report are determined; wherein, the image quality index includes at least one of structural similarity, information entropy, and contrast. The image credibility score of the third medical image is determined based on the image quality index value and the semantic consistency detection result. If the image credibility score is less than a preset threshold, the parameters of the image fusion model are adjusted, and the second medical image is fused again.
8. The method according to claim 7, characterized in that, The method further includes: If the image credibility score is greater than a preset threshold, a third medical image is output, and an interpretation report of the third medical image is generated; wherein, the interpretation report includes at least one of the image processing strategy, image credibility score, and target region in the third medical image; the target region is determined based on the semantic consistency detection result.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the medical image processing method according to any one of claims 1-8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the medical image processing method as described in any one of claims 1-8.