Image processing method and device
By using the image analysis model to perform fusion enhancement processing on multimodal images, the problem of low image analysis efficiency when the modality is missing is solved, and efficient image analysis is achieved in the case of modality missing.
Patent Information
- Application Number
- CN202111154668.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2041-09-29
AI Technical Summary
Existing image processing methods are inefficient when modalities are missing and cannot effectively perform image analysis.
The multimodal images are fused and enhanced through the image analysis model, and the feature extraction sub-model and the modal perception sub-model are used to extract features and fuse the images to be analyzed, generating a target image that enhances the distribution area of the analysis object.
This enables effective image analysis even when the modality is missing, thereby improving the processing efficiency of the image processing method.
Smart Images

Figure CN113850794B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method and device. Background Art
[0002] With the rapid development of computer graphics and image processing technology, computer-based technology has been widely used in the medical field. Through computer analysis and processing of various diagnostic images, the accuracy of medical staff's diagnosis has been effectively improved. At present, commonly used computer-aided diagnosis systems usually perform corresponding enhancement processing on magnetic resonance imaging (MRI) or electronic computed tomography (CT) images to enable doctors to quickly determine the location of the lesion. For example, when current computer-aided diagnosis systems determine liver tumors, they usually adopt a dual-modality solution, that is, collecting enhanced CT images when the contrast agent is in the liver veins and CT images when the contrast agent is in the liver arteries over a period of time for analysis. Since the enhanced CT images when the contrast agent is in the liver veins and the CT images when the contrast agent is in the liver arteries can well complement each other's information, it is helpful to better diagnose liver tumors.
[0003] However, if there is only one modality of image information or more than two modalities of image information, the above-mentioned dual-modality solution cannot be used for image processing. A new corresponding image processing method needs to be redeveloped, resulting in low efficiency of the current image processing method.
[0004] Application Contents
[0005] To solve the above technical problems, the embodiments of the present application hope to provide an image processing method and device. The technical solution of the present application is implemented as follows:
[0006] In a first aspect, an image processing method is provided, the method comprising:
[0007] Obtaining a first number of images to be analyzed; wherein each of the images to be analyzed corresponds to a different target modality of the target photographed object;
[0008] Through the image analysis model, the first number of images to be analyzed are fused and enhanced to obtain a first target image; wherein, the first target image is used to enhance the display of the distribution area of the analysis object in the first number of images to be analyzed, and the analysis object belongs to the photographed object. The image analysis model is trained by sample images corresponding to a second number of different sample modalities, the first number is less than or equal to the second number, and the target modality belongs to the sample modality.
[0009] Optionally, performing fusion enhancement processing on the first number of images to be analyzed using an image analysis model to obtain a first target image includes:
[0010] If the first number is 1, based on the feature extraction sub-models corresponding to the first number of images to be analyzed of different target modalities in the image analysis model, the corresponding images to be analyzed are processed to obtain the first target image; wherein, the feature extraction sub-model corresponding to each of the images to be analyzed can characterize the correlation relationship between the second number of images to be analyzed of different sample modalities.
[0011] Optionally, performing fusion enhancement processing on the first number of images to be analyzed using an image analysis model to obtain a first target image includes:
[0012] If the first number is greater than or equal to 2 and less than or equal to the second number, processing the corresponding images to be analyzed by using the feature extraction sub-models in the image analysis model corresponding to the first number of images to be analyzed of different target modalities to obtain a first number of reference images; wherein the feature extraction sub-model corresponding to each of the images to be analyzed is capable of characterizing the correlation relationship between the second number of images to be analyzed of different sample modalities;
[0013] The first target image is obtained by performing image processing on the first number of reference images through the modality perception sub-model in the image analysis model.
[0014] Optionally, the modality perception sub-model is used to perform image fusion processing on reference images of at least two different modalities, or the modality perception sub-model is used to perform image fusion processing on reference images of at least two different sample modalities after feature enhancement processing.
[0015] Optionally, when the modality perception sub-model is used to perform feature enhancement processing on reference images of at least two different sample modalities and then perform image fusion processing, performing image processing on the first number of reference images using the modality perception sub-model in the image analysis model to obtain the first target image includes:
[0016] Performing image fusion processing on the first number of reference images through the modality perception sub-model to obtain a target fused image;
[0017] Determining a similarity coefficient between each of the reference images and the target fused image using the modality perception sub-model to obtain the first number of similarity coefficients;
[0018] The first target image is obtained by performing feature enhancement processing on the first number of the similarity coefficients and the first number of the reference images through the modal perception sub-model.
[0019] Optionally, performing feature enhancement processing on the first number of the similarity coefficients and the first number of the reference images by the modality perception sub-model to obtain the first target image includes:
[0020] Performing feature enhancement processing on each of the reference images using the corresponding similarity coefficient using the modality perception sub-model to obtain the first number of sub-feature images;
[0021] The first target image is obtained by performing image fusion processing on the first number of sub-feature images through the modal perception sub-model.
[0022] Optionally, the method further includes:
[0023] Obtaining a third number of groups of sample images and a third number of marked positions for the analysis object in the third number of groups of sample images; wherein each group of sample images includes the second number of sample images corresponding to different sample modalities;
[0024] Determine the image model to be trained;
[0025] The image model to be trained is trained using the third number of groups of sample images and the third number of marked positions to obtain the image analysis model.
[0026] Optionally, the using the third number of groups of sample images and the third number of marked positions to perform model training on the image model to be trained to obtain the image analysis model includes:
[0027] Performing fusion enhancement processing on the region where the analysis object is located in each of the third number of groups of the sample images using the image model to be trained, to obtain the third number of second target images;
[0028] Based on the third number of groups of sample images, the third number of the second target images and the third number of the marked positions, model training is performed on the image model to be trained to obtain the image analysis model.
[0029] Optionally, the performing model training on the image model to be trained based on the third number of groups of sample images, the third number of second target images, and the third number of marked positions to obtain the image analysis model includes:
[0030] Determine a loss value between each sample image in each target group and the corresponding mark position to obtain first loss values corresponding to the third number of target group sample images; wherein the first loss value corresponding to each target group sample image includes the second number of loss values;
[0031] Determine a loss value between each second target image corresponding to the target group sample image and the corresponding marking position, to obtain the third number of second loss values corresponding to the target group sample image;
[0032] The first loss values corresponding to the third number of target group sample images and the third number of second loss values corresponding to the target group sample images are reversely transmitted in the image model to be trained to continuously train the parameters of the image model to be trained to obtain the image analysis model.
[0033] In a second aspect, an image processing device is provided, comprising: an acquisition unit and a model processing unit; wherein:
[0034] The obtaining unit is configured to obtain a first number of images to be analyzed, wherein each of the images to be analyzed corresponds to a different target modality of the target photographed object;
[0035] The model processing unit is used to perform fusion enhancement processing on the first number of images to be analyzed through the image analysis model to obtain a first target image; wherein, the first target image is used to enhance the display of the distribution area of the analysis object in the first number of images to be analyzed, and the analysis object belongs to the photographed object. The image analysis model is trained by sample images corresponding to a second number of different sample modalities, the first number is less than or equal to the second number, and the target modality belongs to the sample modality.
[0036] The embodiments of the present application provide an image processing method and apparatus. After obtaining a first number of images to be analyzed, the first number of images to be analyzed are subjected to fusion and enhancement processing by an image analysis model to obtain a first target image. Thus, because the first number can be less than the second number of all inputs to the image analysis model, the image analysis model is used to perform enhanced fusion processing on the first number of images to be analyzed in the missing modality to obtain a first target image for the analysis object. This solves the problem of current image processing methods being unable to perform image analysis when the modality is missing. This method allows analysis to be performed using a unified image processing method even when the modality is missing, thereby improving the processing efficiency of the image processing method. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1A flowchart of an image processing method provided in an embodiment of the present application;
[0038] Figure 2 A flowchart of another image processing method provided in an embodiment of the present application;
[0039] Figure 3 A flowchart of another image processing method provided in an embodiment of the present application;
[0040] Figure 4 A flowchart of another image processing method provided in an embodiment of the present application;
[0041] Figure 5 A schematic diagram of a model structure provided in an embodiment of the present application;
[0042] Figure 6 An arterial phase liver image provided in an embodiment of the present application;
[0043] Figure 7 A venous phase liver image provided in an embodiment of the present application;
[0044] Figure 8 A schematic diagram of a first target image provided in an embodiment of the present application;
[0045] Figure 9 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;
[0046] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0048] The embodiment of the present application provides an image processing method, referring to Figure 1 As shown, the method is applied to an electronic device, and the method includes the following steps:
[0049] Step 101: Obtain a first number of images to be analyzed.
[0050] Each image to be analyzed corresponds to a different target modality of the target photographed object.
[0051] In an embodiment of the present application, the electronic device is a device with computing and analysis capabilities, such as a computer, a server, or a smart mobile device. The first number of images to be analyzed are images captured using different shooting methods for the same target subject and requiring analysis. In other words, one image to be analyzed corresponds to one modality, and thus the first number of images to be analyzed corresponds to the first number of modalities. The images to be analyzed can be images requiring analysis in various fields, such as images in the medical field requiring enhanced display of the area where the analysis subject is located.
[0052] For example, a patient's liver is imaged using CT to obtain a venous phase image when the contrast agent is in the vein and an arterial phase image when the contrast agent is in the artery. In this case, the venous phase image and arterial phase image obtained can be called two-modal images.
[0053] Step 102: Perform fusion enhancement processing on a first number of images to be analyzed using an image analysis model to obtain a first target image.
[0054] Among them, the first target image is used to enhance the display of the distribution area of the analysis object in the first number of images to be analyzed, the analysis object belongs to the photographed object, the image analysis model is trained by sample images corresponding to the second number of different sample modalities, the first number is less than or equal to the second number, and the target modality belongs to the sample modality.
[0055] In an embodiment of the present application, the image analysis model is an analysis model obtained by training with a large number of sample images including a second number of different sample modalities. In one analysis process, the maximum input object of the image analysis model is the second number of different sample modalities of images to be analyzed, and the minimum can be 1 sample modality of images to be analyzed. Since the image analysis model is obtained by training with a large number of sample images including a second number of different sample modalities, the parameters corresponding to each sample modality in the image analysis model are determined by other sample modalities, and it has the characteristic of supplementing the features of other sample modalities. In this way, when the first number of input images to be analyzed is less than the second number, the image analysis model can also supplement the missing feature information of the sample modality, and then perform fusion enhancement processing on the first number of images to be analyzed to obtain the first target image.
[0056] An embodiment of the present application provides an image processing method that, after obtaining a first number of images to be analyzed, performs fusion and enhancement processing on the first number of images to be analyzed using an image analysis model to obtain a first target image. Thus, because the first number can be less than the second number of all inputs to the image analysis model, the image analysis model performs enhanced fusion processing on the first number of images to be analyzed in the missing modality to obtain a first target image for the analysis object. This solves the problem of current image processing methods being unable to perform image analysis when the modality is missing, and achieves a method that can still perform analysis using a unified image processing method even when the modality is missing, thereby improving the processing efficiency of the image processing method.
[0057] Based on the foregoing embodiments, an embodiment of the present application provides an image processing method, which is applied to an electronic device and includes the following steps:
[0058] Step 201: Obtain a first number of images to be analyzed.
[0059] Each image to be analyzed corresponds to a different target modality of the target photographed object.
[0060] In the embodiment of the present application, the images to be analyzed are liver CT images acquired by CT as an example for description. It is assumed that a first number of liver CT images to be analyzed are currently acquired, thereby obtaining a first number of images to be analyzed.
[0061] Step 202: Perform fusion enhancement processing on a first number of images to be analyzed using an image analysis model to obtain a first target image.
[0062] Among them, the first target image is used to enhance the display of the distribution area of the analysis object in the first number of images to be analyzed, the analysis object belongs to the photographed object, the image analysis model is trained by sample images corresponding to the second number of different sample modalities, the first number is less than or equal to the second number, and the target modality belongs to the sample modality.
[0063] In this embodiment of the present application, the corresponding image analysis model is a qualified model obtained by training a large number of liver CT images, including the second number. Assuming the second number is 2, the corresponding two sample modalities may be arterial phase liver CT image modalities and venous phase liver CT image modalities, respectively. In this case, the first number may be 1 or 2, typically determined by the specific number of sample modal images acquired.
[0064] Both the first quantity and the second quantity can be determined by actual application scenarios and are not specifically limited here.
[0065] Based on the above embodiments, in other embodiments of the present application, refer to Figure 2As shown, step 202 can be implemented by step 202a:
[0066] Step 202a: If the first number is 1, based on the feature extraction sub-models corresponding to the first number of images to be analyzed of different target modalities in the image analysis model, process the corresponding images to be analyzed to obtain a first target image.
[0067] The feature extraction sub-model corresponding to each image to be analyzed can characterize the correlation between the second number of images to be analyzed of different sample modalities.
[0068] In an embodiment of the present application, when the first number is 1, a feature extraction sub-model corresponding to the target modality of an obtained image to be analyzed is determined from the image analysis model, and image feature extraction processing is performed on the image to be analyzed by determining the obtained feature extraction sub-model to obtain a first target image.
[0069] Since the image analysis model is obtained by training a large number of sample images including the second number of sample modalities, the relevant parameters in the feature extraction sub-model of each sample modality obtained can not only contain the information of its own sample modality, but also learn the information of the remaining sample modalities. In this way, when the image analysis model is used to analyze an image to be analyzed of a target modality, the effect close to that when the image to be analyzed of the second number of sample modalities is input can be achieved.
[0070] Based on the above embodiments, in other embodiments of the present application, refer to Figure 3 As shown, step 202 can be implemented by steps 202b to 202c:
[0071] Step 202b: If the first number is greater than or equal to 2 and less than or equal to the second number, the corresponding images to be analyzed are processed by the feature extraction sub-model corresponding to the first number of images to be analyzed of different target modalities in the image analysis model to obtain a first number of reference images.
[0072] The feature extraction sub-model corresponding to each image to be analyzed can characterize the correlation between the second number of images to be analyzed of different sample modalities.
[0073] In an embodiment of the present application, when the first number is greater than or equal to 2 and less than or equal to the second number, a feature extraction sub-model corresponding to the target modality of each image to be analyzed is determined from the image analysis model, and then the feature extraction sub-model corresponding to the target modality of each image to be analyzed is used to perform image feature extraction processing on the corresponding image to be analyzed to obtain a feature image corresponding to each image to be analyzed, that is, a reference image, and thus a first number of reference images can be obtained.
[0074] Step 202c: Perform image processing on a first number of reference images using the modal perception sub-model in the image analysis model to obtain a first target image.
[0075] In an embodiment of the present application, after obtaining a first number of reference images, the first number of reference images are processed by the modal perception sub-model in the image analysis model to achieve feature enhancement and image fusion processing to obtain a first target image.
[0076] Based on the foregoing embodiments, in other embodiments of the present application, the modality perception sub-model is used to perform image fusion processing on reference images of at least two different modalities, or the modality perception sub-model is used to perform image fusion processing after feature enhancement processing on reference images of at least two different sample modalities.
[0077] In an embodiment of the present application, the modal perception sub-model in the image analysis model can directly perform image fusion processing on a first number of reference images, enhance the common features in the reference images, and weaken the non-common features, thereby obtaining a first target image.
[0078] The modal perception sub-model in the image analysis model can also perform feature enhancement processing on the first number of reference images respectively, and then perform image fusion processing on the first number of reference images after feature enhancement processing to obtain a first target image.
[0079] Based on the foregoing embodiment, in other embodiments of the present application, when the modality perception sub-model is used to perform feature enhancement processing on reference images of at least two different sample modalities and then perform image fusion processing, step 202c can be implemented by steps a11 to a13:
[0080] Step a11: perform image fusion processing on the first number of reference images through the modal perception sub-model to obtain a target fused image.
[0081] In an embodiment of the present application, the modality perception sub-model performs image fusion processing on a first number of reference images extracted by the feature extraction sub-module to obtain a target fused image.
[0082] Step a12: Determine the similarity coefficient between each reference image and the target fused image using the modality perception sub-model to obtain a first number of similarity coefficients.
[0083] In an embodiment of the present application, after obtaining the target fused image, the modality perception sub-model calculates a similarity coefficient between each reference image and the target fused image, obtaining a corresponding offset coefficient for each reference image, thereby obtaining a first number of similarity coefficients. The modality perception sub-model may calculate the similarity coefficient between each reference image and the target fused image using a similarity coefficient calculation method, such as a histogram matching method or a hash algorithm. The modality perception sub-model may also employ some neural network model algorithm to achieve the similarity coefficient between each reference image and the target fused image.
[0084] Among them, when the modal perception sub-model adopts a neural network model algorithm to implement the similarity coefficient between each reference image and the target fusion image, the corresponding modal perception sub-model can be implemented by at least two cascaded convolutional layers. In the at least two cascaded convolutional layers, each convolutional layer before the last convolutional layer includes an instance normalization layer and a leaky rectified linear unit layer. When the modal perception sub-model is implemented by two cascaded convolutional layers, the size of the convolution kernel of the first convolutional layer that processes the first reference analysis image can be, for example, 3*3*3, and the size of the convolution kernel of the second convolutional layer that processes the processing result of the leaky rectified linear unit layer after the first convolutional layer can be, for example, 1*1*1.
[0085] Step a13: perform feature enhancement processing on the first number of similarity coefficients and the first number of reference images through the modal perception sub-model to obtain a first target image.
[0086] In an embodiment of the present application, after the modal perception sub-model determines the similarity coefficient of each reference image, it multiplies the similarity coefficient of each reference image with each corresponding reference image, so that the modal perception sub-model performs feature enhancement processing on the corresponding reference image through each similarity coefficient to obtain the first target image.
[0087] Based on the above embodiment, in other embodiments of the present application, step a13 can be implemented through steps a131 to a132:
[0088] Step a131: Perform feature enhancement processing on each reference image using the corresponding similarity coefficient through the modal perception sub-model to obtain a first number of sub-feature images.
[0089] In an embodiment of the present application, the modal perception sub-model calculates the product between each reference image and the corresponding similarity coefficient, that is, performs weighted processing on each reference image to implement feature enhancement processing for each reference image, obtains the sub-feature image corresponding to each reference image, and then obtains the first number of sub-feature images.
[0090] Step a132: Perform image fusion processing on the first number of sub-feature images through the modal perception sub-model to obtain a first target image.
[0091] In an embodiment of the present application, the modal perception sub-model performs superposition processing on the obtained first number of sub-feature images, that is, implements image fusion processing, thereby obtaining a first target image.
[0092] Based on the above embodiments, in other embodiments of the present application, refer to Figure 4 As shown, before the electronic device executes step 201, it is also used to execute steps 203 to 205:
[0093] Step 203: Obtain a third number of groups of sample images and a third number of marked positions for the analysis object in the third number of groups of sample images.
[0094] Each group of sample images includes a second number of sample images corresponding to different sample modalities.
[0095] In an embodiment of the present application, the third number is usually the minimum number of samples required for model training, and the marked position is the position of the analysis object in the sample image. Since each group of sample images is obtained by shooting the same object under different sample modalities, in the same group of sample images, the marked position corresponding to the analysis object is one.
[0096] Step 204: Determine the image model to be trained.
[0097] Step 205 : Using a third number of groups of sample images and a third number of marked positions, perform model training on the image model to be trained to obtain an image analysis model.
[0098] In an embodiment of the present application, a third number of groups of sample images and a third number of marked positions are used to perform model training on the image model to be trained, and the weight coefficients in the image model to be trained are adjusted and modified so that the loss value between the model obtained after training and the corresponding marked position is less than a preset threshold value, and the model obtained after training is determined to be an image analysis model.
[0099] It should be noted that steps 203 to 205 can also be implemented as an independent embodiment. That is, before executing step 201, the image model to be trained is trained as an independent embodiment to obtain an image analysis model so that the trained image analysis model can be directly called later.
[0100] Based on the above embodiment, in other embodiments of the present application, step 205 can be implemented by steps 205a to 205b:
[0101] Step 205a: Using the image model to be trained, perform fusion enhancement processing on the areas where the analysis objects are located in the third number of groups of sample images to obtain a third number of second target images.
[0102] In the embodiment of the present application, the specific implementation process of step 205a can be implemented by referring to the specific implementation process of steps 202b to 202c, steps a11 to a13 and steps a131 to a132, and will not be described in detail here.
[0103] Step 205b: Based on a third number of groups of sample images, a third number of second target images, and a third number of marked positions, the image model to be trained is trained to obtain an image analysis model.
[0104] In an embodiment of the present application, a third number of groups of sample images, a third number of second target images and a third number of marked positions are used to perform model training on the image model to be trained, determine the weight coefficients in the expected model to be trained, and obtain an image analysis model.
[0105] It should be noted that, from the third number group of sample images, a group of sample images are obtained in sequence to obtain the target group sample images; the target group sample images are fused and enhanced by the image model to be trained to obtain a second target image; based on the target group sample image, the second target image corresponding to the target group sample image, and the marked position corresponding to the target group sample image, the corresponding loss value is calculated; if the loss value is less than the preset loss threshold, the loss value is back-propagated in the image model to be trained, the parameters in the image model to be trained are updated, and the image model to be trained is updated to the image model to be trained after the updated parameters; continue to obtain the next group of sample images adjacent to the target group sample image from the third array of sample images, and update them to the target group sample images, or randomly obtain a group of sample images from the third number group of sample images, and update them to the target group sample images, and repeat this until the calculated loss value is less than the preset loss threshold and the corresponding image model to be trained is the image analysis model.
[0106] Based on the above embodiment, in other embodiments of the present application, step 205b can be implemented by steps b11 to b13:
[0107] Step b11: determine the loss value between each sample image in the target group of sample images and the corresponding marking position, and obtain a first loss value corresponding to the target group of sample images.
[0108] The first loss value corresponding to the target group sample image includes a second number of loss values.
[0109] In an embodiment of the present application, a loss function is used to calculate the loss value between each sample image in the target group sample image and the corresponding marking position to obtain the loss value of each sample image in the target group sample image, thereby obtaining the first loss value corresponding to the target group sample image.
[0110] Step b12: Determine the loss value between the second target image corresponding to the target group sample image and the corresponding marking position to obtain the second loss value corresponding to the target group sample image.
[0111] In an embodiment of the present application, a loss function is used to calculate the loss value between the second target image corresponding to the target group sample image and the corresponding marking position to obtain the second loss value corresponding to the target group sample image.
[0112] Step b13: The first loss value corresponding to the target group sample image and the second loss value corresponding to the target group sample image are reversely transmitted in the image model to be trained to continuously train the parameters of the image model to be trained to obtain the image analysis model.
[0113] In an embodiment of the present application, the cumulative sum of the first loss values corresponding to the target group sample images is determined, the product of the cumulative sum and the preset weight coefficient is calculated, and the sum of the product and the second loss value corresponding to the target group sample images is calculated to obtain the loss value corresponding to the target group sample images. If the loss value is less than or equal to the preset loss threshold, the corresponding image model to be trained can be determined as the image analysis model. If the loss value is greater than the preset loss threshold, the loss value is reversely transmitted to the image model to be trained, the parameters of the image model to be trained are updated to obtain the first image analysis model, and the image model to be trained is updated to the first image analysis model, and the operations corresponding to step 205a and steps b11 to b13 are repeated until the image analysis model is obtained.
[0114] Based on the above embodiments, the present application embodiment provides an implementation process of an image processing method for liver CT scan images. Correspondingly, a model structure diagram for training a training image model can be referred to as follows: Figure 5 As shown, it includes: an arterial phase liver CT input node 31, a venous phase liver CT input node 32, an arterial phase feature extraction sub-model 33, a venous phase feature extraction sub-model 34, a modal perception sub-model 35, a target image output node 36 and a loss calculation node 37. The modal perception sub-model 35 includes a first image fusion module 351, an arterial phase similarity coefficient analysis module 352, a venous phase similarity coefficient analysis module 353, and a second image fusion module 354. The loss calculation node 37 includes: an arterial phase loss value calculation module 371, a venous phase loss value calculation module 372, a joint loss value calculation module 373 and a comprehensive loss value calculation module 374.
[0115] Among them, in the model training stage, the information transmission process when a group of sample images, namely the first arterial phase liver image and the first venous phase liver image of the same case, are used for model training is explained as an example: the first arterial phase liver image is input to the arterial phase liver CT input node 31, and the first venous phase liver image is input to the venous phase liver CT input node 32; the arterial phase liver CT input node 31 sends the first arterial phase liver image to the arterial phase feature extraction sub-model 33, the arterial phase feature extraction sub-model 33 performs feature extraction on the first arterial phase liver image to obtain the first arterial phase reference image, similarly, the venous phase feature extraction sub-model 34 performs feature extraction on the first venous phase liver image to obtain the first venous phase reference image; the first image fusion module 351 performs image fusion processing on the first arterial phase reference image and the first venous phase reference image to obtain the first The first target fusion image is obtained; the arterial phase similarity coefficient analysis module 352 calculates the similarity coefficient of the first target fusion image and the first arterial phase reference image to obtain the first arterial similarity coefficient. Similarly, the venous phase similarity coefficient analysis module 353 calculates the similarity coefficient of the first target fusion image and the first venous phase reference image to obtain the first venous similarity coefficient. The second image fusion module 354 uses the first arterial similarity coefficient to perform feature enhancement processing on the first arterial phase reference image to obtain the first arterial sub-feature image. Similarly, the second image fusion module 354 uses the first venous similarity coefficient to perform feature enhancement processing on the first venous phase reference image to obtain the first venous sub-feature image. Then, the first arterial sub-feature image and the first venous sub-feature image are subjected to image fusion processing to obtain the first target image. The arterial phase loss value calculation module 371 calculates the loss value of the first target fusion image based on the first arterial phase liver image X. ap and the liver tumor marker position Y to calculate the first artery loss value L1 = L intra (Y|X ap ;W ap ), W ap is the corresponding parameter coefficient in the arterial phase feature extraction sub-model 33 and the arterial phase similarity coefficient analysis module 352 in the model to be trained. Similarly, the venous phase loss value calculation module 372 is based on the first venous phase liver image X vp and the liver tumor marker position Y to calculate the first venous loss value L2 = L intra (Y|X vp ;W vp ), W vp = L is the corresponding parameter coefficient in the venous phase feature extraction sub-model 34 and the venous phase similarity coefficient analysis module 353 in the model to be trained. The joint loss value calculation module 373 calculates the first joint loss value L3=L based on the first target image X and the liver tumor mark position Y. joint (Y|X; W), W={W ap , W vp}, the comprehensive loss module 374 calculates the final loss value using the following formula.
[0116]
[0117] If the final loss value is greater than the preset loss threshold, the final loss value is transmitted back to Figure 5 In the model to be trained shown in FIG, the parameters of the arterial phase feature extraction sub-model 33, the venous phase feature extraction sub-model 34, and the modality perception sub-model 35 are updated, and the updated Figure 5 The model to be trained shown repeats the above process until it is determined that the final loss value obtained through training is less than or equal to the preset loss threshold, and the corresponding model to be trained is an image analysis model.
[0118] Wherein: the arterial phase feature extraction sub-model 33 and the venous phase feature extraction sub-model 34 can both be implemented using a fully convolutional network (FCN).
[0119] The process of determining the similarity coefficient by the arterial phase similarity coefficient analysis module 352 and the venous phase similarity coefficient analysis module 353 can be expressed by the following formula:
[0120] A i =δ(f a ([F dual ; F i ];θ i )), i=ap,vp.
[0121] Among them, δ is a sigmoid function, θ represents f a The learned parameters are composed of two cascaded convolutional layers. The first convolutional layer can be a 3×3×3 convolution kernel, and the second convolutional layer can be a 1×1×1 convolution kernel. Each convolutional layer is followed by an instance normalization and a leaky rectified linear unit. dual Used to represent the first target fusion image, F ap Used to represent the first arterial phase reference image or F vp First venous phase reference image, A ap is the similarity coefficient of the first arterial phase, A vp is the similarity coefficient of the first venous phase. The convolution operation is used to model the correlation between the discriminative bimodal information and each modality feature.
[0122] The implementation process of the second image fusion module 354 can be expressed by the following formula:
[0123]
[0124] Among them, A apis the first artery similarity coefficient, A vp is the first vein similarity coefficient, F ap is the first arterial phase reference image, F vp This is the first venous phase reference image.
[0125] In this way, the first image fusion module 351 can obtain the first target fused image through convolution. Although the first target fused image encodes the arterial and venous information of the liver tumor, it inevitably introduces redundant noise from each sample modality during liver tumor segmentation. To address this redundant noise, a venous phase similarity coefficient analysis module 353 and an arterial phase similarity coefficient analysis module 352 are proposed through an attention mechanism to calculate the influence of each sample modality, adaptively weighing the contribution of each sample modality and visually interpreting it. Furthermore, during the calculation of the final loss value, the arterial phase loss value calculation module 371 encourages each branch to learn to distinguish features specific to the arterial phase, the venous phase loss value calculation module 372 encourages each branch to learn to distinguish features specific to the venous phase, and the joint loss value calculation module 373 encourages each branch to learn from each other to maintain the commonality between high-level features, thereby better integrating multimodal information. In this way, through the above loss value determination method for model training, a combination of cross-entropy loss and slice loss is implemented as the segmentation loss, effectively reducing the impact of uneven tumor data distribution.
[0126] Assumptions based on Figure 5 After the image analysis model shown is trained to obtain an image analysis model, when the first number of images to be analyzed is 1, illustratively, when the image to be analyzed is the second arterial phase liver image to be analyzed, the second arterial phase liver image is input to the arterial phase feature extraction sub-model 33 via the arterial phase liver CT input node 31. The arterial phase feature extraction sub-model 33 performs feature extraction on the second arterial phase liver image to obtain a second arterial phase reference image. Due to the lack of the venous phase liver image, the second arterial phase reference image is directly output from the target image output node 36 to obtain the first target image corresponding to the second arterial phase liver image. When the image to be analyzed is the second venous phase liver image to be analyzed, the implementation process can be similar to that of the image to be analyzed for the second arterial phase liver image to be analyzed, and will not be further described here.
[0127] When the first number is equal to the second number, which is 2, the image to be analyzed is as follows Figure 6 The third arterial phase liver images shown and Figure 7When the third venous phase liver image is shown, the third arterial phase liver image is input to the arterial phase liver CT input node 31, and the third venous phase liver image is input to the venous phase liver CT input node 32; the arterial phase liver CT input node 31 sends the third arterial phase liver image to the arterial phase feature extraction sub-model 33, the arterial phase feature extraction sub-model 33 performs feature extraction on the third arterial phase liver image to obtain a third arterial phase reference image, similarly, the venous phase feature extraction sub-model 34 performs feature extraction on the third venous phase liver image to obtain a third venous phase reference image; the first image fusion module 351 performs image fusion processing on the third arterial phase reference image and the third venous phase reference image to obtain a second target fused image; the arterial phase similarity coefficient analysis module 352 calculates the similarity coefficient on the second target fused image and the third arterial phase reference image to obtain a second arterial phase reference image. Pulse similarity coefficient, similarly, the venous phase similarity coefficient analysis module 353 calculates the similarity coefficient of the second target fusion image and the third venous phase reference image to obtain the second venous similarity coefficient; the second image fusion module 354 uses the second arterial similarity coefficient to perform feature enhancement processing on the third arterial phase reference image to obtain a second arterial sub-feature image. Similarly, the second image fusion module 354 uses the second venous similarity coefficient to perform feature enhancement processing on the third venous phase reference image to obtain a second venous sub-feature image, and then the second image fusion module 354 performs image fusion processing on the second arterial sub-feature image and the second venous sub-feature image, thereby obtaining a first target image corresponding to the third arterial phase liver image and the third venous phase liver image; the first target image corresponding to the third arterial phase liver image and the third venous phase liver image is output through the target image output node 36, which can be referred to Figure 8 As shown, Figure 8 The area filled with diagonal lines is the highlighted liver tumor area.
[0128] In this way, the above-mentioned image analysis model obtained can handle multimodal segmentation problems and missing modality problems without any modification, thereby improving processing efficiency; each single-modality model can implicitly utilize bimodal information by learning from other models, that is, since the parameters in the arterial phase feature extraction submodel and the venous phase feature extraction submodel are determined by bimodal information, in the absence of other modalities, better segmentation results can be obtained based on the parameters in the arterial phase feature extraction submodel or the venous phase feature extraction submodel, that is, by combining the characteristics and commonalities of each modality, a better multimodal segmentation effect can be obtained through the cooperation of all specific modality models.
[0129] It should be noted that, for the description of the same steps and contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.
[0130] An embodiment of the present application provides an image processing method that, after obtaining a first number of images to be analyzed, performs fusion and enhancement processing on the first number of images to be analyzed using an image analysis model to obtain a first target image. Thus, because the first number can be less than the second number of all inputs to the image analysis model, the image analysis model performs enhanced fusion processing on the first number of images to be analyzed in the missing modality to obtain a first target image for the analysis object. This solves the problem of current image processing methods being unable to perform image analysis when the modality is missing, and achieves a method that can still perform analysis using a unified image processing method even when the modality is missing, thereby improving the processing efficiency of the image processing method.
[0131] Based on the above embodiments, the present application provides an image processing device, referring to Figure 9 As shown, the image processing device 4 includes: an acquisition unit 41 and a model processing unit 42; wherein:
[0132] An obtaining unit 41 is configured to obtain a first number of images to be analyzed, wherein each image to be analyzed corresponds to a different target modality of the target photographed object;
[0133] The model processing unit 42 is used to perform fusion enhancement processing on a first number of images to be analyzed through an image analysis model to obtain a first target image; wherein the first target image is used to enhance the display of the distribution area of the analysis object in the first number of images to be analyzed, the analysis object belongs to the photographed object, and the image analysis model is trained by sample images corresponding to a second number of different sample modalities, the first number is less than or equal to the second number, and the target modality belongs to the sample modality.
[0134] Based on the above embodiment, in other embodiments of the present application, the model processing unit 42 includes: a feature extraction sub-model module; wherein:
[0135] A feature extraction sub-model module is used to process the corresponding image to be analyzed based on the feature extraction sub-models corresponding to the first number of images to be analyzed of different target modalities in the image analysis model if the first number is 1, so as to obtain a first target image; wherein the feature extraction sub-model corresponding to each image to be analyzed can characterize the correlation relationship between the second number of images to be analyzed of different sample modalities.
[0136] Based on the above embodiment, in other embodiments of the present application, the model processing unit 42 includes: a feature extraction sub-model module and a modality perception sub-model module; wherein:
[0137] a feature extraction sub-model module, configured to, if the first number is greater than or equal to 2 and less than or equal to the second number, process the corresponding images to be analyzed using the feature extraction sub-models in the image analysis model corresponding to the first number of images to be analyzed of different target modalities to obtain a first number of reference images; wherein the feature extraction sub-model corresponding to each image to be analyzed is capable of characterizing the correlation between the second number of images to be analyzed of different sample modalities;
[0138] The modal perception sub-model module is used to perform image processing on a first number of reference images through the modal perception sub-model in the image analysis model to obtain a first target image.
[0139] Based on the foregoing embodiments, in other embodiments of the present application, the modality perception sub-model is used to perform image fusion processing on reference images of at least two different modalities, or the modality perception sub-model is used to perform image fusion processing after feature enhancement processing on reference images of at least two different sample modalities.
[0140] Based on the foregoing embodiment, in other embodiments of the present application, when the modality perception sub-model is used to perform feature enhancement processing on reference images of at least two different sample modalities and then perform image fusion processing, the modality perception sub-model is specifically used to implement the following steps:
[0141] Performing image fusion processing on a first number of reference images through a modality perception sub-model to obtain a target fused image;
[0142] Determining a similarity coefficient between each reference image and the target fused image using a modality perception sub-model to obtain a first number of similarity coefficients;
[0143] A first target image is obtained by performing feature enhancement processing on a first number of similarity coefficients and a first number of reference images through a modal perception sub-model.
[0144] Based on the foregoing embodiment, in other embodiments of the present application, the modality perception sub-model implementation step may be implemented by performing feature enhancement processing on the first number of similarity coefficients and the first number of reference images using the modality perception sub-model to obtain a first target image, which may be specifically implemented by the following steps:
[0145] Performing feature enhancement processing on each reference image using a corresponding similarity coefficient through the modality perception sub-model to obtain a first number of sub-feature images;
[0146] The first number of sub-feature images are subjected to image fusion processing through the modal perception sub-model to obtain a first target image.
[0147] Based on the foregoing embodiment, in other embodiments of the present application, the image processing apparatus further includes a determination unit and a model training unit; wherein:
[0148] The obtaining unit is further configured to obtain a third number of groups of sample images and a third number of marked positions for the analysis object in the third number of groups of sample images; wherein each group of sample images includes a second number of sample images corresponding to different sample modalities;
[0149] A determination unit, used to determine an image model to be trained;
[0150] The model training unit is used to use a third number of groups of sample images and a third number of marked positions to perform model training on the image model to be trained to obtain an image analysis model.
[0151] Based on the foregoing embodiment, in other embodiments of the present application, the model training unit is specifically configured to implement the following steps:
[0152] Performing fusion enhancement processing on the areas where the analysis objects are located in the third number of groups of sample images using the image model to be trained, to obtain a third number of second target images;
[0153] Based on a third number of groups of sample images, a third number of second target images, and a third number of marked positions, the image model to be trained is trained to obtain an image analysis model.
[0154] Based on the foregoing embodiment, in other embodiments of the present application, the model training unit is configured to implement the step of performing model training on the image model to be trained based on the third number of groups of sample images, the third number of second target images, and the third number of marked positions to obtain the image analysis model, which can be specifically implemented by the following steps:
[0155] Determine a loss value between each sample image in the target group sample image and the corresponding marked position to obtain a first loss value corresponding to the target group sample image; wherein the first loss value corresponding to the target group sample image includes a second number of loss values;
[0156] Determine a loss value between a second target image corresponding to the target group sample image and a corresponding marked position to obtain a second loss value corresponding to the target group sample image;
[0157] The first loss value corresponding to the target group sample image and the second loss value corresponding to the target group sample image are reversely transmitted in the image model to be trained to continuously train the parameters of the image model to be trained to obtain the image analysis model.
[0158] It should be noted that the information interaction process and description between the units and modules in this embodiment can be referred to in detail. Figures 1 to 4 The interactive process in the image processing method shown will not be described in detail here.
[0159] An embodiment of the present application provides an image processing device that, after obtaining a first number of images to be analyzed, performs fusion and enhancement processing on the first number of images to be analyzed using an image analysis model to obtain a first target image. Thus, because the first number can be less than the second number of all inputs to the image analysis model, the image analysis model performs enhanced fusion processing on the first number of images to be analyzed in the missing modality to obtain a first target image for the analysis object. This solves the problem of current image processing methods being unable to perform image analysis when the modality is missing, and achieves a method that can still perform analysis using a unified image processing method even when the modality is missing, thereby improving the processing efficiency of the image processing method.
[0160] Based on the above embodiments, the embodiments of the present application provide an electronic device which can be applied to Figures 1 to 4 In the image processing method provided in the corresponding embodiment, refer to Figure 10 As shown, the electronic device 5 may include: a processor 51, a memory 52 and a communication bus 53, wherein:
[0161] A communication bus 53 is used to implement communication between the processor 51 and the memory 52;
[0162] The processor 51 is used to execute the image processing program stored in the memory 52 to achieve the following Figures 1 to 4 The implementation process of the image processing method provided in the corresponding embodiment will not be repeated here.
[0163] Based on the above embodiments, the embodiments of the present application provide a computer-readable storage medium, referred to as a storage medium for short, which can be applied to Figures 1 to 4 In the method provided in the corresponding embodiment, the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the following Figures 1 to 4 The implementation process of the method provided in the corresponding embodiment will not be repeated here.
[0164] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0165] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0166] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0168] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. An image processing method, comprising: Obtaining a first number of images to be analyzed; wherein each of the images to be analyzed corresponds to a different target modality of the target photographed object; Through the image analysis model, the first number of images to be analyzed are fused and enhanced to obtain a first target image; wherein, the first target image is used to enhance the display of the distribution area of the analysis object in the first number of images to be analyzed, and the analysis object belongs to the photographed object. The image analysis model is trained by sample images corresponding to a second number of different sample modalities, the first number is less than or equal to the second number, and the target modality belongs to the sample modality; the image analysis model includes feature extraction sub-models corresponding to the images to be analyzed of different target modalities; the feature extraction sub-model corresponding to each of the images to be analyzed can characterize the correlation relationship between the second number of images to be analyzed of different sample modalities.
2. The method according to claim 1, wherein the step of performing fusion enhancement processing on the first number of images to be analyzed using an image analysis model to obtain a first target image comprises: If the first number is 1, based on the feature extraction sub-models corresponding to the first number of images to be analyzed of different target modalities in the image analysis model, the corresponding images to be analyzed are processed to obtain the first target image.
3. The method according to claim 1, wherein the step of performing fusion enhancement processing on the first number of images to be analyzed using an image analysis model to obtain a first target image comprises: If the first number is greater than or equal to 2 and less than or equal to the second number, processing the corresponding images to be analyzed of the first number of different target modalities using feature extraction sub-models in the image analysis model to obtain a first number of reference images; The first target image is obtained by performing image processing on the first number of reference images through the modality perception sub-model in the image analysis model.
4. According to the method according to claim 3, the modality perception sub-model is used to perform image fusion processing on reference images of at least two different modalities, or the modality perception sub-model is used to perform feature enhancement processing on reference images of at least two different sample modalities and then perform image fusion processing.
5. The method according to claim 4, wherein, when the modality perception sub-model is used to perform feature enhancement processing on reference images of at least two different sample modalities and then perform image fusion processing, performing image processing on the first number of reference images using the modality perception sub-model in the image analysis model to obtain the first target image comprises: Performing image fusion processing on the first number of reference images through the modality perception sub-model to obtain a target fused image; Determining a similarity coefficient between each of the reference images and the target fused image using the modality perception sub-model to obtain the first number of similarity coefficients; The first target image is obtained by performing feature enhancement processing on the first number of the similarity coefficients and the first number of the reference images through the modal perception sub-model.
6. The method according to claim 5, wherein the step of performing feature enhancement processing on the first number of similarity coefficients and the first number of reference images using the modality perception sub-model to obtain the first target image comprises: Performing feature enhancement processing on each of the reference images using the corresponding similarity coefficient using the modality perception sub-model to obtain the first number of sub-feature images; The first target image is obtained by performing image fusion processing on the first number of sub-feature images through the modal perception sub-model.
7. The method according to any one of claims 1 to 6, further comprising: Obtaining a third number of groups of sample images and a third number of marked positions for the analysis object in the third number of groups of sample images; wherein each group of sample images includes the second number of sample images corresponding to different sample modalities; Determine the image model to be trained; The image model to be trained is trained using the third number of groups of sample images and the third number of marked positions to obtain the image analysis model.
8. The method according to claim 7, wherein the step of using the third number of groups of sample images and the third number of marked positions to perform model training on the image model to be trained to obtain the image analysis model comprises: Performing fusion enhancement processing on the regions where the analysis objects are located in the third number of groups of sample images using the image model to be trained, to obtain the third number of second target images; Based on the third number of groups of sample images, the third number of the second target images and the third number of the marked positions, model training is performed on the image model to be trained to obtain the image analysis model.
9. The method according to claim 8, wherein the step of training the image model to be trained based on the third number of groups of sample images, the third number of second target images, and the third number of marked positions to obtain the image analysis model comprises: Determining a loss value between each sample image in the target group sample image and the corresponding marked position to obtain a first loss value corresponding to the target group sample image; wherein the first loss value corresponding to the target group sample image includes the second number of loss values; Determine a loss value between the second target image corresponding to the target group sample image and the corresponding marked position to obtain a second loss value corresponding to the target group sample image; The first loss value corresponding to the target group sample image and the second loss value corresponding to the target group sample image are reversely transmitted in the image model to be trained to continuously train the parameters of the image model to be trained to obtain the image analysis model.
10. An image processing device, comprising: Obtaining units and model processing units; wherein: The obtaining unit is configured to obtain a first number of images to be analyzed, wherein each of the images to be analyzed corresponds to a different target modality of the target photographed object; The model processing unit is used to perform fusion enhancement processing on the first number of images to be analyzed through an image analysis model to obtain a first target image; wherein, the first target image is used to enhance the display of the distribution area of the analysis object in the first number of images to be analyzed, and the analysis object belongs to the photographed object. The image analysis model is trained by sample images corresponding to a second number of different sample modalities, the first number is less than or equal to the second number, and the target modality belongs to the sample modality; the image analysis model includes feature extraction sub-models corresponding to images to be analyzed of different target modalities; the feature extraction sub-model corresponding to each of the images to be analyzed can characterize the correlation relationship between the second number of images to be analyzed of different sample modalities.
Citation Information
Patent Citations
Training method and device and detection method and device for human face living body detection model, and electronic equipment
CN111597918A
Image detection method, device and equipment and computer readable storage medium
CN112767303A
Training method of image fusion model, image fusion method and electronic equipment
CN113052025A