A method and device for cross-modal image generation and detection

Through cross-modal image generation and detection methods, multimodal image registration and fusion are used to generate and detect multimodal images, solving the problem of high-cost acquisition of multimodal images and achieving efficient tumor detection.

CN115272167BActive Publication Date: 2025-07-29UNIV OF SCI & TECH BEIJING +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210545711.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-07-29
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

The prior art requires the acquisition of multiple modal images for detection in cross-modal medical images, which leads to high cost and inconvenience in time to detect tumors. The image generation is mainly used for data enhancement and not directly used for detection.

Method used

By acquiring image sets of multiple modalities for registration and fusion, using cross-modal image generation model and object detection model for training, generating and detecting multimodal fusion images, achieving end-to-end model optimization.

Benefits of technology

It reduces the medical cost of patients, and can achieve key target detection using only a single modal image, improving the accuracy and efficiency of tumor detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272167B_ABST
    Figure CN115272167B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for cross-modal image generation and detection, belonging to the fields of image processing and artificial intelligence technologies, capable of realizing cross-modal generation of medical images and completing target detection; the method includes: S1. Obtain image sets of two or more modalities and perform registration; S2. Perform fusion on the image sets to obtain a first multi-modal fusion image; S3. Use the first multi-modal fusion image and its corresponding single-modal images to train a cross-modal image generation model; S4. Input the single-modal image corresponding to the first multi-modal fusion image into the cross-modal image generation model to obtain a second multi-modal fusion image; S5. Input the second multi-modal fusion image and the corresponding first multi-modal fusion image into a target detection model for training; S6. Perform end-to-end optimization on the model to obtain a final cross-modal image generation model and a target detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and artificial intelligence, and particularly to a method and device for cross-modal image generation and detection. Background Art

[0002] Cross-modal image processing and analysis is an important research direction in the field of intelligent medical research. It observes and analyzes the same area by integrating imaging methods with different imaging principles, providing data support for key target recognition. Observation modalities include Computed Tomography (CT), Positron Emission Computed Tomography (PET), and Magnetic Resonance Imaging (MRI), etc.

[0003] Generally, to detect and identify a certain key target, it is necessary to collect and process multiple modal images for analysis. For example, cervical cancer screening often mainly uses CT and PET images. Since the imaging mechanisms of medical images in different modalities are different, the degree of lesion information they reflect is also different. Doctors need to collect and compare the CT images and PET images of patients to accurately detect the location of tumors in the images. When using deep learning methods for image recognition, its accuracy depends on the degree of lesion information reflected by medical images in different modalities. Among them, the accuracy of PET / CT image recognition is the highest, the accuracy of CT image recognition is the lowest, and the accuracy of PET image recognition is between the two. However, in actual diagnosis, PET imaging detection items are too expensive, bringing huge economic costs to patients and also hindering the timely detection and prevention of tumors.

[0004] With the breakthrough progress of artificial intelligence theory and computer vision technology in the field of image processing, deep learning has gradually become the mainstream method for image generation and target detection. However, currently, most methods use image generation methods as data augmentation in target detection training and are not used for cross-modal image generation and detection.

[0005] Therefore, it is necessary to study a method and device for cross-modal image generation and detection to address the deficiencies of the existing technology and solve or alleviate one or more of the above problems. Summary of the Invention

[0006] In view of this, the present invention provides a method and device for cross-modal image generation and detection, which can realize cross-modal generation of medical images and complete target detection.

[0007] On the one hand, the present invention provides a method for cross-modal image generation and detection, and the steps of the method include:

[0008] S1. Obtain image sets of two or more modalities and perform registration;

[0009] S2. Fuse the registered image sets to obtain a first multi-modal fusion image;

[0010] S3. Use the first multi-modal fusion image and its corresponding single-modal images to train a cross-modal image generation model, and obtain a trained cross-modal image generation model;

[0011] S4. Input the single-modal images corresponding to the first multi-modal fusion image into the trained cross-modal image generation model to obtain a second multi-modal fusion image;

[0012] S5. Input the second multi-modal fusion image and its corresponding first multi-modal fusion image into an object detection model for training to obtain a trained object detection model;

[0013] S6. Perform end-to-end tuning on the trained cross-modal image generation model and object detection model to obtain the final cross-modal image generation model and object detection model. When in use, the cross-modal image generation model and object detection model are automatically used jointly.

[0014] The first multi-modal fusion images in steps S3 and S4 can be the same image set, different image sets, or partially overlapping.

[0015] For the above aspects and any possible implementation manners, a further implementation manner is provided. The specific content of step S1 includes:

[0016] Obtain a number of first-modal images and a number of second-modal images, and pair them one by one;

[0017] Perform image sharpening, Gaussian blur, feature extraction, and edge detection processing on each pair of images in sequence, and then perform correction using a correction algorithm to obtain a registered image set with the same size.

[0018] For the above aspects and any possible implementation manners, a further implementation manner is provided. The content of fusing the image set in step S2 includes:

[0019] Convert all modality images in the image set into grayscale images;

[0020] Perform linear gray value transformation on the grayscale images corresponding to the first modality image to the N-th modality image respectively, and perform a concatenation operation on the transformed grayscale images corresponding to different modality images in the channel direction. The obtained concatenated image is used as the first multi-modal fusion image; N represents the number of modalities of the images in the image set, which is an integer greater than or equal to 2; one grayscale image corresponds to one channel, that is, a single channel. Performing concatenation in the channel direction mainly refers to the superposition of multiple grayscale images.

[0021] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The content of fusing the image set further includes: professionals annotate the concatenated image to obtain a ground truth bounding box for the abnormal region, and use the multi-modal fusion image with the ground truth bounding box as the first multi-modal fusion image. This first multi-modal fusion image can be used in step S5 and step S6.

[0022] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The content of the linear gray value transformation includes: multiplying all gray values by a transformation coefficient, and then compressing the value range of the gray values to a suitable interval.

[0023] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The transformation coefficients of the grayscale images corresponding to the first modality image and the second modality image are both 0.5.

[0024] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The cross-modal image generation model in step S3 can be various types of generative models, including generative adversarial networks, flow models, variational autoencoders, etc.

[0025] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The object detection model in step S5 can be various types of object detection models, including anchor-based single-stage models (such as Yolov5x model, RetinaNet model, etc.) and two-stage models (such as Faster RCNN model, Cascade RCNN model, etc.), anchor-free models (such as CenterNet model, Fcos model, etc.), and so on.

[0026] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The first modality image is a CT image, and the second modality image is a PET image.

[0027] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The correction algorithm includes image scaling, image cropping, and / or image zero-padding.

[0028] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The content of the end-to-end optimization in step S6 includes: inputting a first-modal image into a trained cross-modal image generation model to obtain a second multi-modal fusion image, bilinearly interpolating and magnifying the second multi-modal fusion image to the same pixel size as the second-modal image, and then inputting it into a trained object detection model to obtain a detection result; backpropagating the loss of the detection result to the cross-modal image generation model to complete the end-to-end training.

[0029] The loss refers to the loss of the above detection result relative to the standard result. The standard result is specifically: the result after professional personnel annotate the abnormal region true value annotation box for the first multi-modal fusion image.

[0030] On the other hand, the present invention provides a device for cross-modal image generation and detection. The device includes:

[0031] A feature extraction unit for separately extracting features from each modal image;

[0032] An image registration unit for separately performing edge detection on each modal image and using a correction algorithm to achieve the registration of different modal images;

[0033] An image fusion unit for fusing the registered different modal images to obtain a first multi-modal fusion image;

[0034] A style transfer unit for training the cross-modal image generation model during the model construction stage; and for generating a second multi-modal fusion image by inputting a single-modal image during the model usage stage;

[0035] An object detection unit for training the object detection model during the model construction stage; and for performing object detection on the second multi-modal fusion image generated by the style transfer unit during the model usage stage;

[0036] A tuning unit for implementing end-to-end tuning between the cross-modal image generation model and the object detection model. As Figure 5 shown.

[0037] On yet another aspect, the present invention provides a device for cross-modal image generation and detection, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The device is characterized in that when the processor executes the computer program, it implements the steps of any of the above methods.

[0038] Compared with the prior art, one of the above technical solutions has the following advantages or beneficial effects: The solution of the present invention can avoid the current situation where it is necessary to collect two modalities of images to detect the target area. Only a single modality image can be used to achieve key target detection through an image generation and detection model, thereby reducing the cost of multi-modal image observation and preparation;

[0039] Another one of the above technical solutions has the following advantages or beneficial effects: The method of the present invention only uses single-modal CT images to generate multi-modal PET / CT images, reducing the medical cost of patients; at the same time, through the detection network, the generated PET / CT multi-modal images are used as detection targets for multi-modal lesion detection of cervical cancer, achieving the goal of obtaining PET / CT images and their lesion detection results only using CT images, and reducing the medical cost of patients;

[0040] Another one of the above technical solutions has the following advantages or beneficial effects: The present invention solves the problem that in the field of cervical cancer medical image generation, most of the generated images are only used for data augmentation to train detection models and have not directly used the generated images as detection targets.

[0041] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned technical effects simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 is a flowchart of a cross-modal image generation and detection method provided by an embodiment of the present invention;

[0044] Figure 2 is a comparison diagram of an adaptive multi-modal image fusion method with a single-modal CT image and a single-modal PET image provided by an embodiment of the present invention;

[0045] Figure 3 is a schematic flowchart of an adaptive multi-modal image fusion method provided by an embodiment of the present invention;

[0046] Figure 4 is a network structure diagram of a cross-modal image generation and detection method for cervical cancer medical images provided by an embodiment of the present invention;

[0047] Figure 5It is a schematic diagram of a cervical cancer medical image cross-modal image generation and detection device provided by an embodiment of the present invention;

[0048] Figure 6 It is the experimental result of object detection for different modal images provided by the present invention;

[0049] Figure 7 It is a visualization result diagram of experimental comparison provided by an embodiment of the present invention. Detailed implementation manners

[0050] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0051] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] Aiming at the deficiencies of the prior art, the present invention provides a method for cross-modal image generation and detection. Taking cervical cancer screening as an example, it can generate PET-CT multi-modal fusion images from CT modal images, and then apply the generated PET-CT modal images to object detection, so as to achieve the goal of obtaining the generated PET-CT fusion images and PET-CT image recognition and analysis only based on CT images, thereby reducing the cost of collecting PET images in practical applications.

[0053] The specific content of the method of the present invention includes:

[0054] Use an image feature extraction algorithm to extract the features of all images in the CT medical image set and the PET medical image set, and use an edge detection algorithm to detect the outer edges of the lesion features of the CT image and the PET image respectively;

[0055] Use a correction method to register the images in the CT image and the PET image by one or more of image cropping, image zero-value filling, and image scaling;

[0056] Use a feature fusion method to first convert the registered CT image and PET image into grayscale images. Perform a linear transformation on the grayscale value of the grayscale image of the PET image to compress the grayscale value range, and copy it to the second channel of the RGB image; then perform a linear transformation on the grayscale value of the grayscale image of the CT image to compress the grayscale value range, and superimpose it with the linearly transformed grayscale image of the PET image, and copy it to the third channel of the RGB image; the pixel values of the first channel of the RGB image are all zero;

[0057] An adaptive multi-scale feature extraction method is adopted. During the model training process, high weights are assigned to the pixels in the lesion area and added to the loss function, while low weights are assigned to the pixels in the non-lesion area and added to the loss function, enabling the model to pay more attention to the features of the lesion area.

[0058] An image generation method is adopted. The real CT image and the annotation of the real CT image are input into the generative adversarial network. The model learns the features of the CT image and the annotation during training to generate a PET-CT image.

[0059] A target detection method is adopted. The generated PET-CT image and the real PET-CT image (which can be the image obtained by fusing the PET image and the CT image as mentioned above) are input into the target detection model, and its label is the label of the real PET-CT image, enabling the model to learn the joint distribution of the generated PET-CT image and the real PET-CT image during training and detect the lesion area of the generated PET-CT image.

[0060] The cross-modal image generation and detection method of the present invention includes a model construction stage and a model usage stage. The specific process of the model construction stage is as Figure 1 shown and includes:

[0061] S101, Obtain N CT images and N PET images taken by medical devices, pair the CT images and the PET images one by one, and adopt a series of image feature extraction methods. First, adopt an image sharpening method, and then adopt a Gaussian blur method to extract the paired CT image features and PET image features. Then, use the canny edge detection method to detect the outer edges of the CT image and the PET image respectively. After obtaining the outer edges of the paired CT image and PET image, adopt a correction method. Calculate the rectangular frame that can enclose this outer edge using the obtained outer edge for calculating the scaling coefficient. Since the size of the CT image is 512 * 512 pixels and the size of the PET image is 128 * 128 pixels, crop and scale the PET image so that the non-zero value features in the CT image and the PET image are aligned, and perform zero-padding on the area that is less than 512 * 512 to make the filled size consistent with the size of the CT image. Finally, complete the registration and obtain a multi-modal image dataset.

[0062] S102. First, convert the CT image and the PET image after image registration into grayscale images. Then, perform a linear transformation on the grayscale values of the grayscale image of the PET image, multiply all its grayscale values by 0.5 to compress the range of grayscale values, and place it in the second channel (i.e., the G channel) of the RGB image. Next, perform a linear transformation on the grayscale values of the grayscale image of the CT image, multiply all its grayscale values by 0.5 to compress the range of grayscale values, and superimpose it on the PET image after the linear transformation of the grayscale values, and place it in the third channel (i.e., the B channel) of the RGB image. There is nothing on the first channel (i.e., the R channel), and all pixel values are zero. The RGB image composed of the aforementioned R, G, and B channels is used as the result of multimodal image fusion.

[0063] S103. Input N CT images and N PET-CT images into the style transfer model (i.e., the cross-modal image generation model) to train the style transfer model. For the input of the discriminator of the style transfer model, the images generated by the generator and the true annotations of the lesion regions are used. The region within the annotation box is defined as the high-weight region, and the region outside the annotation box is defined as the low-weight region. An adaptive multi-scale feature extraction method is adopted.

[0064] The expression for cross-modal image generation is as follows:

[0065]

[0066] Among them, p i represents the i-th pixel in the image, a low refers to the weight of the non-lesion region, and a high refers to the weight of the lesion region. For the pixels in the non-lesion region, low weights are assigned and added to the loss function. For the pixels in the lesion region, high weights are assigned and added to the loss function, so that the model pays more attention to the features of the lesion region. Adopt the cross-modal image conversion method, input the real CT image and the annotation of the real CT image into the style transfer model, the real label is the PET-CT image, and the model learns the features and annotations of the CT image during training to generate a multimodal PET-CT image.

[0067] The cross-modal image generation model in the present invention can be a neural network model based on deep learning.

[0068] S104. Train the object detection model. Input the generated PET-CT image and the real PET-CT image into the object detection model to train it. Its label is the label of the real PET-CT image, so that the model learns the joint distribution of the generated PET-CT image and the real PET-CT image during training and can detect the lesion region of the generated PET-CT image.

[0069] S105. End-to-end optimize the trained style transfer model and object detection model to obtain the final cross-modal image generation and detection model. Specifically, input the original single-modal CT image with a size of 512 pixels * 512 pixels into the cross-modal image generation network, and then output a PET-CT multi-modal image with a size of 512 pixels * 512 pixels. Bilinearly interpolate and enlarge the output PET-CT multi-modal image with a size of 512 pixels * 512 pixels to a size of 1024 pixels * 1024 pixels, and input it into the object detection model. Backpropagate the loss of object detection to the generator for end-to-end training.

[0070] As Figure 2 shown, it is a comparison diagram of the adaptive multi-modal image fusion method provided by an embodiment of the present invention with a single-modal CT image and a single-modal PET image. The first column is the single-modal CT image, which reflects the physiological structure information; the second column is the single-modal PET image, which reflects the tissue metabolism information; the third column is the fusion result obtained by the adaptive multi-modal image fusion method of the present invention, which can reflect both the physiological structure information and the tissue metabolism information in one image.

[0071] As Figure 3 shown, it is a schematic flowchart of the adaptive multi-modal image fusion method provided by an embodiment of the present invention. Process the CT single-modal image and the PET single-modal image respectively. The first row processes the CT single-modal image, and the second row processes the PET single-modal image. Specifically, Figure 3 the "sharpening" operation shown in it refers to performing image sharpening on the image; the "Gaussian Blur" operation refers to performing Gaussian blur on the image; the "EdgeExtraction" operation refers to performing canny edge detection on the image to extract the edges; the "Addition" operation refers to adding the pixel values at the corresponding positions of two images, and for pixels with pixel values exceeding 255, set their pixel values equal to 255.

[0072] Figure 3The "Location" operation shown in [Figure] refers to finding a rectangular box for each of the CT image and the PET image, such that the rectangular box can completely enclose the non-zero value features in the figure, and the area of the rectangular box is as small as possible. Specifically, starting from the center of the image, search for the row or column position indices of the minimum number of non-zero value features in the up, down, left, and right four directions respectively. For the up and down directions, use the row position index, that is, traverse all the pixels in the row at the current position, count the number of non-zero value pixels, and stop the search if it is less than a certain threshold. Then determine that this row is one side of the rectangular box. After completing the operations in the up and down directions, two sides of the rectangular box can be obtained. For the left and right directions, use the column position index, that is, traverse all the pixels in the column at the current position, count the number of non-zero value pixels, and stop the search if it is less than a certain threshold. Then determine that this column is one side of the rectangular box. After completing the operations in the left and right directions, two sides of the rectangular box can be obtained. After completing the operations in the four directions, a complete rectangular box can be obtained.

[0073] Figure 3 The "Calculate offset and ratio" operation shown in [Figure] refers to calculating the width and height of the rectangular box obtained after performing the "Location" operation on the CT single-modal image, calculating the width and height of the rectangular box obtained after performing the "Location" operation on the PET single-modal image, and calculating the ratios of the area of the rectangular box in the PET image to the area of the rectangular box in the CT image in the x and y directions. Then magnify the PET according to the ratios to align the features in the rectangular boxes of the two modal images, thus completing the image registration.

[0074] Figure 3 The "Weight Fusion" operation shown in [Figure] refers to first converting the CT image and the PET image after image registration into grayscale images, and placing them on different channels of the RGB image respectively. Among them, perform a linear transformation on the gray values of the PET image, multiply all its gray values by 0.5 to compress the gray value range, and place it on the second channel of the RGB image. Then perform a linear transformation on the gray values of the CT image, multiply all its gray values by 0.5 to compress the gray value range, and superimpose it with the PET image after the gray value linear transformation, and place it on the third channel of the RGB image to obtain the result of multi-modal image fusion.

[0075] As Figure 4As shown, it is the network structure diagram of the cervical cancer medical image cross-modal image generation and detection method provided by the embodiment of the present invention; among them, "Backbone" is the feature extraction layer in the object detection network (the object detection network refers to the object detection model described above), which is composed of multiple convolutional layers and aggregates at different image fine-grained levels to obtain the features of the image; "Head" is the detection task branch in the object detection network, which is composed of multiple convolutional layers, and its main function is to generate object bounding boxes and predict their respective categories.

[0076] In the cervical cancer medical image cross-modal image generation and detection method, the PET image and the CT image are fused, and a cervical cancer cross-modal image generation network based on deep learning is constructed and the cervical cancer cross-modal image generation network is trained; using the trained cervical cancer cross-modal image generation network, N CT images of the test set are used as inputs, and N generated PET-CT images are output.

[0077] In the specific implementation manner of the foregoing cervical cancer multi-modal medical image fusion method, further, the registration of the N cervical cancer CT image sequences and the N cervical cancer PET image sequences includes: determining the edge of the irregular feature region of the first CT image and determining the edge of the irregular feature region of the first PET image; determining the external rectangular frame of the irregular feature region of the first CT image and determining the external rectangular frame of the irregular feature region of the first PET image; using image processing methods to perform several operations such as image scaling and image cropping on the CT image and the PET image, and aligning the external rectangular frame of the irregular feature region of the first CT image with the external rectangular frame of the irregular feature region of the first PET image. The remaining N-1 CT images and N-1 PET images are also processed and operated in the same manner as above.

[0078] In this embodiment, the image processing methods adopted include: image scaling, image cropping, and image zero-value filling.

[0079] In the specific implementation manner of the foregoing cervical cancer multi-modal medical image fusion method, further, in the fusion method, the registered multi-modal images are converted into grayscale images and stitched according to the channel dimension. The grayscale value of the CT image multiplied by 0.5 plus the grayscale value of the PET image multiplied by 0.5 is placed on the B channel of the three RGB channels of the image, and the grayscale value of the ET image multiplied by 0.5 is placed on the G channel of the three RGB channels of the image.

[0080] In the specific implementation manner of the foregoing fundus cervical cancer cross-modal image generation method, further, the construction of the cervical cancer cross-modal image generation network based on deep learning and the training of the cervical cancer cross-modal image generation network include:

[0081] Obtain a pre-set training data set, input the training data set into the cervical cancer cross-modal image generation network, and use the adaptive moment estimation optimizer to train the cervical cancer cross-modal image generation network until the error between the CT image and the generated PET-CT image is less than a preset threshold, obtaining the trained cervical cancer cross-modal image generation network. The training data set includes: CT images, PET-CT images, and annotations of the lesion areas of the CT images.

[0082] In this embodiment, the cervical cancer cross-modal image generation network uses the Pix2Pix network, which includes a generator and a discriminator. Its generator is a U-Net network: The generator includes an encoding stage and a decoding stage; the encoding stage includes: 5 feature extraction modules, each feature extraction module includes: N convolutional modules, and the N convolutional modules are used to hierarchically extract image features. The convolutional module includes: an activation function layer and a convolution operation; the decoding stage includes: N-1 skip connection operations and N-1 transposed convolution modules. Among them, each transposed convolution module includes: an activation function layer, a transposed convolution operation, and a normalization layer. The discriminator includes N convolutional modules, and the N convolutional modules are used to hierarchically extract image features. The convolutional module includes: an activation function layer and a convolution operation;

[0083] In this embodiment, in the encoding stage of the generator, the CT image is input into the feature extraction branch. Each feature extraction branch extracts the image features of the corresponding level through M convolutional modules, and then conveys them to the decoding stage; the decoding stage is used to restore the received image features to the original image size.

[0084] In this embodiment, assume that each feature extraction branch in the encoding stage of the generator includes: 5 convolutional modules, and each convolutional module includes 1 downsampling operation; then in the encoding stage, the CT image feature extraction branch can be used to extract the image features of the corresponding level through 5 convolutional modules, and then convey them to the decoding stage.

[0085] In this embodiment, assume that the decoding stage of the generator includes: 4 skip connection operations, 4 transposed convolution modules, 1 transposed convolution operation, and 1 convolution operation. Among them, each transposed convolution module includes: 2 activation function layers and 1 transposed convolution operation, and each transposed convolution operation doubles the feature size.

[0086] In this embodiment, the discriminator in the generator includes 6 convolutional modules. Each convolutional module is composed of a convolutional layer, a normalization layer, and an activation function layer. The first 5 convolutional modules are used to hierarchically extract image features, and the output of the last convolutional module is a feature vector of 1 channel, indicating the degree to which the discriminator predicts true or false for the input.

[0087] In this embodiment, the generator of the cervical cancer cross-modal image generation network is a U-shaped network-like structure.

[0088] In this embodiment, the target detection network adopted is a one-stage model Yolov5, which mainly includes a backbone part, a neck part, and a head part. The backbone is a feature extraction layer with continuous downsampling, aggregating at different image fine-grain levels to obtain image features; the neck consists of a series of convolutional layers, and its function is to send the image features to the prediction layer; the main function of the head is to generate target boxes and predict their belonging categories.

[0089] In this embodiment, the training process is divided into two processes: training the image generation network and training the target detection network. The training process of the image generation network is as follows: Input the single-modal CT image into the image generation network for encoding and feature extraction, and then through decoding, convert the extracted features into a multi-modal PETCT image to obtain an output value. The output value is the multi-modal PETCT image, which is compared with the real multi-modal PETCT image, and the loss is calculated using a loss function. The obtained loss is backpropagated to the generator to enable the generator to update the network parameters and complete one training process. This process is repeated until the loss between the output of the generator and the real multi-modal PETCT image reaches the minimum.

[0090] Then, for the trained generator, input the single-modal CT image, and use the generated multi-modal PETCT image and the real multi-modal PETCT image together as inputs to train the target detection network. Calculate the loss with the target detection result of the real multi-modal PETCT image and backpropagate it to the target detection network to complete one training process. This process is repeated until the loss between the output of the target detection network and the target detection result of the real multi-modal PETCT image reaches the minimum. Obtain the target detection result of the generated multi-modal PETCT image.

[0091] As Figure 6 shown, it is the preliminary experimental result of this embodiment. Specifically, CT in the table represents the target detection result for the single-modal CT image, and CT->PETCT represents the target detection result for the PET-CT multi-modal image data generated from the single-modal CT image data of this application (that is, the virtual PETCT image obtained after inputting the single-modal CT image into the trained style transfer model of this application). This experiment selected two target detection models based on the Yolov5x model and the Faster RCNN model respectively for detection, so as to obtain Figure 6The comparison results of the accuracy rate. The SGDMomentum optimizer is adopted, the initial learning rate is 0.01, the input image size is 1024 pixels * 1024 pixels, the number of training rounds is 60, and the evaluation index adopted is AP50 (average precision). For the experiment on CT images, both the training set and the test set are real data (the training here refers to training the object detection model and then performing detection); for the experiment on the virtual PET-CT multimodal image data generated from CT single-modal images, the training set is a 1:1 mixture of real PET-CT data and the generated PET-CT data, and the test set is the generated PET-CT multimodal image data. It should be noted that in this experiment, the original CT images and original PET images corresponding to the training set and the test set in the two types of images are the same, and the number of each set is also the same. The purpose is to obtain experimental results with as single influencing factor as possible. Through the model method proposed by the present invention, the tumor recognition accuracy rate of converting CT into PETCT is better than the accuracy rate of only using CT for analysis and recognition, which corroborates the effectiveness of the method of the present invention. In addition, limited by the data, the experimental results can only reflect the effectiveness of the method of the present invention, but the technical effects that can be achieved by the method of the present invention cannot be judged based on these experimental results. When the amount of the training set is sufficient, the detection accuracy rate of the present invention can be further improved.

[0092] As Figure 7 shown, it is the experimental comparison visualization result of this embodiment. Specifically, Figure 7 in the first column is the annotation of the real CT single-modal image, the second column is the model prediction of the real CT single-modal image, the third column is the annotation of the real PETCT single-modal image, and the fourth column is the model prediction of the PET-CT multimodal image generated from the CT single-modal image. The annotation here is performed by professionals (such as doctors) to obtain the ground truth bounding box of the abnormal area.

[0093] The above has introduced in detail a method and device for cross-modal image generation and detection provided by an embodiment of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

[0094] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used in the embodiments of the present invention and the appended claims is only an associative relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the situation where A exists alone, the situation where A and B exist simultaneously, and the situation where B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented in whole or in part in the form of a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

Claims

1. A method for cross-modal image generation and detection, characterized in that, The steps of the method include: S1. Obtain image sets of two or more modalities and perform registration; S2. Fuse the registered image sets to obtain a first multi-modal fusion image; S3. Use the first multi-modal fusion image and a single-modal image of one of the two or more modal images corresponding thereto to train a cross-modal image generation model to obtain a trained cross-modal image generation model; S4. Input the single-modal image of one of the two or more modal images corresponding to the first multi-modal fusion image into the trained cross-modal image generation model to obtain a second multi-modal fusion image; S5. Input the second multi-modal fusion image and the first multi-modal fusion image corresponding thereto into a target detection model for training to obtain a trained target detection model; S6. Perform end-to-end tuning on the trained cross-modal image generation model and the target detection model to obtain a final cross-modal image generation model and a target detection model.

2. The method for cross-modal image generation and detection according to claim 1, wherein The specific content of step S1 includes: Obtain a plurality of first-modal images and a plurality of second-modal images and pair them one by one; Successively perform image sharpening, Gaussian blurring, feature extraction, and edge detection processing on each pair of images, and then use a correction algorithm for correction to obtain a registered image set with the same size.

3. The method for cross-modal image generation and detection according to claim 1, wherein, The content of fusing the image set in step S2 includes: Convert the images in the image set into grayscale images; Perform a linear transformation on the grayscale value of the grayscale image corresponding to the first-modal image and place it on the G channel of the RGB image; Perform a linear transformation on the grayscale value of the grayscale image corresponding to the second-modal image, and superimpose it on the grayscale image corresponding to the first-modal image after the linear transformation of the grayscale value, and place it on the B channel of the RGB image; Use an RGB image composed of an R channel with all pixel values being zero and the G channel and B channel after placing the images as the first multi-modal fusion image.

4. The method for cross-modal image generation and detection according to claim 3, wherein, The content of the linear transformation of the grayscale value includes: multiplying all grayscale values by a transformation coefficient, and then compressing the value range of the grayscale values to a suitable interval.

5. The method for cross-modal image generation and detection according to claim 4, wherein The transformation coefficient values of the grayscale images corresponding to the first-modal image and the second-modal image are both 0.

5.

6. The method for cross-modal image generation and detection according to claim 1, wherein, The cross-modal image generation model in step S3 is a generative adversarial network, a flow model, or a variational autoencoder.

7. The method for cross-modal image generation and detection according to claim 1, characterized in that The target detection model in step S5 is an anchor-based single-stage model, an anchor-based two-stage model, or an anchor-free model.

8. The method for cross-modal image generation and detection according to claim 2, characterized in that, The first-modal image is a CT image, and the second-modal image is a PET image.

9. The method for cross-modal image generation and detection according to claim 2, wherein The correction algorithm includes image scaling, image cropping, and / or image zero-value filling.

10. An apparatus for cross-modal image generation and detection, characterized in that, The device includes: A feature extraction unit for respectively extracting features from each modal image; An image registration unit for respectively performing edge detection on each modal image and using a correction algorithm to achieve registration of different modal images; An image fusion unit for fusing the registered different modal images to obtain a first multi-modal fusion image; A style transfer unit, which is used to train a cross-modal image generation model using a first multi-modal fusion image and a single-modal image among the corresponding single-modal images of each modality during the model construction phase; and is used to generate a second multi-modal fusion image by inputting the single-modal image among the single-modal images of each modality during the model usage phase. An object detection unit, which is used to train an object detection model using a second multi-modal fusion image and a first multi-modal fusion image during the model construction phase; and is used to perform object detection on the second multi-modal fusion image generated by the style transfer unit during the model usage phase. An optimization unit, which is used to achieve end-to-end optimization between the cross-modal image generation model and the object detection model.

Citation Information

Patent Citations

  • Ophthalmological multi-modal image retrieval method and device, server and storage medium

    CN111428072A

  • Semi-supervised multi-mode nuclear magnetic resonance image synthesis method based on coarse-to-fine learning

    CN114170118A