Image processing method and device, equipment and medium

Through the multimodal image processing model, lesion images are extracted from the original images of different modes, and combined with the information of the two modes, the problem of low segmentation accuracy of single mode images is solved, and lesion area recognition and extraction with higher accuracy is achieved.

CN120279044APending Publication Date: 2025-07-08BINZHOU TRADITIONAL CHINESE MEDICINE HOSPITAL (BINZHOU INTEGRATED TRADITIONAL CHINESE & WESTERN MEDICINE RES CENT)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350983.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, when using convolutional neural network models alone for image segmentation or segmentation of single-modal images, there is a problem of poor model reusability and high cost, resulting in low image recognition and segmentation accuracy.

Method used

The multimodal image processing model is used to extract lesion images from the original images of two different modes. Through normalization processing, image acquisition and feature fusion, the multimodal multi-scale transformer and encoder decoder structure is used to improve the recognition and extraction accuracy of lesion areas.

Benefits of technology

The extraction accuracy of lesion areas is improved, and the problem of low recognition and segmentation accuracy during single-modal image segmentation is solved, thereby achieving more accurate lesion areas identification and extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279044A_ABST
    Figure CN120279044A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, equipment and a medium, and the method comprises the steps: obtaining a first original image and a second original image of a target part, and carrying out the normalization processing of the images, and obtaining a to-be-used image; according to the image extraction size and the image acquisition step length, performing image acquisition on the to-be-used image to obtain a plurality of to-be-applied image groups; for the to-be-applied image group, inputting the first sub-image and the second sub-image into a multi-modal image processing model to obtain a target sub-image and a to-be-used evaluation identifier; determining a target evaluation identifier of the original image based on the target sub-image and the to-be-used evaluation identifier; and based on the target evaluation identifier of the original image, extracting a lesion image corresponding to the lesion from the original image. According to the technical scheme provided by the embodiment of the invention, the lesion image corresponding to the lesion is extracted from the original images of the two different modalities according to the multi-modal image processing model, and the lesion area is identified and extracted more accurately by combining the information of the two different modalities, so that the extraction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing, and in particular, to an image processing method, apparatus, device, and medium. Background Art

[0002] With the improvement of the scientific and technological level, image segmentation technology has gradually emerged in the field of image processing. Image segmentation technology can extract regions of interest, thereby separating the regions of interest from the background. For example, in medical image analysis, for images taken based on CT or MR, image segmentation technology can be used to identify and separate regions of interest in the images.

[0003] Currently, the image segmentation methods mainly used are: using a convolutional neural network model alone for image segmentation, or segmenting images under a single-modal image. However, the image expression information corresponding to different-modal original images is different. Therefore, when extracting lesion images from single-modal original images, models adapted to different modalities need to be trained to process the data corresponding to different modalities, resulting in poor model reusability and high costs. Summary of the Invention

[0004] Embodiments of the present disclosure provide an image processing method, apparatus, device, and medium to extract lesion-corresponding lesion images from two different-modal original images according to a multi-modal image processing model, and improve the accuracy of lesion tissue recognition and extraction in images.

[0005] In a first aspect, embodiments of the present disclosure provide an image processing method, which includes:

[0006] Obtain a first original image of a target part in a first modality and a second original image in a second modality, and perform normalization processing on the first original image to obtain a first image to be used, and perform normalization processing on the second original image to obtain a second image to be used;

[0007] According to the image extraction size and the image acquisition step length, sequentially perform image acquisition on the first image to be used and the second image to be used to obtain a plurality of image groups to be applied, where the image groups to be applied include a first sub-image in the first modality and a second sub-image in the second modality;

[0008] For the plurality of image groups to be applied, input the first sub-image and the second sub-image in the image group to be applied into a pre-trained multi-modal image processing model to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image, where the to-be-used evaluation identifier is used to characterize the probability of the existence of a lesion in the target sub-image;

[0009] Determine the target evaluation identifier corresponding to each pixel point in the first original image or the second original image based on the target subgraph corresponding to the multiple image groups to be applied and the corresponding evaluation identifier to be used;

[0010] Extract the lesion image corresponding to the lesion from the first original image or the second original image based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image;

[0011] Wherein, the lesion is located in the target part.

[0012] In a second aspect, an embodiment of the present invention further provides an image processing apparatus, the apparatus includes:

[0013] An image to be used determination module, configured to obtain a first original image of a target part in a first modality and a second original image in a second modality, perform normalization processing on the first original image to obtain a first image to be used, and perform normalization processing on the second original image to obtain a second image to be used;

[0014] An image group to be applied determination module, configured to sequentially perform image acquisition on the first image to be used and the second image to be used according to an image extraction size and an image acquisition step length to obtain a plurality of image groups to be applied, wherein the image group to be applied includes a first subgraph in the first modality and a second subgraph in the second modality;

[0015] A target subgraph determination module, configured to input the first subgraph and the second subgraph in the image group to be applied into a pre-trained multi-modal image processing model for the plurality of image groups to be applied, to obtain a target subgraph and a corresponding evaluation identifier to be used for the target subgraph, wherein the evaluation identifier to be used is used to characterize the probability of the existence of a lesion in the target subgraph;

[0016] A target evaluation identifier determination module, configured to determine the target evaluation identifier corresponding to each pixel point in the first original image or the second original image based on the target subgraph corresponding to the multiple image groups to be applied and the corresponding evaluation identifier to be used;

[0017] A lesion image extraction module, configured to extract the lesion image corresponding to the lesion from the first original image or the second original image based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image; wherein, the lesion is located in the target part.

[0018] In a third aspect, an embodiment of the present invention further provides an electronic device, the electronic device includes:

[0019] One or more processors;

[0020] A storage device for storing one or more programs,

[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method according to any one of the embodiments of the present invention.

[0022] In a fourth aspect, an embodiment of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the image processing method according to any one of the embodiments of the present invention when executed by a computer processor.

[0023] The technical solution of the embodiment of the present disclosure first obtains a first original image of a target part in a first modality and a second original image in a second modality, normalizes the first original image to obtain a first image to be used, normalizes the second original image to obtain a second image to be used, and then, according to the image extraction size and the image acquisition step length, sequentially performs image acquisition on the first image to be used and the second image to be used to obtain a plurality of image groups to be applied, where the image groups to be applied include a first sub-image in the first modality and a second sub-image in the second modality. Then, for the plurality of image groups to be applied, the first sub-image and the second sub-image in the image group to be applied are input into a pre-trained multi-modal image processing model to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image, where the to-be-used evaluation identifier is used to characterize the probability of the presence of a lesion in the target sub-image. Further, based on the target sub-images corresponding to the plurality of image groups to be applied and the corresponding to-be-used evaluation identifiers, a target evaluation identifier corresponding to each pixel point in the first original image or the second original image is determined. Finally, based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image, a lesion image corresponding to the lesion is extracted from the first original image or the second original image, which solves the problem of low image recognition and segmentation accuracy caused by using a convolutional neural network model alone for image segmentation or segmenting images in a single-modal image in the prior art. According to the multi-modal image processing model of the embodiment of the present invention, a lesion image corresponding to a lesion is extracted from two different-modal original images, and by combining the information of the two different modalities, the lesion area is more accurately recognized and extracted, achieving the effect of improving the extraction accuracy of the lesion area. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the introduced drawings are only the drawings of a part of the embodiments to be described by the present invention, rather than all the drawings. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0025] Figure 1 It is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure;

[0026] Figure 2 It is a schematic diagram of image acquisition provided by an embodiment of the present disclosure;

[0027] Figure 3 It is a schematic diagram of a multi-modal image processing model provided by an embodiment of the present disclosure;

[0028] Figure 4 It is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure;

[0029] Figure 5 It is a schematic structural diagram of an image processing apparatus provided by an embodiment of the present invention;

[0030] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0031] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, rather than all the structures.

[0032] Embodiment 1

[0033] Before introducing the technical solutions provided by the embodiments of the present disclosure, an exemplary description of the application scenario can be given first. The technical solutions provided by the embodiments of the present disclosure can be applied to a scenario where the lesion images corresponding to the lesions in two different-modal original images are accurately extracted based on a multi-modal image processing model. For example, for Computed Tomography (CT) images and Magnetic Resonance (MR) images, CT images provide good bone contrast, while MR images have advantages in soft tissue resolution. At this time, the CT images and MR images can be used as the original images. Based on the technical solutions of the embodiments of the present disclosure, the advantages of the image information of the two different-modal images of CT images and MR images can be combined to more accurately extract the lesion area.

[0034] Figure 1It is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of determining a lesion image corresponding to a lesion based on two different modalities of original images. This method can be executed by an image processing device, which can be implemented in the form of software and / or hardware. The hardware can be an electronic device such as a server, and the electronic device can execute the image processing method provided by this technical solution.

[0035] As Figure 1 shown, the method includes:

[0036] S110. Obtain a first original image of a target part in a first modality and a second original image of the target part in a second modality, perform normalization processing on the first original image to obtain a first image to be used, and perform normalization processing on the second original image to obtain a second image to be used.

[0037] Among them, the target part refers to the part included in the image that needs to be processed currently. For example, if the liver needs to be processed, the captured image may include the liver. At this time, the liver is the target part, and the captured image is the original image. Optionally, the first modality can be the CT modality, and the second modality can be the MR modality. The first original image refers to an image of the target part obtained by CT shooting. The second original image refers to an image of the target part obtained by MR shooting.

[0038] Optionally, for all pixel points of the first original image and the second original image, the normalization processing can be understood as scaling the pixel value data of all pixel points into a specific interval to facilitate subsequent analysis and processing of the pixel value data of the image. The image obtained after normalizing the first original image is called the first image to be used. The image obtained after normalizing the second original image is called the second image to be used.

[0039] In this embodiment, retrieve the first pixel mean and the first pixel variance corresponding to the target part in the first modality determined in advance; for all pixel points in the first original image, process the pixel values of the pixel points based on the first pixel mean and the first pixel variance to obtain the first image to be used.

[0040] It should be noted that for the first original image of the target part captured by CT, the first pixel mean is the mean corresponding to the target part determined in advance, denoted by X mean to represent, the first pixel variance

[0041] is the variance corresponding to the target part determined in advance, denoted by X std to represent. The formula for processing the pixel values of the pixel points based on the first pixel mean and the first pixel variance is:

[0042]

[0043] Among them, X 原图像 (where i, j, k are the pixel values of all pixel points in the first original image, X 归一化后 (where i, j, k are the pixel values of all pixel points in the first image to be used.

[0044] It should be noted that X mean and X std are related to the target part, and the corresponding X mean and X std values for different target parts may vary. For example, assuming the target part is the liver, by calculating the average value of the pixel values of the pixel points corresponding to liver tumors, the mean X corresponding to liver tumors determined based on historical data can be obtained. By calculating the variance of the pixel values of the pixel points corresponding to liver tumors, the variance X corresponding to the liver determined based on historical data can be obtained

[0045] If the X std value for the target part of the liver determined based on historical data is 9 and the X mean value is 4, then when performing image segmentation on liver tumors subsequently, the value of X std is 9, and the value of X mean is 4. std

[0046] In this embodiment, the second pixel mean and the second pixel variance of the target part in the second modality stored in advance are retrieved; for all pixel points in the second original image, the pixel values of the pixel points are processed based on the second pixel mean and the second pixel variance to obtain the second image to be used.

[0047] It should be noted that for the second original image of the target part taken by MR, the calculation methods of the second pixel mean and the second pixel variance are the same as those of the first pixel mean and the second pixel variance, which will not be elaborated here. The calculation formula for obtaining the second image to be used is similar to the calculation formula for obtaining the first image to be used, which will not be elaborated here.

[0048] Specifically, for the part that needs to be recognized and the lesion extracted currently, first, the CT image and the MR image of the target part taken are obtained. After normalizing the CT image and the MR image based on the formula, the first image to be used and the second image to be used after the normalization process are obtained.

[0049] S120. According to the image extraction size and the image acquisition step length, the first image to be used and the second image to be used are sequentially subjected to image acquisition to obtain a plurality of image groups to be applied, where the image groups to be applied include the first sub - image in the first modality and the second sub - image in the second modality.​

[0050] It should be noted that the image extraction size is the specific size information of the images collected when collecting the first image to be used and the second image to be used by a certain method, usually including the length, width and height of the image. For example, the image extraction size can be 1080*1080*1080 pixels. The image acquisition step is the distance that the operation window moves on the image when collecting the first image to be used and the second image to be used. For example, when performing image acquisition, if the step can be half of the extraction size pixels, it means that each time the acquisition window moves half of the extraction size pixels to the right on the image. The setting of the step can ensure that different regions of the image are covered in consecutive acquisition operations, so that more features or details can be extracted from the image. In image acquisition, a larger step can reduce the number of acquisition windows, but more image details will be lost; while a smaller step will increase the number of acquisition windows, but more image details will be retained.

[0051] Among them, after performing image acquisition on the first image to be used in the first modality according to the image extraction size and the image acquisition step, the obtained image is called the first sub-image; after performing image acquisition on the second image to be used in the second modality according to the image extraction size and the image acquisition step, the obtained image is called the second sub-image. The set of the first sub-image and the second sub-image is called the image group to be applied.

[0052] Specifically, for the first image to be used and the second image to be used, perform image acquisition on the first image to be used and the second image to be used according to the same image extraction size and image acquisition step. All the sub-images collected based on the first image to be used are the first sub-images, and all the sub-images collected based on the second image to be used are the second sub-images. The sum of the first sub-images and the second sub-images is called the image group to be applied.

[0053] Exemplarily, as Figure 2 shown, assume that the pixels of the first image to be used and the second image to be used at this time are 9*9*9, the image extraction size is 2*2*2, and the image acquisition step is 1. As Figure 2 shown in (A) below, the first first sub-image or second sub-image is the blue area. As Figure 2 shown in (B) below, after collecting the image according to the acquisition step, the second first sub-image or second sub-image is the blue area. According to the image extraction size of 2*2*2 and the image acquisition step of 1, perform image acquisition on the first image to be used and the second image to be used in sequence, and multiple image groups to be applied can be obtained.

[0054] S130. For multiple image groups to be applied, input the first sub-image and the second sub-image in the image group to be applied into a pre-trained multi-modal image processing model to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image, where the to-be-used evaluation identifier is used to characterize the probability of the presence of lesions in the target sub-image.

[0055] It should be noted that, as Figure 3 shown, the core structure of the multi-modal image processing model includes two parts: an encoder and a decoder. In the encoder, first, the CT image and the MR image of W*H*D*1 are processed through a convolutional layer to obtain an image of W*H*D*C, and the feature fusion of the two images of W*H*D*C is performed through a multi-modal multi-scale transformer (Multi-modal Multi-scale Transformers for Deepfake Detection, M2TR) to obtain an image of W*H*D*C. Then, features are gradually extracted through a series of convolutional layers and downsampling to reduce the spatial resolution of the image. In the decoder, first, multi-scale features of the image are extracted through ASPP pooling to improve the robustness and accuracy of the segmentation result. Then, the size of the image is gradually restored through a series of convolutional layers, upsampling, and M2TR operations, and the feature maps in the encoder are fused to generate a final image of W*H*D*1.

[0056] It should be noted that the M2TR module adopts a self-attention mechanism to perform feature fusion on the feature information of the two input images. The specific formula is as follows:

[0057]

[0058] where Q and K are the input feature information respectively, d K is the dimension size of the feature information, and conv is a convolutional operation. It should be noted that the softmax function is a normalized exponential function, which is mainly used to "compress" a K-dimensional vector containing arbitrary real numbers into another K-dimensional real vector, so that the range of each element is between (0,1), and the sum of all elements is 1.

[0059] Among them, the target sub-image refers to the image output after inputting the first sub-image and the second sub-image into the multi-modal image processing model. The to-be-used evaluation identifier refers to the pixel value of each pixel point in the target sub-image, which is used to characterize the probability of the presence of lesions corresponding to each pixel point in the target sub-image.

[0060] In this embodiment, the first sub-image and the second sub-image in the image group to be applied are input into the multi-modal image processing model to obtain a target sub-image after fusion processing of the first sub-image and the second sub-image, and a to-be-used evaluation identifier corresponding to each pixel point in the target sub-image.

[0061] Specifically, after obtaining the first sub-image in the first modality and the second sub-image in the second modality in the image group to be applied, the first sub-image and the second sub-image are input into a pre-trained multi-modal image processing model, and the output result is the target sub-image after the fusion processing of the first sub-image and the second sub-image. The size of the target sub-image is the same as that of the first sub-image and the second sub-image. The pixel value of each pixel point in the target sub-image is the evaluation identifier to be used, and the evaluation identifier to be used represents the probability that the pixel point in the target sub-image is a lesion.

[0062] S140. Based on the target sub-images corresponding to multiple image groups to be applied and the corresponding evaluation identifiers to be used, determine the target evaluation identifier corresponding to each pixel point in the first original image or the second original image.

[0063] Among them, the target evaluation identifier is determined after calculating all the evaluation identifiers to be used based on a calculation formula. The target evaluation identifier represents the pixel value of each pixel point in the first original image or the second original image, and the target evaluation identifier represents the probability that each pixel point is a lesion finally.

[0064] In this embodiment, based on the target sub-images corresponding to multiple image groups to be applied, determine at least one evaluation identifier to be used corresponding to each pixel point; through at least one evaluation identifier to be used for the same pixel point, determine the target evaluation identifier corresponding to the pixel point in the first original image or the second original image.

[0065] It should be noted that after obtaining at least one evaluation identifier to be used corresponding to each pixel point, for each pixel point, taking the average value of all the evaluation identifiers to be evaluated existing for this pixel point can obtain the target evaluation identifier corresponding to this pixel point in the first original image or the second original image. For example, for a certain pixel point, there are 8 target sub-images of this pixel point, then the evaluation identifiers to be used corresponding to this pixel point are 8, and the average value obtained by calculating the 8 evaluation identifiers to be used is the target evaluation identifier of this pixel point.

[0066] Specifically, for each pixel point, all the target sub-images related to it and the evaluation identifiers to be used corresponding to all the target sub-images can be obtained. For each pixel point, taking the average value of the evaluation identifiers to be used existing for it can obtain the target evaluation identifier of each pixel point, and the target evaluation identifier represents the probability that each pixel point is a lesion finally.

[0067] S150. Based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image, extract the lesion image corresponding to the lesion from the first original image or the second original image.

[0068] Specifically, after obtaining the target evaluation identifier corresponding to each pixel in the first original image or the second original image, perform binary processing on the target evaluation identifier according to a preset threshold. For the target evaluation identifier greater than the preset threshold, set the corresponding pixel as the foreground area, that is, the lesion area; for the target evaluation identifier less than the preset threshold, set the corresponding pixel as the background area, that is, the non-lesion area. After binary processing, the final lesion image corresponding to the lesion can be extracted by removing small false positive areas through methods such as connected component analysis.

[0069] The technical solution of the embodiment of the present disclosure is as follows: First, obtain the first original image of the target part in the first modality and the second original image in the second modality, perform normalization processing on the first original image to obtain the first image to be used, perform normalization processing on the second original image to obtain the second image to be used, and then, according to the image extraction size and the image acquisition step, sequentially perform image acquisition on the first image to be used and the second image to be used to obtain a plurality of image groups to be applied, where the image group to be applied includes the first sub-image in the first modality and the second sub-image in the second modality. Then, for the plurality of image groups to be applied, input the first sub-image and the second sub-image in the image group to be applied into the pre-trained multi-modal image processing model to obtain the target sub-image and the corresponding evaluation identifier to be used, where the evaluation identifier to be used is used to characterize the probability of the presence of a lesion in the target sub-image. Further, based on the target sub-images and the corresponding evaluation identifiers to be used corresponding to the plurality of image groups to be applied, determine the target evaluation identifier corresponding to each pixel in the first original image or the second original image. Finally, based on the target evaluation identifier corresponding to each pixel in the first original image or the second original image, extract the lesion image corresponding to the lesion from the first original image or the second original image, which solves the problem of low image recognition and segmentation accuracy caused by using a convolutional neural network model alone for image segmentation or segmenting images in a single-modal image in the prior art. In the embodiment of the present invention, according to the multi-modal image processing model, the lesion image corresponding to the lesion is extracted from two different-modal original images, and by combining the information of the two different modalities, the lesion area can be more accurately recognized and extracted, achieving the effect of improving the accuracy of lesion area extraction.

[0070] Embodiment 2

[0071] Figure 4 It is a schematic flowchart of an image processing method provided by an embodiment of the present invention. On the basis of the foregoing embodiment, the embodiment of the present disclosure describes the steps of obtaining the multi-modal image processing model, and the specific implementation manner can refer to the technical solution of this embodiment. The same or corresponding technical terms as those in the above embodiment will not be described herein again.

[0072] As Figure 4As shown in the figure, the method specifically includes the following steps:

[0073] S210. Obtain a plurality of first sample data, where the first sample data includes at least one first sample image of a preset lesion site in the first modality and a second sample image in the second modality.

[0074] Among them, the preset lesion site refers to the site where the lesion is currently extracted. For example, the preset lesion site can be parts such as the kidney, liver, spleen, and pancreas. The first sample image refers to an image of the preset lesion site obtained by CT scanning. The second sample image refers to an image of the preset lesion site obtained by MR scanning. It should be noted that the first sample image and the second sample image are images with the pixel points already matched.

[0075] Specifically, for a certain preset lesion site, images obtained by CT and MR scanning respectively can be obtained, that is, the first sample data of this preset site is obtained. After obtaining the first sample images and the second sample images of a plurality of preset lesion sites, a plurality of first sample data can be obtained.

[0076] S220. For the plurality of first sample data, based on the first annotation region of the first sample image and the second annotation region corresponding to the second sample image in the first sample data, determine the first sample pixel mean and the first sample pixel variance of the first annotation region, and the second sample pixel mean and the second sample pixel variance of the second annotation region.

[0077] Among them, the first annotation region refers to the lesion region in the first sample image. Optionally, the lesion region can be the region where the tumor is located. The first sample pixel mean refers to the pixel mean of all pixel points in the lesion region of the first sample image, and the first sample pixel variance refers to the pixel variance of all pixel points in the lesion region of the first sample image; the second annotation region refers to the lesion region in the second sample image, the second sample pixel mean refers to the pixel mean of all pixel points in the lesion region of the second sample image, and the second sample pixel variance refers to the pixel variance of all pixel points in the lesion region of the second sample image.

[0078] Specifically, for each preset lesion site, after obtaining the first sample data of each preset lesion site, the lesion region in the first sample image of this preset lesion site, that is, the first annotation region, can be obtained, and the first sample pixel mean and the first sample pixel variance corresponding to the first annotation region can be calculated; the lesion region in the second sample image of this preset lesion site, that is, the second annotation region, can be obtained, and the second sample pixel mean and the second sample pixel variance corresponding to the second annotation region can be calculated.

[0079] S230. Determine the first image of the first sample image based on the first sample image, the first sample pixel variance, and the first sample pixel mean in the first sample data; and determine the second image of the second sample image based on the second sample image, the second sample pixel variance, and the second sample pixel mean in the first sample data.

[0080] It should be noted that the first image refers to the result of the normalization process of the first sample image. The calculation formula for the pixel value of each pixel point in the first image is:

[0081]

[0082] where X mean refers to the first sample pixel mean, and X std refers to the first sample pixel variance. Subtract the first sample pixel mean from the pixel value of each pixel point in the first sample image and then divide by the first sample pixel variance to obtain the pixel value of each pixel point in the first image.

[0083] The calculation process of the second image is the same as that of the first image and will not be elaborated here.

[0084] Specifically, based on the calculation formula for the pixel value of each pixel point in the first image, the first image can be obtained through the first sample image, the first sample pixel variance, and the first sample pixel mean; and based on the calculation formula for the pixel value of each pixel point in the second image, the second image can be obtained through the second sample image, the second sample pixel variance, and the second sample pixel mean.

[0085] S240. Determine the target sample data for training the multi-modal image processing model based on the first image, the second image of each first sample data, and the actual annotation identifier of each pixel point.

[0086] Among them, determine whether each pixel point in the first image or the second image is a lesion. If the pixel point is a lesion, the actual annotation identifier can be 1; if the pixel point is not a lesion, the actual annotation identifier can be 0. The target sample data refers to the first image, the second image of each first sample data, and the actual annotation identifier of each pixel point.

[0087] Specifically, for a certain first image and a second image, for the pixel points that are lesion regions in both the first image and the second image, the actual annotation identifier of the pixel points can be 1; for the pixel points that are not lesion regions in both the first image and the second image, the actual annotation identifier of the pixel points can be 0; for the pixel points at the lesion edge region or that are lesion regions in a single image, the pixel points can be annotated based on experience. The actual annotation identifier of each annotated pixel point, the first image, and the second image are used as the target sample data for training the multi-modal image processing model for subsequent model training.

[0088] S250. For the target sample data, the first image and the second image in the target sample data are segmented respectively according to the image extraction size and the image acquisition step length to obtain multiple groups of segmented images. Among them, each group of segmented images includes the first segmented image and the second segmented image corresponding to the same image region. The first segmented image is an image in the first modality, and the second segmented image is an image in the second modality.

[0089] Among them, after image acquisition of the first image in the first modality according to the image extraction size and the image acquisition step length, the obtained image is called the first segmented image; after image acquisition of the second image in the second modality according to the image extraction size and the image acquisition step length, the obtained image is called the second segmented image. The set of the first segmented image and the second segmented image is called a group of segmented images.

[0090] Specifically, for the first image in the first modality and the second image in the second modality in the target sample data, the first image and the second image are segmented respectively according to the preset image extraction size and the image acquisition step length. After segmenting the first image, the first segmented image is obtained; after segmenting the second image, the second segmented image is obtained, and the image regions of the first segmented image and the second segmented image correspond to each other. All the corresponding first segmented images and second segmented images form a group of segmented images.

[0091] S260. For multiple groups of segmented images, the first segmented image and the second segmented image in the group of segmented images are input into the multi-modal image processing model to be trained, and a fused sub-image that fuses the first segmented image and the second segmented image, and the prediction identifier corresponding to each pixel point in the fused sub-image are output.

[0092] Among them, after inputting the first segmented image and the second segmented image corresponding to the same image region into the multi-modal image processing model to be trained, the obtained image is the fused sub-image. It should be noted that the pixel value corresponding to each pixel point in the fused sub-image is the prediction identifier.

[0093] Specifically, for the first and second segmented images in multiple groups of segmented images, both are input into the multi-modal image processing model to be trained, and multiple fused sub-images and the prediction labels corresponding to each fused sub-image can be obtained.

[0094] S270. By processing the prediction labels of all pixel points, the labels to be applied for the pixel points are obtained.

[0095] Specifically, after obtaining multiple fused sub-images and the prediction labels corresponding to each fused sub-image, for the prediction labels of the same pixel points, the average value is calculated to obtain the labels to be applied for each pixel point.

[0096] S280. Based on the actual annotation labels of each pixel point in the first and second sub-images and the labels to be applied for all pixel points, the loss value is determined to correct the model parameters in the multi-modal image processing model to be trained based on the loss value.

[0097] It should be noted that the training parameters can be set to default values before training the multi-modal image processing model to be trained. When training the multi-modal image processing model to be trained, the training parameters in the model can be corrected based on the output results of the multi-modal image processing model to be trained. That is to say, the target multi-modal image processing model can be obtained by correcting the loss function in the multi-modal image processing model to be trained. There is a corresponding loss value for each image to be processed, and this loss value is determined based on the actual annotation labels of each processed image and the labels to be applied for all pixel points.

[0098] Specifically, after inputting the first and second segmented images in the segmented image group into the multi-modal image processing model to be trained, the multi-modal image processing model to be trained can obtain the prediction labels corresponding to the first and second segmented images in the segmented image group. By processing the prediction labels of all pixel points, the labels to be applied for the pixel points are obtained. Based on the actual annotation labels of each image to be processed and the labels to be applied for all pixel points, the loss value of the image to be processed can be determined, and the model parameters in the multi-modal image processing model to be trained can be corrected using the backpropagation method.

[0099] S290. Taking the convergence of the loss function in the multi-modal image processing model to be trained as the training objective, the multi-modal image processing model is obtained.

[0100] Specifically, the training error of the loss function, that is, the loss parameter, can be used as the condition for detecting whether the loss function has reached convergence. For example, whether the training error is less than a preset error, or whether the error change trend tends to be stable, or whether the current number of iterations is equal to the preset number. If it is detected that the convergence condition is reached, for example, the training error of the loss function reaches less than the preset error or the error change tends to be stable, it indicates that the multi-modal image processing model to be trained is trained, and at this time, the iterative training can be stopped. If it is detected that the current convergence condition is not reached, the first sample data can be further obtained to train the multi-modal image processing model until the training error of the loss function is within the preset range. When the training error of the loss function reaches convergence, the multi-modal image processing model to be trained can be used as the multi-modal image processing model.

[0101] The technical solution of the embodiment of the present disclosure is as follows: First, obtain the first sample data. Then, for multiple first sample data, based on the first annotation area of the first sample image and the second annotation area corresponding to the second sample image in the first sample data, determine the first sample pixel mean and the first sample pixel variance of the first annotation area, and the second sample pixel mean and the second sample pixel variance of the second annotation area. Then, based on the first sample image, the first sample pixel variance, and the first sample pixel mean in the first sample data, determine the first image of the first sample image; and, based on the second sample image, the second sample pixel variance, and the second sample pixel mean in the first sample data, determine the second image of the second sample image. Further, based on the first image, the second image of each first sample data, and the actual annotation identifier of each pixel point, determine the target sample data for training the multi-modal image processing model. Then, for the target sample data, segment the first image and the second image in the target sample data according to the image extraction size and the image acquisition step length respectively to obtain multiple segmented image groups. For multiple segmented image groups, input the first segmented image and the second segmented image in the segmented image group into the multi-modal image processing model to be trained, and output a fused sub-image that fuses the first segmented image and the second segmented image, and the prediction identifier corresponding to each pixel point in the fused sub-image. Further, by processing the prediction identifiers of all pixel points, obtain the to-be-applied identifier of the pixel points. Finally, based on the actual annotation identifier of each pixel point in the first sub-image and the second sub-image and the to-be-applied identifier of all pixel points, determine the loss value, so as to correct the model parameters in the multi-modal image processing model to be trained, and use the convergence of the loss function in the multi-modal image processing model to be trained as the training target to obtain the multi-modal image processing model.

[0102] It solves the problem of low image recognition and segmentation accuracy caused by using a convolutional neural network model alone for image segmentation or segmenting images under a single-modal image in the prior art, and obtains the effect of improving the extraction accuracy of the lesion area.

[0103] Embodiment III

[0104] Figure 5 It is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As shown in the figure, the device includes: a to-be-used image determination module 310, a to-be-applied image group determination module 320, a target sub-image determination module 330, a target evaluation identifier determination module 340, and a lesion image extraction module 350.

[0105] The to-be-used image determination module 310 is configured to obtain a first original image of a target part in a first modality and a second original image in a second modality, and perform normalization processing on the first original image to obtain a first to-be-used image, and perform normalization processing on the second original image to obtain a second to-be-used image; the to-be-applied image group determination module 320 is configured to sequentially perform image acquisition on the first to-be-used image and the second to-be-used image according to an image extraction size and an image acquisition step length to obtain a plurality of to-be-applied image groups, wherein each to-be-applied image group includes a first sub-image in the first modality and a second sub-image in the second modality; the target sub-image determination module 330 is configured to input the first sub-image and the second sub-image in each to-be-applied image group into a pre-trained multi-modal image processing model for the plurality of to-be-applied image groups to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image, wherein the to-be-used evaluation identifier is used to characterize the probability of the existence of a lesion in the target sub-image; the target evaluation identifier determination module 340 is configured to determine a target evaluation identifier corresponding to each pixel point in the first original image or the second original image based on the target sub-images corresponding to the plurality of to-be-applied image groups and the corresponding to-be-used evaluation identifiers; the lesion image extraction module 350 is configured to extract a lesion image corresponding to the lesion from the first original image or the second original image based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image; wherein, the lesion is located in the target part.

[0106] The technical solution of the embodiment of the present disclosure first obtains the first original image of the target part in the first modality and the second original image in the second modality, normalizes the first original image to obtain the first image to be used, normalizes the second original image to obtain the second image to be used, and then, according to the image extraction size and the image acquisition step size, sequentially performs image acquisition on the first image to be used and the second image to be used to obtain a plurality of image groups to be applied, wherein the image group to be applied includes the first sub-image in the first modality and the second sub-image in the second modality. Then, for the plurality of image groups to be applied, the first sub-image and the second sub-image in the image group to be applied are input into a pre-trained multi-modal image processing model to obtain the target sub-image and the to-be-used evaluation identifier corresponding to the target sub-image, wherein the to-be-used evaluation identifier is used to characterize the probability of the presence of a lesion in the target sub-image. Further, based on the target sub-images corresponding to the plurality of image groups to be applied and the corresponding to-be-used evaluation identifiers, the target evaluation identifier corresponding to each pixel point in the first original image or the second original image is determined. Finally, based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image, the lesion image corresponding to the lesion is extracted from the first original image or the second original image, which solves the problem of low image recognition and segmentation accuracy caused by using a convolutional neural network model alone for image segmentation or segmenting images in a single modality image in the prior art. In the embodiment of the present invention, according to the multi-modal image processing model, the lesion image corresponding to the lesion is extracted from two original images of different modalities, and the information of the two different modalities is combined to more accurately identify and extract the lesion area, achieving the effect of improving the extraction accuracy of the lesion area.

[0107] On the basis of the above technical solutions, the image-to-be-used determination module 310 includes: retrieving the first pixel mean and the first pixel variance corresponding to the target part in the first modality in advance; for all pixel points in the first original image, processing the pixel values of the pixel points based on the first pixel mean and the first pixel variance to obtain the first image to be used; retrieving the second pixel mean and the second pixel variance of the target part in the second modality stored in advance; for all pixel points in the second original image, processing the pixel values of the pixel points based on the second pixel mean and the second pixel variance to obtain the second image to be used.

[0108] On the basis of the above technical solutions, the target sub-image determination module 330 includes: inputting the first sub-image and the second sub-image in the image group to be applied into the multi-modal image processing model to obtain the target sub-image after fusing the first sub-image and the second sub-image, and the to-be-used evaluation identifier corresponding to each pixel point in the target sub-image.

[0109] Based on the above technical solutions, the target evaluation identifier determination module 340 includes determining at least one to-be-used evaluation identifier corresponding to each pixel point based on the target subgraphs corresponding to multiple to-be-applied image groups; and determining the target evaluation identifier corresponding to the pixel point in the first original image or the second original image by means of at least one to-be-used evaluation identifier of the same pixel point.

[0110] Based on the above technical solutions, the device further includes a first sample data acquisition module, a pixel mean determination module, an image determination module, a target sample data determination module, a segmented image group determination module, a fused subgraph output module, a to-be-applied identifier determination module, a loss value determination module, and a multi-modal image processing model determination module.

[0111] The first sample data acquisition module is configured to acquire a plurality of first sample data, where the first sample data includes at least one first sample image of a preset lesion site in the first modality and a second sample image in the second modality;

[0112] The pixel mean determination module is configured to, for a plurality of the first sample data, determine the first sample pixel mean and the first sample pixel variance of the first annotation region and the second sample pixel mean and the second sample pixel variance of the second annotation region based on the first annotation region of the first sample image and the second annotation region corresponding to the second sample image in the first sample data;

[0113] The image determination module is configured to determine the first image of the first sample image based on the first sample image, the first sample pixel variance, and the first sample pixel mean in the first sample data; and determine the second image of the second sample image based on the second sample image, the second sample pixel variance, and the second sample pixel mean in the first sample data;

[0114] The target sample data determination module is configured to determine the target sample data for training the multi-modal image processing model based on the first image, the second image of each first sample data, and the actual annotation identifier of each pixel point;

[0115] The segmented image group determination module is configured to segment the first image and the second image in the target sample data according to the image extraction size and the image acquisition step length respectively for the target sample data, to obtain a plurality of segmented image groups, where the segmented image groups include a first segmented image and a second segmented image corresponding to the same image region, the first segmented image is an image in the first modality, and the second segmented image is an image in the second modality;

[0116] The fused sub - graph output module is used to input the first segmented image and the second segmented image in the segmented image groups into the multi - modal image processing model to be trained for multiple segmented image groups, output a fused sub - graph that fuses the first segmented image and the second segmented image, and the predicted label corresponding to each pixel point in the fused sub - graph;

[0117] The to - be - applied label determination module is used to obtain the to - be - applied label of the pixel points by processing the predicted labels of all pixel points;

[0118] The loss value determination module is used to determine the loss value based on the actual annotation label of each pixel point in the first sub - graph and the second sub - graph and the to - be - applied labels of all pixel points, so as to correct the model parameters in the multi - modal image processing model to be trained based on the loss value;

[0119] The multi - modal image processing model determination module is used to take the convergence of the loss function in the multi - modal image processing model to be trained as the training goal to obtain the multi - modal image processing model.

[0120] On the basis of the above technical solutions, the first modality is the CT modality, and the second modality is the MR modality.

[0121] The image processing device provided by the embodiments of the present disclosure can execute the image processing method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0122] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0123] Embodiment 4

[0124] Figure 6 It is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. The following refers to Figure 6 , which shows a schematic structural diagram of an electronic device (such as Figure 6 the terminal device or server in) 500 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in - vehicle terminals (such as in - vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0125] As Figure 6 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An editing / output (I / O) interface 505 is also connected to the bus 504.

[0126] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wirelesly to exchange data. Although Figure 6 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.

[0127] Particularly, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.

[0128] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0129] The electronic device provided by the embodiments of the present disclosure and the base station frequency point determination method provided by the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0130] Example 5

[0131] An embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the image processing method provided in the above embodiment is implemented.

[0132] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0133] In some embodiments, the server may communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0134] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device.

[0135] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:

[0136] Obtain a first original image of the target part in a first modality and a second original image in a second modality, normalize the first original image to obtain a first image to be used, and normalize the second original image to obtain a second image to be used;

[0137] According to the image extraction size and the image acquisition step length, sequentially perform image acquisition on the first image to be used and the second image to be used to obtain a plurality of image groups to be applied, wherein each image group to be applied includes a first sub-image in the first modality and a second sub-image in the second modality;

[0138] For the plurality of image groups to be applied, input the first sub-image and the second sub-image in the image group to be applied into a pre-trained multi-modal image processing model to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image, wherein the to-be-used evaluation identifier is used to characterize the probability of the existence of a lesion in the target sub-image;

[0139] Based on the target sub-images corresponding to the plurality of image groups to be applied and the corresponding to-be-used evaluation identifiers, determine the target evaluation identifier corresponding to each pixel point in the first original image or the second original image;

[0140] Based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image, extract the lesion image corresponding to the lesion from the first original image or the second original image;

[0141] Wherein, the lesion is located in the target part.

[0142] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0144] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Wherein, the name of the unit does not constitute a limitation on the unit itself in some cases.

[0145] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0148] Furthermore, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0149] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An image processing method, characterized in that, Including: Obtain a first original image of a target part in a first modality and a second original image in a second modality, and perform normalization processing on the first original image to obtain a first image to be used, and perform normalization processing on the second original image to obtain a second image to be used; According to the image extraction size and the image acquisition step length, sequentially perform image acquisition on the first image to be used and the second image to be used to obtain a plurality of image groups to be applied, where the image group to be applied includes a first sub-image in the first modality and a second sub-image in the second modality; For the plurality of image groups to be applied, input the first sub-image and the second sub-image in the image group to be applied into a pre-trained multi-modal image processing model to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image, where the to-be-used evaluation identifier is used to characterize the probability of the presence of a lesion in the target sub-image; Based on the target sub-images corresponding to the plurality of image groups to be applied and the corresponding to-be-used evaluation identifiers, determine the target evaluation identifier corresponding to each pixel point in the first original image or the second original image; Based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image, extract the lesion image corresponding to the lesion from the first original image or the second original image; Wherein, the lesion is located in the target part.

2. The method according to claim 1, characterized in that The performing normalization processing on the first original image to obtain a first image to be used includes: Retrieve the first pixel mean value and the first pixel variance corresponding to the target part in the first modality determined in advance; For all pixel points in the first original image, process the pixel value of the pixel point based on the first pixel mean value and the first pixel variance to obtain the first image to be used; Correspondingly, the performing normalization processing on the second original image to obtain a second image to be used includes: Retrieve the second pixel mean value and the second pixel variance of the target part in the second modality stored in advance; For all pixel points in the second original image, process the pixel value of the pixel point based on the second pixel mean value and the second pixel variance to obtain the second image to be used.

3. The method according to claim 1, characterized in that, The inputting the first sub-image and the second sub-image in the image group to be applied into a pre-trained multi-modal image processing model to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image includes: Input the first sub-image and the second sub-image in the image group to be applied into the multi-modal image processing model to obtain a target sub-image after fusion processing of the first sub-image and the second sub-image, and a to-be-used evaluation identifier corresponding to each pixel point in the target sub-image.

4. The method according to claim 1, characterized in that, The determining the target evaluation identifier corresponding to each pixel point in the first original image or the second original image based on the target sub-images corresponding to the plurality of image groups to be applied and the corresponding to-be-used evaluation identifiers includes: Based on the target sub-images corresponding to a plurality of the image groups to be applied, determine at least one to-be-used evaluation identifier corresponding to each pixel point; Determine the target evaluation identifier corresponding to the pixel in the first original image or the second original image by using at least one evaluation identifier to be used for the same pixel point.

5. The method according to claim 1, wherein The method further includes: Obtain a plurality of first sample data, where the first sample data includes at least one first sample image of a preset lesion site in the first modality and a second sample image in the second modality; For the plurality of first sample data, based on the first annotation region of the first sample image and the second annotation region corresponding to the second sample image in the first sample data, determine the first sample pixel mean and the first sample pixel variance of the first annotation region, and the second sample pixel mean and the second sample pixel variance of the second annotation region; Based on the first sample image, the first sample pixel variance, and the first sample pixel mean in the first sample data, determine the first image of the first sample image; and, based on the second sample image, the second sample pixel variance, and the second sample pixel mean in the first sample data, determine the second image of the second sample image; Based on the first image, the second image of each first sample data, and the actual annotation identifier of each pixel point, determine the target sample data for training the multi-modal image processing model.

6. The method according to claim 5, wherein After obtaining a plurality of target sample data, the method further includes: For the target sample data, segment the first image and the second image in the target sample data according to the image extraction size and the image acquisition step length respectively, to obtain a plurality of segmented image groups, where the segmented image groups include a first segmented image and a second segmented image corresponding to the same image region, the first segmented image is an image in the first modality, and the second segmented image is an image in the second modality; For the plurality of segmented image groups, input the first segmented image and the second segmented image in the segmented image group into the multi-modal image processing model to be trained, and output a fused sub-image that fuses the first segmented image and the second segmented image, and a prediction identifier corresponding to each pixel point in the fused sub-image; Process the prediction identifiers of all pixel points to obtain the application identifier to be used for the pixel point; Based on the actual annotation identifier of each pixel point in the first sub-image and the second sub-image and the application identifier to be used for all pixel points, determine a loss value, and correct the model parameters in the multi-modal image processing model to be trained based on the loss value; Take the convergence of the loss function in the multi-modal image processing model to be trained as the training target to obtain the multi-modal image processing model.

7. The method according to claim 1, characterized in that The first modality is the CT modality, and the second modality is the MR modality.

8. An image processing apparatus, characterized in that, It includes: An image to be used determination module, configured to obtain a first original image of a target part in the first modality and a second original image in the second modality, perform normalization processing on the first original image to obtain a first image to be used, and perform normalization processing on the second original image to obtain a second image to be used; An image group to be applied determination module, configured to sequentially perform image acquisition on the first image to be used and the second image to be used according to an image extraction size and an image acquisition step length, to obtain a plurality of image groups to be applied, where each image group to be applied includes a first sub-image in the first modality and a second sub-image in the second modality; A target sub-image determination module, configured to, for the plurality of image groups to be applied, input the first sub-image and the second sub-image in the image group to be applied into a pre-trained multi-modal image processing model, to obtain a target sub-image and a to-be-used evaluation identifier corresponding to the target sub-image, where the to-be-used evaluation identifier is used to characterize the probability of the existence of a lesion in the target sub-image; A target evaluation identifier determination module, configured to determine a target evaluation identifier corresponding to each pixel point in the first original image or the second original image based on the target sub-images corresponding to the plurality of image groups to be applied and the corresponding to-be-used evaluation identifiers; A lesion image extraction module, configured to extract a lesion image corresponding to a lesion from the first original image or the second original image based on the target evaluation identifier corresponding to each pixel point in the first original image or the second original image; where the lesion is located in the target part.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to execute the image processing method according to any one of claims 1-7.