Focus segmentation model training method and device

Through the multi-sample training method, combined with the lesion segmentation model, distance information prediction model and loss value correction, the problems of high labeling cost and low label utilization in the existing technology are solved, and the segmentation accuracy of the lesion segmentation model in small sample scenarios is improved.

CN120451739APending Publication Date: 2025-08-08TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510540284.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing lesion segmentation model based on U-Net architecture or its variants has high labeling cost, low label utilization rate during training, and insufficient segmentation accuracy in small sample scenarios.

Method used

Using multiple training samples, including the first image set, the second image set and the third image set, the first model is used to predict lesion segmentation, the second model is used to predict distance information, and the third model is used to predict distance information, and the loss value is determined based on the image domain set, the prediction segmentation image, the label domain set, and the distance information, the model parameters are corrected, and unnecessary models are removed to obtain the target lesion segmentation model.

Benefits of technology

It reduces the annotation cost, improves the label utilization rate, and improves the generalization ability of the model in small sample scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451739A_ABST
    Figure CN120451739A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a focus segmentation model training method and apparatus. The method comprises the steps of obtaining a plurality of training samples; inputting the training sample into a pre-constructed focus segmentation model, and outputting an image domain set, a prediction segmentation image and a label domain set; determining three loss values based on the image domain set, the prediction segmentation image, the label domain set, the first distance information, the second distance information and the lesion segmentation image, and correcting model parameters in the lesion segmentation model based on the loss values; and taking the focus segmentation model obtained when the iteration condition is satisfied as a to-be-used segmentation model, and removing a second model and a third model in the to-be-used segmentation model to obtain a target focus segmentation model for processing the input image. According to the technical scheme, the focus segmentation model is trained according to the first image set, the second image set and the third image set, the target focus segmentation model is obtained, and the generalization ability of the model in a small sample scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and more particularly to a method and apparatus for training a lesion segmentation model. Background Art

[0002] With the advancement of science and technology, the field of image processing has gradually developed the use of constructed segmentation models to separate regions of interest from the background. For example, in medical image analysis, after pre-processing images acquired through CT or MR imaging, they can be input into a pre-built segmentation model to train the trained segmentation model. The trained segmentation model can then accurately identify and separate pathological areas in the medical image, such as tumors and lesions.

[0003] Currently, pre-built segmentation models are primarily based solely on the U-Net architecture, employing a symmetrical encoder-decoder structure connected by numerous skip connections to fuse information from different layers. Alternatively, segmentation models are constructed based on variations and extensions of the U-Net architecture. For example, integrating the concept of residual networks, residual connections are introduced into the convolutional layers of the U-Net architecture to help address the vanishing gradient problem in deep networks. However, training segmentation models based solely on the U-Net architecture or its variants requires a large number of high-quality annotated training samples. Furthermore, in medical images obtained through tomography, the image variability between layers is low. Therefore, training segmentation models using the U-Net architecture or its variants is subject to high annotation costs and low label utilization. Furthermore, applying the trained segmentation models to segmentation results in low segmentation accuracy. Summary of the Invention

[0004] The embodiments of the present disclosure provide a method and device for training a lesion segmentation model, so as to achieve the effects of reducing labeling costs, improving label utilization, and enhancing the generalization ability of the model in small sample scenarios when training the lesion segmentation model.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for training a lesion segmentation model, the method comprising:

[0006] Acquiring a plurality of training samples, wherein the training samples are composed of a plurality of first 2D images extracted from the same 3D image, the training samples including a first image set, a second image set, and a third image set, the first image set including a second 2D image and a lesion segmentation image corresponding to the second 2D image, the second image set including a plurality of third 2D images including the second 2D image, and first distance information between a level of the plurality of third 2D images in the 3D image and a level of the second 2D image, the third image set including a plurality of first segmented images corresponding to the plurality of first 2D images including the lesion segmentation image, and second distance information between a level of the plurality of first segmented images in the 3D image and a level of the lesion segmentation image;

[0007] For the multiple training samples, the training samples are input into a pre-built lesion segmentation model, and an image domain set, a predicted segmentation image, and a label domain set are output, wherein the lesion segmentation model includes a first model for performing lesion segmentation prediction on the first image set, a second model for performing distance information prediction on the second image set, and a third model for performing distance information prediction on the third image set, the first model includes multiple data processing layers, and the output result of at least one data processing layer is used to correct model parameters in the second model and the third model, the image domain set includes a first predicted distance from the third 2D image to the second 2D image predicted by the second model, and the label domain set includes a second predicted distance from the multiple second segmented images to the second 2D image predicted by the third model;

[0008] Determining, based on the image domain set, the predicted segmented image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model, and correcting model parameters in the lesion segmentation model based on the loss values; wherein the first loss value is determined based on at least the second loss value and the third loss value;

[0009] The lesion segmentation model obtained when the iteration condition is satisfied is used as the segmentation model to be used, and the second model and the third model in the segmentation model to be used are removed to obtain a target lesion segmentation model for processing the input image.

[0010] In a second aspect, an embodiment of the present invention further provides a training device for a lesion segmentation model, the device comprising:

[0011] a plurality of training sample acquisition modules, configured to acquire a plurality of training samples, wherein the training samples are composed of a plurality of first 2D images extracted from the same 3D image, the training samples comprising a first image set, a second image set, and a third image set, the first image set comprising a second 2D image and a lesion segmentation image corresponding to the second 2D image, the second image set comprising a plurality of third 2D images including the second 2D image, and first distance information between the plurality of third 2D images at a level in the 3D image and the second 2D image, the third image set comprising a plurality of first segmented images corresponding to the plurality of first 2D images including the lesion segmentation image, and second distance information between the plurality of first segmented images at a level in the 3D image and the lesion segmentation image;

[0012] a predicted segmented image output module, configured to input the plurality of training samples into a pre-built lesion segmentation model, and output an image domain set, a predicted segmented image, and a label domain set, wherein the lesion segmentation model includes a first model for predicting lesion segmentation for the first image set, a second model for predicting distance information for the second image set, and a third model for predicting distance information for the third image set, the first model includes a plurality of data processing layers, and an output result of at least one data processing layer is used to correct model parameters in the second model and the third model, the image domain set includes a first predicted distance from the third 2D image to the second 2D image predicted by the second model, and the label domain set includes a second predicted distance from the plurality of second segmented images to the second 2D image predicted by the third model;

[0013] a model parameter correction module, configured to determine, based on the image domain set, the predicted segmented image, the label domain set, the first distance information, the second distance information, and the second 2D image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model, and to correct model parameters in the lesion segmentation model based on the loss values; wherein the first loss value is determined based on at least the second loss value and the third loss value;

[0014] The target lesion segmentation model determination module is used to use the lesion segmentation model obtained when the iteration condition is met as the segmentation model to be used, and remove the second model and the third model in the segmentation model to be used to obtain the target lesion segmentation model for processing the input image.

[0015] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising:

[0016] one or more processors;

[0017] a storage device for storing one or more programs,

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the lesion segmentation model as described in any one of the embodiments of the present invention.

[0019] In a fourth aspect, an embodiment of the present invention further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the training method for a lesion segmentation model as described in any one of the embodiments of the present invention.

[0020] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, the computer program implements the training method of the lesion segmentation model as described in any one of the embodiments of the present invention.

[0021] The technical solution of the embodiment of the present disclosure is as follows: first, a plurality of training samples are obtained. Then, for the plurality of training samples, the training samples are input into a pre-constructed lesion segmentation model, and an image domain set, a predicted segmentation image, and a label domain set are output. Furthermore, based on the image domain set, the predicted segmentation image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model are determined, and the model parameters in the lesion segmentation model are corrected based on the loss value. Finally, the lesion segmentation model obtained when the iteration condition is met is used as the segmentation model to be used, and the second model and the third model in the segmentation model to be used are removed to obtain a target lesion segmentation model for processing the input image. This solves the problem in the prior art that after constructing a lesion segmentation model based only on the U-Net architecture or a variant of the U-Net architecture, there is a high labeling cost and a low label utilization rate when training based on the constructed lesion segmentation model. When training based on a lesion segmentation model, the embodiment of the present invention includes not only a first model constructed based on a variant of the U-Net architecture, but also a second model for predicting distance information for a second image set, and a third model for predicting distance information for a third image set. The training method based on the first, second, and third models can achieve the effects of reducing labeling costs, increasing label utilization, and improving the model's generalization ability in small sample scenarios when training the lesion segmentation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings introduced here are only drawings of some embodiments to be described in the present invention, and are not all drawings. A person of ordinary skill in the art can also derive other drawings based on these drawings without inventive effort.

[0023] Figure 1 is a flowchart of a training method for a lesion segmentation model provided by an embodiment of the present disclosure;

[0024] Figure 2 1 is a flow chart of another method for training a lesion segmentation model provided by an embodiment of the present invention;

[0025] Figure 3 is a flow chart of the lesion segmentation model provided in an embodiment of the present invention;

[0026] Figure 4 1 is a schematic structural diagram of a training device for a lesion segmentation model provided by an embodiment of the present invention;

[0027] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0029] Before introducing the technical solutions provided by the embodiments of the present disclosure, an example description of the application scenarios can be given. The technical solutions provided by the embodiments of the present disclosure can be applied to the scenario where, when performing image segmentation tasks, the model is trained based on the constructed lesion segmentation model and multiple training samples to obtain a target lesion segmentation model that can process the input image. For example, for training samples obtained based on computer tomography (CT) images, the training samples can be input into a pre-constructed lesion segmentation model based on the technical solutions of the embodiments of the present disclosure to train a target lesion segmentation model for completing subsequent image segmentation tasks. It should be noted that the boundary area of the prostate is usually closely adjacent to the surrounding tissues (such as the bladder and rectum) and presents fuzzy features in 3D images, which brings difficulties to the lesion segmentation task in the prostate image. Obtaining multiple first segmented images corresponding to multiple first 2D images requires rich professional knowledge and high labeling costs, which may result in a limited number of first segmented images. There are significant differences in the morphology of the prostate area of different people, which increases the difficulty of training the model for segmenting the prostate image. Therefore, the preset segmentation area in the embodiment of the present invention is the prostate area. Based on the technical solutions of the disclosed embodiments, the lesion segmentation model includes not only a first model constructed based on a variant of the U-Net architecture, but also a second model for predicting distance information for a second image set, and a third model for predicting distance information for a third image set. The training method based on the first, second, and third models can reduce labeling costs, increase label utilization, and enhance the model's generalization capabilities in small sample size scenarios when training the lesion segmentation model.

[0030] Example 1

[0031] Figure 1 It is a flow chart of a training method for a lesion segmentation model provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation where model training is performed based on a constructed lesion segmentation model and multiple training samples to obtain a target lesion segmentation model that can process the input image. The method can be executed by a training device for the lesion segmentation model, which can be implemented in the form of software and / or hardware. The hardware can be a mobile electronic device, which can execute the training method for the lesion segmentation model provided by this technical solution.

[0032] like Figure 1 As shown, the method includes:

[0033] S110: Obtain multiple training samples.

[0034] The training samples are composed of multiple first 2D images extracted from the same 3D image, and the training samples include a first image set, a second image set, and a third image set. The first image set includes a second 2D image and a lesion segmentation image corresponding to the second 2D image. The second image set includes multiple third 2D images, and the multiple third 2D images include the second 2D image, as well as first distance information between the levels of the multiple third 2D images in the 3D image and the levels of the second 2D images. The third image set includes multiple first segmented images corresponding to the multiple first 2D images, and the first segmented images include the lesion segmentation image, as well as second distance information between the levels of the multiple first segmented images in the 3D image and the lesion segmentation image.

[0035] 3D images refer to images of a target area captured using 3D imaging technology. A 3D image is a three-dimensional visual representation of the human body's internal structures and organs. 3D images can present images of the body's internal structures and organs in three spatial dimensions (length, width, and height). 3D images provide richer and more three-dimensional visual information, demonstrating the three-dimensional relationships of target areas, preventing information loss, and increasing the image's usability. 3D images can be acquired using a variety of medical imaging technologies, each of which uses different physical principles to capture detailed information about the human body. 3D images can be based on CT scans, which use X-rays to scan at different angles and then use computer processing to generate volumetric data, which is then reconstructed into a 3D image. For example, if the target area is the prostate, a CT image of the prostate can provide a detailed display of the prostate and its surrounding structures. 3D images can also be based on MRI scans, which use strong magnetic fields and radiofrequency pulses to obtain high-resolution images of internal soft tissues and reconstruct them into 3D views. When the target site is the prostate, MRI images of the prostate can generate high-resolution 3D images of the prostate. 3D images can also be based on 3D ultrasound images, which use multiple ultrasound probes to collect data from different angles and combine computer processing to produce a stereoscopic image. When the target site is the prostate, 3D ultrasound images of the prostate can capture three-dimensional data of the prostate, providing more comprehensive anatomical information. 3D images can also be based on positron emission tomography (PET), which uses radioactive tracer distribution and imaging technology to generate functional 3D images. When the target site is the prostate, PET images of the prostate can provide detailed information about the metabolic activity of prostate lesions. It should be noted that the resulting 3D images can reveal the size, location, and morphology of lesions such as tumors, cysts, or inflammation within the target site. For example, a 3D image can be a CT-based image of a prostate containing a tumor.

[0036] It should be noted that 2D images are the basic building blocks of 3D images. By acquiring and combining multiple 2D images with spatial positioning, computer processing and reconstruction techniques can be used to combine 3D images with a sense of three-dimensionality. A 2D image refers to an image displayed on a plane, which contains only two spatial dimensions (length and width). In the field of medical imaging, 2D images are usually presented as cross-sectional, coronal or sagittal slice images of the internal structure of the human body. The continuity and consistency between slices are ensured by clear spatial positioning parameters (such as slice thickness, interval, position coordinates). These 2D images provide detailed views of tissues, organs and lesions at specific levels and are the basic units for constructing 3D images. It should be noted that there is a close layering and combination relationship between 3D images and 2D images. The construction of a 3D image can be based on a series of continuous 2D images acquired from different positions. These 2D images usually have a certain thickness and interval, covering the entire or partial area of a human body part. 3D images can be generated by combining and superimposing multiple 2D images according to their spatial position information through computer algorithms (such as reconstruction algorithms and volume rendering technology) to generate a 3D image with a three-dimensional sense. 2D images and 3D images can achieve complementary perspectives. 2D images provide high-resolution local detail information, which helps to accurately identify and analyze the structure of a specific area. 3D images integrate multiple 2D images to display the spatial relationship and overall morphology of the entire structure, which facilitates comprehensive understanding and evaluation. In an embodiment of the present invention, for the same 3D image, its basic constituent unit, that is, each layer of 2D image, can be referred to as a first 2D image. For example, assuming that the number of layers of the CT image obtained by CT imaging of the prostate area is 320, then each of the 320 layers can be referred to as a first 2D image, and there are 320 first 2D images that can be extracted from the 3D image.

[0037] Among them, the second 2D image refers to the first 2D image of a certain layer. For example, assuming that the number of layers of the CT image obtained by taking a CT image of the prostate area is 320, at this time, assuming that the first 2D image of the 30th layer is extracted, then the first 2D image of the 30th layer is the second 2D image. The lesion segmentation image corresponding to the second 2D image is an image that converts the second 2D image into a binary label (0 / 1). For the second 2D image, first, determine the area that needs to be segmented (such as a tumor or blood vessel), and then, you can use manual annotation or automatic threshold segmentation tools to annotate the second 2D image, save the area that needs to be segmented as 1, and set the remaining background areas to 0, and finally generate a lesion segmentation image corresponding to the second 2D image. After obtaining the first image set, the first image set can be used as a training sample for training the first model.

[0038] It should be noted that the third 2D image refers to the first 2D image randomly selected from the set of first 2D images. It should be noted that all the third 2D images in the second image set include the second 2D image. The purpose of this is to ensure that the network corresponding to the second image set can help the network corresponding to the first image set to improve the learning effect and accuracy of the network. Among them, for each third 2D image, the distance between the level of each third 2D image in the 3D image and the level of the second 2D image in the 3D image is calculated, and the set of distances is called the first distance information. Optionally, assuming that the level of the second 2D image in the 3D image is the Pth layer, and assuming that the level of the third 2D image in the 3D image is the Kth layer, then the first distance calculation formula d=exp(-(KP)) can be used. 2 ) calculates the distance between the level of each third 2D image in the 3D image and the level of the second 2D image in the 3D image. Because for 3D images, closer first 2D images have more similar features and higher image similarity, each value in the first distance information can indirectly represent the similarity between each third 2D image and the second 2D image. After obtaining the second image set, the second image set can be used as training samples for training the second model.

[0039] It should be noted that the first segmented image refers to the lesion segmentation image corresponding to the first 2D image. Each first 2D image can be converted into an image with a binary label (0 / 1), and the converted image can be called a first segmented image. The number of first segmented images corresponds to the first 2D image, and each first 2D image has a first segmented image corresponding to it. It should be noted that the first segmented image includes the lesion segmentation image. The purpose of doing so is to ensure that the network corresponding to the third image set can help the network corresponding to the first image set to improve the learning effect and accuracy of the network. Among them, for each first segmented image, the distance between the level of each first segmented image in the 3D image and the level of the lesion segmentation image in the 3D image is calculated, and the set of distances is called the second distance information. Optionally, assuming that the level of the lesion segmentation image in the 3D image is the Mth layer, and assuming that the level of the first segmented image in the 3D image is the Lth layer, then the second distance calculation formula d=exp(-(LM)) can be used. 2 ) Calculate the distance between the level of each first segmented image in the 3D image and the level of the lesion segmented image in the 3D image. Because for 3D images, first segmented images with closer distances have more similar features and higher image similarity, each value in the second distance information can indirectly represent the similarity between each first segmented image and the lesion segmented image. After obtaining the third image set, the third image set can be used as training samples for training the third model.

[0040] In this embodiment, before obtaining multiple training samples, the method also includes: obtaining at least one sample pair, wherein the sample pair includes a 3D sample image corresponding to a preset part and a 3D segmentation image of the preset part corresponding to the 3D sample image; for at least one sample pair, randomly selecting a first position from the Z axis of the 3D sample image and the 3D segmentation image of the sample pair to obtain a first 2D image corresponding to the first position and a lesion segmentation image to be used; and determining multiple training samples based on the first 2D image and the corresponding lesion segmentation image to be used.

[0041] Optionally, the preset location is the prostate. The 3D sample image is obtained using scanning technologies such as CT or MRI, and displays the detailed three-dimensional structure of the prostate. The 3D segmented image of the preset location is a 3D segmented image corresponding to the 3D sample image. The 3D segmented image of the preset location refers to an image in which the target region in the 3D sample image is annotated. Optionally, the target region may be tumor tissue in the prostate. The 3D segmented image of the preset location is typically obtained by manual annotation by an expert, or by automatic segmentation of the 3D sample image using an algorithm. Each pixel in the 3D segmented image of the preset location is labeled as a lesion region (labeled as 1) or a non-lesion region (labeled as 0). It should be noted that the 3D segmented image of the preset location differs only in pixel values; the rest of the image dimensions and other aspects are identical to the 3D sample image. When training a model based on machine learning, in order to prevent overfitting, improve the model's generalization ability, and enhance transfer learning capabilities, multiple 3D sample images corresponding to the preset location and 3D segmented images of the preset location corresponding to the 3D sample image are generally obtained, i.e., at least one sample pair is obtained.

[0042] It should also be noted that after obtaining the 3D sample image corresponding to the preset part and the 3D segmentation image of the preset part corresponding to the 3D sample image, the same data augmentation processing can be performed on the 3D sample image and the 3D segmentation image. The purpose is to expand the training samples and improve the generalization ability of the model by performing the same transformation on the 3D sample image and the 3D segmentation image corresponding to the preset part. Optionally, the 3D sample image and the 3D segmentation image can be rotated, and new samples can be generated by rotating the image around the center point of the image. Rotation is usually a fixed-angle operation, and the rotation angle can be positive or negative, usually common angles such as 90°, 180°, 270°, or any angle; the 3D sample image and the 3D segmentation image can be translated to displace the entire image horizontally or vertically. Through the translation transformation, the position of the image will change, but the content of the image will not be changed. Optionally, you can first select the number of pixels to translate, for example, by randomly selecting the translation amount in the x- and y-axis directions, and then use the translation function to translate the 3D sample image and 3D segmentation image. You can also perform elastic transformations on the 3D sample image and 3D segmentation image, creating a more natural effect by nonlinearly deforming the image. This is achieved by adding a certain degree of random displacement to each pixel in the image. Elastic transformations are very effective for enhancing model robustness. Optionally, you can create a random field (usually Gaussian noise) and apply it to each pixel in the image. Then, you can adjust the position of each pixel in the image using the displacement field to perform elastic transformations on the image. You can also add noise to the 3D sample image and 3D segmentation image, adding noise to the image to increase the robustness of the model and prevent overfitting when processing real-world data. Common noise types include Gaussian noise and salt and pepper noise. Optionally, you can add random noise that follows a Gaussian distribution to the pixel values of the image, or randomly select some pixels in the image and set them to extreme black or white values. The above data augmentation methods can help improve the generalization ability of the model and avoid overfitting, especially when there are insufficient training samples.

[0043] For the 3D sample image and 3D segmentation image of a sample pair, the image is typically composed of three dimensions: the X-axis is the horizontal axis, representing the left-right direction of the image; the Y-axis is the vertical axis, representing the up-down direction of the image; and the Z-axis is the interslice direction, representing the spatial depth of the image, typically corresponding to the axial direction of the human body from head to toe (or back to abdomen). In CT or MR images, the Z-axis corresponds to continuous cross-sectional slices. For example, the Z-axis of a prostate CT scan can extend from the fifth lumbar vertebra or sacral promontory to 1-2 cm below the ischial tuberosity, ensuring coverage of the bladder base, the entire prostate, and adjacent pelvic floor structures. The Z-axis of a prostate CT scan is approximately 10-15 cm in total length and can contain 20-150 slices of 2D images. For example, a 3D sample image of a prostate CT scan may be 512*512*150, representing 150 slices of 2D images along the Z-axis of the 3D sample image and 3D segmentation image. The first position refers to a position point randomly determined within the range of the Z axis, and the 2D image corresponding to the first position is the first 2D image. For example, the first position can be the 80th layer of the 3D sample image of a prostate CT. In this case, the first 2D image is the 2D image corresponding to the 80th layer of the 3D sample image of the CT. Because the Z axis of the 3D sample image and the 3D segmentation image are consistent, after determining the first 2D image corresponding to the first position, the image corresponding to the first position in the 3D segmentation image of the preset part is extracted to obtain the lesion segmentation image to be used. Moreover, different first 2D images and lesion segmentation images to be used can be determined based on different first positions, that is, multiple training samples can be determined. The main purpose of determining multiple training samples is to help train the deep learning model and improve the training effect of the model.

[0044] It should be noted that all first 2D images can be normalized using window width and window position. The purpose of normalization is to map the grayscale values of the first 2D images to a certain range. Optionally, the window width and window position can be used to keep the pixel values in the first 2D images between 0 and 1. The window position determines the brightness of the image and defines the grayscale midpoint of a selected range in the image, which can be used to control the overall brightness of the image. The window width defines the grayscale range displayed in the image and controls the image contrast. A larger window width results in a wider grayscale range and more detail; a smaller window width results in higher contrast and less detail. Specifically, when processing the first 2D image using window width and window position, first, an appropriate window width and window position are selected. The selection of window width and window position generally depends on the characteristics of the medical image and the area to be observed. Then, based on the window width and window position, the minimum and maximum grayscale values of the image are calculated, and the pixel values are linearly mapped to the range of 0 to 1. Finally, this calculation is performed for each pixel in the image to obtain a normalized image. The contrast and brightness of the image can be adjusted by adjusting the window width and window level. By adjusting these parameters, the visibility of specific areas can be enhanced, making the key areas of the medical image (such as lesions) clearer. After obtaining the normalized first 2D image, the normalized first 2D image and the lesion segmentation image to be used can be resized to ensure that the image size remains consistent when input into the model. For example, we can use the resizing technique to resize the normalized first 2D image and the lesion segmentation image to be used to a fixed size C*W*D. Where C is the number of channels of the image, W is the number of columns of the image, and D is the number of rows of the image. Specifically, after the image is loaded, the image can be resized using interpolation methods such as nearest neighbor interpolation and bilinear interpolation to resize each image to C*W*D. Finally, the image is resized. For color images, the number of channels is usually kept at 3, and for grayscale images, the number of channels is kept at 1. In addition, the width and height of the image can be achieved by adjusting the pixel size of the image. In some cases, you may want to maintain the aspect ratio of an image. This can be done by scaling it while resizing, but preserving its original proportions. Optionally, you can add padding to maintain consistent image dimensions. Padding areas can be filled with the background color or black. This resizing allows the image to be handled uniformly within the model.

[0045] Optionally, after obtaining the first 2D image and the corresponding lesion segmentation image to be used of at least one sample pair, the method further includes: randomly selecting a second 2D image from multiple first 2D images, and calling the lesion segmentation image to be used corresponding to the second 2D image as the lesion segmentation image, so as to obtain a first image set based on the second 2D image and the lesion segmentation image; randomly selecting a preset number of third 2D images from multiple first 2D images, and constructing a second image set based on the third 2D images and the second 2D image in the first image set; randomly selecting a preset number of first segmentation images from multiple lesion segmentation images to be used, and determining a third image set based on the first segmentation images and the lesion segmentation images in the first image set.

[0046] The value of the preset number can be determined based on factors such as the size of the dataset, the capacity and complexity of the network architecture, the stability of training, and the time constraints of computing resources. For example, the preset number can be 10. In this case, the number of 2D images in the second image set is 11, and the number of annotated segmented images in the third image set is 11.

[0047] In this embodiment, first distance information between the third 2D image in the second image set and the second 2D image in the Z-axis direction of the 3D image is obtained, and the first distance information is arranged according to the arrangement information of the third 2D image in the second image set to obtain a first distance sequence, so as to update the second image set based on the first distance sequence; second distance information between the first segmented image in the third image set and the lesion segmented image in the Z-axis direction of the 3D image is obtained, and the second distance information is arranged according to the arrangement information of the first segmented image in the third image set to obtain a second distance sequence, so as to update the third image set based on the second distance sequence.

[0048] The arrangement information of the third 2D images in the second image set refers to a method that describes the order and position of these randomly selected third 2D images relative to the first 2D image in space. The arrangement information may include the layer number, order, or alignment of the 2D images in a certain direction. The first distance sequence refers to the reordering of these distance information according to the arrangement information of the third 2D images (i.e., their order in the 3D image) after calculating the first distance information to form an ordered distance sequence. The purpose of the arrangement information is to rearrange these first distance information according to the order of the third 2D images in space. For example, assuming that the arrangement information of the randomly selected third 2D images is the 5th, 10th, 15th, 45th, 50th, 53rd, 67th, 73rd, 74th, and 75th layers, the calculated first distance information is 0.69, 0.70, 0.76, 0.86, 0.98, 0.97, 0.89, 0.80, 0.75, and 0.70. In this case, if the layer numbers of the third 2D image are arranged in ascending order, the calculated first distance information should be arranged in the same order, thereby obtaining a first distance sequence. The arrangement information of the first segmented images in the third image set refers to the method that describes the spatial order and position of these randomly selected third 2D images relative to the lesion segmentation image to be used. The arrangement information may include the layer number, order, or alignment of the lesion segmentation image to be used in a certain direction. The second distance sequence refers to the reordering of these distance information based on the arrangement information of the first segmented images (i.e., their order in the 3D segmentation image of the preset part) after the second distance information is calculated, to form an ordered distance sequence. The purpose of the arrangement information is to rearrange the second distance information according to the spatial order of the first segmented images. If the layer numbers of the first segmented images are arranged in ascending order, the calculated second distance information should be arranged in the same order, thereby obtaining a second distance sequence.

[0049] Specifically, multiple sample pairs are obtained, each comprising a 3D sample image corresponding to a preset location and a 3D segmented image of the preset location corresponding to the 3D sample image. For each sample, a first position is randomly selected from the Z-axis of the 3D sample image and the 3D segmented image of the sample pair, and the first 2D image corresponding to the first position is determined as the second 2D image. For the 3D segmented image of the preset location, a lesion segmented image to be used at the same first position on the Z-axis is determined as the lesion segmented image. The obtained second 2D image and the lesion segmented image corresponding to the second 2D image are used as a first image set. For all first 2D images in the same 3D image, a preset number of first 2D images are randomly selected as third 2D images. After first distance information is obtained using a first distance calculation formula, the first distance information is arranged according to the arrangement information of the third 2D images to obtain a first distance sequence. The set of the third 2D image, the second 2D image in the first image set, and the first distance sequence is used as a second image set. For the 3D segmented image of the preset location corresponding to the same 3D sample image, a preset number of lesion segmented images to be used are randomly selected. After obtaining the second distance information using the second distance calculation formula, the second distance information is arranged according to the arrangement information of the lesion segmentation images to be used to obtain a second distance sequence. A set of the selected preset number of lesion segmentation images to be used, the lesion segmentation images, and the second distance sequence is used as a third image set.

[0050] S120 . For a plurality of training samples, input the training samples into a pre-built lesion segmentation model, and output an image domain set, a predicted segmentation image, and a label domain set.

[0051] Among them, the lesion segmentation model is a deep learning model used to identify and separate lesion areas (such as tumor areas) in medical images (such as images of the prostate area taken by CT). The lesion segmentation model is constructed to automatically detect and segment possible lesion areas from the input medical image. It should be noted that the lesion segmentation model at this time is a pre-built lesion segmentation model to be trained. In addition, the lesion segmentation model includes three models, namely the first model, the second model, and the third model.

[0052] Among them, the lesion segmentation model includes a first model for predicting lesion segmentation for the first image set, a second model for predicting distance information for the second image set, and a third model for predicting distance information for the third image set. The first model includes multiple data processing layers, and the output results of at least one data processing layer are used to correct the model parameters in the second model and the third model. The image domain set includes the first predicted distance from the third 2D image to the second 2D image predicted by the second model, and the label domain set includes the second predicted distances from multiple second segmented images to the second 2D image predicted by the third model.

[0053] It should be noted that the first model, as part of the lesion segmentation model, is a model that predicts lesion segmentation for the first image set. The input of the first model is the first image set, that is, the second 2D image and the lesion segmentation image to be used corresponding to the second 2D image. The output of the first model is a predicted segmentation image, which shows the area that the first model considers to be a lesion. The predicted segmentation image refers to the segmentation result image generated by the first model after the input second 2D image is processed by the first model. In the context of medical image processing, the predicted segmentation image is generally used to display the first model's prediction of the lesion area. Each pixel in the predicted segmentation image is marked as a lesion area (represented by 1) or a non-lesion area (represented by 0).

[0054] Among them, the second model, as part of the lesion segmentation model, is a model that predicts distance information for the second image set. The input of the second model is the second image set, that is, a preset number of selected third 2D images, the second 2D images in the first image set, and the first distance sequence. The output of the second model is an image domain set, which shows the set of distances between the levels of all third 2D images in the 3D image and the levels of the second 2D images as considered by the second model. The image domain set refers to the first predicted distance generated by the second model after the second model processes the input second image set.

[0055] Among them, the third model, as part of the lesion segmentation model, is a model for predicting distance information for the third image set. The input of the third model is the third image set, that is, a preset number of lesion segmentation images to be used, lesion segmentation images, and a second distance sequence. The output of the third model is a label domain set, which shows the set of distances between the level of all lesion segmentation images to be used in the 3D segmentation image and the level of the lesion segmentation image as considered by the third model. The label domain set refers to the second predicted distance generated by the third model after the third model processes the input third image set.

[0056] It should be noted that the multiple data processing layers in the first model refer to the layers in the first model used to process input data, extract features, or transform data representations. These layers perform different levels of processing and transformation on the raw data, i.e., the first dataset, to ultimately output the predicted segmented image.

[0057] Specifically, after obtaining multiple training samples, the first image set in the training samples is input into the first model of the pre-built lesion segmentation model, the second image set in the training samples is input into the second model of the pre-built lesion segmentation model, and the third image set in the training samples is input into the third model of the pre-built lesion segmentation model. The three models in the pre-built lesion segmentation model are trained, the first model outputs a predicted segmentation image for lesion segmentation prediction of the first image set, the second model outputs an image domain set for distance information prediction of the second image set, and the third model outputs a label domain set for distance information prediction of the third image set. The three models in the pre-built lesion segmentation model do not exist completely independently, and the output results of the multiple data processing layers in the first model will modify the model parameters in the second model and the third model.

[0058] S130. Based on the image domain set, the predicted segmentation image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, determine a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model, and correct the model parameters in the lesion segmentation model based on the loss values.

[0059] The first loss value is determined based on at least the second loss value and the third loss value.

[0060] In this embodiment, the second loss value is determined based on the first predicted distance in the image domain set and the first distance sequence in the second image set; the third loss value is determined based on the second predicted distance in the label domain set and the second distance sequence in the third image set; the image loss value is determined based on the lesion segmentation image and the predicted segmentation image, and the first loss value is determined based on the image loss value, the second loss value and the third loss value; the parameters of the first model are corrected based on the first loss value, the parameters of the second model are corrected based on the second loss value, and the parameters of the third model are corrected based on the third loss value.

[0061] It should be noted that the second loss value is a numerical value used during the second model training process to measure the difference between the first predicted distance output by the second model and the actual first distance information. The second loss value reflects the error of the second model in the current training step. The second loss value is used to guide the optimization process of the second model, with the goal of minimizing this value to improve the performance of the second model. To achieve this goal, a loss function is typically used to quantify this error, and the second loss value is used as the optimization target to guide the learning process of the second model. Optionally, for the distance prediction task corresponding to the second model, a mean squared error function can be used as the loss function to measure the difference between the first predicted distance and the actual first distance sequence. The mean squared error function is used to calculate the difference between the first predicted distance in the image domain set and the actual first distance sequence, thereby obtaining the second loss value for all images. The smaller the second loss value, the smaller the difference between the first predicted distance and the actual first distance sequence in the image domain set, and the better the prediction performance of the second model. It should also be noted that after the second loss value is calculated, the backpropagation algorithm can be used to calculate the gradient of the second loss value relative to the model parameters of the second model. Then, the parameters of the second model can be updated through the optimization algorithm to reduce the second loss value and improve the performance of the second model, so that the second model can more accurately predict the first predicted distance from the third 2D image to the second 2D image in future training.

[0062] It should be noted that the third loss value is a numerical value used during the training process of the third model to measure the difference between the second predicted distance output by the third model and the true second distance information. The third loss value reflects the error of the third model in the current training step. The third loss value is used to guide the optimization process of the third model, with the goal of minimizing this value to improve the performance of the third model. To achieve this goal, a loss function is typically used to quantify this error, and the third loss value is used as the optimization target to guide the learning process of the third model. Optionally, for the distance prediction task corresponding to the third model, a mean squared error function can be used as the loss function to measure the difference between the second predicted distance and the true second distance sequence. Using the mean squared error function, the difference between the second predicted distance in the label domain set and the true second distance sequence is calculated, thereby obtaining the third loss value for all images. The smaller the third loss value, the smaller the difference between the second predicted distance in the label domain set and the true second distance sequence, and the better the prediction performance of the third model. It should also be noted that after the third loss value is calculated, the back propagation algorithm can be used to calculate the gradient of the third loss value relative to the model parameters of the third model. Then, the parameters of the third model can be updated through the optimization algorithm to reduce the third loss value and improve the performance of the third model, so that the third model can more accurately predict the second predicted distance from multiple second segmented images to the second 2D image, that is, to the lesion segmentation image in future training.

[0063] The image loss value is a numerical value used during the training process of the first model to measure the difference between the predicted segmentation image output by the first model and the true labeled lesion segmentation image. The image loss value reflects the error of the first model in the current training step. The first loss value is a loss value obtained by combining the image loss value, the second loss value, and the third loss value. The first loss value is used to guide the optimization process of the first model, with the goal of minimizing this value to improve the performance of the first model. To achieve this goal, a loss function is typically used during the training process of the first model to quantify the error, and the resulting loss value is used as the optimization target to guide the learning process of the first model. Optionally, for the image segmentation task corresponding to the first model, a cross-entropy loss function can be used as the loss function to measure the difference between the predicted segmentation image and the true labeled lesion segmentation image. The cross-entropy loss function calculates the difference between the predicted value of each pixel in the predicted segmentation image and the true labeled lesion segmentation image, thereby obtaining the image loss value. The smaller the image loss value, the smaller the difference between the predicted segmentation image and the true labeled lesion segmentation image, and the better the prediction effect of the first model. It should also be noted that after the loss value of the entire image is calculated, the image loss value, the second loss value, and the third loss value can be combined to obtain a first loss value. Optionally, a weighted sum method is usually used to calculate the first loss value based on a given weight coefficient using the image loss value, the second loss value, and the third loss value. After obtaining the first loss value, the backpropagation algorithm can be used to calculate the gradient of the first loss value relative to the model parameters of the first model, and then the parameters of the first model can be updated through the optimization algorithm to reduce the first loss value and improve the performance of the first model, so that the first model can more accurately predict the lesion area in future training. Through the image loss value, the second loss value, and the third loss value, the first model can optimize multiple objectives at the same time, so that the first model performs better when processing different tasks.

[0064] Specifically, the first, second, and third image sets are input into the first, second, and third models, respectively, for model training. During model training, a second loss value is determined based on the first predicted distance output by the second model, the first predicted distance, and the first distance sequence. A third loss value is determined based on the second predicted distance output by the third model, the second predicted distance, and the second distance sequence. An image loss value is determined based on the predicted segmentation image output by the third model, the lesion segmentation image, and the predicted segmentation image. A first loss value is determined based on the image loss value, the second loss value, and the third loss value. After obtaining the first, second, and third loss values, the parameters of the first model are modified based on the first loss value, the parameters of the second model are modified based on the second loss value, and the parameters of the third model are modified based on the third loss value.

[0065] S140 , using the lesion segmentation model obtained when the iteration condition is satisfied as the segmentation model to be used, and removing the second model and the third model in the segmentation model to be used to obtain a target lesion segmentation model for processing the input image.

[0066] The "segmentation model to be used" refers to the model that can be directly used for lesion segmentation tasks after the model is trained and meets the iteration conditions. The segmentation model to be used has undergone a certain degree of training and optimization and can complete the lesion segmentation task on the input image. The target lesion segmentation model refers to the final lesion segmentation model after being streamlined. The target lesion segmentation model is a model that focuses on the optimal part of the lesion segmentation task by removing unnecessary second and third model components from the obtained segmentation model to be used.

[0067] It should be noted that during the training process, the segmentation model to be used will go through multiple iterations, and each iteration will adjust the model parameters until it reaches a predetermined performance standard or meets specific iteration conditions. The iteration conditions can be determined based on the model's loss value, accuracy, or other evaluation indicators. If the segmentation model to be used meets the predetermined performance standard (such as the loss value drops to a low enough level, or the segmentation accuracy reaches a certain threshold), it can be used as the segmentation model to be used.

[0068] Specifically, after the model is trained, the lesion segmentation model obtained when the iteration conditions are met is used as the segmentation model to be used. When generating the target lesion segmentation model, the second model and the third model in the segmentation model to be used are removed. Therefore, the target lesion segmentation model only retains the part directly related to lesion segmentation, that is, the first model. Removing the second model and the third model in the segmentation model to be used can remove unnecessary complexity and parts that interfere with the segmentation results, making the model more concise and efficient, and also making the target lesion segmentation model more focused on accurate identification of lesion areas.

[0069] The technical solution of the embodiment of the present disclosure is as follows: first, a plurality of training samples are obtained. Then, for the plurality of training samples, the training samples are input into a pre-constructed lesion segmentation model, and an image domain set, a predicted segmentation image, and a label domain set are output. Furthermore, based on the image domain set, the predicted segmentation image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model are determined, and the model parameters in the lesion segmentation model are corrected based on the loss value. Finally, the lesion segmentation model obtained when the iteration condition is met is used as the segmentation model to be used, and the second model and the third model in the segmentation model to be used are removed to obtain a target lesion segmentation model for processing the input image. This solves the problem in the prior art that after constructing a lesion segmentation model based only on the U-Net architecture or a variant of the U-Net architecture, there is a high labeling cost and a low label utilization rate when training based on the constructed lesion segmentation model. When training based on a lesion segmentation model, the embodiment of the present invention includes not only a first model constructed based on a variant of the U-Net architecture, but also a second model for predicting distance information for a second image set, and a third model for predicting distance information for a third image set. The training method based on the first, second, and third models can achieve the effects of reducing labeling costs, increasing label utilization, and improving the model's generalization ability in small sample scenarios when training the lesion segmentation model.

[0070] Example 2

[0071] Figure 2 This is a flowchart of a training method for a lesion segmentation model provided by an embodiment of the present invention. Based on the previous embodiment, this method provides a detailed description of inputting training samples into a pre-built lesion segmentation model, outputting an image domain set, a predicted segmented image, and a label domain set. For specific implementation methods, please refer to the technical solution of this embodiment. Technical terms that are identical or corresponding to those in the above embodiments are not repeated here.

[0072] like Figure 2 As shown, the method specifically includes the following steps:

[0073] S210: Acquire multiple training samples.

[0074] S220 . After inputting the training samples into the lesion segmentation model, the first image set is processed based on the first model, the second image set is processed based on the second model, and the third image set is processed based on the third model.

[0075] In this embodiment, based on the input-output relationship, the first model includes a residual convolutional network layer, at least one target network layer, a residual convolutional network layer and an output layer, and the target network layer includes a residual convolutional network layer and an attention layer; based on the input-output relationship, the second model and the third model include multiple residual convolutional network layers and a regressor layer; the output of the target network layer of the first model at the same level is respectively connected to the input of the residual convolutional network in the second model and the third model to construct a lesion segmentation model.

[0076] Among them, such as Figure 3As shown, based on the input-output relationship, the first model includes a residual convolutional network layer, at least one target network layer, a residual convolutional network layer, and an output layer, and the target network layer includes a residual convolutional network layer and an attention layer. It should be noted that the residual convolutional network layer is a variant of the convolutional neural network. Each convolution layer in the residual convolutional network layer operates on the input image by applying a convolution kernel (or filter) to extract features. These convolutional layers are responsible for identifying different local patterns (such as edges, textures, shapes, etc.) from the input image. The residual convolutional network layer can be used to solve the gradient disappearance problem in the neural network by introducing residual connections, making the network deeper while avoiding training difficulties; the target network layer is part of the entire network architecture of the first model, and the target network layer is specifically used to learn and optimize specific target tasks. In tasks such as lesion segmentation, the target network layer is responsible for focusing on and refining important information related to the segmentation task from the features extracted by the convolutional layer. The target network layer consists of a residual convolutional network layer and an attention layer. This structure enables the network to accurately focus on areas related to the target task (such as lesion area segmentation) while extracting features; the attention layer is a mechanism that enables the model to selectively focus on different parts of the input when processing input data. The attention mechanism helps the model focus on the most important information for the final prediction by assigning different weights to different parts of the input. In image segmentation tasks, lesion areas may be very small and inconspicuous, so special attention should be paid to these areas while ignoring other irrelevant parts. The attention layer calculates the importance weight of each pixel or area, allowing the model to focus on key areas during training and inference. The attention layer makes the model more efficient in feature extraction and can focus on the most important areas, thereby improving the accuracy of the segmentation task; the output layer is the last layer of the first model, responsible for converting the final feature map of the first model into the final prediction result. In image segmentation tasks, the output layer is usually a pixel-level classifier used to generate a classification label for each pixel (such as lesion or non-lesion). Optionally, the output layer outputs whether each pixel belongs to the lesion area (1) or the non-lesion area (0); the regressor layer is a network layer specifically used for regression tasks. The main task of the regressor layer is to convert the features extracted from the previous layer (such as the residual convolutional network layer) into a continuous output value. In the regressor layer, a linear activation function (or no activation function) is usually used because the output of the regression task is a continuous value. If the target output range is within a specific interval, a sigmoid or tanh activation function can be used for output constraint, but usually the output of the regressor layer is a direct numerical value. In an embodiment of the present invention, the regressor layer is usually located at the end of the second model and the third model, and the continuous numerical output is predicted by the regressor layer to directly predict the first prediction distance and the numerical value of the first prediction distance.

[0077] Specifically, such as Figure 3As shown, based on the input-output relationship, the first model includes a residual convolutional network layer, at least one target network layer, a residual convolutional network layer, and an output layer, wherein the target network layer includes a residual convolutional network layer and an attention layer. Based on the input-output relationship, the second model and the third model include multiple residual convolutional network layers and a regressor layer. During the downsampling process of the first model, the number of residual convolutional layers of the first model and the residual convolutional layers of the second model and the third model are the same and correspond. The target network layer of the first model and the residual convolutional network layers of the second model and the third model are the same and correspond. The output of the target network layer of the first model, located at the same level, is connected to the input of the residual convolutional network in the second model and the third model, respectively, to finally construct a lesion segmentation model. After the lesion segmentation model is constructed, the training sample is first input into the lesion segmentation model. Then, the first image set is processed based on the first model, the second image set is processed based on the second model, and the third image set is processed based on the third model.

[0078] S230. When the attention layer in the target network layer in the first model outputs an intermediate result, the intermediate result is sent to the residual convolutional network layer in the second model and the third model that is at the same level as the target network layer, so that the second model and the third model learn the output parameters in the first model.

[0079] The intermediate result refers to the Q value calculated by convolution at each attention layer in the first model, i.e., the query value. The query value is a vector used to "ask" the model for certain information about the input data, indicating the current attention paid to other elements.

[0080] Specifically, such as Figure 3As shown, during the model training process, after the training sample is input into the lesion segmentation model, the first image set is processed based on the first model, the second image set is processed based on the second model, and the third image set is processed based on the third model. In the second model and the third model, all the residual convolution layers except the first residual convolution layer can calculate the K value (key value) and the V value (value) respectively by convolution. For each K, Q and V value calculated, the calculation process remains consistent. Therefore, in an embodiment of the present invention, the output of the target network layer of the first model is connected to the input of the residual convolution network in the second model and the third model respectively based on being located at the same level as an example for explanation. Optionally, after obtaining a set of K, Q and V values, first, the correlation between Q and K is calculated, and the similarity is usually calculated using dot product. Then, a weighted sum is performed, that is, the corresponding V is weighted by the result of the similarity, and the K pair with higher similarity will be given a greater weight. Finally, the weighted value V is weighted and summed to obtain a new representation, which is the output result. After obtaining the output result, it is calculated through a residual convolution layer to restore it to the same dimension as the previous residual convolution layer. After the above processing, the output result of the data processing layer in the first model can be used to correct the model parameters in the second and third models. Figure 3 As shown, during the model training process, the output results of each residual convolution layer in the first model, that is, the weight results, can be shared with each corresponding residual convolution layer in the second model. It should be noted that the weight of the residual convolution layer usually refers to the parameters learned by the residual convolution layer, and the weight of the residual convolution layer usually includes parameters such as convolution kernel, bias, and weight matrix. Optionally, the weight of each residual convolution layer in the first model can be shared with each corresponding residual convolution layer in the second model by sharing the convolution kernel.

[0081] S240: outputting an image domain set based on the second model, outputting a label domain set based on the third model, and outputting a predicted segmented image based on the first model;

[0082] The first prediction distance in the image domain is used to characterize the similarity between the third 2D image and the second 2D image, and the second prediction distance in the label domain set is used to characterize the similarity between the first segmented image and the lesion segmented image.

[0083] It should be noted that because the first predicted distance is calculated based on the levels of each third 2D image and the levels of the second 2D image, and for 3D images, the similarity graphs between images of closely spaced levels are higher, while the similarity graphs between images of distant levels are lower. Therefore, the image domain set is generally used to represent the predicted similarity between all third 2D images and the second 2D image in the second model. Each element in the image domain set, namely the first predicted distance, represents a predicted distance value and also represents the similarity between the third 2D image and the second 2D image. Because the second predicted distance is calculated based on the levels of the first segmented image and the levels of the lesion segmented image, and for 3D segmented images, the similarity graphs between images of closely spaced levels are higher, while the similarity graphs between images of distant levels are lower. Therefore, the label domain set is generally used to represent the predicted similarity between the first segmented image and the lesion segmented image in the third model. Each element in the label domain set, namely the second predicted distance, represents a predicted distance value and also represents the similarity between the first segmented image and the lesion segmented image.

[0084] Specifically, after inputting the sample image into the constructed lesion segmentation model, the second model can output data representing the similarity between the third 2D image and the second 2D image, namely the first predicted distance. All first predicted distances are referred to as the image domain set. The third model can also output data representing the similarity between the first segmented image and the lesion segmented image, namely the second predicted distance. All second predicted distances are referred to as the label domain set. The first model can then output the final predicted segmented image, thus completing the predictive segmentation task.

[0085] S250. Based on the image domain set, the predicted segmentation image, the label domain set, the first distance information, the second distance information and the lesion segmentation image, determine the first loss value for correcting the first model, the second loss value for correcting the second model, and the third loss value for correcting the third model, and correct the model parameters in the lesion segmentation model based on the loss values.

[0086] S260: Use the lesion segmentation model obtained when the iteration condition is satisfied as the segmentation model to be used, and remove the second model and the third model in the segmentation model to be used to obtain a target lesion segmentation model for processing the input image.

[0087] In this embodiment, an image to be segmented including a preset part is input into a target lesion segmentation model, and a target image segmented from the lesion part is output.

[0088] Among them, the image to be segmented refers to the input image that needs to be segmented, which can be an image of the prostate area based on CT or the like. It aims to extract a specific area (such as the lesion area) from the image to be segmented. The image to be segmented contains all the information that needs to be processed by the segmentation model to be used, and is the image passed to the lesion segmentation model after preprocessing. The target image refers to the result image output by the lesion segmentation model when the image to be segmented is input into the target lesion segmentation model. The target image represents the lesion area predicted by the target lesion segmentation model in the image to be segmented. As the final output of the segmentation task, the target image can be a binary image in which the position and shape of the lesion area are marked. It should be noted that the size of the target image is usually the same as that of the image to be segmented, maintaining spatial alignment.

[0089] Optionally, the target image is a target image segmented from the image to be segmented of the prostate area.

[0090] Specifically, when the lesion segmentation model meets the iteration conditions during training, the current lesion segmentation model is used as the target segmentation model. After obtaining the target segmentation model, the second and third models in the target segmentation model are removed to obtain the target lesion segmentation model used to process the input image. After the image to be segmented, including the prostate and other areas, is input into the obtained target lesion segmentation model, the target lesion segmentation model outputs a target image of the prostate segmentation area.

[0091] The technical solution of the embodiment of the present disclosure is, first, to obtain a plurality of training samples. Then, after the training samples are input into the lesion segmentation model, the first image set is processed based on the first model, the second image set is processed based on the second model, and the third image set is processed based on the third model. When the attention layer in the target network layer in the first model outputs an intermediate result, the intermediate result is sent to the residual convolutional network layer in the second model and the third model that is at the same level as the target network layer, so that the second model and the third model learn the output parameters in the first model. An image domain set is output based on the second model, a label domain set is output based on the third model, and a predicted segmentation image is output based on the first model. Furthermore, based on the image domain set, the predicted segmentation image, the label domain set, the first distance information, the second distance information and the second 2D image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model are determined, and the model parameters in the lesion segmentation model are corrected based on the loss values. Finally, the lesion segmentation model obtained when the iteration conditions are met is used as the target segmentation model, and the second and third models in the target segmentation model are removed to obtain the target lesion segmentation model for processing the input image. The three models in the lesion segmentation model process the three image sets separately, and the second and third models learn the output parameters of the first model. This ensures that the second and third models inherit the learning capabilities of the first model, and that the first model assists in the training of the second and third models. The three models in the lesion segmentation model learn collaboratively, making the overall model performance more stable. The first loss value of the first model is corrected by the output image domain set and first distance information of the second model; the third loss value of the third model is corrected by the output label domain set and second distance information of the third model; and the second loss value of the second model is corrected by the output predicted segmentation image, lesion segmentation image, first loss value, and third loss value of the first model. Through information exchange between the models and dynamic adjustment of the loss function, mutual promotion and optimization are achieved, thereby improving the overall segmentation effect, fully leveraging the advantages of each model, enhancing overall segmentation performance, improving segmentation accuracy, and reducing the risk of overfitting. The sharing of features and outputs between models promotes knowledge transfer and improves each model's understanding of lesion characteristics.

[0092] Example 3

[0093] As an optional embodiment of the above embodiment, its specific implementation can be combined with Figure 3 And the following text to understand.

[0094] In this embodiment, data import and preprocessing steps are first performed. Optionally, for a CT image of the prostate region, a 2D image is first extracted from this 3D image, and a lesion segmentation image is obtained after manual or algorithmic annotation. Next, a certain number of 2D images can be randomly extracted from this 3D image, and the distance between the level of each 2D image in the CT image and the level of the first extracted 2D image is calculated. Furthermore, a certain number of 2D annotated images can be randomly extracted from the annotated images corresponding to this 3D image, and the distance between the level of each 2D annotated image in the CT image and the level of the first extracted 2D image is calculated.

[0095] Data augmentation can be performed on the input training set samples. It is important to note that the data augmentation steps for the labeled images in the samples must be consistent with the data augmentation steps for the 2D images. Furthermore, all extracted 2D images must be normalized based on their annotations. All sample images must be resized to maintain the same dimensions of C*W*D.

[0096] The model is trained on three parallel tasks. One is the conventional segmentation task, which is to predict the result of the input and compare the predicted value with the true label; the other is self-supervised learning with labeled images, which is to compare the current input to be predicted with several labels to find the corresponding true label; the last one is self-supervised learning with original images, which is to compare the current input to be predicted with several original images to find the corresponding original images with the same content.

[0097] Specifically, for model segmentation training, the model of the disclosed embodiment is based on the UNet structure, incorporating self-attention into each convolutional network of the segmentation network and employing cross-attention in the regression network for contrastive learning. The segmentation network of this solution consists of an encoder and a decoder, with the encoder having L encoding layers and the decoder having L-1 decoding layers. The encoders of the segmentation network and the regression network of this solution both use the same architecture.

[0098] The specific encoding process is as follows: the input of the encoding layer, the input of the jth encoding layer is the output of the j-1th encoding layer, that is, the size of the input feature is C j-1 *W j-1 *D j-1 ; Perform convolution calculation. The jth encoding layer first calculates the input features using the residual convolution network, and the output size is C j *W j *D j , where W j It's W j-1Half of D j It's D j-1 Half of the self-attention; self-attention calculation, for the result of convolution calculation, calculate self-attention Q, K and V, and update the features according to Q, K and V to get the result of self-attention, whose size is C j *W j *D j ; Convolution calculation, the result of self-attention calculation is input into the residual convolution network again for feature calculation, and the output feature size is C j *W j *D j At this point, the jth encoding layer of the segmentation network is calculated; and so on until the last encoding layer is calculated. The output of the last encoding layer is then used as the input of the decoder.

[0099] The specific decoding process is as follows: the input of the decoding layer, the input of the jth decoding layer is the output of the j+1th decoding layer (feature size is C j+1 *W j+1 *D j+1 ) and the output of the j-th layer encoder (feature size is C j *W j *D j ). It should be noted that the order of the decoding layer is exactly the opposite of the encoding layer; the convolution calculation is similar to the encoding layer. For the input features, the decoding layer first performs residual convolution calculation on them; the self-attention calculation is similar to the encoding layer. For the convolved features, the decoding layer calculates self-attention on them; the convolution calculation is similar to the encoding layer. Finally, the attention features are calculated using the residual convolution network, and the output is the final output of the decoding layer; and so on. The final decoding output is the prediction result of the model.

[0100] To conduct self-supervised learning of labels, the output of the encoder of the segmentation network is not only used for decoding of the segmentation network, but also for self-supervised learning with the labels. Therefore, in addition to the real label image, the input of the self-supervised learning network also includes C-1 2D label images. Therefore, the input of the network is these C label images, and the output of the network is C numerical values, that is, regression values. Each numerical value is between 0 and 1, representing the model's confidence in each label. The larger the value, the more similar the model believes that the corresponding label image is to the segmentation image of the current original image. Finally, the true label of the network, that is, the true regression value is the distance between each standard image and the true label image (between 0 and 1). For example, the true label is the Kth image in the 3D image, and a certain label image is the Pth image in the 3D image, so the distance between the image and the true label is d = exp(-(KP) 2), the distance is the regression label value of the label image; for self-supervised input, the input of the jth encoding layer is the output of the j-1th encoding layer, that is, the size of the input feature is C j-1 *W j-1 *D j-1 ; Convolution calculation, the jth encoding layer first uses the residual convolution network to calculate the input features, and the output size is C j *W j *D j , where W j It's W j-1 Half of D j It's D j-1 Half of the cross attention, input the self-attention Q of the j-th encoder of the segmentation network, calculate K and V according to the output of the convolution calculation, and then calculate the cross attention, and the final output feature size is C j *W j *D j ; Convolution calculation, the result of the cross attention output is input into the residual convolution network again for feature calculation, and the output feature size is C j *W j *D j At this point, the calculation of the j-th encoding layer of the self-supervised learning network is completed.

[0101] Regression calculations input the features of the last encoding layer into a regressor, which outputs C regression values, representing the model's confidence in the C labeled images. Specifically, they represent the degree of similarity between each labeled image and its corresponding true labeled image. Regression is used instead of classification because if a labeled image is very close to the true labeled image, such as the K+1th image of a 3D labeled image, the difference between the two images is very small. The classification network's binary output of either true or false can lead to unstable training.

[0102] In addition to being used for segmentation network decoding, the segmentation network's encoder output undergoes self-supervised learning with the original image. Therefore, in addition to the true labeled image, the self-supervised learning network's input also includes C-1 2D original images. The network's input is these C original images, and its output is C numerical values, or regression values. Each value, between 0 and 1, represents the model's confidence in each original image. The larger the value, the more similar the model believes the corresponding original image is to the current image. Finally, the network's true label, or true regression value, is the distance between each standard image and the true labeled image (between 0 and 1). This self-supervised learning network is similar to label self-supervised learning, replacing the labeled image with the original image. However, the parameters of this self-supervised learning network and the segmentation network are shared, as both are used to encode the original image.

[0103] For loss calculation, the regression network uses mean square error, and the segmentation network uses cross entropy function to calculate the regression loss and segmentation loss respectively.

[0104] Model update, regression loss and segmentation loss are used to update the segmentation network parameters, and regression loss is also used to update the parameters of the self-supervised learning network.

[0105] Model selection: After the model training is completed, the validation set is used to select a model with the best segmentation result from the trained models as the final model.

[0106] Model prediction, input a new data into the final segmentation network, and output the result predicted by the model.

[0107] According to the technical solution of the embodiment of the present disclosure, after data is imported, the model is trained according to three parallel tasks, one is a conventional segmentation task, the other is self-supervised learning with labeled images, and the last is self-supervised learning with original images. The segmentation task model is based on the structure of UNet, and self-attention is added to each convolutional network of the segmentation network, while cross attention is used in the regression network for contrastive learning. The segmentation network consists of an encoder and a decoder. The encoder has L encoding layers and the decoder has L-1 decoding layers. The encoders of the segmentation network and the regression network use the same architecture. The model is updated based on the loss calculation, and after the model training is completed, the validation set is used to select a model with the best segmentation result from the trained models as the final model. After the final model is determined, model prediction is performed, a new data is input into the final segmentation network, and the model prediction result is output. This can achieve the goal of reducing the labeling cost, improving the utilization rate of labels, and improving the generalization ability of the model in small sample scenarios during model training, and fully utilizing the advantages of each model to improve the overall segmentation performance.

[0108] Example 4

[0109] Figure 4 This is a structural diagram of a training device for a lesion segmentation model provided by an embodiment of the present disclosure. As shown in the figure, the device includes: multiple training sample acquisition modules 310, a predicted segmentation image output module 320, a model parameter correction module 330 and a target lesion segmentation model determination module 340.

[0110] A plurality of training sample acquisition modules are configured to acquire a plurality of training samples, wherein the training samples are composed of a plurality of first 2D images extracted from the same 3D image, the training samples include a first image set, a second image set, and a third image set, the first image set includes a second 2D image and a lesion segmentation image corresponding to the second 2D image, the second image set includes a plurality of third 2D images and the plurality of third 2D images include the second 2D image, and first distance information between the plurality of third 2D images and the second 2D image in the 3D image layer, the third image set includes the plurality of third 2D images and the plurality of third 2D images. A plurality of first segmented images corresponding to a first 2D image and including the lesion segmentation image in the first segmented image, and second distance information between the plurality of first segmented images and the lesion segmentation image level in the 3D image level; a predicted segmented image output module, for inputting the plurality of training samples into a pre-built lesion segmentation model, and outputting an image domain set, a predicted segmented image and a label domain set, wherein the lesion segmentation model includes a first model for performing lesion segmentation prediction on the first image set, a second model for performing distance information prediction on the second image set, and a second model for performing distance information prediction on the lesion segmentation model. a third model for predicting distance information based on the third image set, wherein the first model includes multiple data processing layers, and the output result of at least one data processing layer is used to correct model parameters in the second model and the third model; the image domain set includes the first predicted distance from the third 2D image to the second 2D image predicted by the second model, and the label domain set includes the second predicted distance from the multiple second segmented images to the second 2D image predicted by the third model; a model parameter correction module is used to determine a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model based on the image domain set, the predicted segmented image, the label domain set, the first distance information, the second distance information, and the second 2D image, and to correct the model parameters in the lesion segmentation model based on the loss values; wherein the first loss value is determined based on at least the second loss value and the third loss value; and a target lesion segmentation model determination module is used to use the lesion segmentation model obtained when the iteration condition is satisfied as the segmentation model to be used, and remove the second model and the third model from the segmentation model to be used to obtain a target lesion segmentation model for processing the input image.

[0111] The technical solution of the embodiment of the present disclosure is as follows: first, a plurality of training samples are obtained. Then, for the plurality of training samples, the training samples are input into a pre-constructed lesion segmentation model, and an image domain set, a predicted segmentation image, and a label domain set are output. Furthermore, based on the image domain set, the predicted segmentation image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model are determined, and the model parameters in the lesion segmentation model are corrected based on the loss value. Finally, the lesion segmentation model obtained when the iteration condition is met is used as the segmentation model to be used, and the second model and the third model in the segmentation model to be used are removed to obtain a target lesion segmentation model for processing the input image. This solves the problem in the prior art that after constructing a lesion segmentation model based only on the U-Net architecture or a variant of the U-Net architecture, there is a high labeling cost and a low label utilization rate when training based on the constructed lesion segmentation model. When training based on a lesion segmentation model, the embodiment of the present invention includes not only a first model constructed based on a variant of the U-Net architecture, but also a second model for predicting distance information for a second image set, and a third model for predicting distance information for a third image set. The training method based on the first, second, and third models can achieve the effects of reducing labeling costs, increasing label utilization, and improving the model's generalization ability in small sample scenarios when training the lesion segmentation model.

[0112] On the basis of the above technical solutions, the device also includes: a lesion segmentation model construction module, which is used to, based on the input-output relationship, the first model includes a residual convolutional network layer, at least one target network layer, a residual convolutional network layer and an output layer, and the target network layer includes a residual convolutional network layer and an attention layer; based on the input-output relationship, the second model and the third model include multiple residual convolutional network layers and a regressor layer; the output of the target network layer of the first model at the same level is respectively connected to the input of the residual convolutional network in the second model and the third model to construct the lesion segmentation model.

[0113] Based on the above technical solutions, the device also includes: a 2D image acquisition module for acquiring at least one sample pair, wherein the sample pair includes a 3D sample image corresponding to a preset part and a 3D segmentation image of the preset part corresponding to the 3D sample image; for the at least one sample pair, randomly selecting a first position from the Z axis of the 3D sample image and the 3D segmentation image of the sample pair to obtain a first 2D image corresponding to the first position and a lesion segmentation image to be used; based on the first 2D image and the corresponding lesion segmentation image to be used, determining the multiple training samples.

[0114] Based on the above technical solutions, the device further includes an image set acquisition module, which includes a first image set acquisition submodule, a second image set acquisition submodule, and a third image set acquisition submodule.

[0115] a first image set acquisition submodule, configured to randomly select a second 2D image from the plurality of first 2D images, and retrieve a to-be-used lesion segmentation image corresponding to the second 2D image as the lesion segmentation image, so as to obtain the first image set based on the second 2D image and the lesion segmentation image;

[0116] a second image set acquisition submodule, configured to randomly select a preset number of third 2D images from the plurality of first 2D images, and construct the second image set based on the third 2D images and the second 2D images in the first image set;

[0117] The third image set acquisition submodule is configured to randomly select a preset number of first segmented images from the plurality of lesion segmentation images to be used, and determine the third image set based on the first segmented images and the lesion segmentation images in the first image set.

[0118] On the basis of the above technical solutions, the device further includes: an image set updating module. The image set updating module includes: a second image set updating submodule and a third image set updating submodule.

[0119] a second image set updating submodule, configured to obtain first distance information between the third 2D image in the second image set and the second 2D image in the Z-axis direction of the 3D image, and arrange the first distance information according to arrangement information of the third 2D image in the second image set to obtain a first distance sequence, so as to update the second image set based on the first distance sequence;

[0120] The third image set updating submodule is used to obtain the second distance information between the first segmented image in the third image set and the lesion segmentation image in the Z-axis direction of the 3D image, and arrange the second distance information according to the arrangement information of the first segmented image in the third image set to obtain a second distance sequence, so as to update the third image set based on the second distance sequence.

[0121] On the basis of the above technical solutions, the predicted segmented image output module 320 includes: an image set processing submodule, an output parameter learning submodule and a predicted segmented image output submodule.

[0122] an image set processing submodule, configured to, after inputting the training samples into the lesion segmentation model, process the first image set based on the first model, process the second image set based on the second model, and process the third image set based on the third model;

[0123] an output parameter learning submodule, configured to, when the attention layer in the target network layer in the first model outputs an intermediate result, send the intermediate result to the residual convolutional network layer in the second model and the third model that is located at the same level as the target network layer, so that the second model and the third model learn the output parameters in the first model;

[0124] A predicted segmented image output submodule is configured to output the image domain set based on the second model, output the label domain set based on the third model, and output the predicted segmented image based on the first model; wherein the first predicted distance in the image domain is used to characterize the similarity between the third 2D image and the second 2D image, and the second predicted distance in the label domain set is used to characterize the similarity between the first segmented image and the lesion segmented image.

[0125] On the basis of the above technical solutions, the model parameter correction module 330 includes: a second loss determination submodule, a third loss determination submodule, a first loss determination submodule and a parameter correction submodule.

[0126] a second loss determination submodule, configured to determine a second loss value based on the first predicted distance in the image domain set and the first distance sequence in the second image set;

[0127] a third loss determination submodule, configured to determine a third loss value based on the second predicted distance in the label domain set and the second distance sequence in the third image set;

[0128] a first loss determination submodule, configured to determine an image loss value based on the lesion segmentation image and the predicted segmentation image, and determine a first loss value based on the image loss value, the second loss value, and the third loss value;

[0129] A parameter correction submodule is used to correct the parameters of the first model based on the first loss value, correct the parameters of the second model based on the second loss value, and correct the parameters of the third model based on the third loss value.

[0130] On the basis of the above technical solutions, the device further includes: a target image output module, which is used to input the image to be segmented including the preset part into the target lesion segmentation model and output the target image segmented from the lesion part.

[0131] On the basis of the above technical solution, the preset part is the prostate part, and the target image is the target image segmented from the image to be segmented of the prostate part.

[0132] The training device for the lesion segmentation model provided in the embodiments of the present disclosure can execute the training method for the lesion segmentation model provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0133] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0134] Example 5

[0135] Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 5 , which shows an electronic device (eg Figure 5 The terminal device in the embodiments of the present disclosure may include, but is not limited to, a mobile terminal such as a mobile phone, a laptop computer, a digital broadcast receiver, a PDA (personal digital assistant), a PAD (tablet computer), a PMP (portable multimedia player), an in-vehicle terminal (such as an in-vehicle navigation terminal), and the like. Figure 5 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0136] like Figure 5 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.

[0137] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0138] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0139] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0140] The electronic device provided in the embodiment of the present disclosure and the training method of the lesion segmentation model provided in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0141] Example 6

[0142] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method for the lesion segmentation model provided in the above embodiment.

[0143] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0144] In some embodiments, the server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0145] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0146] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0147] Acquiring a plurality of training samples, wherein the training samples are composed of a plurality of first 2D images extracted from the same 3D image, the training samples including a first image set, a second image set, and a third image set, the first image set including a second 2D image and a lesion segmentation image corresponding to the second 2D image, the second image set including a plurality of third 2D images including the second 2D image, and first distance information between a level of the plurality of third 2D images in the 3D image and a level of the second 2D image, the third image set including a plurality of first segmented images corresponding to the plurality of first 2D images including the lesion segmentation image, and second distance information between a level of the plurality of first segmented images in the 3D image and a level of the lesion segmentation image;

[0148] For the multiple training samples, the training samples are input into a pre-built lesion segmentation model, and an image domain set, a predicted segmentation image, and a label domain set are output, wherein the lesion segmentation model includes a first model for performing lesion segmentation prediction on the first image, a second model for performing distance information prediction on the second image set, and a third model for performing distance information prediction on the third image set, the first model includes multiple data processing layers, and the output result of at least one data processing layer is used to correct model parameters in the second model and the third model, the image domain set includes a first predicted distance from the third 2D image to the second 2D image predicted by the second model, and the label domain set includes a second predicted distance from the multiple second segmented images to the second 2D image predicted by the third model;

[0149] Determining, based on the image domain set, the predicted segmented image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model, and correcting model parameters in the lesion segmentation model based on the loss values; wherein the first loss value is determined based on at least the second loss value and the third loss value;

[0150] The lesion segmentation model obtained when the iteration condition is satisfied is used as the segmentation model to be used, and the second model and the third model in the segmentation model to be used are removed to obtain a target lesion segmentation model for processing the input image.

[0151] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0153] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0154] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0155] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0157] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0158] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A training method for a lesion segmentation model, characterized in that: include: Acquiring a plurality of training samples, wherein the training samples are composed of a plurality of first 2D images extracted from the same 3D image, the training samples including a first image set, a second image set, and a third image set, the first image set including a second 2D image and a lesion segmentation image corresponding to the second 2D image, the second image set including a plurality of third 2D images including the second 2D image, and first distance information between a level of the plurality of third 2D images in the 3D image and a level of the second 2D image, the third image set including a plurality of first segmented images corresponding to the plurality of first 2D images including the lesion segmentation image, and second distance information between a level of the plurality of first segmented images in the 3D image and a level of the lesion segmentation image; For the multiple training samples, the training samples are input into a pre-built lesion segmentation model, and an image domain set, a predicted segmentation image, and a label domain set are output, wherein the lesion segmentation model includes a first model for performing lesion segmentation prediction on the first image set, a second model for performing distance information prediction on the second image set, and a third model for performing distance information prediction on the third image set, the first model includes multiple data processing layers, and the output result of at least one data processing layer is used to correct model parameters in the second model and the third model, the image domain set includes a first predicted distance from the third 2D image to the second 2D image predicted by the second model, and the label domain set includes a second predicted distance from the multiple second segmented images to the second 2D image predicted by the third model; Determining, based on the image domain set, the predicted segmented image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model, and correcting model parameters in the lesion segmentation model based on the loss values; wherein the first loss value is determined based on at least the second loss value and the third loss value; The lesion segmentation model obtained when the iteration condition is satisfied is used as the segmentation model to be used, and the second model and the third model in the segmentation model to be used are removed to obtain a target lesion segmentation model for processing the input image.

2. The method according to claim 1, characterized in that Also includes: Constructing the lesion segmentation model; The constructing of the lesion segmentation model comprises: According to the input-output relationship, the first model includes a residual convolutional network layer, at least one target network layer, a residual convolutional network layer, and an output layer, wherein the target network layer includes a residual convolutional network layer and an attention layer; The second model and the third model include a plurality of residual convolutional network layers and a regressor layer according to the input-output relationship; The output of the target network layer of the first model at the same level is connected to the input of the residual convolutional network in the second model and the third model respectively to construct the lesion segmentation model.

3. The method according to claim 1, characterized in that Before obtaining a plurality of training samples, the method further includes: Acquire at least one sample pair, wherein the sample pair includes a 3D sample image corresponding to a preset part and a 3D segmented image of the preset part corresponding to the 3D sample image; For the at least one sample pair, randomly selecting a first position from the Z axis of the 3D sample image and the 3D segmentation image of the sample pair to obtain a first 2D image corresponding to the first position and a lesion segmentation image to be used; The plurality of training samples are determined based on the first 2D image and the corresponding lesion segmentation image to be used.

4. The method according to claim 3, characterized in that After obtaining the first 2D image of the at least one sample pair and the corresponding lesion segmentation image to be used, the method further includes: randomly selecting a second 2D image from the plurality of first 2D images, and retrieving a to-be-used lesion segmentation image corresponding to the second 2D image as the lesion segmentation image, so as to obtain the first image set based on the second 2D image and the lesion segmentation image; randomly selecting a preset number of third 2D images from the plurality of first 2D images, and constructing the second image set based on the third 2D images and the second 2D images in the first image set; A preset number of first segmented images are randomly selected from the plurality of lesion segmented images to be used, and the third image set is determined based on the first segmented images and the lesion segmented images in the first image set.

5. The method according to claim 4, characterized in that The method further comprises: obtaining first distance information between the third 2D image in the second image set and the second 2D image in the Z-axis direction of the 3D image, and arranging the first distance information according to arrangement information of the third 2D image in the second image set to obtain a first distance sequence, and updating the second image set based on the first distance sequence; Obtain second distance information between the first segmented image in the third image set and the lesion segmented image in the Z-axis direction of the 3D image, and arrange the second distance information according to the arrangement information of the first segmented image in the third image set to obtain a second distance sequence, so as to update the third image set based on the second distance sequence.

6. The method according to claim 1, characterized in that The step of inputting the training sample into a pre-built lesion segmentation model and outputting an image domain set, a predicted segmentation image, and a label domain set comprises: After inputting the training samples into the lesion segmentation model, processing the first image set based on the first model, processing the second image set based on the second model, and processing the third image set based on the third model; When the attention layer in the target network layer in the first model outputs an intermediate result, the intermediate result is sent to the residual convolutional network layer in the second model and the third model that is located at the same level as the target network layer, so that the second model and the third model learn the output parameters in the first model; The image domain set is output based on the second model, the label domain set is output based on the third model, and the predicted segmented image is output based on the first model; wherein the first predicted distance in the image domain is used to characterize the similarity between the third 2D image and the second 2D image, and the second predicted distance in the label domain set is used to characterize the similarity between the first segmented image and the lesion segmented image.

7. The method according to claim 1, characterized in that The determining, based on the image domain set, the predicted segmented image, the label domain set, the first distance information, the second distance information, and the lesion segmentation image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model, and correcting model parameters in the lesion segmentation model based on the loss values, includes: determining a second loss value based on the first predicted distance in the image domain set and the first distance sequence in the second image set; determining a third loss value based on the second predicted distance in the label domain set and the second distance sequence in the third image set; Determining an image loss value based on the lesion segmentation image and the predicted segmentation image, and determining a first loss value based on the image loss value, the second loss value, and the third loss value; Parameters of the first model are corrected based on the first loss value, parameters of the second model are corrected based on the second loss value, and parameters of the third model are corrected based on the third loss value.

8. The method according to claim 1, characterized in that The method further comprises: The image to be segmented including the preset part is input into the target lesion segmentation model, and the target image segmented from the lesion part is output.

9. The method according to claim 8, characterized in that The preset part is a prostate part, and the target image is a target image segmented from the image to be segmented of the prostate part.

10. A training device for a lesion segmentation model, characterized in that: include: a plurality of training sample acquisition modules, configured to acquire a plurality of training samples, wherein the training samples are composed of a plurality of first 2D images extracted from the same 3D image, the training samples comprising a first image set, a second image set, and a third image set, the first image set comprising a second 2D image and a lesion segmentation image corresponding to the second 2D image, the second image set comprising a plurality of third 2D images including the second 2D image, and first distance information between the plurality of third 2D images at a level in the 3D image and the second 2D image, the third image set comprising a plurality of first segmented images corresponding to the plurality of first 2D images including the lesion segmentation image, and second distance information between the plurality of first segmented images at a level in the 3D image and the lesion segmentation image; a predicted segmented image output module, configured to input the plurality of training samples into a pre-built lesion segmentation model, and output an image domain set, a predicted segmented image, and a label domain set, wherein the lesion segmentation model includes a first model for predicting lesion segmentation for the first image set, a second model for predicting distance information for the second image set, and a third model for predicting distance information for the third image set, the first model includes a plurality of data processing layers, and an output result of at least one data processing layer is used to correct model parameters in the second model and the third model, the image domain set includes a first predicted distance from the third 2D image to the second 2D image predicted by the second model, and the label domain set includes a second predicted distance from the plurality of second segmented images to the second 2D image predicted by the third model; a model parameter correction module, configured to determine, based on the image domain set, the predicted segmented image, the label domain set, the first distance information, the second distance information, and the second 2D image, a first loss value for correcting the first model, a second loss value for correcting the second model, and a third loss value for correcting the third model, and to correct model parameters in the lesion segmentation model based on the loss values; wherein the first loss value is determined based on at least the second loss value and the third loss value; The target lesion segmentation model determination module is used to use the lesion segmentation model obtained when the iteration condition is met as the segmentation model to be used, and remove the second model and the third model in the segmentation model to be used to obtain the target lesion segmentation model for processing the input image.