Construction method and training method of image segmentation model and image processing method
By introducing channel attention layer and forward reverse processing order into the image segmentation model, the problems of high annotation cost and low segmentation accuracy of the U-Net architecture model are solved, and a more efficient medical image segmentation effect is achieved.
Patent Information
- Application Number
- CN202510540264.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
When the existing image segmentation model is built on the U-Net architecture or its variants, there are problems such as high annotation cost, low label utilization rate and low segmentation accuracy, especially in medical images.
When building an image segmentation model, a channel attention layer is introduced, and the residual convolution network and up-down sampling layer are combined with forward and reverse processing orders to reduce redundant information and improve model generalization capabilities and accuracy.
It effectively reduces the annotation cost of model training, improves the accuracy and generalization ability of image segmentation, and performs excellently in medical image segmentation tasks.
Smart Images

Figure CN120451738A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular to a method for constructing and training an image segmentation model and an image processing method. Background Art
[0002] With the improvement of science and technology, the field of image processing has gradually developed the use of constructed image segmentation models to separate the region of interest from the background.
[0003] Currently constructed image segmentation models are primarily based solely on the U-Net architecture, employing a symmetrical encoder-decoder structure connected by numerous skip connections to fuse information from different layers. Alternatively, segmentation models are constructed based on variations and extensions of the U-Net architecture. For example, integrating the principles of residual networks, residual connections are introduced into the convolutional layers of the U-Net architecture to help address the vanishing gradient problem in deep networks. However, training segmentation models based solely on the U-Net architecture or its variants requires a large number of high-quality annotated training samples. Furthermore, in medical images obtained through tomography, the image variability between layers is low. Therefore, image segmentation models constructed using the U-Net architecture or its variants suffer from high annotation costs and low label utilization during training. Furthermore, applying the trained image segmentation models to segmentation results in low segmentation accuracy. Summary of the Invention
[0004] The embodiments of the present disclosure provide a method for constructing, training, and processing an image segmentation model to effectively reduce redundant information and improve the generalization ability and accuracy of the image segmentation model.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for constructing an image segmentation model, the method comprising:
[0006] Constructing an image segmentation model including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer according to the input-output relationship;
[0007] According to the forward processing order of the at least one first target network layer and the reverse processing order of the at least one second target network layer, connecting the output of the first target network layer in the same processing order to the input of the second target network layer;
[0008] Among them, the first target network layer includes a residual convolutional network layer, a channel attention layer and a downsampling layer, the second target network layer includes a residual convolutional network layer, a channel attention layer and an upsampling layer, and the number of layers of the first target network layer is the same as the number of layers of the second target network layer.
[0009] In a second aspect, an embodiment of the present disclosure provides a method for training an image segmentation model, the method comprising:
[0010] Acquiring a plurality of training samples, wherein the training samples include at least two consecutive 2D sample images extracted from 3D images of the same subject and the same lesion site, and lesion annotated images corresponding to the at least two consecutive 2D images;
[0011] For the plurality of training samples, inputting at least two consecutive 2D sample images in the training samples into the image segmentation model to be trained, and outputting predicted labeled images corresponding to the at least two consecutive 2D images;
[0012] Processing the lesion annotated images and the corresponding predicted annotated images of the at least two consecutive 2D images based on the loss function in the image segmentation model to be trained to determine a target loss value;
[0013] The model parameters in the image segmentation model to be trained are modified based on the target loss value, and the model obtained when the loss function converges is used as the image segmentation model.
[0014] In a third aspect, the present disclosure further provides an image processing method, the method comprising:
[0015] Acquiring at least one 2D image including a target lesion site, wherein the at least one 2D image is extracted based on a 3D image of the target lesion site;
[0016] The at least one 2D image is input into a pre-trained image segmentation model, and a lesion segmentation image corresponding to the at least one 2D image is output.
[0017] In a fourth aspect, an embodiment of the present invention further provides a device for constructing an image segmentation model, the device comprising:
[0018] An image segmentation model construction module is used to construct an image segmentation model including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer according to the input-output relationship;
[0019] a second target network layer access module configured to access the output of the first target network layer in the same processing order to the input of the second target network layer according to the forward processing order of the at least one first target network layer and the reverse processing order of the at least one second target network layer;
[0020] Among them, the first target network layer includes a residual convolutional network layer, a channel attention layer and a downsampling layer, the second target network layer includes a residual convolutional network layer, a channel attention layer and an upsampling layer, and the number of layers of the first target network layer is the same as the number of layers of the second target network layer.
[0021] In a fifth aspect, an embodiment of the present invention provides a training device for an image segmentation model, the device comprising:
[0022] a training sample acquisition module, configured to acquire a plurality of training samples, wherein the training samples include at least two consecutive 2D sample images extracted from 3D images of the same subject and the same lesion site, and lesion annotated images corresponding to the at least two consecutive 2D images;
[0023] a predicted and annotated image output module, configured to input, for the plurality of training samples, at least two consecutive 2D sample images in the training samples into the image segmentation model to be trained, and output predicted and annotated images corresponding to the at least two consecutive 2D images;
[0024] a target loss value determination module, configured to process the lesion annotated images and the corresponding predicted annotated images of the at least two consecutive 2D images based on the loss function in the image segmentation model to be trained to determine a target loss value;
[0025] An image segmentation model determination module is used to modify the model parameters in the image segmentation model to be trained based on the target loss value, and use the model obtained when the loss function converges as the image segmentation model.
[0026] In a sixth aspect, an embodiment of the present invention provides an image processing device, the device comprising:
[0027] a 2D image acquisition module, configured to acquire at least one 2D image including a target lesion site, wherein the at least one 2D image is extracted based on a 3D image of the target lesion site;
[0028] The lesion segmentation image output module is used to input the at least one 2D image into a pre-trained image segmentation model and output a lesion segmentation image corresponding to the at least one 2D image.
[0029] In a seventh aspect, an embodiment of the present invention further provides an electronic device, comprising:
[0030] one or more processors;
[0031] a storage device for storing one or more programs,
[0032] When the one or more programs are executed by the one or more processors, the one or more processors implement the image segmentation model construction method, training method or image processing method as described in any one of the embodiments of the present invention.
[0033] In an eighth aspect, an embodiment of the present invention further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute an image segmentation model construction method, training method or image processing method as described in any one of the embodiments of the present invention.
[0034] In the ninth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the image segmentation model construction method, training method or image processing method as described in any one of the embodiments of the present invention.
[0035] The technical solution of the embodiment of the present disclosure first constructs an image segmentation model based on the input-output relationship, including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer. Each first target network layer in the image segmentation model includes a residual convolution network layer, a channel attention layer, and a downsampling layer. Each second target network layer in the image segmentation model includes a residual convolution network layer, a channel attention layer, and an upsampling layer. The number of layers in the first target network layer is the same as the number of layers in the second target network layer. In the constructed image segmentation model, based on the forward processing order of all first target network layers and the reverse processing order of all second target network layers, the outputs of the first target network layers with the same processing order are connected to the inputs of the second target network layers. This solves the problems of high labeling cost and low label utilization during model training in the prior art when constructing image segmentation models based solely on the U-Net architecture or its variants, as well as the problem of low segmentation accuracy when applying the trained image segmentation model for segmentation. When constructing an image segmentation model based on the input-output relationship, the embodiment of the present invention introduces a channel attention layer into the model structure to effectively reduce redundant information, thereby achieving the effect of improving the generalization ability and accuracy of the image segmentation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings introduced here only illustrate some of the embodiments to be described by the present invention, and are not exhaustive. A person skilled in the art can derive other drawings based on these drawings without inventive effort.
[0037] Figure 1 This is a flowchart of a method for constructing an image segmentation model provided by an embodiment of the present disclosure;
[0038] Figure 2 This is a flow chart of an image segmentation model provided by an embodiment of the present invention;
[0039] Figure 3 This is a flow chart of a channel attention layer provided by an embodiment of the present invention;
[0040] Figure 4 1 is a flow chart of a method for training an image segmentation model provided by an embodiment of the present invention;
[0041] Figure 5 is a flow chart of an image processing method provided by an embodiment of the present invention;
[0042] Figure 6 1 is a schematic structural diagram of a device for constructing an image segmentation model provided by an embodiment of the present invention;
[0043] Figure 7 Schematic diagram of the structure of a training device for an image segmentation model provided by an embodiment of the present invention;
[0044] Figure 8 is a structural diagram of an image processing device provided by an embodiment of the present invention;
[0045] Figure 9 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0047] Before introducing the technical solutions provided by the embodiments of the present disclosure, an example description of the application scenarios can be given first. The technical solutions provided by the embodiments of the present disclosure can be applied to scenarios in which an image segmentation model is constructed, the constructed image segmentation model is trained, and the trained model is applied. For example, for an image of the prostate part acquired based on a computer tomography (CT) image, if one wants to obtain a lesion segmentation image of the image based on the model, an image segmentation model can be constructed based on the technical solutions of the embodiments of the present disclosure. After the image segmentation model is constructed, the model can be trained based on multiple training samples to obtain a trained image segmentation model. After obtaining the trained image segmentation model, a lesion segmentation image of a 2D image can be output for at least one 2D image of the input prostate part.
[0048] It should be noted that the 2D image can be extracted from the 3D image of the target lesion site. The target lesion site can be the prostate site. The 3D image refers to an image of the target lesion site taken based on 3D imaging technology. 3D images can present images of the internal structures and organs of the human body in three spatial dimensions (length, width and height). The acquisition of 3D images can rely on a variety of medical imaging technologies, which capture detailed information about the human body through different physical principles. Optionally, the 3D image can be an image taken based on CT, or an image taken based on MRI. As the basic building block of the 3D image of the target lesion site, the 2D image is an image extracted from the 3D image. By acquiring and combining multiple 2D images with spatial positioning, a 3D image with a three-dimensional sense can be combined using computer processing and reconstruction technology. A 2D image refers to an image displayed on a plane, which only contains two spatial dimensions (length and width).
[0049] Based on the technical solution of the embodiment of the present disclosure, the image segmentation model includes at least a channel attention layer. Based on the constructed and trained image segmentation model, a lesion segmentation image corresponding to the 2D image can be obtained, which effectively reduces redundant information and improves the generalization ability and accuracy of the image segmentation model.
[0050] Example 1
[0051] Figure 1 It is a flow chart of a method for constructing an image segmentation model provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to situations based on constructing an image segmentation model. The method can be executed by a device for constructing an image segmentation model. The device can be implemented in the form of software and / or hardware. The hardware can be a mobile electronic device, which can execute the method for constructing an image segmentation model provided by this technical solution.
[0052] like Figure 1 As shown, the method includes:
[0053] S110. Construct an image segmentation model including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer based on the input-output relationship.
[0054] Among them, the first target network layer includes a residual convolutional network layer, a channel attention layer and a downsampling layer, the second target network layer includes a residual convolutional network layer, a channel attention layer and an upsampling layer, and the number of layers of the first target network layer is the same as the number of layers of the second target network layer.
[0055] It should be noted that the image segmentation model, as a deep learning model, aims to accurately locate and segment the lesion area of 2D images from images taken based on medical images (such as CT, MRI, etc.).
[0056] Generally speaking, the input layer's primary responsibility is to receive and preprocess 2D images, converting them into a format and range suitable for image segmentation models. This step is crucial for ensuring that the subsequent first target network layer can effectively extract features. Residual convolutional networks (RCNs) utilize residual connections and residual blocks. In each residual block, the input is directly added to the convolutional layer output. This makes it easier to learn the "residual" rather than directly learning the mapping between input and output, thus avoiding the training difficulties associated with increased depth. The channel attention layer dynamically adjusts the weights of each channel, allowing the image segmentation model to adaptively focus on important feature channels and suppress unimportant ones. This layer provides flexibility to image segmentation models, allowing them to adaptively adjust their attention to different channels based on the task and input data. This adaptability enables the image segmentation model to automatically focus on key information in different image processing scenarios. By suppressing the influence of irrelevant channels and reducing the interference of redundant features, the channel attention layer improves the generalization ability of the image segmentation model. This is particularly important in complex image segmentation tasks, helping the model more accurately identify and segment target regions.
[0057] The downsampling layer can be used to reduce the spatial dimensions of an image (such as its width and height) to extract more abstract features and reduce computational complexity. Downsampling can be achieved through pooling or convolutional layers. The upsampling layer can expand or restore the spatial dimensions of an image, typically used to generate high-resolution output or restore feature maps to their original size. Upsampling can be achieved through methods such as transposed convolution or interpolation. The output layer is the final layer of the lesion segmentation model, and its function is to convert the final feature map of the image segmentation model into the model's prediction results.
[0058] Specifically, in order to segment the 2D image extracted from the 3D image of the target lesion site and to segment the tumor area and other areas in the 2D image, an image segmentation model can be constructed first. Figure 2 The constructed image segmentation model may include an input layer, a first target network layer consisting of multiple residual convolutional network layers, channel attention layers and downsampling layers, a second target network layer consisting of multiple residual convolutional network layers, channel attention layers and upsampling layers, a convolutional network layer and an output layer.
[0059] S120: According to a forward processing order of at least one first target network layer and a reverse processing order of at least one second target network layer, connect the output of the first target network layer in the same processing order to the input of the second target network layer.
[0060] It should be noted that, in the embodiment of the present invention, when constructing the image segmentation model, different layers in the network can be combined to achieve the transmission of context information and efficient feature extraction through forward and reverse processing sequences. Figure 2 , the forward processing order refers to the layer-by-layer processing order from the input layer to at least one first target network layer. After the forward processing order, the resolution of the image gradually decreases, the features gradually abstract, and finally a feature map based on deep feature representation is obtained. Figure 2 , the reverse processing order refers to the layer-by-layer processing order from at least one second target network layer to the output layer. After the reverse processing order, the resolution of the image is gradually restored, more detailed local information is restored, and finally a high-resolution output image is generated. In the image segmentation model, the output of each first target network layer will form a mapping relationship with the input of each second target network layer during the processing process, and their order is consistent. The same processing order means that the first target network layer and the second target network layer are processed in the corresponding hierarchical order. That is to say, the output of the first target network layer of the Lth layer will be connected to the input of the Lth layer of the second target network layer. The same processing order is maintained from the input image, through the downsampling of the first target network layer, until the upsampling process of the second target network layer ends.
[0061] In this embodiment, based on the input-output relationship, a channel feature extraction layer, a node feature calculation layer, a graph attention network layer, and an attention value layer are constructed.
[0062] Among them, the channel feature extraction layer includes at least two channel extraction units; the node feature calculation layer includes a node feature calculation unit and a similarity matrix calculation unit; the node feature calculation unit includes at least two node feature calculation subunits; the node feature calculation subunit includes at least a first multilayer perceptron; the similarity matrix calculation unit includes at least two node relationship feature calculation subunits and a similarity matrix calculation subunit; the node relationship feature calculation subunit includes at least a second multilayer perceptron; the graph attention network layer includes at least a third multilayer perceptron; the attention value layer includes an attention value extraction unit and a first channel feature calculation unit; the first multilayer perceptron, the second multilayer perceptron and the third multilayer perceptron are all composed of multiple fully connected layers.
[0063] It should be noted that according to the input-output relationship, Figure 3 The channel attention layer includes a channel feature extraction layer, a node feature calculation layer, a graph attention network layer, and an attention value layer. The channel feature extraction layer includes at least two channel extraction units, each of which can extract channel feature information for its corresponding channel from the first output feature graph. The node feature calculation unit includes at least two node feature calculation subunits. Based on the channel feature information extracted for each channel in the first output feature graph, each node feature calculation subunit can use a first multilayer perceptron to calculate the node features for its corresponding channel. Based on the channel feature information extracted for each channel in the first output feature graph, the similarity matrix calculation unit in the node feature calculation layer can ultimately output a similarity matrix. The similarity matrix calculation unit includes at least two node relationship feature calculation subunits. Based on the channel feature information extracted for each channel in the first output feature graph, each node relationship feature calculation subunit can use a second multilayer perceptron to calculate the node relationship features for its corresponding channel. Based on the node relationship features of each channel, the similarity matrix calculation subunit in the similarity matrix calculation unit can be used to calculate the similarity matrix. Based on the similarity matrix and node features, the third multilayer perceptron in the graph attention network layer can be used to calculate the attention matrix. The attention value layer includes at least an attention value extraction unit, based on which elements in the attention matrix can be extracted. The attention value layer includes at least a first channel feature calculation unit, based on which first channel feature calculation unit, elements in the attention matrix, and the first output feature map can obtain the first channel feature.
[0064] Optionally, the channel extraction unit may extract and process at least two features in the first output feature map to obtain at least two second channel features.
[0065] The first output feature map is the feature map output by the residual convolutional network layer in the first target network layer of the same level or the residual convolutional network layer in the second target network layer of the same level. The second channel feature is represented by a matrix consisting of the width and height of each channel in the first output feature map. The number of second channel features is consistent with the number of channels in the first output feature map.
[0066] Specifically, for each channel in the first output feature map, the channel extraction unit may extract and process each channel feature in the first output feature map, and ultimately obtain second channel features having the same number of channels as the first output feature map. Furthermore, each second channel feature may be represented by a matrix consisting of the width and height of each channel in the first output feature map.
[0067] Optionally, at least two second channel features are stretched to obtain at least two first vectors; characteristic elements in at least two first vectors are randomly discarded to obtain at least two second vectors; and at least two second vectors are input into a first multilayer perceptron to output at least two node features.
[0068] The first vector is the vector obtained by stretching the second channel feature, and the first vector is a one-dimensional vector. For example, if the second channel feature corresponding to a channel is represented by a matrix of size W*D, the second channel feature can be stretched into a one-dimensional vector of length W*D, namely the first vector.
[0069] Among them, the characteristic element refers to the element in the first vector. The second vector refers to the vector obtained after randomly discarding the characteristic elements in the first vector. For example, the Dropout technology can be used to randomly discard the characteristic elements in the first vector. Specifically, the probability of Dropout refers to the proportion of randomly discarded characteristic elements. For example, if the probability is equal to 0.5, then there is a 50% probability that each characteristic element in the first vector will be discarded, that is, set to 0. In order to achieve random discarding, a random binary mask vector of the same size as the first vector can be generated. Each element in the mask is 0 or 1. For the elements at each position in the random binary mask vector, the probability can be set to 0, and the remaining elements can be set to 1. Finally, the random binary mask vector can be multiplied by the characteristic elements in the first vector to implement the discarding operation.
[0070] It should be noted that the node feature is the output result after the second vector is input into the first multilayer perceptron. Each node feature can be used to represent the characteristic information of each second channel feature. To calculate the node feature, the second vector can be nonlinearly transformed and learned by the first multilayer perceptron. Optionally, the second vector can be input into the first multilayer perceptron, and the second vector can be calculated through the fully connected layer in the first multilayer perceptron to finally obtain the node feature. The number of node features is consistent with the number of second vectors.
[0071] Optionally, at least two second vectors are input into a second multilayer perceptron to output at least two node relationship features; at least two node relationship features are processed based on a distance function to output a first similarity value corresponding to every two node relationship features; an average similarity value is obtained by averaging all the first similarity values; for each first similarity value, the average similarity value is subtracted to obtain a second similarity value; the second similarity value is set to 0 if it is a non-positive value to obtain a third similarity value; and based on the third similarity value, a similarity matrix is output.
[0072] The elements in the similarity matrix represent the third similarity values between the second vector pairs at corresponding positions.
[0073] It should be noted that the node relationship feature is the output result after the second vector is input into the second multilayer perceptron. The node relationship feature of a certain channel can be used to represent the characteristic information between this second vector and all other second vectors. To calculate the node relationship feature, the second vector can be nonlinearly transformed and learned by the second multilayer perceptron. Optionally, the second vector can be input into the second multilayer perceptron, and the second vector can be calculated through the fully connected layer in the second multilayer perceptron to finally obtain the node relationship feature.
[0074] The distance function is a function for calculating the similarity between any two node relationship features. The first similarity value is a calculation result obtained by calculating the distance function for any two node relationship features.
[0075] Optionally, after obtaining all node relationship features, if each node relationship feature is represented as a vector of length K, that is, each node relationship feature is K-dimensional. All node relationship features can be organized into a matrix of shape (C, K), where C is the number of second vectors and K is the dimension of each node relationship feature. The first similarity value can be calculated based on cosine similarity using matrix multiplication. The calculated first similarity value represents the similarity relationship between each pair of second vectors, providing a more efficient way to capture the similarity relationship between different channels.
[0076] The average similarity value is the result obtained by averaging all the first similarities. Each second similarity value is the result obtained by subtracting the average similarity value from each first similarity value. For each obtained second similarity value, if the second similarity value is greater than 0, then the second similarity value corresponding to the second vector pair is retained. If the second similarity value is not greater than 0, then the second similarity value corresponding to the second vector pair is set to 0. The updated value of the second similarity value becomes the third similarity value between the second vector pairs. The similarity matrix refers to the matrix composed of the third similarity values.
[0077] Optionally, for at least one positive element in the similarity matrix, extract the first row information and the first column information corresponding to the positive element; concatenate the node features corresponding to the first row information and the node features corresponding to the first column information to obtain at least one third vector; input the at least one third vector into a third multilayer perceptron to obtain at least one first element; and output the attention matrix based on the first element.
[0078] Among them, the first row information is the row corresponding to the positive element, the first column information is the column corresponding to the positive element, and the element in the attention matrix represents the first element at the corresponding position.
[0079] It should be noted that positive elements refer to elements in the similarity matrix whose values are positive. For positive elements, the node features corresponding to the number of rows of positive elements and the node features corresponding to the number of columns of positive elements are concatenated, and the resulting vector is called the third vector. Concatenation means combining two node features into a longer vector. For example, assuming the length of the node features corresponding to the number of rows of positive elements is D, and the length of the node features corresponding to the number of columns of positive elements is also D, then the length of the resulting third vector after concatenation is 2D. The third multilayer perceptron is a neural network composed of multiple fully connected layers. After the third vector is input into the third multilayer perceptron, it learns the mapping through the weights between the multiple fully connected layers, ultimately outputting a single value, which is the first element. The number of first elements matches the number of positive elements in the similarity matrix. For example, assuming the resulting third vector after concatenation is [0.1, 0.2, 0.4, 0.5], after calculation by the third multilayer perceptron, the final output first element may be 0.75.
[0080] It should be noted that the first element represents the strength of the relationship between channels and represents the attention value between the two channels. Similarly, the value of the first element corresponding to all positive elements can be calculated, and finally the attention matrix can be obtained. The element in the Pth row and Qth column of the attention matrix represents the first element value between the Pth second channel feature and the Qth second channel feature, which is also the attention value between the Pth second channel feature and the Qth second channel feature.
[0081] Optionally, at least one first element corresponding to at least one second row of information in the attention matrix is extracted; at least one first element in the second row of information and at least one second channel feature in the first output feature map of the residual convolutional network layer in the first target network layer at the corresponding position or the residual convolutional network layer in the second target network layer are multiplied and added to obtain a first channel feature.
[0082] The second row of information is the row corresponding to the attention matrix. For a certain first channel feature, the result of multiplying the second channel feature in the same row of the attention matrix with the second channel feature in the same row of the first output feature map is the first channel feature. Taking the Pth second channel feature as an example, the Pth row of the attention matrix is extracted, and then the Pth first output feature map is obtained. The first element in the Pth row of the attention matrix and the corresponding position of the Pth first output feature map are multiplied and added to obtain the first channel feature.
[0083] The technical solution of the embodiment of the present disclosure first constructs an image segmentation model based on the input-output relationship, including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer. Each first target network layer in the image segmentation model includes a residual convolution network layer, a channel attention layer, and a downsampling layer. Each second target network layer in the image segmentation model includes a residual convolution network layer, a channel attention layer, and an upsampling layer. The number of layers in the first target network layer is the same as the number of layers in the second target network layer. In the constructed image segmentation model, based on the forward processing order of all first target network layers and the reverse processing order of all second target network layers, the outputs of the first target network layers with the same processing order are connected to the inputs of the second target network layers. This solves the problems of high labeling cost and low label utilization during model training in the prior art when constructing image segmentation models based solely on the U-Net architecture or its variants, as well as the problem of low segmentation accuracy when applying the trained image segmentation model for segmentation. When constructing an image segmentation model based on the input-output relationship, the embodiment of the present invention introduces a channel attention layer into the model structure to effectively reduce redundant information, thereby achieving the effect of improving the generalization ability and accuracy of the image segmentation model.
[0084] Example 2
[0085] Figure 4 This is a flowchart of a method for training an image segmentation model provided by an embodiment of the present invention. Based on the previous embodiment, this method further details how to train a to-be-trained image segmentation model using multiple training samples to obtain an image segmentation model. For specific implementations, please refer to the technical solution of this embodiment. Technical terms that are identical or corresponding to those in the previous embodiment are not repeated here.
[0086] like Figure 4 As shown, the method specifically includes the following steps:
[0087] S210: Obtain multiple training samples.
[0088] The multiple training samples refer to multiple samples for training the image segmentation model.
[0089] It should be noted that the training samples include at least two consecutive 2D sample images extracted from 3D images of the same subject and the same lesion site, as well as lesion annotated images corresponding to the at least two consecutive 2D images. It should be noted that for a particular 2D sample image, the area that needs to be annotated (such as a tumor or blood vessel) can be determined. Then, this 2D sample image can be annotated using methods such as manual annotation or automatic threshold segmentation tools, with the area that needs to be annotated saved as 1 and the remaining background areas set to 0. Finally, a lesion annotated image corresponding to the 2D sample image is generated.
[0090] Specifically, at least two consecutive 2D sample images and lesion annotated images corresponding to the at least two consecutive 2D images can be extracted from the 3D image as training samples for training the image segmentation model. The same data augmentation processing can also be performed on the 3D image and the 3D annotated images corresponding to the 3D image. At least two consecutive 2D sample images and lesion annotated images corresponding to the at least two consecutive 2D images can be extracted from the enhanced 3D image as training samples for training the image segmentation model.
[0091] S220 . For a plurality of training samples, input at least two consecutive 2D sample images in the training samples into the image segmentation model to be trained, and output predicted labeled images corresponding to the at least two consecutive 2D images.
[0092] The image segmentation model to be trained is a model whose model parameters are initial parameters or default parameters. The predicted labeled image is the labeled image output after at least two consecutive 2D sample images in the training sample are input into the image segmentation model to be trained.
[0093] It should be noted that the model parameters in the image segmentation model to be trained do not meet the expected requirements. Therefore, there is a certain difference between the predicted annotation image and the lesion annotation image output based on the model parameters at this time. Therefore, the corresponding error loss value can be determined based on the predicted annotation image and lesion annotation image corresponding to each 2D sample image.
[0094] S230 , processing the lesion annotated images and the corresponding predicted annotated images of at least two consecutive 2D images based on the loss function in the image segmentation model to be trained to determine a target loss value.
[0095] Among them, the loss function is a mathematical formula used to measure the difference between the predicted annotation image of the image segmentation model to be trained and the lesion annotation image. The loss function is usually based on the comparison of the predicted annotation image of the 2D image segmentation (usually a binary image, indicating whether the pixel belongs to the lesion area) with the lesion annotation image. The loss function in the image segmentation model to be trained can be a cross-entropy loss function. The target loss value refers to the loss value that the image segmentation model to be trained is expected to be optimized to during the training process of the image segmentation model to be trained. When optimizing the image segmentation model to be trained, the value of the loss function can be minimized as much as possible, that is, the predicted annotation image of the image segmentation model to be trained is made as close to the lesion annotation image as possible.
[0096] Specifically, we select an appropriate loss function and calculate the loss for each pair of 2D images, the lesion annotation image and the corresponding predicted annotation image. We then sum all the loss values and gradually optimize the weights of the image segmentation model to be trained based on the sum of the loss values. The ultimate goal is to minimize the target loss value while achieving good generalization performance and avoiding overfitting.
[0097] S240. Modify the model parameters in the image segmentation model to be trained based on the target loss value, and use the model obtained when the loss function converges as the image segmentation model.
[0098] When training the image segmentation model to be trained, the model parameters in the model can be corrected based on the output results of the diseased image segmentation model to be trained. That is, the image segmentation model can be obtained by correcting the loss function in the image segmentation model to be trained.
[0099] Specifically, after a 2D image is input into the image segmentation model to be trained, the model can generate a predicted annotated image corresponding to the 2D image. Based on the predicted annotated image and the lesion annotated image, the loss value corresponding to the 2D image can be determined. The model parameters of the image segmentation model to be trained can then be modified using the backpropagation method. The model obtained when the loss function converges is used as the image segmentation model.
[0100] The technical solution of the embodiment of the present disclosure obtains multiple training samples, and for the multiple training samples, inputs at least two continuous 2D sample images in the training samples into the image segmentation model to be trained, and outputs predicted annotation images corresponding to at least two continuous 2D images. Furthermore, the lesion annotation images and corresponding predicted annotation images of at least two continuous 2D images are processed based on the loss function in the image segmentation model to be trained to determine the target loss value. Based on the target loss value, the model parameters in the diseased image segmentation model to be trained are corrected, and the model obtained when the loss function converges is used as the image segmentation model. Training with continuous 2D images can help the model better understand the temporal relationship between images, thereby improving the accuracy of lesion segmentation. Continuous image input enables the model to capture the contextual information between images, so that it can better handle complex structures and lesion locations in the image, and improve the performance of the model in different slices and image noise, making the model more stable in practical applications.
[0101] Example 3
[0102] Figure 5 This is a flowchart of an image processing method provided by an embodiment of the present invention. Based on the previous embodiment, this method provides a detailed description of acquiring a 2D image and obtaining a corresponding lesion segmentation image based on a pre-trained image segmentation model. For specific implementations, please refer to the technical solution of this embodiment. Technical terms that are the same or corresponding to those in the previous embodiment are not repeated here.
[0103] like Figure 5 As shown, the method specifically includes the following steps:
[0104] S310: Acquire at least one 2D image including the target lesion site.
[0105] Wherein, at least one 2D image is extracted based on a 3D image of the target lesion site.
[0106] It should be noted that the 3D image refers to an image of the target lesion site taken based on 3D imaging technology. In this embodiment, the target lesion site is the prostate lesion site.
[0107] S320: Input at least one 2D image into a pre-trained image segmentation model, and output a lesion segmentation image corresponding to the at least one 2D image.
[0108] In this embodiment, for at least one 2D image, the 2D image is input into the image segmentation model, and the 2D image is processed based on the input layer, at least one first target network layer, at least one second target network layer, the residual convolutional network layer and the output layer to output a lesion segmentation image of the 2D image.
[0109] Specifically, at least one 2D image can be acquired based on a 3D image of the prostate area. All acquired 2D images can then be input into a pre-trained image segmentation model. Based on the image segmentation model, a lesion segmentation image corresponding to all input 2D images can be output.
[0110] The technical solution of the disclosed embodiments first acquires at least one 2D image including the target lesion site. Then, the at least one 2D image is input into a pre-trained image segmentation model, which outputs a lesion segmentation image corresponding to the at least one 2D image. The embodiments of the present invention obtain a lesion segmentation image corresponding to the 2D image based on an image segmentation model that includes at least a channel attention layer, effectively reducing redundant information and improving the accuracy of the lesion segmentation image.
[0111] Example 4
[0112] As an optional embodiment of the above embodiment, its specific implementation can be combined with Figure 2 And the following text to understand.
[0113] In this embodiment, first, input data, input image data x of training set sample i (i) and its label y (i) . For x (i) Perform data enhancement, such as rotation, translation, elastic transformation, noise addition, etc. If spatial transformation is involved, its label y (i) The same spatial transformation is also applied to ensure that x (i) and y (i) Corresponding. Need to be from 3D x (i) In the z-axis of the image, a position is randomly selected and c consecutive 2D images at that position are taken out. Each 2D image is regarded as a channel, thus forming a 2D image with c channels (for example, 3 channels when c=3). Similarly, y (i) The corresponding labels are taken out as the labels of 2D input. (i) Normalize the pixels so that their values are between 0 and 1. (i) and y (i) Resize them so that their dimensions are consistent, namely c×w×d.
[0114] The model is trained. The model consists of two parts. The first part is the encoder encoding process, and the second part is the decoder decoding process. During the encoding process, the model will continuously use the convolutional network to downsample, convolve and calculate attention on the features, while in the decoding process, the model will continuously upsample, convolve and calculate attention on the features. In the encoding process, the encoding network has a total of L encoding layers, which are numbered 1, 2, ..., L. Each encoding layer is composed of a residual convolutional network and channel attention. Taking the jth encoding layer as an example, the calculation process is as follows: input feature, the input of the jth encoding layer is the output of the j-1th encoding layer, that is, the size of the feature is c j-1 ×w j-1 ×d j-1 Using the residual convolutional network with stride=2, the input features are convolved and the output features with half the size are output, with a size of c j ×w j ×d j , which is the output w of this convolution j It is w j-1 Half of d j It is d j-1 Calculate the channel attention based on the graph attention network. The calculation of channel attention depends on the graph attention mechanism network. In the graph attention mechanism, attention needs to be calculated based on the nodes and edges of the graph. At this time, this scheme treats each channel as a different node to calculate the node features and the edges of the nodes. Therefore, at this time, there is c j nodes, each node is represented by w j ×d j Represented by. Computing node characteristics, for c j For each node (channel), each node is composed of w j ×d j Therefore, this scheme firstly calculates w j ×d j Stretched into a one-dimensional vector, the representation of each node is obtained. Dropout is used to randomly discard features. For the robustness of the model, the Dropout technology is then used to randomly discard some feature elements of the one-dimensional feature. Node relationship features and node feature calculation, in order to calculate the relationship between each node, that is, the edge between two nodes, this scheme inputs the one-dimensional feature into an MLP network to obtain the node relationship feature. At the same time, in order to calculate the attention between nodes, the features of each node are also required, so the one-dimensional feature needs to be input into another MLP network to obtain the features of each node. Similarity matrix calculation, based on the node relationship feature, the distance function, such as the cosine function, is used to calculate the similarity of the two node relationship features, and finally a size of c is obtained. j ×c jThe similarity matrix S, where the p-th row and q-th column elements of the similarity matrix represent the similarity between the p-th node (channel) and the q-th node (channel). Next, we calculate the mean of the similarity matrix S, subtract the mean from S, and finally apply the ReLU function to the similarity matrix. This is done to remove edges with low similarity, making the similarity matrix sparse, while also reducing noise interference. Graph attention calculation, according to the similarity matrix S, calculate the attention values between connected nodes. Since the similarity matrix S is a sparse matrix, this solution does not need to calculate the attention between every two nodes, but calculates the attention between nodes whose edges are not 0. For example, if the p-th row and q-th column element is greater than 0, the node features of the p-th node and the node features of the q-th node are spliced together to form a one-dimensional vector, and then the MLP network is used to calculate its attention value, that is, the output is a single element. By analogy, the attention between two nodes with edges greater than 0 is calculated according to the similarity matrix S, and finally a matrix of size c is formed. j ×c j The attention matrix A is a matrix in which the elements in the p-th row and q-th column represent the attention values between the p-th node and the q-th node. The channel features are re-represented, and then the features of all channels are updated based on the attention matrix A. Since the nodes in this scheme represent channels, taking the p-th channel as an example, the relationship between this channel and all other channels is represented by the p-th row of the attention matrix A. Specifically, the updated features of the p-th channel are, A p,0 Multiply by the feature of the 0th channel plus A p,1 Multiply by the first channel feature, ..., Multiply by c j After obtaining the updated channel features, the residual convolution network is used to calculate the features again, and the output of the residual convolution network is used as the output of the j-th encoding layer.
[0115] According to the above encoding process, the input is encoded in sequence until it reaches the final L-th encoding layer, and then the output of the L-th encoding layer is input to the L-1-th decoding layer, and the decoding operation is performed in sequence. In the decoding process, the decoding network is also composed of L decoding layers, which are numbered L-1, L-2, ..., 0. The calculation process of the decoding layer is similar to that of the encoding layer. Taking the j-th decoding layer as an example, the decoding process is as follows: Input feature, the input of the j-th decoding layer is the output of the j+1-th decoding layer, that is, the size of the feature is c j+1 ×w j+1 ×d j+1 At the same time, since this solution adopts the UNet network structure, the input of the j-th decoding layer also has the input of the j-th encoding feature, that is, the size is c j ×w j ×dj So c j+1 ×w j+1 ×d j+1 After upsampling the features and c j ×w j ×d j The features of S221 are concatenated in the channel dimension to obtain the input of the final j-th decoding layer. Convolution calculation, the features of S221 are input into the residual convolution network to obtain c j ×w j ×d j The feature output of S222 is obtained. Graph attention calculation inputs the features of S222 into the graph attention network, calculates the attention matrix A, and updates the channel features again based on this matrix. Convolution calculation inputs the output of S223 into the residual convolution network, and uses the output of the residual convolution network as the output of the j-th decoding network. According to the above decoding process, the encoded features are decoded in sequence, and the prediction results are finally output.
[0116] Loss calculation: Calculate the cross-entropy loss based on the predicted results and the true labels, and update the model parameters accordingly. Model selection: After model training is complete, use the validation set to select the best-performing model from the trained models as the final model. Model prediction: Based on the final model, input new test examples into the model and output the model's prediction results.
[0117] The technical solution of the disclosed embodiments proposes a new channel attention mechanism network, which is constructed based on the graph attention mechanism network. The disclosed embodiments treat each channel of the convolved features as a node and use graph attention to calculate the relationship between nodes and the attention between nodes, thereby obtaining the relationship between channels and ultimately updating the features of each channel. Sparse processing is performed when constructing the similarity matrix, which reduces the computational complexity and reduces the impact of noise.
[0118] Example 5
[0119] Figure 6 This is a structural diagram of a device for constructing an image segmentation model provided by an embodiment of the present disclosure. As shown in the figure, the device includes: an image segmentation model construction module 410 and a second target network layer access module 420.
[0120] An image segmentation model construction module is used to construct an image segmentation model including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer and an output layer based on the input-output relationship; a second target network layer access module is used to connect the output of the first target network layer with the same processing order to the input of the second target network layer based on the forward processing order of the at least one first target network layer and the reverse processing order of the at least one second target network layer; wherein, the first target network layer includes a residual convolution network layer, a channel attention layer and a downsampling layer, and the second target network layer includes a residual convolution network layer, a channel attention layer and an upsampling layer, and the number of layers of the first target network layer is the same as the number of layers of the second target network layer.
[0121] The technical solution of the embodiment of the present disclosure first constructs an image segmentation model based on the input-output relationship, including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer. Each first target network layer in the image segmentation model includes a residual convolution network layer, a channel attention layer, and a downsampling layer. Each second target network layer in the image segmentation model includes a residual convolution network layer, a channel attention layer, and an upsampling layer. The number of layers in the first target network layer is the same as the number of layers in the second target network layer. In the constructed image segmentation model, based on the forward processing order of all first target network layers and the reverse processing order of all second target network layers, the outputs of the first target network layers with the same processing order are connected to the inputs of the second target network layers. This solves the problems of high labeling cost and low label utilization during model training in the prior art when constructing image segmentation models based solely on the U-Net architecture or its variants, as well as the problem of low segmentation accuracy when applying the trained image segmentation model for segmentation. When constructing an image segmentation model based on the input-output relationship, the embodiment of the present invention introduces a channel attention layer into the model structure to effectively reduce redundant information, thereby achieving the effect of improving the generalization ability and accuracy of the image segmentation model.
[0122] Based on the above technical solutions, the channel attention layer includes:
[0123] According to the input-output relationship, a channel feature extraction layer, a node feature calculation layer, a graph attention network layer and an attention value layer are constructed; wherein, the channel feature extraction layer includes at least two channel extraction units; the node feature calculation layer includes a node feature calculation unit and a similarity matrix calculation unit; the node feature calculation unit includes at least two node feature calculation subunits; the node feature calculation subunit includes at least a first multilayer perceptron; the similarity matrix calculation unit includes at least two node relationship feature calculation subunits and a similarity matrix calculation subunit; the node relationship feature calculation subunit includes at least a second multilayer perceptron; the graph attention network layer includes at least a third multilayer perceptron; the attention value layer includes an attention value extraction unit and a first channel feature calculation unit; the first multilayer perceptron, the second multilayer perceptron and the third multilayer perceptron are all composed of multiple fully connected layers.
[0124] Based on the above technical solutions, the at least two channel extraction units are used to: extract and process at least two features in the first output feature map to obtain the at least two second channel features; wherein, the first output feature map is the feature map output by the residual convolutional network layer in the first target network layer of the same level or the residual convolutional network layer in the second target network layer of the same level; the second channel feature is represented by a matrix consisting of the width and height of each channel of the first output feature map; the number of the second channel features is consistent with the number of channels in the first output feature map.
[0125] Based on the above technical solutions, the at least two node feature calculation subunits are used to: stretch the at least two second channel features to obtain at least two first vectors; randomly discard the characteristic elements in the at least two first vectors to obtain at least two second vectors; input the at least two second vectors into the first multilayer perceptron to output at least two node features.
[0126] Based on the above technical solutions, the at least two node relationship feature calculation subunits and the similarity matrix calculation subunit are used to: input the at least two second vectors into a second multilayer perceptron to output at least two node relationship features; process the at least two node relationship features based on a distance function to output a first similarity value corresponding to each two node relationship features; obtain an average similarity value by averaging all the first similarity values; subtract the average similarity value from each first similarity value to obtain a second similarity value; set the second similarity value to 0 if it is a non-positive value to obtain a third similarity value; output the similarity matrix based on the third similarity value; wherein the elements in the similarity matrix represent the third similarity values between the second vector pairs at corresponding positions.
[0127] Based on the above technical solutions, the graph attention network layer is used to: for at least one positive element in the similarity matrix, extract the first row information and the first column information corresponding to the positive element; wherein, the first row information is the row corresponding to the positive element, and the first column information is the column corresponding to the positive element; splice the node features corresponding to the first row information and the node features corresponding to the first column information to obtain at least one third vector; input the at least one third vector into the third multi-layer perceptron to obtain at least one first element; based on the first element, output the attention matrix; wherein, the elements in the attention matrix represent the first elements at the corresponding positions.
[0128] Based on the above technical solutions, the attention value extraction unit and the first channel feature calculation unit are used to: extract the at least one first element corresponding to at least one second row of information in the attention matrix; wherein the second row of information is the row corresponding to the attention matrix; multiply the at least one first element in the second row of information and the at least one second channel feature in the first output feature map of the residual convolution network layer in the first target network layer at the corresponding position or the residual convolution network layer in the second target network layer, and then add them to obtain the first channel feature.
[0129] Example 6
[0130] Figure 7 This is a structural diagram of a training device for an image segmentation model provided by an embodiment of the present invention. As shown in the figure, the device includes: a training sample acquisition module 510, a predicted and labeled image output module 520, a target loss value determination module 530 and an image segmentation model determination module 540.
[0131] A training sample acquisition module is used to acquire multiple training samples, wherein the training samples include at least two continuous 2D sample images extracted from 3D images of the same object and the same lesion site, and lesion annotation images corresponding to the at least two continuous 2D images; a predicted annotation image output module, for the multiple training samples, inputs at least two continuous 2D sample images in the training samples into the image segmentation model to be trained, and outputs predicted annotation images corresponding to the at least two continuous 2D images; a target loss value determination module is used to process the lesion annotation images and corresponding predicted annotation images of the at least two continuous 2D images based on the loss function in the image segmentation model to be trained, and determine the target loss value; an image segmentation model determination module is used to correct the model parameters in the image segmentation model to be trained based on the target loss value, and use the model obtained when the loss function converges as the image segmentation model.
[0132] The technical solution of the embodiment of the present disclosure obtains multiple training samples, and for the multiple training samples, inputs at least two continuous 2D sample images in the training samples into the image segmentation model to be trained, and outputs predicted annotation images corresponding to at least two continuous 2D images. Furthermore, the lesion annotation images and corresponding predicted annotation images of at least two continuous 2D images are processed based on the loss function in the image segmentation model to be trained to determine the target loss value. Based on the target loss value, the model parameters in the diseased image segmentation model to be trained are corrected, and the model obtained when the loss function converges is used as the image segmentation model. Training with continuous 2D images can help the model better understand the temporal relationship between images, thereby improving the accuracy of lesion segmentation. Continuous image input enables the model to capture the contextual information between images, so that it can better handle complex structures and lesion locations in the image, and improve the performance of the model in different slices and image noise, making the model more stable in practical applications.
[0133] Example 7
[0134] Figure 8 6 is a structural diagram of an image processing device provided by an embodiment of the present invention. As shown in the figure, the device includes: a 2D image acquisition module 610 and a lesion segmentation image output module 620.
[0135] A 2D image acquisition module is used to acquire at least one 2D image including a target lesion site, wherein the at least one 2D image is extracted based on a 3D image of the target lesion site; a lesion segmentation image output module is used to input the at least one 2D image into a pre-trained image segmentation model and output a lesion segmentation image corresponding to the at least one 2D image.
[0136] Based on the above technical solutions, the lesion segmentation image output module is used to input the at least one 2D image into the image segmentation model, and process the 2D image based on the input layer, at least one first target network layer, at least one second target network layer, the residual convolutional network layer and the output layer, and output the lesion segmentation image of the 2D image.
[0137] The technical solution of the disclosed embodiments first acquires at least one 2D image including the target lesion site. Then, the at least one 2D image is input into a pre-trained image segmentation model, which outputs a lesion segmentation image corresponding to the at least one 2D image. The embodiments of the present invention obtain a lesion segmentation image corresponding to the 2D image based on an image segmentation model that includes at least a channel attention layer, effectively reducing redundant information and improving the accuracy of the lesion segmentation image.
[0138] Example 8
[0139] Figure 9 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 9 , which shows an electronic device (eg Figure 9 The terminal device in the embodiments of the present disclosure may include, but is not limited to, a mobile terminal such as a mobile phone, a laptop computer, a digital broadcast receiver, a PDA (personal digital assistant), a PAD (tablet computer), a PMP (portable multimedia player), an in-vehicle terminal (such as an in-vehicle navigation terminal), and the like. Figure 9 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0140] like Figure 9 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.
[0141] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 9 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0142] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0143] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0144] The electronic device provided in the embodiment of the present disclosure and the image segmentation model construction method, training method or image processing method provided in the above-mentioned embodiment belong to the same inventive concept. Technical details not fully described in this embodiment can be referred to the above-mentioned embodiment, and this embodiment has the same beneficial effects as the above-mentioned embodiment.
[0145] Embodiment 9
[0146] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the image segmentation model construction method, training method or image processing method provided in the above embodiments.
[0147] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0148] In some embodiments, the server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0149] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0150] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0151] Constructing an image segmentation model including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer according to the input-output relationship;
[0152] According to the forward processing order of the at least one first target network layer and the reverse processing order of the at least one second target network layer, connecting the output of the first target network layer in the same processing order to the input of the second target network layer;
[0153] Among them, the first target network layer includes a residual convolutional network layer, a channel attention layer and a downsampling layer, the second target network layer includes a residual convolutional network layer, a channel attention layer and an upsampling layer, and the number of layers of the first target network layer is the same as the number of layers of the second target network layer.
[0154] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0155] Acquiring a plurality of training samples, wherein the training samples include at least two consecutive 2D sample images extracted from 3D images of the same subject and the same lesion site, and lesion annotated images corresponding to the at least two consecutive 2D images;
[0156] For the plurality of training samples, inputting at least two consecutive 2D sample images in the training samples into the image segmentation model to be trained, and outputting predicted labeled images corresponding to the at least two consecutive 2D images;
[0157] Processing the lesion annotated images and the corresponding predicted annotated images of the at least two consecutive 2D images based on the loss function in the image segmentation model to be trained to determine a target loss value;
[0158] The model parameters in the image segmentation model to be trained are modified based on the target loss value, and the model obtained when the loss function converges is used as the image segmentation model.
[0159] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0160] Acquiring at least one 2D image including a target lesion site, wherein the at least one 2D image is extracted based on a 3D image of the target lesion site;
[0161] The at least one 2D image is input into a pre-trained image segmentation model, and a lesion segmentation image corresponding to the at least one 2D image is output.
[0162] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0164] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0165] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0166] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0167] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0168] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0169] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for constructing an image segmentation model, characterized in that: include: Constructing an image segmentation model including an input layer, at least one first target network layer, at least one second target network layer, a residual convolution layer, and an output layer according to the input-output relationship; According to the forward processing order of the at least one first target network layer and the reverse processing order of the at least one second target network layer, connecting the output of the first target network layer in the same processing order to the input of the second target network layer; Among them, the first target network layer includes a residual convolutional network layer, a channel attention layer and a downsampling layer, the second target network layer includes a residual convolutional network layer, a channel attention layer and an upsampling layer, and the number of layers of the first target network layer is the same as the number of layers of the second target network layer.
2. The method according to claim 1, characterized in that The channel attention layer includes: Based on the input-output relationship, the channel feature extraction layer, node feature calculation layer, graph attention network layer, and attention value layer are constructed; Among them, the channel feature extraction layer includes at least two channel extraction units; the node feature calculation layer includes a node feature calculation unit and a similarity matrix calculation unit; the node feature calculation unit includes at least two node feature calculation subunits; the node feature calculation subunit includes at least a first multilayer perceptron; the similarity matrix calculation unit includes at least two node relationship feature calculation subunits and a similarity matrix calculation subunit; the node relationship feature calculation subunit includes at least a second multilayer perceptron; the graph attention network layer includes at least a third multilayer perceptron; the attention value layer includes an attention value extraction unit and a first channel feature calculation unit; the first multilayer perceptron, the second multilayer perceptron and the third multilayer perceptron are all composed of multiple fully connected layers.
3. The method according to claim 2, characterized in that The at least two channel extraction units are configured to: Extracting and processing at least two features from the first output feature map to obtain the at least two second channel features; The first output feature map is a feature map output by the residual convolutional network layer in the first target network layer of the same level or the residual convolutional network layer in the second target network layer of the same level; the second channel feature is represented by a matrix consisting of the width and height of each channel of the first output feature map; the number of the second channel features is consistent with the number of channels in the first output feature map.
4. The method according to claim 2, characterized in that The at least two node feature calculation subunits are configured to: Stretching the at least two second channel features to obtain at least two first vectors; Randomly discarding characteristic elements in the at least two first vectors to obtain at least two second vectors; The at least two second vectors are input into a first multilayer perceptron, and at least two node features are output.
5. The method according to claim 2, characterized in that The at least two node relationship feature calculation subunits and the similarity matrix calculation subunit are used to: Inputting the at least two second vectors into a second multilayer perceptron, and outputting at least two node relationship features; Processing the at least two node relationship features based on a distance function, and outputting a first similarity value corresponding to each two node relationship features; An average similarity value is obtained by averaging all the first similarity values; For each first similarity value, subtract the average similarity value to obtain a second similarity value; Setting the non-positive value of the second similarity value to 0 to obtain a third similarity value; outputting the similarity matrix based on the third similarity value; The elements in the similarity matrix represent third similarity values between the second vector pairs at corresponding positions.
6. The method according to claim 2, characterized in that The graph attention network layer is used to: For at least one positive element in the similarity matrix, extract first row information and first column information corresponding to the positive element; wherein the first row information is the row corresponding to the positive element, and the first column information is the column corresponding to the positive element; Concatenate the node features corresponding to the first row of information and the node features corresponding to the first column of information to obtain at least one third vector; Inputting the at least one third vector into a third multilayer perceptron to obtain at least one first element; Based on the first element, the attention matrix is output; wherein the elements in the attention matrix represent the first elements at the corresponding positions.
7. The method according to claim 2, characterized in that The attention value extraction unit and the first channel feature calculation unit are used to: Extracting the at least one first element corresponding to at least one second row of information in the attention matrix; wherein the second row of information is the row corresponding to the attention matrix; The at least one first element in the second row of information and at least one second channel feature in the first output feature map of the residual convolutional network layer in the first target network layer in the corresponding position or the residual convolutional network layer in the second target network layer are multiplied and added to obtain the first channel feature.
8. A method for training an image segmentation model, comprising the image segmentation model according to claims 1-7, characterized in that: The method further comprises: Acquiring a plurality of training samples, wherein the training samples include at least two consecutive 2D sample images extracted from 3D images of the same subject and the same lesion site, and lesion annotated images corresponding to the at least two consecutive 2D images; For the plurality of training samples, inputting at least two consecutive 2D sample images in the training samples into the image segmentation model to be trained, and outputting predicted labeled images corresponding to the at least two consecutive 2D images; Processing the lesion annotated images and the corresponding predicted annotated images of the at least two consecutive 2D images based on the loss function in the image segmentation model to be trained to determine a target loss value; The model parameters in the image segmentation model to be trained are modified based on the target loss value, and the model obtained when the loss function converges is used as the image segmentation model.
9. An image processing method comprising the image segmentation model according to claims 1 to 7, characterized in that: The method further comprises: Acquiring at least one 2D image including a target lesion site, wherein the at least one 2D image is extracted based on a 3D image of the target lesion site; The at least one 2D image is input into a pre-trained image segmentation model, and a lesion segmentation image corresponding to the at least one 2D image is output.
10. The method according to claim 9, characterized in that Inputting the at least one 2D image into a pre-trained image segmentation model and outputting a lesion segmentation image corresponding to the at least one 2D image includes: For the at least one 2D image, the 2D image is input into the image segmentation model, and the 2D image is processed based on the input layer, at least one first target network layer, the at least one second target network layer, the residual convolutional network layer and the output layer, and a lesion segmentation image of the 2D image is output.