Prostate transrectal ultrasound image intelligent segmentation method and system
By using the regional adaptive Transformer module and jump connection in the prostate transrectal ultrasound image segmentation, the problem of identifying interfering information and prostate shape in the prior art is solved, and segmentation accuracy and model generalization ability are improved.
Patent Information
- Application Number
- CN202510108958.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The prior art has the problem of identifying interfering information and various prostate shapes in transrectal ultrasound image segmentation of the prostate, and the global modeling ability is weak, making it challenging to accurately segment the glands in the face of complex interfering information.
The regional adaptive Transformer (RAT) module is adopted to calculate the scale coefficient τ of the feature sequence to realize adaptive self-attention calculation, and combine jump connection and Patch fusion module to enhance the generalization ability of the model to TRUS image data.
It improves the segmentation accuracy of prostate TRUS images and the generalization ability of the model, reduces the computational complexity, avoids the loss of model flexibility, and maintains good generalization performance when the data set is small.
Smart Images

Figure CN120198447A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of prostate ultrasound image segmentation, and particularly relates to an intelligent segmentation method and system for transrectal ultrasound images of the prostate. Background Art
[0002] Worldwide, prostate cancer is the fourth most common cancer and the second most common cancer in men. Transrectal Ultrasound (TRUS) is one of the effective methods for early screening and diagnosis of prostate cancer. Accurately segmenting the prostate from TRUS images is crucial for computer-aided diagnosis of prostate cancer. However, manual segmentation of the prostate is time-consuming and laborious, and the segmentation results lack objectivity. With the rapid development of convolutional neural networks, many deep learning-based prostate segmentation methods have been proposed successively.
[0003] Traditional segmentation methods are based on the shallow features of the prostate, so they can only extract weak semantic information, which makes it difficult to handle problems such as speckle noise and various artifacts in TRUS images. The segmentation method based on Convolutional Neural Network (CNN) uses the edge features in different channels to enhance the edge features of the prostate, but it has defects in identifying interference information and various different prostate shapes. Although CNN can solve some problems (such as low signal-to-noise ratio and contrast), its global modeling ability is weak. Facing complex interference information, accurately segmenting the gland is challenging. Transformer has a powerful global modeling ability. The segmentation methods improved based on Transformer learn inductive biases in large-scale data, but these methods require additional pre-trained models, which brings huge computational complexity and loss of model flexibility, and cannot generalize the model well when the dataset is small. Summary of the Invention
[0004] The main purpose of the present invention is to overcome the deficiencies of the prior art and provide an intelligent segmentation method and system for transrectal ultrasound images of the prostate, which can accurately segment prostate TRUS images and effectively perform computer-aided diagnosis of prostate cancer.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: One aspect of the present invention provides an intelligent segmentation method for transrectal ultrasound images of the prostate, including the following steps: Encoding stage: Segment the TRUS image to be segmented and project it onto the channels; Perform downsampling, normalization, and linear layer processing on the image projected onto the channels in sequence, and repeat N - 1 times to obtain N feature sequences of different sizes; Calculate the scale coefficient τ of each feature sequence, and perform adaptive self-attention calculation on the generated triples of the feature sequence. Among them, when the scale coefficient τ is lower than the threshold T, global self-attention calculation is performed; when the scale coefficient τ is higher than the threshold T, local self-attention calculation is performed; Pass the adaptive self-attention calculation results of the first N - 1 feature sequences to the corresponding level in the decoding stage through skip connections, and output the adaptive self-attention calculation results of the Nth feature sequence as the encoding results; Decoding stage: Based on the encoding results and the adaptive self-attention calculation results of the first N - 1 feature sequences passed through skip connections, perform sampling and splicing level by level, and reshape them into a feature map; Input the feature map into the upsampling layer and convolutional layer in sequence for feature extraction to obtain the final segmentation result.
[0006] As a preferred technical solution, the calculation of the scale coefficient τ of each feature sequence is specifically: ; Among them, τ is the scale coefficient, used to represent the ratio of the computational complexity of global and local self-attention; H and W are the height and width of the feature map, C is the dimension of the feature map, and M is the window size.
[0007] As a preferred technical solution, the adaptive self-attention calculation is specifically: ; ; Among them, LN represents normalization, A-MSA(·) represents multi-head self-attention calculation, z i-1 represents the features output by the previous layer, represents the features output by the multi-head self-attention of the current layer, MLP(·) represents the calculation performed by the multi-layer perceptron module, and z i is the features output by the multi-layer perceptron module; the multi-head self-attention calculation is specifically: Generate triples {key, query, value} for the feature sequence, and use three learnable parameter matrices W q , W k and W v to calculate the triples {key, query, value}, as follows: Q = z i-1 W q ; K = z i-1 W k ; V = z i-1 W v ; ; where Q, K, and V are the calculated query matrix, key matrix, and value matrix respectively, Softmax(·) is the Softmax activation function, d is the feature dimension, B is the deviation of relative positions, and T is the threshold.
[0008] As a preferred technical solution, when performing the calculation of global self-attention, Q, K, V, B ∈ R HW×d ; when performing the calculation of local self-attention, Q, K, V, B ∈ R (M^2)×d , and M is the window size.
[0009] As a preferred technical solution, the adaptive self-attention calculation results of the encoding result and the first N - 1 feature sequences transmitted through skip connections are sampled and concatenated level by level and reshaped into a feature map. Specifically: The following steps are repeated N - 1 times to obtain the reshaped feature map: Upsample the encoding result, concatenate it with the adaptive self-attention calculation result of the feature sequence of the previous dimension transmitted through skip connections and converted into a sequence format; perform convolution on the concatenated result to match the dimensions.
[0010] Another aspect of the present invention also provides an intelligent segmentation system for transrectal ultrasound images of the prostate, including an encoder and a decoder. The encoder includes an embedding layer and N region adaptive Transformer modules; the decoder includes a Patch fusion module, an upsampling layer, and a convolutional layer connected in sequence; The embedding layer is used to extract N feature sequences of different sizes of the TRUS image to be segmented and connect them to N region adaptive Transformer modules respectively; The region adaptive Transformer module includes a multi-head self-attention and a multi-layer perceptron module connected in sequence, which is used to perform adaptive self-attention calculation on the feature sequence. Among them, the adaptive self-attention calculation results of the first N - 1 feature sequences are transmitted to the corresponding level of the decoder through skip connections, and the adaptive self-attention calculation result of the Nth feature sequence is output as the encoding result; the adaptive self-attention calculation is specifically: calculate the scale coefficient τ of the feature sequence, perform global self-attention calculation when the scale coefficient τ is lower than the threshold T, and perform local self-attention calculation when the scale coefficient τ is higher than the threshold T; The Patch fusion module is used to sample and concatenate the encoding result and the adaptive self-attention calculation results of the first N - 1 feature sequences transmitted through skip connections level by level and reshape them into a feature map; The convolutional layer is used to extract features from the feature map to obtain the final segmentation result.
[0011] As a preferred technical solution, the embedding layer includes a segmentation layer, a linear projection layer, and N-1 feature sequence layers connected in sequence; the segmentation layer is used to evenly segment the TRUS image to be segmented; the linear projection layer is used to project the image obtained by the segmentation layer onto the channels; the feature sequence layer includes a downsampling layer, a normalization layer, and a linear layer, and is used to extract feature sequences of different sizes.
[0012] As a preferred technical solution, the region adaptive Transformer module is configured to perform the following calculations: ; where τ is a scale coefficient used to characterize the ratio of the computational complexity of global and local self-attention; H and W are the height and width of the feature map, C is the dimension of the feature map, and M is the window size; The multi-head self-attention is configured to perform the following calculations: ; ; where LN represents normalization, z i-1 represents the features output by the previous layer, represents the features output by the multi-head self-attention of the current layer; A-MSA(·) represents the multi-head self-attention calculation, Softmax(·) is the Softmax activation function, d is the feature dimension, B is the deviation of the relative position, T is the threshold, Q, K, and V are the query matrix, key matrix, and value matrix calculated respectively, and are calculated by using three learnable parameter matrices W q , W k and W v as follows: Q = z i-1 W q ; K = z i-1 W k ; V = z i-1 W v ; When the scale coefficient τ is lower than the threshold T and the global self-attention calculation is performed, Q, K, V, B ∈ R HW×d ; when the scale coefficient τ is higher than the threshold T and the local self-attention calculation is performed, Q, K, V, B ∈ R (M^2)×d ; The multi-layer perceptron module is configured to perform the following calculations: ; where MLP(·) represents the calculation performed by the multi-layer perceptron module, and z i is the feature output by the multi-layer perceptron module.
[0013] As a preferred technical solution, the Patch fusion module is provided with N - 1 connected in sequence, and each Patch fusion module is configured to perform the following steps: Upsample the encoded result or the feature sequence output by the previous Patch fusion module, and splice it with the adaptive self - attention calculation result of the feature sequence of the previous dimension passed through the skip connection from the encoder and converted into the sequence format; process the spliced result through a convolutional layer to match the dimensions, and obtain the output of the Patch fusion module at this level.
[0014] As a preferred technical solution, the convolutional layers all introduce inductive biases of locality, spatial invariance, and hierarchical structure.
[0015] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) The regional adaptive Transformer (RAT) module of the present invention realizes adaptive operation by calculating the scale coefficient τ of the current sequence complexity, adopts local self - attention for low computational complexity, reduces the computational complexity, and avoids the loss of model flexibility.
[0016] (2) The present invention constructs a hierarchical structure of the corresponding encoder, introduces inductive biases such as locality and spatial invariance in convolutions at different levels, and enhances the generalization ability of the entire model for TRUS image data.
[0017] (3) The present invention uses the Patch fusion module to receive the global information from the encoder as an additional input, avoids the continuous convolution from destroying the context information extracted by the transformer, and ensures that the network can not only reduce the data requirements, but also maintain the ability to recognize interference problems and various prostate shapes. Description of the Drawings
[0018] Figure 1 is the image segmentation flowchart of an intelligent prostate transrectal ultrasound image segmentation method according to an embodiment of the present invention; Figure 2 is the structural schematic diagram of an intelligent prostate transrectal ultrasound image segmentation system according to an embodiment of the present invention; Figure 3 is the schematic diagram of the qualitative result comparison of an intelligent prostate transrectal ultrasound image segmentation system according to an embodiment of the present invention and other segmentation models in the prostate segmentation task. Detailed Embodiments
[0019] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0020] Embodiment 1: As Figure 1 shown, this embodiment provides an intelligent segmentation method for transrectal ultrasound images of the prostate, including the following steps: S1. Encoding stage, as Figure 1 shown in the left half.
[0021] S1.1. In the encoder, the embedding layer evenly divides the TRUS image to be segmented ( ), and then projects it onto the channels through the linear projection layer. The current layer is the first layer. On this basis, downsampling is performed, and after normalization and linear layer processing, the second layer is constructed. According to the above method, through N - 1 downsampling layers, an N - level hierarchical structure is constructed to provide N different - sized feature sequences for subsequent processing to provide global context information.
[0022] S1.2. The present invention realizes adaptive operation through a simple but effective threshold - constraint method. For the input feature sequence, first calculate the scale coefficient τ of each feature sequence, and then perform adaptive self - attention calculation on the generated triplet of the feature sequence. Among them, when the scale coefficient τ is lower than the threshold T, global self - attention calculation is performed, and when the scale coefficient τ is higher than the threshold T, local self - attention calculation is performed to achieve low computational complexity.
[0023] (1) Calculation of the scale coefficient τ of the feature sequence: ; where τ is the scale coefficient, used to characterize the ratio of the computational complexity of global and local self - attention; H and W are the height and width of the feature map, C is the dimension of the feature map, and M is the window size.
[0024] (2) Adaptive self - attention calculation: ; ; where LN represents normalization, z i-1 represents the feature output by the previous layer, Represents the features output by the multi-head self-attention of the current layer; MLP(·) represents the calculation performed by the multi-layer perceptron module in the traditional Vision Transformer (VIT) module, and z i Is the feature output by the multi-layer perceptron module; A-MSA(·) represents the multi-head self-attention calculation, which generates a triple {key, query, value} for the feature sequence and uses three learnable parameter matrices W q 、W k And W v To calculate the triple {key, query, value}, and solve the interference problem in prostate segmentation by establishing the relationship between different objects (distinguishing interfering objects and the prostate) and different pixels within the target (identifying different prostate shapes), as shown in the following formula: Q = z i-1 W q ; K = z i-1 W k ; V = z i-1 W v ; ; Among them, Q, K, and V are the calculated query matrix, key matrix, and value matrix respectively, Softmax(·) is the Softmax activation function, d is the feature dimension, B is the deviation of the relative position, and T is the threshold.
[0025] Furthermore, when the scale coefficient τ is lower than the threshold T and the global self-attention calculation is performed, Q, K, V, B ∈ R HW ×d ; when the scale coefficient τ is higher than the threshold T and the local self-attention calculation is performed, Q, K, V, B ∈ R (M^2)×d .
[0026] S1.3. Pass the adaptive self-attention calculation results of the first N - 1 feature sequences (including global context information) to the corresponding level in the decoding stage through skip connections, and use the adaptive self-attention calculation results of the Nth feature sequence as the encoding result for output.
[0027] S2. In the decoding stage, as Figure 1 Shown in the right half: S2.1. Based on the encoding result and the adaptive self-attention calculation results of the first N - 1 feature sequences passed through skip connections, perform sampling and splicing level by level and reshape them into a feature map. Specifically: Repeat the following steps N - 1 times to obtain the reshaped feature map: Upsample the encoded result and concatenate it with the result of the adaptive self-attention calculation of the feature sequence of the previous dimension passed through the skip connection and transformed into a sequence format; perform convolution on the concatenated result to match the dimensions.
[0028] Through the local dominance of multiple convolutional layers, rich local features can be extracted. However, to avoid the neglect and destruction of global context information by consecutive convolutions and ensure that the network can not only reduce data requirements but also maintain a high ability to identify interference problems and various prostate shapes, the present invention adopts a skip connection method to receive global information from the encoding stage as an additional input during the decoding stage for full fusion, and reduces the fused feature dimension image, outputting a feature map containing rich global information and local information.
[0029] S2.2. Sequentially input the feature map obtained in step S2.1 into an upsampling layer and a convolutional layer for feature extraction to obtain the final segmentation result.
[0030] Furthermore, the convolutional layer used for convolving the concatenated result in step S2.1 and the convolutional layer in step S2.2 both introduce inductive biases of locality, spatial invariance, and hierarchical structure, and the inductive biases of convolutional layers at different levels are different to enhance the generalization ability of the entire model for TRUS image data.
[0031] As a preferred technical solution, N is set to 4.
[0032] Embodiment 2: As Figure 2 shown, in this embodiment, an intelligent segmentation system for transrectal ultrasound images of the prostate is provided, and the system includes an encoder and a decoder.
[0033] (a) The encoder includes an embedding layer and N region adaptive Transformer modules (also referred to as region adaptive converter modules); the decoder includes N - 1 Patch fusion modules (also referred to as block fusion modules), an upsampling layer, and a convolutional layer.
[0034] The embedding layer is used to extract N feature sequences of different sizes of the TRUS image to be segmented and connect them to N region adaptive Transformer modules respectively.
[0035] Furthermore, the embedding layer includes a segmentation layer, a linear projection layer, and N - 1 feature sequence layers connected in sequence; the segmentation layer is used to evenly segment the TRUS image to be segmented; the linear projection layer is used to project the image obtained by the segmentation layer onto the channels; the feature sequence layer includes a downsampling layer, a normalization layer, and a linear layer, and is used to extract feature sequences of different sizes.
[0036] The region adaptive Transformer module includes a multi-head self-attention (A-MSA) and a multi-layer perceptron (MLP) module, which are used to perform adaptive self-attention calculation on the feature sequence. Among them, the adaptive self-attention calculation results of the first N-1 feature sequences are transmitted to the corresponding level of the decoder through skip connections, and the adaptive self-attention calculation result of the Nth feature sequence is output as the encoding result; the specific adaptive self-attention calculation is as follows: calculate the scale coefficient τ of the feature sequence, and perform global self-attention calculation when the scale coefficient τ is lower than the threshold T, and perform local self-attention calculation when the scale coefficient τ is higher than the threshold T.
[0037] Furthermore, the region adaptive Transformer module is configured to perform the following calculations: ; where τ is the scale coefficient, which is used to represent the ratio of the computational complexity of global and local self-attention; H and W are the height and width of the feature map, C is the dimension of the feature map, and M is the window size; Furthermore, the multi-head self-attention is configured to perform the following calculations: ; ; where LN represents normalization, z i-1 represents the features output by the previous layer, represents the features output by the multi-head self-attention of the current layer; A-MSA(·) represents the multi-head self-attention calculation, Softmax(·) is the Softmax activation function, d is the feature dimension, B is the deviation of the relative position, T is the threshold, Q, K, and V are the query matrix, key matrix, and value matrix calculated respectively, and are calculated by using three learnable parameter matrices W q , W k and W v as follows: Q = z i-1 W q ; K = z i-1 W k ; V = z i-1 W v ; When the scale coefficient τ is lower than the threshold T and the global self-attention calculation is performed, Q, K, V, B ∈ R HW×d ; when the scale coefficient τ is higher than the threshold T and the local self-attention calculation is performed, Q, K, V, B ∈ R (M^2)×d ; Furthermore, the multi-layer perceptron module is configured to perform the following calculations: ; where MLP(·) represents the calculation performed by the multi-layer perceptron module, and z i is the feature output by the multi-layer perceptron module.
[0038] (b) The Patch fusion module is used to perform sampling and splicing level by level according to the coding result and the adaptive self-attention calculation result of the first N - 1 feature sequences transmitted through the skip connection, and reshape them into a feature map.
[0039] Furthermore, the Patch fusion module is provided with N - 1 connected in sequence, and each Patch fusion module is configured to perform the following steps: Upsample the coding result or the feature sequence output by the previous Patch fusion module, and splice it with the adaptive self-attention calculation result of the feature sequence of the previous dimension transmitted through the skip connection from the encoder and converted into the sequence format; process the spliced result through a convolutional layer to match the dimensions, and obtain the output of the Patch fusion module at this level.
[0040] The convolutional layer is used to extract features from the feature map to obtain the final segmentation result.
[0041] Furthermore, the convolutional layers all introduce inductive biases of locality, spatial invariance, and hierarchy, and the inductive biases of convolutional layers at different levels are different to enhance the generalization ability of the entire model for TRUS image data.
[0042] As a preferred technical solution, N is set to 4.
[0043] It should be noted here that the system provided in this embodiment is only illustrated by the above division of each functional module. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure is divided into different functional modules to complete all or part of the functions described above. This system can be applied to an intelligent segmentation method for prostate transrectal ultrasound images in the above embodiment.
[0044] Embodiment 3: To enable those skilled in the art to better understand the technical solutions of the present application and the excellent effects achieved, this embodiment uses the clinical prostate TRUS image dataset collected by the First Affiliated Hospital of Jinan University to train and evaluate the intelligent segmentation system for prostate transrectal ultrasound images in the above embodiment.
[0045] The clinical prostate TRUS image dataset contains 1000 TRUS images of 135 patients.
[0046] (1) Experimental running environment, running devices, and training parameter settings.
[0047] In this embodiment, relevant experiments were conducted in ubantu20.04.2, python3.8.12, and pytorch1.10.1. All training processes were run on an Nvidia GeForce RTX 2080Ti GPU with 11GB of memory. The learning rate was set to 0.01, the learning rate scheduler was the "Poly" decay strategy with a parameter power of 0.9, the optimizer was SGD, the momentum was set to 0.99, the weight decay was set to 3e-5, the number of training epochs was 1500, and the batch size was set to 8 / GPU. This embodiment used cross-entropy loss as the loss function.
[0048] (2) Comparison of experimental results.
[0049] In this embodiment, 6 advanced segmentation methods were compared on the dataset to verify the segmentation performance of the present invention (hereinafter referred to as TCF-Net). CNN-based segmentation methods include U-Net, UNet++, and Attention U-net (Att UNet). Transformer-based segmentation methods include UNext, TransUNet, and ScaleFormer.
[0050] (2.1) The comparative experimental results with other advanced segmentation methods are shown in Table 1. The metrics for measuring computational complexity are Params (Parameters) and FLOPs (Floating Point Operations), and the performance metrics and results for evaluating the segmentation model are mIoU (Mean Intersection Over Union), DSC (Dice similarity coefficient), and PA (Pixel Accuracy). Among them, the mean intersection over union (mIoU) metric of the TCF-Net of the present invention is 94.4%, the Dice similarity coefficient (DSC) is 95.5%, and the pixel accuracy (PA) is 97.9%, all of which are better than the 6 segmentation methods used as comparisons.
[0051] ; Table 1. Comparison of computational complexity and segmentation performance between the present invention and other segmentation methods.
[0052] (2.2) Visualization results. As Figure 3As shown, the key areas in the TRUS image are highlighted in a box to better visualize the differences in different segmentation predictions.
[0053] Based on the comparison of the experimental results in (2.1) and (2.2), the effectiveness of the present invention for prostate TRUS image segmentation is verified.
[0054] Example 4: This embodiment provides a storage medium storing a program, which when executed by a processor, implements an intelligent segmentation method for prostate transrectal ultrasound images in the above embodiment, specifically: Encoding stage: Segment the TRUS image to be segmented and project it onto the channels; Perform downsampling, normalization, and linear layer processing on the image projected onto the channels in sequence, and repeat N - 1 times to obtain N feature sequences of different sizes; Calculate the scale coefficient τ of each feature sequence, and perform adaptive self - attention calculation on the generated triplets of the feature sequences. Among them, when the scale coefficient τ is lower than the threshold T, global self - attention calculation is performed, and when the scale coefficient τ is higher than the threshold T, local self - attention calculation is performed; Transmit the adaptive self - attention calculation results of the first N - 1 feature sequences to the corresponding levels in the decoding stage through skip connections, and output the adaptive self - attention calculation result of the Nth feature sequence as the encoding result; Decoding stage: Based on the encoding result and the adaptive self - attention calculation results of the first N - 1 feature sequences transmitted through skip connections, perform sampling and splicing level by level and reshape them into a feature map; Input the feature map into an upsampling layer and a convolutional layer in sequence for feature extraction to obtain the final segmentation result.
[0055] Example 5: This embodiment provides a computer device, which includes: a memory and at least one processor. Instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the instructions in the memory so that the computer device executes an intelligent segmentation method for prostate transrectal ultrasound images as described in the above embodiment.
[0056] It should be understood that each part of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0057] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for intelligent segmentation of prostate transrectal ultrasound images, characterized in that: The steps include: Coding phase: Segment the TRUS image to be segmented and project it onto the channel; The image projected onto the channel is sequentially downsampled, normalized, and processed with a linear layer, and this process is repeated N-1 times to obtain N feature sequences of different sizes. Calculate the scale coefficient τ of each feature sequence, and perform adaptive self-attention calculation on the feature sequence generated triples, where global self-attention calculation is performed when the scale coefficient τ is lower than the threshold T, and local self-attention calculation is performed when the scale coefficient τ is higher than the threshold T; The adaptive self-attention calculation results of the first N-1 feature sequences are passed to the corresponding level of the decoding stage through skip connections, and the adaptive self-attention calculation results of the Nth feature sequence are output as the encoding result; Decoding stage: Based on the encoding results and the adaptive self-attention calculation results of the first N-1 feature sequences transmitted through the jump connection, the samples are spliced level by level and reshaped into a feature map; The feature map is sequentially input into the upsampling layer and the convolution layer for feature extraction to obtain the final segmentation result.
2. The method for intelligent segmentation of prostate transrectal ultrasound images according to claim 1, characterized in that: The calculation of the scale coefficient τ of each feature sequence is specifically: ; Among them, τ is the scale factor, which is used to characterize the ratio of the computational complexity of global and local self-attention; H and W are the height and width of the feature map, C is the dimension of the feature map, and M is the window size.
3. The method for intelligent segmentation of prostate transrectal ultrasound images according to claim 1, characterized in that: The adaptive self-attention calculation is specifically as follows: ; ; Among them, LN represents normalization, A-MSA(·) represents multi-head self-attention calculation, z i-1 Represents the features of the previous layer output, represents the features of the multi-head self-attention output of the current layer, MLP(·) represents the calculation performed by the multi-layer perceptron module, and z i is the feature output by the multi-layer perceptron module; the multi-head self-attention calculation is specifically as follows: Generate triples {key, query, value} for feature sequences and use three learnable parameter matrices W q , W k and W v To calculate the triple {key, query, value}, as follows: Q=z i-1 W q ; F=z i-1 IN k ; V=z i-1 W v ; ; where Q, K, and V are the calculated query matrix, key matrix, and value matrix, respectively, Softmax(·) is the Softmax activation function, d is the feature dimension, B is the deviation of the relative position, and T is the threshold.
4. The method for intelligent segmentation of prostate transrectal ultrasound images according to claim 3, characterized in that: When performing global self-attention calculation, Q,K,V,B∈R HW×d ; When performing local self-attention calculations, Q,K,V,B∈R (M ^2)×d , M is the window size.
5. The method for intelligent segmentation of prostate transrectal ultrasound images according to claim 1, characterized in that: The adaptive self-attention calculation results based on the encoding results and the first N-1 feature sequences transmitted through the skip connection are sampled and spliced level by level and reshaped into a feature map, specifically: Repeat the following steps N-1 times to get the reshaped feature map: The encoded result is upsampled and concatenated with the adaptive self-attention calculation result of the feature sequence of the previous dimension passed through the skip connection and converted into the sequence format; the concatenated result is convolved to match the dimension.
6. A prostate transrectal ultrasound image intelligent segmentation system, comprising an encoder and a decoder, characterized in that: The encoder includes an embedding layer and N region adaptive Transformer modules; the decoder includes a patch fusion module, an upsampling layer and a convolution layer connected in sequence; The embedding layer is used to extract N feature sequences of different sizes of the TRUS image to be segmented, and is respectively connected to N regional adaptive Transformer modules; The regional adaptive Transformer module includes a multi-head self-attention and multi-layer perceptron module connected in sequence, which is used to perform adaptive self-attention calculation on the feature sequence, wherein the adaptive self-attention calculation results of the first N-1 feature sequences are transmitted to the corresponding level of the decoder through a jump connection, and the adaptive self-attention calculation result of the Nth feature sequence is output as the encoding result; the adaptive self-attention calculation is specifically: calculating the scale coefficient τ of the feature sequence, performing global self-attention calculation when the scale coefficient τ is lower than the threshold T, and performing local self-attention calculation when the scale coefficient τ is higher than the threshold T; The Patch fusion module is used to perform sampling and splicing level by level according to the encoding result and the adaptive self-attention calculation results of the first N-1 feature sequences transmitted through the skip connection, and reshape them into a feature map; The convolution layer is used to extract features from the feature map to obtain the final segmentation result.
7. The intelligent segmentation system for prostate transrectal ultrasound images according to claim 6, characterized in that: The embedding layer includes a segmentation layer, a linear projection layer and N-1 feature sequence layers connected in sequence; the segmentation layer is used to evenly segment the TRUS image to be segmented; the linear projection layer is used to project the image obtained by the segmentation layer onto the channel; the feature sequence layer includes a downsampling layer, a normalization layer and a linear layer, which are used to extract feature sequences of different sizes.
8. The intelligent segmentation system for prostate transrectal ultrasound images according to claim 6, characterized in that: The Region Adaptive Transformer module is configured to perform the following computations: ; Where τ is the scale factor, which is used to characterize the ratio of the computational complexity of global and local self-attention; H and W are the height and width of the feature map, C is the dimension of the feature map, and M is the window size; The multi-head self-attention is configured to perform the following computations: ; ; Among them, LN represents normalization, z i-1 Represents the features of the previous layer output, represents the features of the multi-head self-attention output of the current layer; A-MSA(·) represents the multi-head self-attention calculation, Softmax(·) is the Softmax activation function, d is the feature dimension, B is the deviation of the relative position, T is the threshold, Q, K and V are the calculated query matrix, key matrix and value matrix respectively, by using three learnable parameter matrices W q , W k and W v To calculate, as follows: Q=z i-1 W q ; F=z i-1 IN k ; V=z i-1 W v ; When the scale factor τ is lower than the threshold T, the global self-attention calculation is performed, Q,K,V,B∈R HW×d ; When the scale factor τ is higher than the threshold T, the local self-attention calculation is performed, Q,K,V,B∈R (M^2)×d ; The Multilayer Perceptron module is configured to perform the following computations: ; where MLP(·) represents the computation performed by the multilayer perceptron module, z i The features output by the multilayer perceptron module.
9. The intelligent segmentation system for prostate transrectal ultrasound images according to claim 6, characterized in that: The patch fusion module is provided with N-1 sequentially connected ones, and each patch fusion module is configured to perform the following steps: The encoding result or the feature sequence output by the previous Patch fusion module is upsampled and concatenated with the adaptive self-attention calculation result of the feature sequence of the previous dimension from the encoder that is transmitted through the jump connection and converted into the sequence format; the concatenated result is processed through a convolutional layer to match the dimension to obtain the output of the Patch fusion module at this level.
10. The intelligent segmentation system for prostate transrectal ultrasound images according to claim 9, characterized in that: The convolutional layers all introduce locality, spatial invariance, and hierarchical inductive biases.
Citation Information
Patent Citations
Lightweight multi-scale feature fusion real-time image semantic segmentation method and system
CN114445430A
Medical image segmentation model and method based on connection Swin Transform path
CN114912575A
Prostate MRI (Magnetic Resonance Imaging) image segmentation method and system based on adaptive multi-scale Transform optimization
CN115272170A
Prostate segmentation algorithm based on window multi-head self-attention and deep convolution
CN118735941A
Abdomen multi-organ image segmentation method fusing multi-scale features
CN119205824A