Double-task ultrasonic scoliosis image segmentation method based on edge constraint
Through the dual-task ultrasonic scoliosis image segmentation method based on edge constraints, the EDGE-Net network and Gaussian filter combined with a dual attention mechanism is used to solve the problem of inaccurate scoliosis segmentation in small samples and low-quality ultrasonic images, and achieve higher accuracy image segmentation.
Patent Information
- Application Number
- CN202510382542.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
AI Technical Summary
In the case of a small number of samples, it is difficult for the prior art to achieve accurate segmentation of scoliosis in ultrasound images, especially in the case of low noise and image quality, resulting in large measurement errors.
A dual-task ultrasonic scoliosis image segmentation method based on edge constraints is adopted, and feature extraction is performed through the EDGE-Net network, combined with Gaussian filtering and dual attention mechanism, and feature enhancement is used to achieve image segmentation.
In small samples and low-quality ultrasound images, the segmentation accuracy of scoliosis is improved, the noise impact is reduced, and more accurate segmentation results are provided.
Smart Images

Figure CN120355731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision image processing, and more specifically, to a dual-task ultrasonic scoliosis image segmentation method based on edge constraints. Background Art
[0002] Idiopathic scoliosis (AIS) is a condition where the spinal curvature is greater than 10°. School-age children and adolescents are the most common population affected by scoliosis, affecting 0.47%-5.2% of adolescents. During the rapid growth period, the risk of AIS deterioration is high, and severe scoliosis can lead to lung diseases and even affect future life. Full-spine standing X-ray radiography is the gold standard for detecting and monitoring the progression of scoliosis, but the frequency of routine evaluation is usually 3-6 months. The cumulative X-ray radiation dose may lead to an increased incidence of cancer, such as breast cancer, leukemia, etc. Therefore, it is recommended to use low-radiation or non-radiation methods to monitor scoliosis, especially for children and adolescents. Ultrasound is a real-time, non-radiation, cost-effective alternative to X-rays.
[0003] In the traditional clinical method of diagnosing scoliosis using ultrasound, there are mainly two methods for estimating the scoliosis angle by manual measurement. First, in the coronal view, identify the inflection point of the middle dark spinal profile, mark the most tilted vertebra, draw a short line at the inflection point, and measure the curvature angle accordingly. Second: Using the ribs, thoracic transverse processes, and lumbar features of the ultrasound image, after determining the most tilted vertebral area, draw a line through the middle to the most tilted transverse process to evaluate the spinal curvature. However, ultrasound images are prone to poor image quality due to the patient's own or equipment factors, such as noise, speckle, etc. Directly measuring the angle will cause relatively large errors. Therefore, using the segmented image for measurement will be relatively clear and accurate, bringing convenience to both doctors and patients. In the field of deep learning, a large number of samples are generally required for the model to learn and extract features, but in the medical field, issues such as patient privacy and international conventions are involved. Therefore, the number of samples is usually very limited, and the number of labels marked by experts is even less. Therefore, a small number of high-quality samples are crucial, which can not only save costs but also obtain relatively accurate training results.
[0004] Therefore, in view of the above problems, a segmentation algorithm is needed to achieve the complete segmentation of vertebrae with a small number of samples, and at the same time, accurate segmentation of ultrasound images can be achieved according to the corresponding edge constraints. Summary of the Invention
[0005] In view of this, the present invention provides a dual-task ultrasonic scoliosis image segmentation method based on edge constraints, which can achieve the complete segmentation of vertebrae with a small number of samples, and at the same time, accurate segmentation of ultrasound images can be achieved according to the corresponding edge constraints.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A dual-task ultrasound scoliosis image segmentation method based on edge constraint, comprising:
[0008] Preprocess the ultrasound image;
[0009] The preprocessed ultrasound image that meets the requirements is input into the EDGE-Net network for four-layer downsampling, and four-layer feature maps are obtained in sequence. The four-layer feature maps are used as the sampling feature stream and the boundary label stream respectively;
[0010] Take the first three layers of sampling feature maps and the corresponding upsampled feature maps of the first three layers of boundary label streams as inputs, calculate the local mean, local covariance and local variance of the image using a Gaussian filter, and calculate the linearly constrained output feature map according to the local mean, local covariance and local variance of the image;
[0011] Perform dynamic filtering feature enhancement based on the upsampled feature maps of the first three layers of sampling feature streams, the upsampled feature maps of the first three layers of boundary label streams, and the linearly constrained output feature map, and output the dynamically filtered feature enhancement feature map; this step highlights the high-frequency information through pixel-by-pixel subtraction, enhances the low-frequency features and high-frequency features of the feature map respectively, and supplements the missing information in the sampling process. Therefore, it helps to identify the details of image segmentation, highlight the image boundary, and there are also skip connection parts, that is, in the upsampling process of each layer, the corresponding feature maps of downsampling will also be concatenated in the channel dimension to supplement the high-level and global features;
[0012] Based on the fourth layer of sampling feature map and the fourth layer of sampling feature map that has undergone channel direction information separation and global average pooling, obtain the aggregated direction feature map, and use the dual attention mechanism to perform direction feature extraction to obtain the direction feature map.; this step utilizes the directional information in the potential direction space of medical images, and combines the spinal directional information and semantic information to more effectively assist image segmentation;
[0013] Obtain the final segmentation prediction map based on the linearly constrained output feature map, the dynamically filtered feature enhancement feature map, the direction feature map, and the aggregated direction feature map.
[0014] Preferably, performing dynamic filtering feature enhancement based on the upsampled feature maps of the first three layers of sampling feature streams, the upsampled feature maps of the first three layers of boundary label streams, and the linearly constrained output feature map, and outputting the dynamically filtered feature enhancement feature map, specifically including:
[0015] Step 1: The linearly constrained output feature map The upsampled feature maps c1, c2, c3 of the first three layers of the sampling feature stream and the upsampled feature maps e1, e2, e3 of the first three layers of the boundary label stream are used as inputs for dynamic filtering feature enhancement. First, the high-frequency information is filtered using a Gaussian convolution kernel, and the size of the Gaussian convolution kernel is as follows:
[0016]
[0017] Convolve each pixel of the upsampled maps c1, c2, c3 of the first three layers of the sampling feature stream using the Gaussian kernel to smooth the image and enhance the low-frequency features of the image, obtaining the feature maps H F1 、H F2 、H F3 , and use linear constraints to output the feature map Subtract the corresponding feature map H F1 、H F2 、H F3 to obtain the difference feature maps D1, D2, D3. The calculation formula is as follows:
[0018] D m =F GUIm -H Fm
[0019] where m represents the m-th layer of the feature map, and m takes values of 1, 2;
[0020] Step 2: Perform feature enhancement processing on the linearly constrained output feature map . The calculation formulas for enhancing the low-frequency features L1, L2, L3 are as follows:
[0021] L n =Concate{Up[W(L n-1 、L n-2 、L n-3 、L n-4 )]}
[0022] where W is the adaptive average pooling convolution, the convolution kernel sizes are 1x1, 2x2, 3x3, 6x6, and 9x9, Up is the upsampling operation, Concate represents concatenation in the channel dimension, n refers to one of the four divisions of each layer of the low-frequency features L1, L2, L3, and n takes values of 1, 2, 3, 4;
[0023] Enhance the high-frequency features H1, H2, H3. Based on the division of the feature map channels, perform random division, divide each layer of the high-frequency features H1, H2, H3 into four parts, use dilated convolutions with different scales for each part of the high-frequency features, and perform convolution in the form of grouped convolution during the convolution process, and perform concatenation in the channel dimension on the respectively enhanced high-frequency features;
[0024] After enhancing the high-frequency features and low-frequency features, a feature enhancement map is obtained;
[0025] The difference feature maps D1, D2, D3, the first three layers of boundary label upsampling feature maps e1, e2, e3 and feature enhancement maps En1, En2, En3 are spliced in channel dimension to obtain the dynamic filtering feature enhancement map D f1 , D f2 , D f3 .
[0026] Preferably, based on the fourth layer sampling feature map and the fourth layer sampling feature map after channel direction information separation and global average pooling, an aggregated directional feature map is obtained, and a dual attention mechanism is used to extract directional features to obtain the directional feature map, which specifically includes:
[0027] Step a: For the fourth layer sampling feature map F4, initialize the feature list of eight directions: up, down, left, right, upper left, upper right, lower left, and lower right, and extract features according to categories. For each pixel value of each category feature, translate it up, down, left, right, upper left, upper right, lower left, and lower right once, respectively. Use matrix multiplication to calculate whether the category feature after translation and the original category feature still belong to the same category, so as to separate the direction information later. Calculate the directional connectivity loss in turn, perform channel direction information separation, and perform global average pooling to obtain the aggregated directional feature map. Perform the following operations on F4 and the aggregated directional feature map V. pri Perform slicing operations, i.e. F4 and V pri When combined into a pair, the pair of images will be randomly divided into slices according to the channel dimension to form eight groups of slices. The dual attention module is used for them, and then the features of the two groups of slices are fused as a whole. Then, they are spliced into a feature map of separated directional information according to the channel dimension to construct the directional feature map Dir;
[0028] Step b: Perform directional voting on the directional feature map Dir, that is, calculate which direction each pixel of each category feature map has the greatest probability of belonging to, and obtain the accurate directional feature map Dir max , through the directional feature map Dir max Calculate the DICE loss function with the true label and use the direction information to constrain the segmentation:
[0029]
[0030] Among them, TP is the number of pixels whose true label and prediction result are both positive samples, FP is the number of pixels whose true label is negative samples but the prediction result is positive samples, and FN is the number of pixels whose true label is positive samples but the prediction result is negative samples.
[0031] Preferably, the dual attention mechanism includes performing channel attention mechanism and spatial attention mechanism operations:
[0032]
[0033] Among them, represents a learnable weight matrix, represents matrix multiplication, concate represents concatenation in the channel dimension, m represents one of eight parts, Q m , K m , V m represent the query value, key value, and value generated based on the aggregated direction feature map and the fourth-layer downsampled feature map, CAM represents the channel attention mechanism, and represents the spatial attention mechanism.
[0034] Preferably, the calculation formula for the Gaussian kernel of the Gaussian filter is:
[0035]
[0036] Among them, G'(x,y) represents the Gaussian kernel value between two vectors x and y, ||x - y|| 2 represents the square of the Euclidean distance between vectors x and y, and σ is the standard deviation of the Gaussian kernel.
[0037] Preferably, preprocessing the ultrasonic image specifically includes:
[0038] Unifying the size of the ultrasonic image and enhancing the contrast.
[0039] Preferably, unifying the size of the ultrasonic image specifically includes:
[0040] Obtaining the maximum size among all ultrasonic images;
[0041] Padding all ultrasonic images with 0 bits and unifying them to the maximum size;
[0042] Performing proportional reduction on all padded ultrasonic images to obtain a set of ultrasonic images with unified size.
[0043] Preferably, contrast enhancement is performed through the contrast adaptive histogram equalization image processing technique, and the calculation formula is:
[0044] s k = T(r k ) = (L - 1)·CDF(r k )
[0045] Among them, r k is the original gray value, s k is the new gray value after conversion, L is the total number of gray levels, CDF(r k ) is the gray value rk Cumulative distribution function
[0046] According to the above technical solution, compared with the prior art, the present invention discloses a dual-task ultrasonic scoliosis image segmentation method based on edge constraints, which forms a dual-task stream by combining the boundary labels of ultrasonic images, uses linear constraints to guide image segmentation, combines Gaussian filtering at the same time to enhance low-frequency information, obtains edge differences, combines directionality to supervise deep features, and obtains a segmentation image with higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0048] Figure 1 It is a flowchart of a dual-task ultrasonic scoliosis image segmentation method based on edge constraints provided by the present invention.
[0049] Figure 2 It is a structural diagram of the EDGE-Net model provided by the present invention.
[0050] Figure 3 It is a flowchart of feature enhancement provided by the present invention.
[0051] Figure 4 It is a flowchart of direction supervision provided by the present invention.
[0052] Figure 5 It is a schematic diagram of specific pixel translation provided by the present invention. The color is abstractly used to represent the class of the spine. In the figure, green is the same class of spine features, and red is the same class of spine features, which can represent combining direction information. After translation, the original judged class probability P. If it still belongs to the same class, it is recorded as 1, otherwise it is recorded as 0. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0054] The embodiment of the present invention discloses a dual-task ultrasonic scoliosis image segmentation method based on edge constraints, as Figure 1 shown, including:
[0055] Preprocess the ultrasonic image;
[0056] The preprocessed ultrasonic image that meets the requirements is input into the EDGE-Net network for four-layer downsampling, and four-layer downsampling feature maps are obtained in sequence. The four-layer downsampling feature maps are used as the sampling feature stream and the boundary label stream respectively;
[0057] Take the first three-layer downsampling feature maps and the corresponding upsampled feature maps of the first three-layer boundary label streams as inputs, use a Gaussian filter to calculate the local mean, local covariance, and local variance of the image, and calculate the linearly constrained output feature map based on the local mean, local covariance, and local variance of the image;
[0058] Perform dynamic filtering feature enhancement based on the upsampled feature maps of the first three-layer sampling feature streams, the upsampled feature maps of the first three-layer boundary label streams, and the linearly constrained output feature map, and output the dynamically filtered feature enhancement feature map;
[0059] Based on the fourth-layer downsampling feature map and the fourth-layer downsampling feature map after channel direction information separation and global average pooling, obtain the aggregated direction feature map, and use the dual attention mechanism to extract the direction feature to obtain the direction feature map;
[0060] Based on the linearly constrained output feature map, the dynamically filtered feature enhancement feature map, the direction feature map, and the aggregated direction feature map, obtain the final segmentation prediction map.
[0061] The purpose of the present invention is to propose a dual-task technical solution for reducing the negative impact of noise and improving the edge segmentation accuracy for low-quality ultrasonic images with noise. One task calculates the DICE loss mainly based on the sampling feature map, and the other task calculates the loss mainly based on the boundary label.
[0062] Furthermore, the ultrasonic image preprocessing is divided into two steps: image size unification and contrast enhancement. The specific process is as follows:
[0063] Step 1.1.1, for the original ultrasonic image κ i , where the value range of i is the size of the entire dataset. In this embodiment, i = 109, perform size unification. The original images have different sizes, about 640 * 2900, and form a set {κ i \ κ1, κ2,......}, and uniformly normalize them to 640 * 160. Based on this, the specific process is as follows:
[0064] First, find the largest size М among all ultrasonic images, such as 640 * 3200;
[0065] For all {κ i \ κ1, κ2,......}, fill them with 0 bits and unify them to 640 * 3200;
[0066] After that, the filled image {κ i \κ1, κ2,......} is scaled down proportionally, with a width of 640, a scaling factor β = 0.25, a height of 3200, and a scaling factor α = 0.2, to obtain an ultrasound image set {κ i \κ1, κ2,......} with a unified size of 640*160.
[0067] Step 1.1.2: Enhance the contrast of the ultrasound image set {κ i \κ1, κ2,......} obtained in Step 1.1.1 using the Contrast Limited Adaptive Histogram Equalization (CLAHE) image processing technique to improve the local contrast of the image. The enhancement amplitude of the local contrast is restricted by calculating the height of the local histogram for each pixel value, thereby restricting the amplification of noise and over-enhancement of the local contrast. The formula for calculating the pixel value is shown as follows:
[0068] s k = T(r k ) = (L - 1)·CDF(r k )
[0069] where r k is the original gray value, s k is the new gray value after conversion, L is the total number of gray levels (usually 256), and CDF(r k ) is the cumulative distribution function of the gray value r k .
[0070] After the above steps, the preprocessed image Μ k is obtained, where the value of k is still the length of the data set, 109.
[0071] Furthermore, the ultrasound images that meet the requirements after preprocessing are input into the EDGE-Net model. As Figure 2 shown, first, EDGE-Net performs a double-layer convolution on the processed preprocessed image using a 3x3 convolution kernel to obtain shallow features. Then, after four downsamplings, that is, using the max pooling operation to extract features from global to local, four downsampled feature maps are obtained.
[0072] According to the downsampled features, a dual-task flow is constructed, that is, there are two upsampling flows. The four-layer feature maps obtained in one task path are the sampled feature flows, named UP c1 , UP c2 , UP c3 , UP c4 respectively, and the obtained feature maps are named c1, c2, c3, c4. The four-layer feature maps obtained in the other task path are the boundary label flows, named UPe1 , UP e2 , UP e3 , UP e4 Named, the obtained feature maps are named e1, e2, e3, e4. The output of the boundary label stream only provides boundary information and is not used as the main evaluation criterion. Next, different steps are implemented according to different streams:
[0073] First, perform linear constraints on the boundary labels. The boundary labels are manually marked by experts. Use the three downsampled feature maps F1, F2, F3 obtained by successive downsampling above and the three boundary label feature maps e1, e2, e3 obtained by upsampling the boundary label stream as inputs. Use a Gaussian filter to calculate the local mean, local covariance, and local variance of the image. The local window size is r. Use these values to obtain the coefficients mean a and mean b , construct a linear function expression, and mutually constrain to guide the boundary segmentation result to be more accurate. The calculation formula of the Gaussian kernel is as follows:
[0074]
[0075] Among them, G'(x,y) represents the Gaussian kernel value between two vectors x and y, ||x - y|| 2 represents the square of the Euclidean distance between vectors x and y. σ is the standard deviation of the Gaussian kernel, which determines the influence range of data points. A larger σ means a wider influence range for each data point, while a smaller σ means a smaller influence range. This formula describes the similarity measure between two points. When the two points are very close, the value of the kernel function approaches 1; when the two points are far apart, the value of the kernel function approaches 0. Therefore, the Gaussian kernel can effectively measure the similarity between two points in a high-dimensional space, even if these points are not linearly separable in the original input space.
[0076] Finally, after the above linear constraints, calculate the linearly constrained output feature map F GUIl :
[0077] b l = G'(e l , r) (l = 1, 2, 3)
[0078] mean bl = G'(b l , r) (l = 1, 2, 3)
[0079] F GUIl = mean al ·I + mean bl (l = 1, 2, 3)
[0080] Among them, mean alis the linear coefficient obtained from the local mean, local covariance, and local variance calculated by the Gaussian kernel in the above text, mean bl First, according to the boundary feature map e input by the boundary flow l calculate the local mean b, and then use the Gaussian kernel and the local window r to calculate the average value of b to obtain mean bl , where Ι is the boundary label, l is the feature map of the l-th layer, and the boundary label is manually marked by professional doctors.
[0081] In this embodiment, the dynamic filtering enhancement process includes:
[0082] Step 1: Use the linearly constrained output feature map The upsampled feature maps c1, c2, c3 of the first three layers of sampled feature flows and the upsampled feature maps e1, e2, e3 of the first three layers of boundary labels are used as inputs for dynamic filtering feature enhancement. First, use the Gaussian convolution kernel to filter high-frequency information to solve the Gaussian noise and speckle noise existing in the image. The Gaussian kernel size is as follows:
[0083]
[0084] Convolve each pixel of the upsampled feature maps c1, c2, c3 of the first three layers of sampled feature flows with the Gaussian kernel to obtain the low-frequency enhanced feature maps H F1 、H F2 、H F3 , and then use subtract the corresponding feature maps H F1 、H F2 、H F3 to obtain the difference feature maps D1, D2, D3, which are expressed by the formula as follows:
[0085] D m =F GUIm -H Fm
[0086] where m represents the feature map of the m-th layer, and m takes values of 1, 2, 3;
[0087] Step 2: Continue to perform feature enhancement processing on the linearly constrained output feature map For the most original image, information loss is relatively serious after operations such as downsampling and high-frequency information filtering. Therefore, feature enhancement is performed, and the enhanced features are enhanced separately according to low-frequency features and high-frequency features. Here, L1, L2, L3 and H1, H2, H3 are used to represent the low-frequency features and high-frequency features of
[0088] such as Figure 3As shown in the figure, the low-frequency features L1, L2, and L3 are enhanced. The low-frequency feature maps are divided into four parts based on channels. Each part uses adaptive average pooling convolutions of different sizes to enhance the global-to-local low-frequency features. Finally, the original feature maps L1, L2, and L3 are restored using channel concatenation. The above process is shown by the following formula, where W is the adaptive average pooling convolution with kernel sizes of 1x1, 2x2, 3x3, 6x6, and 9x9, Up is upsampling after pooling, and finally Concate is used for concatenation in the channel dimension. Here, n refers to one of the four parts after each layer of low-frequency features L1, L2, and L3 are divided, and the value of n is 1, 2, 3, or 4:
[0089] L n = Concate{Up[W(L n-1 、L n-2 、L n-3 、L n-4 )]}
[0090] The high-frequency features H1, H2, and H3 are enhanced. First, they are also divided into four parts based on channels. For each part of the high-frequency features, dilated convolutions of different scales are used, and during the convolution process, convolution is performed in the form of grouped convolution. The kernel sizes are 3x3, 5x5, 7x7, and the kernel size is 1; finally, the enhanced high-frequency features are concatenated in the channel dimension. After the high-frequency and low-frequency features are enhanced respectively, the enhanced high-frequency and low-frequency features are fused into the linearly constrained output feature map using per-pixel multiplication and pointwise summation to obtain the feature enhancement maps En1, En2, and En3.
[0091] The reason it is called dynamic filtering is that the enhancement of high-frequency and low-frequency features updates the weights with each backpropagation, rather than fixing the weights in advance. Finally, the difference feature maps Dm, the upsampled feature maps of the first three layer boundary labels, and the feature enhancement maps En1, En2, and En3 are concatenated in the channel dimension to integrate the extracted information and obtain the final output dynamic filtering feature enhancement maps D f1 、D f2 、D f3 .
[0092] The above steps only consider semantic features, that is, category features, and ignore the directionality of image features. Especially for spinal ultrasound images, the orientation of the spine in the image reflects category information from the side. If the important influence of direction information is ignored, accurate results cannot be obtained. Therefore, the deepest downsampled feature map F4 in the downsampled features is used as the input for directional relay supervision, as Figure 4 shown, specifically:
[0093] Step a: Initialize the feature lists in eight directions, namely up, down, left, right, upper left, upper right, lower left, and lower right, for F4. Then extract features according to categories. The category of this segmentation task is four, and these features are Lab1, Lab2, Lab3, and Lab4. Using the directional connectivity loss, for each pixel value of each category feature, translate it upward, downward,..., downward right once respectively, and use matrix multiplication to calculate whether the category feature after translation and the original category feature still belong to the same category, which is convenient for separating directional information later. Calculate the directional connectivity loss, perform channel directional information separation, and global average pooling in turn to obtain the aggregated directional feature map V pri , downsample the feature map F4 of the fourth layer and the aggregated directional feature map V pri Perform slicing operations, that is, F4 and V pri Are combined into a pair, and this pair of images will be randomly divided into slices according to the channel dimension, forming eight groups of slices, and the double attention mechanism will be used for each of the 8 groups of slices, and they will be spliced according to the channel dimension to form a feature map for separating directional information, so that the directional feature map Dir can be constructed, as Figure 5 Shown
[0094] Step b: Perform directional voting on the directional feature map Dir, that is, calculate which direction each pixel of each category feature map belongs to with the greatest possibility, and then obtain the accurate directional feature map Dir max , use this determined directional feature map Dir max Calculate the DICE loss function with the true label. The DICE loss is a relatively popular loss function in medical image segmentation, and its loss function expression is shown as follows:
[0095]
[0096] Among them, TP is the number of pixels where both the true label and the prediction result are positive samples, FP is the number of pixels where the true label is a negative sample but the prediction result is a positive sample, and FN is the number of pixels where the true label is a positive sample but the prediction result is a negative sample. Calculate the loss function to guide the forward propagation;
[0097] In addition to using the deepest feature map F4 as the directional relay supervision output, input the aggregated directional feature map V pri And the deepest feature map F4 into the double attention mechanism module. The specific steps are as follows:
[0098] This double attention mechanism module is to better extract information and correspond to the eight directional features. It has two types of feature maps. One type refers to the last upsampled feature map F4 of the sampling feature stream, and the other type is the aggregated directional feature map V pri, it is randomly divided into eight parts by channel, and the attention mechanism is used. This attention mechanism combines the advantages of self-attention mechanism, spatial attention mechanism and channel attention mechanism. First, Q query, K key and V value are generated according to two types of feature maps, then the matrix multiplication is calculated for the respective Q and K, and then the softmax is used for each element of the matrix to obtain the probability value. After that, it is multiplied by V to obtain the attention feature map. Finally, these eight feature maps are concatenated according to the channel dimension. The above process can be expressed by the following formula:
[0099]
[0100] Among them, represents the learnable weight matrix, represents matrix multiplication, concate represents concatenation in the channel dimension, m represents one of the eight parts, Q m 、K m 、V m represent the query, key and value generated according to two types of feature maps, CAM represents the channel attention mechanism, and SAM represents the spatial attention mechanism.
[0101] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0102] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A dual-task ultrasound scoliosis image segmentation method based on edge constraints, characterized in that Including: Preprocess the ultrasonic image; The preprocessed ultrasonic image that meets the requirements is input into the EDGE-Net network for four-layer downsampling to obtain four-layer downsampled feature maps in sequence. The four-layer downsampled feature maps are used as the sampling feature stream and the boundary label stream respectively; Take the first three layers of downsampled feature maps and the corresponding upsampled feature maps of the first three layers of boundary label streams as inputs, use a Gaussian filter to calculate the local mean, local covariance and local variance of the image, and calculate the linearly constrained output feature map according to the local mean, local covariance and local variance of the image; Perform dynamic filtering feature enhancement based on the upsampled feature maps of the first three layers of sampling feature streams, the upsampled feature maps of the first three layers of boundary label streams, and the linearly constrained output feature map, and output the dynamically filtered feature enhancement feature map; Obtain the aggregated direction feature map based on the fourth-layer downsampled feature map and the fourth-layer downsampled feature map after channel direction information separation and global average pooling. Use the dual attention mechanism to perform direction feature extraction to obtain the direction feature map; Obtain the final segmentation prediction map based on the linearly constrained output feature map, the dynamically filtered feature enhancement feature map, the direction feature map, and the aggregated direction feature map.
2. The dual-task ultrasound scoliosis image segmentation method based on edge constraint according to claim 1, wherein, Perform dynamic filtering feature enhancement based on the upsampled feature maps of the first three layers of sampling feature streams, the upsampled feature maps of the first three layers of boundary label streams, and the linearly constrained output feature map, and output the dynamically filtered feature enhancement feature map, specifically including: Step 1: Sequentially take the linear constraint output feature maps F GUI1 , F GUI2 , F GUI3 , the upsampled feature maps c1, c2, c3 of the first three layers of the sampling feature stream and the upsampled feature maps e1, e2, e3 of the first three layers of the boundary label stream as inputs for dynamic filtering feature enhancement. First, use a Gaussian convolution kernel to filter high-frequency information. The size of the Gaussian convolution kernel is as follows: Convolve each pixel of the upsampled feature maps c1, c2, and c3 of the first three layers of sampled features using a Gaussian kernel to smooth the image and enhance the low-frequency features of the image, obtaining the feature map H that filters out the high frequencies F1 、H F2 、H F3 , and use linear constraints to output the feature map F GUI1 、F GUI2 、F GUI3 Subtract the corresponding feature map H F1 、H F2 、H F3 to obtain the difference feature maps D1, D2, and D3. The calculation formulas are as follows: D m = F GUIm - H Fm Where m represents the m-th layer feature map, and m takes values of 1, 2, 3; Step 2: Perform feature enhancement on the linear constraint output feature maps F GUI1 , F GUI2 , F GUI3 Perform feature enhancement processing, where the enhanced features are enhanced separately according to low-frequency features and high-frequency features. The calculation formulas for enhancing the low-frequency features L1, L2, and L3 are as follows: L n = Concat{Up[W(L n-1 、L n-2 、L n-3 、L n-4 )]} Where W is the adaptive average pooling convolution, the convolution kernel sizes are 1x1, 2x2, 3x3, 6x6 and 9x9, Up is the upsampling operation, Concat represents concatenation in the channel dimension, and n refers to one of the four divisions of each layer of low-frequency features L1, L2, L3. The value of n is 1, 2, 3, 4; Enhance the high-frequency features H1, H2, H3. Based on the division of the feature map channels, perform random division. Divide each layer of high-frequency features H1, H2, H3 into four parts. For each part of the high-frequency features, use dilated convolutions with different scales, and perform convolution in the form of grouped convolution during the convolution process. Concatenate the enhanced high-frequency features in the channel dimension; Obtain the feature enhancement maps En1, En2, En3 after enhancing the high-frequency features and low-frequency features; Concatenate the difference feature maps D1, D2, D3, the upsampled feature maps e1, e2, e3 of the first three-layer boundary labels, and the feature enhancement maps En1, En2, En3 in the channel dimension to obtain the output dynamic filtering feature enhancement map D f1 , D f2 , D f3 .
3. A dual-task ultrasonic scoliosis image segmentation method based on edge constraints according to claim 1, characterized in that Obtain the aggregated direction feature map based on the fourth-layer downsampled feature map and the fourth-layer downsampled feature map after channel direction information separation and global average pooling. Use the dual attention mechanism to perform direction feature extraction to obtain the direction feature map, specifically including: Step a: Initialize the feature lists in eight directions, namely up, down, left, right, upper left, upper right, lower left, and lower right, for the fourth-layer downsampled feature map F4. Extract features by category. Using the directional connectivity loss, for each pixel value of each category feature, translate it upward, downward, leftward, rightward, upper leftward, upper rightward, lower leftward, and lower rightward once respectively. Use matrix multiplication to calculate whether the translated category feature and the original category feature still belong to the same category. Calculate the directional connectivity loss, perform channel-direction information separation, and global average pooling in sequence to obtain the aggregated directional feature map. Slice the fourth-layer downsampled feature map F4 and the aggregated directional feature map V pri respectively. For each slice of the aggregated directional feature map and each slice of the fourth-layer downsampled feature map, use the dual attention mechanism to splice them along the channel dimension into a feature map that separates directional information, and construct the directional feature map Dir; Step b: Perform direction voting on the direction feature map Dir, that is, calculate which direction each pixel of each class feature map is most likely to belong to, and obtain the accurate direction feature map Dir max , through the direction feature map Dir max Calculate the DICE loss function with the true label, and use the direction information to constrain the segmentation: Where TP is the number of pixels where both the true label and the prediction result are positive samples, FP is the number of pixels where the true label is negative but the prediction result is positive, and FN is the number of pixels where the true label is positive but the prediction result is negative.
4. A dual-task ultrasound scoliosis image segmentation method based on edge constraints according to claim 3, characterized in that, The dual attention mechanism includes performing channel attention mechanism and spatial attention mechanism operations: Among them, represents a learnable weight matrix, represents matrix multiplication, concat represents concatenation in the channel dimension, m represents one of eight parts, Q m , K m and V m represent query values, key values, and values generated from the aggregated direction feature map and the fourth-layer downsampled feature map, CAM represents the channel attention mechanism, and SAM represents the spatial attention mechanism.
5. A dual-task ultrasound scoliosis image segmentation method based on edge constraints according to claim 1, characterized in that, The calculation formula of the Gaussian kernel of the Gaussian filter is: Among them, G'(x, y) represents the Gaussian kernel value between two vectors x and y, ||x - y|| 2 represents the square of the Euclidean distance between vectors x and y, and σ is the standard deviation of the Gaussian kernel.
6. A method for dual-task ultrasonic scoliosis image segmentation based on edge constraint according to claim 1, characterized in that, Preprocess the ultrasonic image, specifically including: Unify the size of the ultrasonic image and enhance the contrast.
7. A dual-task ultrasonic scoliosis image segmentation method based on edge constraints according to claim 6, characterized in that, Unify the size of the ultrasonic image, specifically including: Obtain the maximum size among all ultrasonic images; Pad all ultrasound images with 0 bits and unify them to the maximum size; Scale down all the padded ultrasound images proportionally to obtain a set of ultrasound images with unified size.
8. A method for dual-task ultrasound scoliosis image segmentation based on edge constraints according to claim 6, characterized in that Perform contrast enhancement through the contrast adaptive histogram equalization image processing technique. The calculation formula is as follows: s k = T(r k ) = (L - 1)·CDF(r k ) where r k is the original gray value, s k is the new converted gray value, L is the total number of gray levels, CDF(r k ) is the cumulative distribution function of the gray value r k .