Feature Alignment Tongue Image Segmentation Method and System Based on Improved UNet++

By introducing feature alignment module and morphological processing layer in UNet++ network, the problem of feature misalignment in tongue image segmentation is solved, and the accuracy and edge smoothness of tongue image segmentation are improved.

CN115719352BActive Publication Date: 2025-07-08UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211426206.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-07-08
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

In the existing tongue image segmentation method, the encoding network and the decoding network fail to effectively consider the relationship between high-level features and low-level features in the feature downsampling and upsampling process, resulting in image features being misaligned, the tongue edges are inaccurate, and the segmentation accuracy is not high.

Method used

Feature alignment network is introduced in the UNet++ network architecture. Feature alignment module is used to supervise feature offsets during downsampling at different depths, and combine the morphological processing layer to optimize segmentation results to solve the feature misalignment problem.

Benefits of technology

It improves the accuracy of tongue image segmentation, ensures smooth edges of tongue, and achieves efficient tongue image segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719352B_ABST
    Figure CN115719352B_ABST
Patent Text Reader

Abstract

The present invention discloses a feature alignment tongue image segmentation method and system based on improved Unet++. The method includes: obtaining an original tongue image, and performing marking and preprocessing on the original tongue image; constructing a feature alignment tongue image segmentation model according to the preprocessed tongue image and performing model training to obtain a trained feature alignment tongue image segmentation model; optimizing the feature alignment tongue image segmentation model through morphological processing to obtain an optimized tongue image segmentation model; and using the optimized tongue image segmentation model to perform segmentation processing on the tongue image to be segmented to obtain a segmentation result. The present invention applies a feature alignment network to the Unet++ network architecture, adds different feature alignment modules to the downsampling processes at different depths, and successively supervises the offset of features in the downsampling process, solves the problem of feature misalignment in the tongue image segmentation process, fully improves the accuracy of tongue image segmentation, and optimizes the segmentation result through a morphological processing layer to obtain a final result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically relates to a feature alignment tongue image segmentation method and system based on improved UNet++. Background Technique

[0002] The "four diagnostic methods" in traditional Chinese medicine are inspection, auscultation and olfaction, inquiry, and palpation. Tongue diagnosis, as an important part of inspection, is an important basis for traditional Chinese medicine diagnosis. With the development of the Internet, the development speed of the diagnosis and treatment plan of Internet + healthcare has accelerated, and gratifying performance has been achieved. Among them, tongue image recognition, as a non-invasive detection, can effectively reduce the pain of the human body and has great significance in auxiliary detection.

[0003] As the basis of tongue image recognition, the accuracy of tongue image segmentation affects the accuracy of tongue image recognition. In the initial stage of development, the segmentation method based on color decomposition and threshold technology has a high segmentation efficiency and can accurately segment the tongue images collected by the tongue image device. Due to complex backgrounds and noise interference, deep learning image segmentation methods have also become important methods for tongue image segmentation.

[0004] The patent disclosure text with the application number "CN201710498517.1" and the name "An Automatic Segmentation Method for Traditional Chinese Medicine Tongue Images Based on Deep Convolutional Neural Networks" provides a convolutional neural network structure. This method includes an offline training stage and an online segmentation stage. This method can be applied to both closed and open tongue image acquisition environments, and can effectively improve the accuracy and robustness of automatic segmentation of traditional Chinese medicine tongue images.

[0005] The technical solution described in the patent with the application number "CN202011347107.5" and the name "A Tongue Image Segmentation Method, Device and Storage Medium" segments the tongue body by inputting each target tongue body image into a trained binary classification semantic segmentation model, and finally outputs the background and tongue body regions of the image.

[0006] However, in the segmentation method based on deep learning, the segmentation accuracy largely depends on a large number of original tongue images and the segmentation labels of the tongue images. The accuracy of classic deep learning network segmentation depends on the dataset captured by standard devices. When segmenting the tongue body for datasets captured by some ordinary devices, such as mobile phones and handheld cameras, the segmentation accuracy will not reach the ideal result. In addition, when the number of datasets is small, the tongue body segmentation accuracy of ordinary deep learning neural networks will also drop significantly.

[0007] In the patent disclosure texts of the application numbers "CN201710498517.1" and "CN202011347107.5", in the encoding network and decoding network proposed, the relationship between high-level features and low-level features during feature downsampling and upsampling processes is not considered, which easily causes the problem of feature misalignment in images. Moreover, in the segmentation result, there are many sharp parts at the edge of the tongue body, which does not match the real tongue body; furthermore, it causes the problem of low accuracy in tongue image segmentation. Summary of the Invention

[0008] The technical problem to be solved by the present invention is that in the existing tongue image segmentation method, in the encoding network and decoding network, the relationship between high-level features and low-level features during feature downsampling and upsampling processes is not considered, which easily causes the problem of feature misalignment in images. Moreover, in the segmentation result, there are many sharp parts at the edge of the tongue body, which does not match the real tongue body; furthermore, it causes the problem of low accuracy in tongue image segmentation.

[0009] The purpose of the present invention is to provide a feature alignment tongue image segmentation method and system based on improved UNet++. The present invention proposes to apply a feature alignment network to the Unet++ network architecture, add different feature alignment modules to the downsampling processes at different depths, and sequentially supervise the offset of features during the downsampling process to solve the feature misalignment in the tongue image segmentation process, fully improve the accuracy of tongue image segmentation, and finally optimize the segmentation result of the network through a morphological processing layer.

[0010] The present invention is realized through the following technical solutions:

[0011] In the first aspect, the present invention provides a feature alignment tongue image segmentation method based on improved UNet++, and the method includes:

[0012] Obtain the original tongue image, and perform labeling and preprocessing on the original tongue image;

[0013] According to the preprocessed tongue image, construct a feature alignment tongue image segmentation model and perform model training to obtain a trained feature alignment tongue image segmentation model;

[0014] Optimize the feature alignment tongue image segmentation model through morphological processing to obtain an optimized tongue image segmentation model;

[0015] Use the optimized tongue image segmentation model to perform segmentation processing on the tongue image to be segmented to obtain a segmentation result.

[0016] Further, performing labeling and preprocessing on the original tongue image includes:

[0017] Label the original tongue image, mark the tongue body area through feature labeling software, and output the mask image of the original tongue image; and divide the original tongue image and the corresponding mask image into a training set and a test set respectively;

[0018] Perform image enhancement preprocessing on the original tongue image. The image enhancement preprocessing includes noise addition processing, image contrast and brightness change processing, and histogram equalization processing.

[0019] Furthermore, the noise addition processing is to add salt-and-pepper noise to the tongue images in the training set of the original tongue image to obtain the tongue image after noise addition processing; the noise addition processing specifically includes the following steps:

[0020] Step A, input an image and customize the signal-to-noise ratio SNR (the value range of the signal-to-noise ratio SNR is between [0, 1]);

[0021] Step B, calculate the number of image pixels SP, and obtain the number of pixels of salt-and-pepper noise NP = SP * (1 - SNR);

[0022] Step C, randomly obtain each pixel position img[i, j] to be added with noise;

[0023] Step D, randomly generate a floating point number between [0, 1];

[0024] Step E, determine whether the floating point number is greater than 0.5. If it is greater than 0.5, specify the pixel value as 255; if it is less than 0.5, specify the pixel value as 0;

[0025] Step F, repeat steps C to E three times to complete NP pixel bold styles;

[0026] Step G, output the tongue image after adding noise.

[0027] Furthermore, the feature-aligned tongue image segmentation model integrates the feature-aligned network into the UNet++ network, and inputs the preprocessed tongue image I ∈ 3×H×W into the convolutional block VGG block of the UNet++ network to obtain the tongue image feature X(0,0), where H represents the height of the input feature tongue image, and W represents the width of the input feature tongue image; different feature alignment modules are added to the downsampling processes at different depths, and the offsets of the tongue image feature X(0,0) in the downsampling process are supervised in turn according to different feature alignment modules;

[0028] Among them, the convolutional block VGG block performs two operations on the input tongue image features of different sizes respectively. The operations include convolutional operation, pooling operation, and activation Relu operation;

[0029] The feature alignment module aligns and fuses the low-level features of the downsampling layer and the high-level features of the upsampling layer to obtain the aligned and fused tongue image features.

[0030] Furthermore, the execution process of the downsampling layer is as follows:

[0031] The entire network in the feature alignment tongue image segmentation model includes two paths: the Unet++ image segmentation path and the feature alignment path; Unet++ image segmentation path: From a horizontal perspective, this path combines multi-scale features from all previous nodes of the current feature node at the same resolution; from a vertical perspective, this path integrates multi-scale features of different resolutions from the previous node of the current feature node, and during use, the depth of the UNet++ network can be dynamically changed, enabling it to improve the segmentation efficiency with a relatively small impact on accuracy.

[0032] Feature alignment path: This path maintains the spatial information of the tongue image feature X(0,0), that is, there is no downsampling or downsampling operation in the second path; this path generates a series of tongue image features {A1, A2, …, Ai-1, Ai} with the same resolution and number of channels as X(0,0) from top to bottom. The tongue image feature Ai is obtained by combining the tongue image feature Ai-1 with the tongue image feature X(i,j) from the Unet++ image segmentation path. The tongue image feature X(i,j) outputs a feature map Ai' with the same resolution as the tongue image feature Ai-1 through the upsampling layer operation; the two high-level tongue image feature maps of the tongue image feature Ai-1 and the feature map Ai' are concatenated together, and a first offset matrix Δ of size H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer BN layer, an activation layer Relu, and a first 3x3 convolutional layer XA , and a second offset matrix Δ of size H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer BN layer, an activation layer Relu, and a second 3x3 convolutional layer A ; and fusing the output results of the first offset matrix Δ XA and the second offset matrix Δ A to obtain the high-level feature map Ai.

[0033] Among them, the three-dimensional matrix Δ represents the direction and distance of moving the pixel at position (h, w). The second offset matrix Δ A is used to align the tongue image feature Ai-1, and the first offset matrix Δ XA is used to align the feature map Ai'.

[0034] Furthermore, the fusion function in the output results of fusing the first offset matrix Δ XA and the second offset matrix Δ A is as follows:

[0035] A i = u(UP(X i ), Δ XA ) + u(A i-1, Δ A ) (3)

[0036] wherein, A0 is equivalent to X(0,0), A i-1 is the input high-resolution feature of the (i-1)th layer, and X i is the low-resolution feature of the ith layer, where UP represents the upsampling operation and u(,) is the alignment function;

[0037] The expression of the u(,) alignment function is:

[0038]

[0039] wherein, Δ 1hw , Δ 2hw respectively represent the horizontal and vertical learning two-dimensional transformation offset amounts at the position (h, w) in the tongue map (i.e., the tongue image) with a size of H×W, and F h′w′ represents the input tongue image feature at the position (h′, w′) in the tongue map, and u hw represents the output offset map at the corresponding position, and || represents the absolute value function.

[0040] Furthermore, the execution process of the upsampling layer is as follows:

[0041] The low-level feature X(i,j) output by the upsampling operation of the UNet++ network passes through the upsampling layer operation in the feature alignment module FA to output a feature map Bi' with the same resolution as the high-level feature Bi+1; the two high-level feature maps Bi+1 and Bi' are concatenated together, and a third offset matrix Δ with a size of H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer BN, an activation layer Relu, and a third 3x3 convolutional layer XB , and a fourth offset matrix Δ with a size of H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer BN, an activation layer Relu, and a fourth 3x3 convolutional layer B ; and the output results of the third offset matrix Δ XB and the fourth offset matrix Δ B are fused to obtain the high-level feature map Bi;

[0042] wherein, the three-dimensional matrix Δ represents the direction and distance of moving the pixel at the position (h, w); the fourth offset matrix Δ B is used to align the high-resolution tongue image feature Bi+1, and the third offset matrix Δ XB is used to align the high-resolution tongue image feature Bi'.

[0043] Furthermore, the morphological processing optimizes the feature alignment tongue image segmentation model, including:

[0044] Align the segmentation result of the feature-aligned tongue image segmentation model, and remove the noise outside the central tongue body of the segmentation result through the morphological reconstruction method to obtain the image after morphological reconstruction;

[0045] Remove the external noise in the tongue image area of the image after morphological reconstruction through opening operation to obtain the image after opening operation;

[0046] Smooth the image after opening operation through closing operation to obtain the final output result image.

[0047] In the second aspect, the present invention further provides a feature-aligned tongue image segmentation system based on the improved UNet++. This system supports the above-mentioned feature-aligned tongue image segmentation method based on the improved UNet++. The system includes:

[0048] An acquisition unit for acquiring the original tongue image;

[0049] A marking and preprocessing unit for marking and preprocessing the original tongue image;

[0050] A segmentation model construction unit for constructing a feature-aligned tongue image segmentation model based on the preprocessed tongue image and performing model training to obtain a trained feature-aligned tongue image segmentation model;

[0051] A morphological processing unit for optimizing the feature-aligned tongue image segmentation model through morphological processing to obtain an optimized tongue image segmentation model;

[0052] A segmentation unit for using the optimized tongue image segmentation model to perform segmentation processing on the tongue image to be segmented to obtain a segmentation result.

[0053] Furthermore, the feature-aligned tongue image segmentation model integrates the feature alignment network into the UNet++ network. The preprocessed tongue image I∈3×H×W is input into the convolutional block VGG block of the UNet++ network to obtain the tongue image feature X(0,0), where H represents the height of the input feature tongue image, and W represents the width of the input feature tongue image. Different feature alignment modules are added to the downsampling processes at different depths, and the offsets of the tongue image feature X(0,0) in the downsampling process are supervised in sequence according to different feature alignment modules;

[0054] Among them, the convolutional block VGG block is used to perform two operations on the input tongue image features of different sizes respectively. The operations include convolutional operation, pooling operation, and activation Relu operation;

[0055] The feature alignment module is used to align and fuse the low-level features of the downsampling layer and the high-level features of the upsampling layer to obtain the aligned and fused tongue image features.

[0056] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0057] The present invention relates to a feature alignment tongue image segmentation method and system based on improved UNet++. The feature alignment network is applied to the Unet++ network architecture, and different feature alignment modules are added to the downsampling processes at different depths to sequentially monitor the offset of features during the downsampling process. Finally, the output result of the entire network is optimized through a morphological processing layer. Therefore, the present invention can effectively avoid artificially setting the depth of the neural network layer, aggregate tongue image semantic features at different scales, form a highly flexible fusion scheme, and can monitor the offset of features during the downsampling process. Finally, it can process the morphology of the output image, solve the feature misalignment in the tongue image segmentation process, and fully improve the accuracy of tongue image segmentation. The present invention is simple to implement and has a relatively high segmentation efficiency, meeting the application requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0059] Figure 1 It is a flowchart of the feature alignment tongue image segmentation method based on improved UNet++ of the present invention.

[0060] Figure 2 It is a flowchart of the FAUNet++ network of the present invention.

[0061] Figure 3 It is a structural diagram of the FAUnet++ network of the present invention.

[0062] Figure 4 It is an architecture diagram of the VGG network of the present invention.

[0063] Figure 5 It is a schematic diagram of the feature alignment module FA of the present invention.

[0064] Figure 6 It is a schematic structural diagram of the feature alignment tongue image segmentation system based on improved UNet++ of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and the drawings. The illustrative embodiments and descriptions thereof of the present invention are only used to explain the present invention and do not limit the present invention.

[0066] In the existing encoding network and decoding network of tongue image segmentation methods, the relationship between high-level features and low-level features during feature downsampling and upsampling is not considered, which easily causes the problem of feature misalignment in images. Moreover, in the segmentation results, there are many sharp parts at the edge of the tongue body, which does not conform to the real tongue body; furthermore, it causes the problem of low accuracy in tongue image segmentation.

[0067] The present invention provides a feature alignment tongue image segmentation method and system based on improved Unet++. The present invention proposes to apply a feature alignment network to the Unet++ network architecture, add different feature alignment modules to the downsampling processes at different depths, and sequentially supervise the offset of features during the downsampling process to solve the feature misalignment in the tongue image segmentation process, fully improve the accuracy of tongue image segmentation, and finally optimize the segmentation result through a morphological processing layer.

[0068] The FAUNet++ network structure proposed by the present invention is as Figure 3 shown: 1. Add a feature alignment module on the basis of the existing Unet++, which solves the misalignment problem in the feature aggregation process caused by downsampling operations and context information fusion, automatically selects an appropriate network depth according to the segmentation task to perform the segmentation task, shares the learning parameters in the encoding stage, and improves the segmentation performance. 2. Use a morphological processing unit to process the network output result, remove the noise and uneven edge problems in the segmentation result of the feature alignment tongue image segmentation model, and fully improve the accuracy of image segmentation.

[0069] Example 1

[0070] As Figure 1 and Figure 2 shown, the feature alignment tongue image segmentation method based on improved Unet++ of the present invention includes:

[0071] Step 1, obtain the original tongue image, and perform marking and preprocessing on the original tongue image; specifically including:

[0072] Step 11, mark the original tongue image, mark the tongue body area through feature marking software, and output the mask image of the original tongue image; and divide the original tongue image and the corresponding mask image into a training set and a test set respectively; specifically, the marking can be manual marking, and the tongue image training set taken by ordinary equipment is marked with the tongue body area through feature marking software.

[0073] Step 12, perform image enhancement preprocessing on the original tongue image, and the image enhancement preprocessing includes noise addition processing, image contrast and brightness change processing, and histogram equalization processing.

[0074] Specifically, noise addition processing: Add salt-and-pepper noise to the tongue images in the training set of the original tongue images to obtain the tongue images after noise addition processing. The noise addition processing specifically includes the following steps:

[0075] Step A, input an image and customize the signal-to-noise ratio SNR (the value range of the signal-to-noise ratio SNR is between [0,1]);

[0076] Step B, calculate the number of image pixel points SP, and obtain the number of pixel points of salt-and-pepper noise NP = SP * (1 - SNR);

[0077] Step C, randomly obtain each pixel position img[i,j] to be added with noise;

[0078] Step D, randomly generate a floating point number between [0,1];

[0079] Step E, determine whether the floating point number is greater than 0.5. If it is greater than 0.5, specify the pixel value as 255. If it is less than 0.5, specify the pixel value as 0;

[0080] Step F, repeat steps C to E three steps to complete NP pixel bold styles;

[0081] Step G, output the tongue image after adding noise.

[0082] Specifically, image contrast and brightness change processing: Adjust the image contrast, perform color changes, simulate images in different scenarios, and the transformation process is as follows:

[0083] g(i,j) = αf(i,j) + β (1)

[0084] f(i,j) represents the input pixel value, g(i,j) represents the output pixel value, where α > 0, β is the gain variable. α mainly controls the strength of the contrast. A value greater than 1 indicates strong contrast, and a value less than 1 indicates weak contrast. β represents the strength of the brightness. The larger the value, the higher the brightness.

[0085] Specifically, histogram equalization processing: Transform the histogram distribution of the image into an approximately uniform distribution through the cumulative distribution function, thereby enhancing the contrast of the image. In order to expand the brightness range of the original image, a mapping function is required to evenly map the pixel values of the original image to the new histogram. The cumulative distribution function can be used to achieve histogram equalization. Since the image is composed of individual pixel points, the histogram equalization of the image is solved through the discrete form of the cumulative distribution function. During the histogram equalization process, the histogram of the original image is transformed into the form of a specified histogram, and the mapping method is:

[0086]

[0087] Among them, Sk Refers to the value after the current gray level is mapped by the cumulative distribution function. n is the total number of pixels in the image, and n j is the number of pixels at the current gray level, and L is the total number of gray levels in the image.

[0088] Finally, according to the mapping relationship, the pixels of each gray level set are subjected to image transformation to obtain the equalized image.

[0089] Step 2: Based on the preprocessed tongue image, construct a feature-aligned tongue image segmentation model and perform model training to obtain a trained feature-aligned tongue image segmentation model;

[0090] The present invention creatively designs a feature-aligned tongue image segmentation model. The feature-aligned tongue image segmentation model integrates a feature alignment network into the UNet++ network, that is, integrates a feature alignment module (FAM) into the Unet++ model. The Unet++ model introduces a built-in set of depth-variable U-Nets, thereby improving the segmentation performance for images of various sizes. And Unet++ redesigned the skip connections in UNet, thereby achieving flexible feature fusion in the decoder. Specifically, during the segmentation process, from a horizontal perspective, multi-scale features from all previous nodes of the current node are combined at the same resolution; from a vertical perspective, multi-scale features of different resolutions are integrated from the previous node of the current node. Finally, during use, the depth of the network can be dynamically changed, enabling it to improve the segmentation efficiency with a relatively small impact on accuracy. Among them, each processing module is a convolutional block VGG block (as Figure 4 shown). The convolutional block VGG block performs two convolutional operations, pooling operations, and activation Relu operations on input feature images of different sizes respectively. The feature alignment module aligns and fuses the low-level features of the downsampling layer and the high-level features of the upsampling layer to obtain the tongue image features after alignment and fusion.

[0091] Specifically, the execution process of the downsampling layer is as follows:

[0092] In the solution of the present invention, the pre - processed tongue image \(I\in3\times H\times W\) is input into the convolutional block VGG block of the UNet++ network to obtain the tongue image feature \(X^{(0,0)}\), where \(H\) represents the height of the input feature tongue image, and \(W\) represents the width of the input feature tongue image; in the feature alignment tongue image segmentation model, the entire network includes two paths: the Unet++ image segmentation path and the feature alignment path; Unet++ image segmentation path: From a horizontal perspective, this path combines multi - scale features from all previous nodes of the current feature node at the same resolution; from a vertical perspective, this path integrates multi - scale features of different resolutions from the previous node of the current feature node, and during use, the depth of the UNet++ network can be dynamically changed to improve the segmentation efficiency with a small impact on accuracy;

[0093] Feature alignment path: This path preserves the spatial information of the tongue image feature \(X^{(0,0)}\), that is, there is no downsampling or downsampling operation in the second path; specifically, this path generates a series of tongue image features \(\{A_1,A_2,A_3,A_4\}\) with the same resolution and number of channels as \(X^{(0,0)}\) from top to bottom. The tongue image feature \(A_i\) is obtained by combining the tongue image feature \(A_{i - 1}\) with the tongue image feature \(X^{(i,j)}\) from the Unet++ image segmentation path. The specific process is as Figure 5 shown. The tongue image feature \(X^{(i,j)}\) outputs a feature map \(A_i'\) with the same resolution as the tongue image feature \(A_{i - 1}\) through the upsampling layer operation; the two high - level tongue image feature maps of the tongue image feature \(A_{i - 1}\) and the feature map \(A_i'\) are concatenated together, and a first offset matrix \(\Delta\) of size \(H\times W\times2\) is generated through a \(1\times1\) convolutional layer, a batch normalization layer BN layer, an activation layer Relu, and a first \(3\times3\) convolutional layer XA , and a second offset matrix \(\Delta\) of size \(H\times W\times2\) is generated through a \(1\times1\) convolutional layer, a batch normalization layer BN layer, an activation layer Relu, and a second \(3\times3\) convolutional layer A ; and the output results of the first offset matrix \(\Delta\) XA and the second offset matrix \(\Delta\) A are fused to obtain the high - level feature map \(A_i\).

[0094] Among them, the three - dimensional matrix \(\Delta\) represents the direction and distance of moving the pixel at position \((h,w)\). The second offset matrix \(\Delta\) A is used to align the tongue image feature \(A_{i - 1}\), and the first offset matrix \(\Delta\) XA is used to align the feature map \(A_i'\).

[0095] Specifically, the fusion function in the output results of fusing the first offset matrix \(\Delta\) XA and the second offset matrix \(\Delta\) A is:

[0096] A i = u(UP(Xi ), Δ XA ) + u(A i-1 , Δ A ) (3)

[0097] Wherein, A0 is equivalent to X(0, 0), A i-1 is the input high-resolution feature of the (i - 1)-th layer, X i is the low-resolution feature of the i-th layer, where UP represents the upsampling operation, and u(,) is the alignment function;

[0098] The expression of the u(,) alignment function is:

[0099]

[0100] Wherein, Δ 1hw , Δ 2hw respectively represent the horizontal and vertical learning two-dimensional transformation offset amounts at the position (h, w) in the tongue map (i.e., the tongue image) of size H×W, F h′w′ represents the input tongue image feature at the position (h′, w′) in the tongue map, u hw represents the output offset map at the corresponding position, and || represents the absolute value function.

[0101] Specifically, the execution process of the upsampling layer is as follows:

[0102] The low-level feature X(i, j) output by the upsampling operation of the UNet++ network passes through the upsampling layer operation in the feature alignment module FA to output a feature map Bi' with the same resolution as the high-level feature Bi+1; the two high-level feature maps Bi+1 and Bi' are concatenated together, and a third offset matrix Δ of size H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer BN, an activation layer Relu, and a third 3x3 convolutional layer XB , and a fourth offset matrix Δ of size H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer BN, an activation layer Relu, and a fourth 3x3 convolutional layer B ; and the output results of fusing the third offset matrix Δ XB and the fourth offset matrix Δ B are used to obtain the high-level feature map Bi;

[0103] Among them, the three-dimensional matrix Δ represents the direction and distance of moving the pixel at the position (h, w); the fourth offset matrix Δ B is used to align the high-resolution tongue image feature Bi+1, and the third offset matrix Δ XB is used to align the high-resolution tongue image feature Bi'.

[0104] In the above technical solution, the Ai module layer in the downsampling process of the present invention and the Bi module layer in the upsampling process always maintain the spatial information of the original input image. The tongue image features in the downsampling process and the tongue image features in the upsampling process are respectively fused through feature alignment into {A1, A2, A3, A4} and {B1, B2, B3}. As the number of network layers deepens, the semantic information in this path gradually enriches, and finally the segmentation result is jointly output by the UNet++ output module and the feature alignment path.

[0105] In addition, the result loss function of the feature alignment tongue image segmentation model adopts BCEDiceLoss, where Diceloss is as shown in formula (5) and BCEloss is as shown in formula (6):

[0106]

[0107] Among them, Z represents the segmentation result image, and Y represents the marked image

[0108] L BCE =-w[y n ·logz n +(1 - y n )·log(1 - z n )] (6)

[0109] z n represents the nth pixel in the segmentation image, y n represents the nth pixel in the marked image, and w is a user-defined weight.

[0110] Finally, the total loss is the weighted sum of the dice loss and the BCE loss, and a and β are user-defined weight parameters.

[0111] L BCE-DiceLoss =aL Dice +βL BCE (7)

[0112] Set the model training parameters, select the Adam optimization algorithm for gradient update, and save the model with the highest accuracy according to the test set.

[0113] Step 3, optimize the feature alignment tongue image segmentation model through morphological processing to obtain an optimized tongue image segmentation model;

[0114] Considering that the segmentation image output by the FAUnet++ model may have noise and a rough tongue, and there are some small "holes" in it, the segmentation result is input into the morphological processing unit for processing. The specific steps are as follows:

[0115] (a) Align the segmentation results of the feature-aligned tongue image segmentation model, and remove the noise outside the central tongue body of the segmentation results through morphological reconstruction to obtain the image after morphological reconstruction;

[0116] Morphological reconstruction is a method that applies morphological transformation. By marking the image, the key regions defined in the template image are maintained. The specific process is as follows:

[0117] In the first step, first calculate the reverse image R of the image I segmented by FAUnet++. To switch the hole region to the foreground and the non-hole region to the background, the formula is as follows:

[0118] R = 1 – I (5)

[0119] Next, perform edge detection on R, where each pixel belongs to the edge composed of the point set E. Then, set the pixels that do not belong to the edge to zero to obtain the image L:

[0120]

[0121] Use L as the marked image for morphological reconstruction. At the same time, use R as the template image. Based on R and L, perform morphological reconstruction on the image R to generate the hole image H:

[0122] H = MR(R, L, s, t) (7)

[0123] Where MR represents morphological reconstruction, s is the size of the structural element, and t is the tolerance that controls the iterative reconstruction operation. The MR algorithm process is as follows:

[0124] Input: template image R, marked image L with size m×n, size s of the structural element, tolerance t = 50;

[0125] 1. Initialize the structural components

[0126] 2. Let p be the number of iterations, and L0 = L. ⊕ represents the dilation operation, and ∩ represents the intersection operation. Execute:

[0127] L p+1 = (L p ⊕ K) ∩ R

[0128] 3. Until

[0129] Finally, overlap the image I and the image H to obtain the image F after filling the holes. The image F after filling the holes is the image after morphological reconstruction:

[0130] F = I + H (8)

[0131] (b) Remove the external noise in the tongue image area of the image after morphological reconstruction through opening operation to obtain the image O after opening operation;

[0132]

[0133] represents a filtering operation, f1 is an opening operation filter, f d1 represents the center point in the filter f1, and T represents the matrix operation of the filter on the image.

[0134] (c) Smooth the image O after opening operation through closing operation to obtain the final output result image C.

[0135]

[0136] represents a filtering operation, f2 is a closing operation filter, f d2 represents the center point in the filter f2, and T represents the matrix operation of the filter on the image.

[0137] Step 4: Use the optimized tongue image segmentation model to segment the tongue image to be segmented to obtain the segmentation result.

[0138] Specifically, read the single image or multiple images to be segmented taken, load the saved optimized tongue image segmentation model, send the image to be segmented into the feature-aligned tongue image segmentation model, output the mask image, and process the mask image through the morphological processing layer to obtain the final result.

[0139] Compared with the prior art, the feature-aligned tongue image segmentation method based on improved UNet++ of the present invention can effectively avoid artificially setting the depth of the neural network layer, aggregate the semantic features of tongue images at different scales, form a highly flexible fusion scheme, and can supervise the feature offset in the downsampling process, solve the feature misalignment in the tongue image segmentation process, fully improve the accuracy of tongue image segmentation, and finally can process the morphology of the output image and optimize the segmentation result. The present invention is simple to implement, has a high segmentation efficiency, and meets the application requirements.

[0140] Embodiment 2

[0141] As Figure 6 shown, the difference between this embodiment and Embodiment 1 is that this embodiment further provides a feature-aligned tongue image segmentation system based on improved UNet++, and this system supports the feature-aligned tongue image segmentation method based on improved UNet++ described in Embodiment 1; this system includes:

[0142] An acquisition unit for acquiring the original tongue image;

[0143] A marking and preprocessing unit for marking and preprocessing the original tongue image;

[0144] A segmentation model construction unit for constructing a feature-aligned tongue image segmentation model based on the preprocessed tongue image and performing model training to obtain a trained feature-aligned tongue image segmentation model;

[0145] A morphological processing unit for optimizing the feature-aligned tongue image segmentation model through morphological processing to obtain an optimized tongue image segmentation model;

[0146] A segmentation unit for using the optimized tongue image segmentation model to perform segmentation processing on the tongue image to be segmented to obtain a segmentation result.

[0147] As a further implementation, the feature-aligned tongue image segmentation model integrates a feature alignment network into the UNet++ network. The preprocessed tongue image I ∈ 3×H×W is input into the convolutional block VGG block of the UNet++ network to obtain the tongue image feature X(0,0), where H represents the height of the input feature tongue image and W represents the width of the input feature tongue image. Different feature alignment modules are added to the downsampling processes at different depths, and the offsets of the tongue image feature X(0,0) in the downsampling process are supervised in sequence according to different feature alignment modules;

[0148] Among them, the convolutional block VGG block is used to perform two operations on input tongue image features of different sizes respectively. The operations include convolutional operation, pooling operation, and activation Relu operation;

[0149] The feature alignment module is used to align and fuse the low-level features of the downsampling layer and the high-level features of the upsampling layer to obtain the aligned and fused tongue image features.

[0150] The execution processes of each unit can be carried out according to the flow steps of the feature-aligned tongue image segmentation method based on the improved UNet++ described in Embodiment 1, and will not be elaborated one by one in this embodiment.

[0151] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0152] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a means for realizing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for realizing the functions specified in one or more of the blocks.

[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means realizes the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for realizing the functions specified in one or more of the blocks.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for realizing the functions specified in one or more of the blocks.

[0155] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An improved UNet++-based feature alignment tongue image segmentation method, characterized in that, The method includes: Obtaining an original tongue image, and performing labeling and preprocessing on the original tongue image; Constructing a feature alignment tongue image segmentation model based on the preprocessed tongue image and performing model training to obtain a trained feature alignment tongue image segmentation model; Optimizing the feature alignment tongue image segmentation model through morphological processing to obtain an optimized tongue image segmentation model; Using the optimized tongue image segmentation model to perform segmentation processing on the tongue image to be segmented to obtain a segmentation result; The feature alignment tongue image segmentation model integrates a feature alignment network into the UNet++ network. The preprocessed tongue image I∈3×H×W is input into the convolutional block VGG block of the UNet++ network to obtain tongue image features X(0,0), where H represents the height of the input feature tongue image, and W represents the width of the input feature tongue image; different feature alignment modules are added to the downsampling processes at different depths, and the offsets of the tongue image features X(0,0) in the downsampling process are supervised in sequence according to different feature alignment modules; Among them, the convolutional block VGG block performs two operations on the input tongue image feature images of different sizes, and the operations include convolutional operation, pooling operation, and activation operation; The feature alignment module aligns and fuses the low-level features of the downsampling layer and the high-level features of the upsampling layer to obtain aligned and fused tongue image features; The execution process of the downsampling layer is: The entire network in the feature alignment tongue image segmentation model includes two paths: the Unet++ image segmentation path and the feature alignment path; the Unet++ image segmentation path: from a horizontal perspective, this path combines the multi-scale features from all previous nodes of the current feature node at the same resolution; from a vertical perspective, this path integrates the multi-scale features of different resolutions from the previous node of the current feature node, and during use, it can dynamically change the depth of the UNet++ network; The feature alignment path: This path maintains the spatial information of the tongue image features X(0,0). This path generates a series of tongue image features {A1, A2, …, Ai-1, Ai} with the same resolution and number of channels as X(0,0) from top to bottom. The tongue image feature Ai is obtained by combining the tongue image feature Ai-1 with the tongue image feature X(i,j) from the Unet++ image segmentation path. The tongue image feature X(i,j) outputs a feature map Ai' with the same resolution as the tongue image feature Ai-1 through the upsampling layer operation; connecting the tongue image feature Ai-1 and the feature map Ai', and generating a first offset matrix of size H×W×2 through a 1x1 convolutional layer, a batch normalization layer, an activation layer, and a first 3x3 convolutional layer, and generating a second offset matrix of size H×W×2 through a 1x1 convolutional layer, a batch normalization layer, an activation layer, and a second 3x3 convolutional layer; and fusing the output results of the first offset matrix and the second offset matrix to obtain a high-level feature map Ai; Among them, the second offset matrix is used to align the tongue image feature Ai-1, and the first offset matrix is used to align the feature map Ai'; 2. The feature alignment tongue image segmentation method based on the improved UNet++ according to claim 1, characterized in that Performing labeling and preprocessing on the original tongue image includes: Mark the original tongue image, mark the tongue body area through feature marking software, and output the mask image of the original tongue image; and divide the original tongue image and the corresponding mask image into a training set and a test set respectively; Perform image enhancement preprocessing on the original tongue image, and the image enhancement preprocessing includes noise addition processing, image contrast and brightness change processing, and histogram equalization processing.

3. The feature alignment tongue image segmentation method based on the improved UNet++ according to claim 2, wherein The noise addition processing is to add salt and pepper noise to the tongue images in the training set of the original tongue image to obtain the tongue image after noise addition processing; the noise addition processing specifically includes the following steps: Step A, input an image and define the signal-to-noise ratio SNR; Step B, calculate the number of image pixels SP, and obtain the number of pixels of salt and pepper noise NP = SP * (1 - SNR); Step C, randomly obtain each pixel position img[i, j] to be added with noise; Step D, randomly generate a floating point number between [0, 1]; Step E, determine whether the floating point number is greater than 0.

5. If it is greater than 0.5, specify the pixel value as 255, and if it is less than 0.5, specify the pixel value as 0; Step F, repeat steps C to E three steps to complete the bold style of NP pixels; Step G, output the tongue image after adding noise.

4. The feature alignment tongue image segmentation method based on the improved UNet++ according to claim 1, wherein, The fusion first offset matrix Δ XA and the second offset matrix Δ A The fusion function in the output result is: A i = u(UP(X i ), Δ XA ) + u(A i-1 , Δ A ) wherein, A0 is equivalent to X(0,0), and A i-1 is the input high-resolution feature of the (i-1)th layer, and X i is the low-resolution feature of the ith layer, where UP represents the upsampling operation, and u(,) is the alignment function; The expression of the u(,) alignment function is: where, Δ 1hw , Δ 2hw respectively represent the horizontal and vertical learning two-dimensional transformation offsets at the position (h, w) in the tongue image of size H×W, F h′w′ represents the input tongue image feature at the position (h′, w′) in the tongue image, u hw represents the output offset map at the corresponding position, and || represents the absolute value function.

5. The method for tongue image segmentation based on improved UNet++ for feature alignment according to claim 1, wherein The execution process of the upsampling layer is: The low-level feature X(i, j) output by the upsampling operation of the UNet++ network is output through the upsampling layer operation in the feature alignment module to obtain a feature map Bi' with the same resolution as the high-level feature Bi+1; connect the two high-level feature maps Bi+1 and Bi' together, and generate a third offset matrix of size H×W×2 through a 1x1 convolutional layer, a batch normalization layer, an activation layer, and a third 3x3 convolutional layer, and generate a fourth offset matrix of size H×W×2 through a 1x1 convolutional layer, a batch normalization layer, an activation layer, and a fourth 3x3 convolutional layer; And fuse the output results of the third offset matrix and the fourth offset matrix to obtain the high-level feature map Bi; Among them, the fourth offset matrix Δ B is used to align the high-resolution tongue image feature Bi+1, and the third offset matrix Δ XB is used to align the high-resolution tongue image feature Bi'; The high-resolution tongue image feature Bi+1 is the high-level feature map Bi+1, and the high-resolution tongue image feature Bi' is the high-level feature map Bi'.

6. The feature alignment tongue image segmentation method based on improved UNet++ according to claim 1, wherein The optimization of the feature alignment tongue image segmentation model through morphological processing includes: Use the morphological reconstruction method to remove the noise outside the central tongue body of the segmentation result of the feature alignment tongue image segmentation model to obtain the image after morphological reconstruction; Remove the external noise in the tongue image area of the image after morphological reconstruction through opening operation to obtain the image after opening operation; Smooth the image after opening operation through closing operation to obtain the final output result image.

7. Feature alignment tongue image segmentation system based on improved UNet++, characterized in that The system includes: An acquisition unit for acquiring the original tongue image; A marking and preprocessing unit for marking and preprocessing the original tongue image; A segmentation model construction unit for constructing a feature alignment tongue image segmentation model based on the preprocessed tongue image and performing model training to obtain a trained feature alignment tongue image segmentation model; A morphological processing unit for optimizing the feature-aligned tongue image segmentation model through morphological processing to obtain an optimized tongue image segmentation model; A segmentation unit for segmenting a tongue image to be segmented using the optimized tongue image segmentation model to obtain a segmentation result; The feature-aligned tongue image segmentation model integrates a feature alignment network into the UNet++ network. The preprocessed tongue image I ∈ 3×H×W is input into the convolutional block VGG block of the UNet++ network to obtain the tongue image feature X(0,0), where H represents the height of the input feature tongue image and W represents the width of the input feature tongue image. Different feature alignment modules are added to the downsampling processes at different depths, and the offsets of the tongue image feature X(0,0) in the downsampling process are supervised sequentially according to different feature alignment modules; Among them, the convolutional block VGG block is used to perform two operations on input tongue image features of different sizes, and the operations include convolutional operation, pooling operation, and activation operation; The feature alignment module is used to align and fuse the low-level features of the downsampling layer and the high-level features of the upsampling layer to obtain the aligned and fused tongue image features; The execution process of the downsampling layer is as follows: The entire network in the feature-aligned tongue image segmentation model includes two paths: the Unet++ image segmentation path and the feature alignment path. The Unet++ image segmentation path: From a horizontal perspective, this path combines the multi-scale features from all previous nodes of the current feature node at the same resolution. From a vertical perspective, this path integrates the multi-scale features of different resolutions from the previous node of the current feature node, and during use, it can dynamically change the depth of the UNet++ network; The feature alignment path: This path preserves the spatial information of the tongue image feature X(0,0). This path generates a series of tongue image features {A1, A2, …, Ai-1, Ai} with the same resolution and number of channels as X(0,0) from top to bottom. The tongue image feature Ai is obtained by combining the tongue image feature Ai-1 with the tongue image feature X(i,j) from the Unet++ image segmentation path. The tongue image feature X(i,j) outputs a feature map Ai' with the same resolution as the tongue image feature Ai-1 through the upsampling layer operation. The tongue image feature Ai-1 is concatenated with the feature map Ai', and a first offset matrix of size H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer, an activation layer, and a first 3x3 convolutional layer. A second offset matrix of size H×W×2 is generated through a 1x1 convolutional layer, a batch normalization layer, an activation layer, and a second 3x3 convolutional layer. And the output results of the first offset matrix and the second offset matrix are fused to obtain the high-level feature map Ai; Among them, the second offset matrix is used to align the tongue image feature Ai-1, and the first offset matrix is used to align the feature map Ai'.

Citation Information

Patent Citations

  • Deep convolutional neural network-based traditional Chinese medicine tongue image automatic segmentation method

    CN107316307A

  • Tongue image segmentation method and device and storage medium

    CN112489053A

  • Continuous multi-frame image super-resolution reconstruction method based on multi-scale motion compensation framework and recursive learning

    CN112102163A

  • Image fusion method and device, electronic equipment and storage medium

    CN113255756A