Tongue picture segmentation method
By improving the U-Net network to a dual-branch encoder network and introducing a jump connection network of the attention mechanism and Vision Transformer encoder, the problem of low accuracy of tongue image segmentation in the prior art is solved, and a higher accuracy of tongue image segmentation is achieved.
Patent Information
- Application Number
- CN202510234736.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-10
AI Technical Summary
The existing U-Net networks extracted features in tongue-like segmentation are not accurate enough, resulting in low accuracy in tongue-like segmentation.
Based on U-Net, the encoder network is improved as a dual-branch encoder network, the bottleneck layer network of attention mechanism and the jump connection network of the Vision Transformer encoder are introduced, and the capability of feature extraction and multi-scale feature fusion is enhanced.
It improves the accuracy and accuracy of tongue image segmentation, and enhances the model's ability to learn and express tongue image features.
Smart Images

Figure CN120125601A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tongue image segmentation, and specifically, to a tongue image segmentation method. Background Art
[0002] Tongue diagnosis is an important part of the four diagnostic methods in traditional Chinese medicine (observation, auscultation and olfaction, interrogation, and palpation). It is a diagnostic method that judges the health status of the human body by observing the tongue image (the shape, color, texture, etc. of the tongue body and the tongue coating). The theoretical basis of tongue diagnosis stems from the holistic concept and the theory of zang-fu organs and meridians in traditional Chinese medicine, believing that the tongue is closely connected to the internal organs, and its changes can reflect the internal pathological state of the body.
[0003] The key to tongue diagnosis lies in "observing the tongue image", mainly including two aspects: the tongue proper and the tongue coating.
[0004] The tongue proper refers to the color, shape, and movement of the tongue body itself. A healthy tongue proper is generally light red. If the color of the tongue proper is pale, red and crimson, cyanotic and purple, etc., it can reflect the situation of qi and blood circulation and yin-yang imbalance in the human body. The shape of the tongue, such as a plump and large tongue body, a thin and emaciated tongue, cracks, etc., can reflect the function of the spleen and stomach and the condition of internal water and dampness in the body. The movement of the tongue includes whether the tongue is flexible, whether there is tremor, deviation, etc.
[0005] The tongue coating refers to the thin layer covering on the tongue surface, which is produced by the steaming of stomach qi. The thickness of the tongue coating reflects the severity of pathogenic factors and the depth of the disease location. Moist tongue coating indicates sufficient body fluid, while dryness suggests deficiency of body fluid or damage to body fluid by heat pathogens. In terms of color, the normal tongue coating is a thin white coating, and yellow coating, gray coating or black coating may indicate heat syndromes, cold-dampness or serious diseases.
[0006] Tongue diagnosis is widely used in clinical practice, with the characteristics of non-invasive, intuitive, and fast. It can assist in differentiating the nature of diseases, judging the severity of the condition, etc. Modern research shows that the tongue image is closely related to the body's microcirculation, metabolic state and immune function, providing a scientific basis for tongue diagnosis. For example, a dark red tongue proper may be related to microcirculation disorders, and a yellow and greasy tongue coating may indicate digestive function disorders.
[0007] Therefore, the combination of traditional Chinese medicine tongue diagnosis and modern science and technology is of great importance. By using imaging technology, etc., tongue images can be collected quickly and accurately to obtain information about the tongue. Relevant medical staff, etc., can judge diseases based on the information of the tongue.
[0008] Tongue image segmentation is an important technology in the field of tongue diagnosis. It segments the tongue image to accurately extract the tongue body and tongue coating regions, thereby providing a reliable basis for tongue diagnosis analysis. The tongue image is an important basis for traditional Chinese medicine diagnosis and can reflect the functional state of the body's zang-fu organs and disease changes.
[0009] The emergence of tongue image segmentation technology makes it possible to intelligentize tongue diagnosis. Through deep learning and computer vision algorithms, tongue image segmentation can automatically identify the contour of the tongue body, separate regions such as the tongue proper, tongue coating, and background, laying a foundation for subsequent feature extraction and analysis.
[0010] In the prior art, such as the Chinese patent application document with the publication number CN 117765566 A, whose name is an end-to-end tongue image recognition method based on a multi-task network, which segments tongue images based on the U-Net network. However, the features extracted by the existing U-Net network are usually not precise enough, resulting in a low accuracy of tongue image segmentation.
[0011] To solve the above existing problems, people have been seeking an ideal technical solution. Summary of the Invention
[0012] Based on this, it is necessary to provide a tongue image segmentation method for the above technical problems, providing a new idea for constructing a tongue image segmentation model.
[0013] To achieve the above object, the first aspect of the present invention provides a tongue image segmentation method, which includes: Step 1, obtaining an image of a human tongue; Step 2, performing image segmentation and label annotation on the image of the human tongue to establish a data set of the image of the human tongue; Step 3, building a U-Net segmentation model for the image of the human tongue, the U-Net segmentation model including a dual-branch encoder network, a decoder network, a bottleneck layer network introducing an attention mechanism, a skip connection network introducing a Vision Transformer encoder, and an output network; The dual-branch encoder includes a dilated convolution branch and a normal convolution branch arranged in parallel. The dilated convolution branch includes a plurality of cascaded dilated convolution layers, and the normal convolution branch includes a plurality of cascaded normal convolution units. The normal convolution unit includes a normal convolution layer and a max pooling layer, and there are skip connections between each level structure of the dilated convolution branch and the normal convolution branch; The decoder network includes a plurality of cascaded transposed convolution layers. The topmost transposed convolution layer is connected to the output network through a 1*1 convolution layer, and the remaining transposed convolution layers are respectively connected to the output network through upsampling layers; Step 4, training the U-Net segmentation model using the data set of the image of the human tongue to obtain a trained U-Net segmentation model; Step 5, inputting the image of the human tongue to be segmented into the trained U-Net segmentation model to obtain a segmented tongue image.
[0014] Based on the above, the skip connection network includes a Vision Transformer encoder, and the Vision Transformer encoder includes a projection head and a plurality of Transformer encoding modules. Each Transformer encoding module implements the residual connection of the ordinary convolutional layer in the dual-branch encoder and the transposed convolutional layer in the decoder network as a residual connection block.
[0015] Based on the above, the Transformer encoding module includes two composite blocks. The first composite block includes a Layer Norm layer, a multi-head attention mechanism, and a residual connection; the second composite block includes a Layer Norm layer, a multi-layer perceptron activated by GeLu and ReLu activation functions respectively, and a residual connection; a Dropout layer is added after each composite block.
[0016] Based on the above, the bottleneck layer network includes a 3×3 dilated convolutional layer, a 3×3 convolutional layer, and a hybrid attention layer. The 3×3 dilated convolutional layer is connected to the dilated convolution branch, the 3×3 convolutional layer is connected to the 3×3 dilated convolutional layer and the ordinary convolution branch respectively, and the hybrid attention layer is connected to the 3×3 convolutional layer and the decoder respectively.
[0017] Based on the above, the hybrid attention layer includes a CBAM attention module and an SE attention mechanism module connected in parallel and added together.
[0018] Based on the above, the transposed convolutional layer includes 1 2×2 transposed convolutional layer and 2 3×3 ordinary convolutional layers, and the output network is an upsampler.
[0019] To achieve the above object, a second aspect of the present invention provides a tongue image segmentation device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned tongue image segmentation method are implemented.
[0020] To achieve the above object, a third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned tongue image segmentation method are implemented.
[0021] To achieve the above object, a fourth aspect of the present invention provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned tongue image segmentation method are implemented.
[0022] The beneficial effects of the present invention are as follows: Based on U-Net, the encoder network of the present invention is changed to a dual-branch encoder network. The dual-branch encoder includes an atrous convolution branch and a normal convolution branch arranged in parallel, enabling the model to extract features from multiple scales, enhancing the model's feature extraction ability, enabling the model to more accurately capture the features of tongue images on a similar color background, and thus improving the accuracy of tongue image segmentation. The deconvolution layer at the topmost layer is connected to the output network through a 1*1 convolution layer, which helps to finely adjust and optimize the segmentation result and improve the overall performance of the tongue image segmentation task. The remaining deconvolution layers are respectively connected to the output network through upsampling layers, ensuring the restoration of the image size and the global information, helping to improve the image reconstruction quality, and thus improving the tongue image segmentation accuracy. The bottleneck layer network introducing the attention mechanism can automatically select and weight important features, enhance the model's learning and expression ability of key features, etc., and thus improve the tongue image segmentation accuracy. The skip connection network introducing the Vision Transformer encoder can improve the feature representation ability, enhance multi-scale feature fusion, etc., and can improve the segmentation effect in the tongue image segmentation task. Description of the Drawings
[0023] Figure 1 It is a simplified structural schematic diagram of the U-Net segmentation model of the present invention. Figure 2 It is an external view schematic diagram of the tongue diagnosis instrument of the present invention. Figure 3 It is a schematic diagram of the tongue image segmentation result of the present invention. Detailed Embodiments
[0024] The following is a further detailed description of the technical solutions of the present invention through specific embodiments.
[0025] The terms "first", "second", "third", "fourth", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application.
[0026] This embodiment gives a specific implementation of a tongue image segmentation method, and the method includes: Step 1, obtain a human tongue image.
[0027] In some embodiments, images of the tongue surface, tongue back, etc. of multiple people are collected through a tongue diagnosis instrument and a camera as human tongue images.
[0028] Such as Figure 2As shown, the human tongue image is collected by the YM-SA-I tongue diagnosis instrument. In some embodiments, the device simulates a natural light environment, the illuminance at the collection port is approximately 3600 lx, and its error is less than ±8%; the color temperature in the collection environment is in the range of 5000K to 5800K; the color rendering index of the device light source is greater than 95; the imaging resolution is not less than 5 lp / mm; the radiance at the maximum illuminance in the spectral range of 300nm to 2500nm does not exceed 350W / m2, and the effective ultraviolet radiation illuminance does not exceed 0.008W / m2; the imaging device accurately restores colors, and after imaging the colors on the color card, the color difference of each color in the CIE LAB color space is less than 20; the second-generation relative distortion of the imaging device does not exceed 5%.
[0029] Step 2: Perform image segmentation and label annotation on the human tongue image to establish a dataset of human tongue images.
[0030] In some embodiments, the Labelme software is used to uniformly perform image segmentation and label annotation on the collected human tongue images (tongue images); The Labelme software is a software for label annotation of pictures. Specifically, a polygon box connected by multiple points is used for label annotation. Generally, the size of a single tongue image is 1728×1100, the number of polygon points is more than 130, and the label is saved in the json file format. Finally, a label dataset of a full json file is obtained.
[0031] In some embodiments, a data augmentation method is adopted to expand the dataset of tongue images, and the data augmentation method includes rotating, cropping, enhancing contrast, and injecting noise into the human tongue image, etc.
[0032] For example, in terms of geometric transformation: rotate the image by a certain angle; move the image horizontally or vertically to simulate the scenario of camera position change; adjust the size of the image; flip horizontally and vertically. In terms of color space transformation: increase or decrease the brightness of the image; adjust the contrast of the image; change the hue of the image; change the saturation of the color. In terms of image deformation: randomly crop sub-regions from the image; stretch and distort the image, etc. In terms of noise injection: simulate random noise in the sensor or transmission process; randomly insert black and white pixel points into the image to simulate the situation of image damage; simulate the noise caused by photon counting errors.
[0033] Step 3: Build a U-Net segmentation model for human tongue images, and the U-Net segmentation model includes a double-branch encoder network, a decoder network, a bottleneck layer network introducing an attention mechanism, a skip connection network introducing a Vision Transformer encoder, and an output network; Such as Figure 1As shown, the dual-branch encoder includes a dilated convolution branch and a normal convolution branch arranged in parallel. The dilated convolution branch includes multiple cascaded dilated convolution layers, and the normal convolution branch includes multiple cascaded normal convolution units. The normal convolution unit includes a normal convolution layer and a max pooling layer. There are skip connections between each level of the dilated convolution branch and the normal convolution branch. The decoder network includes multiple cascaded transposed convolution layers. The topmost transposed convolution layer is connected to the output network through a 1*1 convolution layer, and the remaining transposed convolution layers are respectively connected to the output network through upsampling layers.
[0034] In some embodiments, the bottleneck layer network includes a 3×3 dilated convolution layer, a 3×3 convolution layer, and a hybrid attention layer. The 3×3 dilated convolution layer is connected to the dilated convolution branch, the 3×3 convolution layer is respectively connected to the 3×3 dilated convolution layer and the normal convolution branch, and the hybrid attention layer is respectively connected to the 3×3 convolution layer and the decoder.
[0035] In some embodiments, the transposed convolution layer includes 1 2×2 transposed convolution layer and 2 3×3 normal convolution layers, and the output network is an encoder.
[0036] In some embodiments, the dilated convolution branch, the normal convolution branch, and the decoder are all four-level structures; the size of the convolution kernel of the dilated convolution layer in the dilated convolution branch is 3×3; the size of the convolution kernel of the normal convolution layer in the normal convolution branch is 3×3; the size of the pooling window of the max pooling layer is 2×2.
[0037] In some embodiments: as Figure 3 shown, the convolutions of the dual-branch encoder network are all double convolutions. For example, the convolution of each level of the dilated convolution branch is a double convolution. The dilation rate of the dilated convolution in the dilated convolution branch is 2; the strides of the two dilated convolutions in the initial double convolution of the dilated convolution branch are both 1; the stride of the first dilated convolution in each of the remaining double convolutions of the dilated convolution branch is 2, and the stride of the second dilated convolution is 1. In the dual-branch encoder network, the dilated convolution adapts to the image size output by the normal convolution. The strides of the normal convolutions in the normal convolution branch are all 1, the padding is 1, and there is no bias. After the double convolution in the normal convolution branch increases the number of channels, 2×2 max pooling is used for size reduction.
[0038] In some embodiments, the initial scale of the image is 0.5, and the image size is reduced to 864×550 and input into the model for convolution.
[0039] In some embodiments, the padding of the dilated convolution in the dilated convolution branch is 2. The pooling layer is not used, and only double dilated convolutions are adopted to vary both the size and the number of channels simultaneously. Through parameter adaptation, the output sizes of the ordinary convolution and the dilated convolution in the same layer can be made the same, so that the skip connections of each layer of the two branches can splice the features, and the spliced results are jointly input into the decoder part.
[0040] In some embodiments, the hybrid attention layer includes a CBAM (Convolutional Block Attention Module) attention module and an SE (Squeeze-and-Excitation) attention mechanism module that are added in parallel.
[0041] Among them, the formulas for CBAM attention and SE attention are as follows: Among them, X is the input, with a size of (B, C, H, W), where B represents the batch size, C represents the number of channels of the image, H represents the height, and W represents the width; the [... ; …] operation is the channel splicing operation; and respectively represent the Sigmoid activation function and the ReLU activation function, and respectively represent the fully connected layer for channel compression and the fully connected layer for channel restoration; is the weighted operation.
[0042] In some embodiments, the skip connection network includes a Vision Transformer encoder, and the Vision Transformer encoder includes a projection head and multiple Transformer encoding modules. Each Transformer encoding module implements the residual connection between the ordinary convolution layer in the double-branch encoder and the transposed convolution layer in the decoder network as a residual connection block.
[0043] In some embodiments, the Transformer encoding module includes two composite blocks. The first composite block includes a Layer Norm layer, a multi-head attention mechanism, and a residual connection; the second composite block includes a Layer Norm layer, a multi-layer perceptron activated by the GeLu and ReLu activation functions respectively, and a residual connection; a Dropout layer is added after each composite block.
[0044] In some embodiments, the Vision Transformer encoder part includes positional encoding at the front. The picture is divided into multiple small patches, and all positional encodings use dynamic encoding. The patch sizes on the four-layer skip connections of the dual-branch encoder are 32, 16, 8, and 4 respectively.
[0045] In some embodiments, before the Vision Transformer encoder, it is necessary to convert the four-dimensional tensor of the image into a three-dimensional tensor. Therefore, a form of positional encoding is adopted, and the image is divided into multiple patches, and positional encoding is performed separately to emphasize the positional features of the image patches.
[0046] The following is the formula for positional encoding: Among them, pos represents the position of the patch in the sequence, i represents the index in the embedding dimension, and the size of the feature dimension is D. Control the change frequencies of different dimensions so that different dimensions have different sensitivities to positions. In traditional positional encoding, an additional 0th patch is separated for dynamic encoding. In the present invention, the adopted embedding has no 0th patch number, and the rest of the patches all use dynamic encoding. Since the Vision Transformer encoder is located on different layers, the patch sizes are not fixed either. For different input sizes and patch-divided images, the positional encoding automatically adjusts the length, and there is no need to define fixed encodings in advance for all possible inputs.
[0047] The formula for multi-head attention is as follows: Among them, Q, K, and V represent query, key, and value respectively. is the learned weight that maps the channel dimension to the dimension of each head. ; is the attention score matrix, which represents the similarity between the query at each position and all keys; h represents the number of attention heads; after the final concatenation, it is projected back to the original number of channels through
[0048] In some embodiments, after the output of each layer of the decoder, an upsampling step is added to change the image into the final two-channel image, and the size of the two-channel image is the same as the size of the original image.
[0049] It should be noted that in the feedforward network, the attention mechanism, positional encoding, multi-head attention linear fully connected layer in the Vision Transformer encoder, and MLP (Multilayer Perceptron) all may result in negative numbers in the graphic matrix. However, in segmentation, the required result is a binary image, and negative numbers will cause improper processing of the final Sigmoid function, leading to model degradation. Therefore, the ReLU activation function is added at the corresponding positions to eliminate negative numbers.
[0050] In some embodiments, the U-Net segmentation model uses the Pytorch environment built based on Python, the optimizer is the Adam optimizer (Adaptive Moment Estimation), and the loss function used is the combined loss composed of BCEWithLogitsLoss and Dice Loss. The calculation formulas of the loss functions are as follows: The BCEWithLogitsLoss formula combines the Sigmoid function and binary cross-entropy. Among them, represents the multiplication sign, N is the number of samples in the binary classification problem, and each sample corresponds to a binary label , and the probability of each sample to be predicted is , is the Sigmoid function, and log is the natural logarithm. In DiceLoss, and respectively represent the label value and predicted value of pixel i, and N is the total number of pixel points, which is equal to the number of pixels in a single image multiplied by the batch size.
[0051] In some embodiments, the tongue image dataset collected by the tongue diagnosis instrument is randomly divided into a training set and a validation set according to a ratio of 9:1. In addition, 300 tongue images are taken out separately as a test set for testing the accuracy rate. Among them, the test set does not participate in gradient update and is only used to test the accuracy rate of the model on other datasets in each round. By comparing the accuracy rate of the test set inside the model, the model with the best epoch is selected.
[0052] Step four, training the U-Net segmentation model using the dataset of the human tongue images to obtain a trained U-Net segmentation model; As Figure 3 shown, Figure 3 is a schematic diagram of the tongue image segmentation result of the present invention. Step five, inputting the human tongue image to be segmented into the trained U-Net segmentation model to obtain the segmented tongue image.
[0053] The embodiment of the present application also provides a tongue image segmentation device, including: An image acquisition module, configured to acquire a human tongue image; A dataset construction module, configured to perform image segmentation and label annotation on the human tongue image to establish a dataset of the human tongue image; A model training module, configured to build a U-Net segmentation model for the human tongue image, where the U-Net segmentation model includes a dual-branch encoder network, a decoder network, a bottleneck layer network introducing an attention mechanism, a skip connection network introducing a Vision Transformer encoder, and an output network; The dual-branch encoder includes a dilated convolution branch and a normal convolution branch arranged in parallel. The dilated convolution branch includes multiple cascaded dilated convolution layers, and the normal convolution branch includes multiple cascaded normal convolution units. The normal convolution unit includes a normal convolution layer and a max pooling layer. There are skip connections between each level of the dilated convolution branch and the normal convolution branch; The decoder network includes multiple cascaded transposed convolution layers. The topmost transposed convolution layer is connected to the output network through a 1*1 convolution layer, and the remaining transposed convolution layers are respectively connected to the output network through upsampling layers; A tongue image segmentation module, configured to input the human tongue image to be segmented into the trained U-Net segmentation model to obtain the segmented tongue image.
[0054] The embodiment of the present application also provides a tongue image segmentation device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a tongue image segmentation method disclosed in the embodiment of the present application are implemented.
[0055] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a tongue image segmentation method disclosed in the embodiment of the present application are implemented.
[0056] The embodiment of the present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of a tongue image segmentation method disclosed in the embodiment of the present application are implemented.
[0057] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0058] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0059] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices, storage media, and program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0060] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0061] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A tongue image segmentation method, characterized in that: include: Step 1, obtaining a human tongue image; Step 2: perform image segmentation and labeling on the human tongue image to establish a data set of human tongue images; Step 3: Build a U-Net segmentation model for human tongue images, where the U-Net segmentation model includes a dual-branch encoder network, a decoder network, a bottleneck layer network with an attention mechanism, a skip connection network with a Vision Transformer encoder, and an output network; The dual-branch encoder includes a hole convolution branch and a common convolution branch arranged in parallel, the hole convolution branch includes a plurality of cascaded hole convolution layers, the common convolution branch includes a plurality of cascaded common convolution units, the common convolution unit includes a common convolution layer and a maximum pooling layer, and each level of the hole convolution branch and the common convolution branch is skip-connected; The decoder network includes a plurality of cascaded deconvolution layers, the top deconvolution layer is connected to the output network through a 1*1 convolution layer, and the remaining deconvolution layers are connected to the output network through upsampling layers respectively; Step 4, using the data set of the human tongue image to train the U-Net segmentation model to obtain a trained U-Net segmentation model; Step 5: Input the human tongue image to be segmented into the trained U-Net segmentation model to obtain the segmented tongue image.
2. A tongue image segmentation method according to claim 1, characterized in that: The jump connection network includes a Vision Transformer encoder, which includes a projection head and multiple Transformer encoding modules, each of which is used as a residual connection block to implement the residual connection of the ordinary convolution layer in the dual-branch encoder and the deconvolution layer in the decoder network.
3. A tongue image segmentation method according to claim 2, characterized in that: The Transformer encoding module includes two composite blocks, the first composite block includes a Layer Norm layer, a multi-head attention mechanism and a residual connection; the second composite block includes a Layer Norm layer, a multi-layer perceptron activated by GeLu and ReLu activation functions respectively, and a residual connection; and a Dropout layer is added after each composite block.
4. A tongue image segmentation method according to any one of claims 1 to 3, characterized in that: The bottleneck layer network includes a 3×3 dilated convolution layer, a 3×3 convolution layer and a mixed attention layer, the 3×3 dilated convolution layer is connected to the dilated convolution branch, the 3×3 convolution layer is respectively connected to the 3×3 dilated convolution layer and the ordinary convolution branch, and the mixed attention layer is respectively connected to the 3×3 convolution layer and the decoder.
5. A tongue image segmentation method according to claim 4, characterized in that: The hybrid attention layer includes a CBAM attention module and a SE attention mechanism module added in parallel.
6. A tongue image segmentation method according to claim 1, characterized in that: The deconvolution layer includes one 2×2 transposed convolution layer and two 3×3 normal convolution layers, and the output network is an encoder.
7. A tongue image segmentation device, characterized in that: include: An image acquisition module, used for acquiring an image of a human tongue; A data set building module is used to segment and label human tongue images and build a data set of human tongue images; A model training module, used to build a U-Net segmentation model for human tongue images, wherein the U-Net segmentation model includes a dual-branch encoder network, a decoder network, a bottleneck layer network introducing an attention mechanism, a skip connection network introducing a Vision Transformer encoder, and an output network; The dual-branch encoder includes a hole convolution branch and a common convolution branch arranged in parallel, the hole convolution branch includes a plurality of cascaded hole convolution layers, the common convolution branch includes a plurality of cascaded common convolution units, the common convolution unit includes a common convolution layer and a maximum pooling layer, and each level of the hole convolution branch and the common convolution branch is skip-connected; The decoder network includes a plurality of cascaded deconvolution layers, the top deconvolution layer is connected to the output network through a 1*1 convolution layer, and the remaining deconvolution layers are connected to the output network through upsampling layers respectively; The tongue image segmentation module is used to input the human tongue image to be segmented into the trained U-Net segmentation model to obtain the segmented tongue image.
8. A tongue image segmentation device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of a tongue image segmentation method described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a tongue image segmentation method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of a tongue image segmentation method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Tongue picture end-to-end identification method based on multi-task network
CN117765566A