Tongue picture segmentation method and device based on semantic segmentation model, equipment and medium
By adopting a semantic segmentation model method in tongue image segmentation, using the combination of parallel network layer and multi-layer perceptron, the problem of poor segmentation of tongue edges and details is solved, and more accurate tongue image segmentation is achieved.
Patent Information
- Application Number
- CN202510126449.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-23
AI Technical Summary
The existing tongue image segmentation method has losses in the edge of the tongue, tongue detail and segmentation accuracy, resulting in poor segmentation of tongue image images in patients.
Using a semantic segmentation model based method, a semantic segmentation model is established by designing an encoder including cascaded multiple parallel network layers and a decoder of a cascaded multiple multi-layer perceptrons to segment the patient's tongue image from the patient's facial image.
It realizes the precise segmentation of the patient's tongue image from the patient's facial image, improves the segmentation accuracy of the tongue edge and details, and meets the application needs of traditional Chinese medicine tongue diagnosis.
Smart Images

Figure CN120032128A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a tongue image segmentation method, device, equipment and medium based on a semantic segmentation model. Background Art
[0002] Tongue diagnosis is an important part of the four diagnostic methods of traditional Chinese medicine, namely, observation, auscultation, inquiry and palpation. It is used to understand the internal health status of the patient by observing the shape, color, texture and changes in the tongue coating. Traditional tongue diagnosis is easily affected by factors such as the subjective experience of the physician and the lighting in the clinic, and the results of tongue diagnosis are not objective and accurate enough.
[0003] To solve this problem, tongue segmentation methods were first applied to remove the face and lips in the patient's facial image, retain the tongue in the patient's facial image, and segment the patient's tongue image from the patient's facial image for intelligent tongue diagnosis. However, the existing tongue segmentation methods are prone to problems such as loss of tongue edges, tongue details, and segmentation accuracy, resulting in poor segmentation of patient tongue images. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a tongue image segmentation method, device, equipment and medium based on a semantic segmentation model, so as to achieve the technical effect of accurately segmenting the patient's tongue image.
[0005] In a first aspect, an embodiment of the present application provides a tongue image segmentation method based on a semantic segmentation model, comprising:
[0006] Acquire a facial image of the patient;
[0007] The patient's facial image is input into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model; wherein the semantic segmentation model includes an encoder and a decoder, the encoder includes a plurality of cascaded parallel network layers, and the decoder includes a plurality of cascaded multi-layer perceptrons.
[0008] In the above implementation process, a semantic segmentation model is established by designing an encoder including multiple cascaded parallel network layers and a decoder including multiple cascaded multi-layer perceptrons. The semantic segmentation model is used to segment the patient's tongue image from the patient's facial image. The parallel computing power of multiple parallel network layers and the nonlinear learning ability of multiple multi-layer perceptrons can be utilized to accurately segment the patient's tongue image from the patient's facial image.
[0009] Furthermore, each of the multiple parallel network layers includes a residual network module and a Transformer module.
[0010] In the above implementation process, by selecting the residual network module and the Transformer module to design the parallel network layer, it is possible to ensure that the global features and detail features of the tongue part are completely extracted from the patient's facial image, thereby achieving accurate segmentation of the patient's tongue image.
[0011] Furthermore, the Transformer module includes an overlapping block merging unit and multiple feature extraction units, and each of the multiple feature extraction units includes a self-attention network and a hybrid feedforward neural network based on the ECA attention mechanism.
[0012] In the above implementation process, by designing a feature extraction unit including a self-attention network and a hybrid feedforward neural network based on the ECA attention mechanism, selecting an overlapping block merging unit and multiple feature extraction units to construct a Transformer module, and using the optimized Transformer module to establish a semantic segmentation model, the semantic segmentation model can be combined with the self-attention mechanism and the ECA attention mechanism to more efficiently extract global features and detail features from the patient's facial images, which is conducive to improving the efficiency of patient tongue image segmentation.
[0013] Furthermore, some of the multiple multilayer perceptrons are target multilayer perceptrons, and the target multilayer perceptron includes a multi-level feature aggregation network.
[0014] In the above implementation process, by adding a multi-level feature aggregation network to some multi-layer perceptrons in the decoder, problems such as loss of segmentation accuracy in the upsampling process can be effectively avoided, thereby more accurately segmenting the patient's tongue image.
[0015] Furthermore, the semantic segmentation model is established based on the Segformer model.
[0016] In the above implementation process, by improving the Segformer model to establish a semantic segmentation model, the semantic segmentation model can be quickly established while ensuring the model performance, which is beneficial to improving the efficiency of patient tongue image segmentation.
[0017] Furthermore, before inputting the patient's facial image into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model, the method further includes:
[0018] intercepting a region of interest in the patient's facial image as a target patient's facial image;
[0019] The patient facial image is updated to the target patient facial image.
[0020] In the above implementation process, by intercepting the region of interest in the patient's facial image as the target patient's facial image and updating the patient's facial image to the target patient's facial image, it is possible to remove the region images in the patient's facial image that are irrelevant to tongue image segmentation, intercept the region of interest image in the patient's facial image for subsequent tongue image segmentation, thereby effectively reducing the processing pressure of the semantic segmentation model and facilitating the improvement of the patient tongue image segmentation efficiency.
[0021] Further, before inputting the patient's facial image into a pre-established semantic segmentation model to obtain the patient tongue image output by the semantic segmentation model, it further includes:
[0022] Performing image preprocessing on the patient's facial image; wherein, the image preprocessing includes image filtering and / or image enhancement.
[0023] In the above implementation process, by first performing one or more of image filtering and image enhancement and other image preprocessing on the patient's facial image, and then inputting the preprocessed patient's facial image into a pre-established semantic style model, it is possible to improve the image quality of the patient's facial image, thereby ensuring that the semantic segmentation model accurately segments the patient tongue image and facilitating the improvement of the patient tongue image segmentation efficiency.
[0024] In a second aspect, an embodiment of the present application provides a tongue image segmentation device based on a semantic segmentation model, including:
[0025] An image acquisition module, configured to acquire a patient's facial image;
[0026] An image segmentation module, configured to input the patient's facial image into a pre-established semantic segmentation model to obtain the patient tongue image output by the semantic segmentation model; wherein, the semantic segmentation model includes an encoder and a decoder, the encoder includes a plurality of cascaded parallel network layers, and the decoder includes a plurality of cascaded multi-layer perceptrons.
[0027] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the above-described method is implemented.
[0028] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium including a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the above-described method. Description of the Drawings
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0030] Figure 1 A schematic diagram of a flow chart of a tongue image segmentation method based on a semantic segmentation model provided in the first embodiment of the present application;
[0031] Figure 2 This is a structural diagram of an ECA attention mechanism of an optional embodiment of the first embodiment of the present application;
[0032] Figure 3 A schematic structural diagram of a Segformer model according to another optional embodiment of the first embodiment of the present application;
[0033] Figure 4 A schematic diagram of the structure of a semantic segmentation model according to another optional embodiment of the first embodiment of the present application;
[0034] Figure 5 A schematic diagram of the structure of a tongue image segmentation device based on a semantic segmentation model provided in the second embodiment of the present application;
[0035] Figure 6 A schematic diagram of the structure of an electronic device provided in the third embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.
[0037] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0038] Tongue diagnosis is an important part of the four diagnostic methods of traditional Chinese medicine, namely, observation, auscultation, inquiry and palpation. It is used to understand the internal health status of the patient by observing the shape, color, texture and changes in the tongue coating. Traditional tongue diagnosis is easily affected by factors such as the subjective experience of the physician and the lighting in the clinic, and the results of tongue diagnosis are not objective and accurate enough.
[0039] To solve this problem, the tongue image segmentation method was applied to remove the face and lips of the patient's facial image, retain the tongue part of the patient's facial image, and segment the patient's tongue image from the patient's facial image for intelligent tongue diagnosis.
[0040] In related technologies, semantic segmentation models based on CNN (Convolutional Neural Network), such as Unet series models or DeepLab series models, are mainly used to segment the patient's tongue image from the patient's facial image. However, the CNN-based semantic segmentation model is difficult to effectively extract the global contextual semantic information in the patient's facial image, and is prone to problems such as judgment errors or partial missing of the tongue in actual application scenarios.
[0041] Compared with traditional convolutional neural networks, Transformer has a wider field of view and can more effectively extract global features. Although Transformer-based semantic segmentation models, such as the Segformer model, have good segmentation effects in other fields, they rely on the Positional Embedding structure to supplement the position information and cannot accurately transmit the position information of each patch in the patient's facial image. They are prone to problems such as loss of tongue edge details. In addition, the Transformer-based semantic segmentation model has a large number of parameters, which is not conducive to deployment and application.
[0042] In addition, most existing semantic segmentation models adopt an encoding-decoding structure, and the feature map needs to be upsampled during the decoding process to restore the feature map to the same size as the patient's facial image, which often leads to problems such as loss of segmentation accuracy.
[0043] In the specific application scenario of TCM tongue segmentation, the edge of the tongue is similar in color to the oral cavity, palate, etc., and detailed information such as tooth marks on the edge of the tongue is also an important part of tongue diagnosis. Therefore, the semantic segmentation model needs to pay more attention to the detailed information of the edge of the tongue.
[0044] Obviously, the application of existing tongue image segmentation methods is prone to problems such as loss of tongue edges, tongue details and segmentation accuracy, resulting in poor segmentation of patient tongue images and making it difficult to meet the application needs of traditional Chinese medicine tongue diagnosis.
[0045] To this end, the present application proposes a tongue image segmentation method based on a semantic segmentation model. The semantic segmentation model is established by designing an encoder including multiple cascaded parallel network layers and a decoder including multiple cascaded multi-layer perceptrons. The semantic segmentation model is used to segment the patient's tongue image from the patient's facial image. The parallel computing power of multiple parallel network layers and the nonlinear learning ability of multiple multi-layer perceptrons can be utilized to accurately segment the patient's tongue image from the patient's facial image.
[0046] Combine the following Figure 1 A tongue image segmentation method based on a semantic segmentation model provided in the first embodiment of the present application is described. The tongue image segmentation method based on a semantic segmentation model provided in the first embodiment of the present application can be executed by a user terminal.
[0047] Please see Figure 1 , Figure 1 The first embodiment of the present application provides a tongue image segmentation method based on a semantic segmentation model, comprising steps S101 to S102:
[0048] S101, obtaining a facial image of a patient;
[0049] S102. Input the patient's facial image into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model; wherein the semantic segmentation model includes an encoder and a decoder, the encoder includes a plurality of cascaded parallel network layers, and the decoder includes a plurality of cascaded multi-layer perceptrons.
[0050] As an example, according to actual application requirements, multiple parallel network layers are pre-designed, the multiple parallel network layers are cascaded to construct an encoder, and multiple multi-layer perceptrons are designed, the multiple multi-layer perceptrons are cascaded to construct a decoder, and based on the encoding-decoding structure, a semantic segmentation model is established according to the constructed encoder and decoder.
[0051] Parallel Networks refers to the process of processing different aspects of data simultaneously by setting up multiple sub-networks or processing paths in parallel in the neural network architecture. This design can utilize the computing power of multiple sub-networks or processing paths to enhance the model's ability to learn features from different perspectives and ultimately integrate these features to make more accurate predictions.
[0052] Multilayer Perceptron (MLP) is a feedforward artificial neural network model that maps multiple input data sets to a single output data set. Multilayer Perceptron can learn any complex nonlinear function, making it perform well in dealing with complex problems. In practical applications, lightweight multilayer perceptrons can be selected.
[0053] When the user needs to perform tongue segmentation, the user can use an image acquisition device to acquire the patient's facial image, and input the patient's facial image into the user terminal, so that the user terminal acquires the patient's facial image.
[0054] After obtaining the patient's facial image, the user terminal calls the pre-established semantic segmentation model and inputs the patient's facial image into the semantic segmentation model. At this time, the encoder in the semantic segmentation model will completely extract the global features and detail features of the tongue from the patient's facial image. The decoder in the semantic segmentation model will restore the patient's tongue image based on the global features and detail features of the tongue, thereby obtaining the patient's tongue image output by the semantic segmentation model.
[0055] By pre-designing an encoder including multiple cascaded parallel network layers and designing a decoder including multiple cascaded multi-layer perceptrons to establish a semantic segmentation model, the parallel computing capabilities of the multiple parallel network layers can be utilized to completely extract the global features and detail features of the tongue part from the patient's facial image, and the nonlinear learning capabilities of the multiple multi-layer perceptrons can be utilized to accurately restore the global features and detail features of the tongue part into the patient's tongue image, thereby achieving accurate segmentation of the patient's tongue image from the patient's facial image.
[0056] The embodiment of the present application establishes a semantic segmentation model by designing an encoder including multiple cascaded parallel network layers and a decoder including multiple cascaded multi-layer perceptrons. The semantic segmentation model is used to segment the patient's tongue image from the patient's facial image. The parallel computing power of multiple parallel network layers and the nonlinear learning ability of multiple multi-layer perceptrons can be used to accurately segment the patient's tongue image from the patient's facial image.
[0057] In an optional embodiment, each of the multiple parallel network layers includes a residual network module and a Transformer module.
[0058] As an example, according to actual application requirements, a residual network module and a Transformer module are selected, the residual network module and the Transformer module are connected, and a parallel network layer is designed. According to this operation, multiple parallel network layers are designed, and multiple parallel network layers are cascaded to construct an encoder.
[0059] Residual Network (ResNet) is a convolutional neural network. The characteristics of residual network are that it is easy to optimize and can improve accuracy by increasing the depth. The residual block inside it uses skip connections. This design allows low-level detail information to be directly passed to high-level layers without having to go through stacked layers of nonlinear transformations. In the back-propagation process, the skip connection ensures that the gradient can be smoothly passed back from the output layer to the input layer, alleviating the gradient vanishing problem caused by increasing the depth in the deep neural network, thereby ensuring that the model can effectively learn local detail features.
[0060] Transformer is a deep learning network architecture based on the self-attention mechanism. Transformer realizes parallel processing of input sequences through the self-attention mechanism, greatly improving the training speed and reasoning efficiency of the model, and more effectively capturing long-distance dependencies, allowing the model to understand the overall structure, which is particularly important for tasks such as processing complex images.
[0061] By selecting the residual network module and the Transformer module to design the parallel network layer, we can fully utilize the ability of the residual network module to capture detail features and the ability of the Transformer module to effectively focus on global features based on the self-attention mechanism, and completely extract the global features and detail features of the tongue from the patient's facial image, which helps to accurately segment the patient's tongue image from the patient's facial image.
[0062] The embodiment of the present application designs parallel network layers by selecting residual network modules and Transformer modules, which can ensure that the global features and detail features of the tongue part are completely extracted from the patient's facial image, thereby achieving accurate segmentation of the patient's tongue image.
[0063] In an optional embodiment, the Transformer module includes an overlapping block merging unit and multiple feature extraction units, each of the multiple feature extraction units includes a self-attention network and a hybrid feedforward neural network based on the ECA attention mechanism.
[0064] As an example, considering that the model's feature learning and extraction capabilities based on the self-attention mechanism are still relatively limited and it is difficult to fully meet the application requirements of the specific application scenario of TCM tongue image segmentation, the Transformer module can be optimized.
[0065] According to the actual application requirements, the overlapping block merging unit, the self-attention network and the hybrid feedforward neural network are selected, and the ECA attention mechanism is introduced in the hybrid feedforward neural network. After obtaining the self-attention network and the hybrid feedforward neural network based on the ECA attention mechanism, the self-attention network and the hybrid feedforward neural network based on the ECA attention mechanism are connected to design a feature extraction unit. According to this operation, multiple feature extraction units are designed, the overlapping block merging unit and multiple feature extraction units are connected, and the Transformer module is constructed, thereby constructing an encoder.
[0066] Overlap Patch Merging is a technique used to process images or feature maps, especially when CNN or Transformer is involved. This technique merges image patches with overlapping areas to improve the model's understanding of spatial information, enhance feature representation, reduce redundancy, and improve the final prediction performance.
[0067] The Mixed Feedforward Network (Mix-FFN) is an important component in the Segformer model, responsible for feature mixing and extraction. The original Mixed Feedforward Network consists only of convolutional neural networks.
[0068] The original Transformer module usually includes overlapping block merging units, multiple self-attention networks, and multiple hybrid feedforward neural networks.
[0069] The SE channel attention mechanism (Squeeze-and-Excitation) is a typical attention mechanism that adjusts the feature weights of each channel by learning the nonlinear interactions between features. However, the fully connected layer in the SE channel attention mechanism will greatly increase the number of parameters and computational complexity of the model. To solve this problem, the ECA (Efficient Channel Attention) attention mechanism is proposed. Different from the SE channel attention mechanism, the ECA attention mechanism uses a 1×1 convolutional layer instead of a fully connected layer, which not only reduces the number of parameters of the model, but also enables the model to achieve better performance while maintaining low computational complexity. The structure of the ECA attention mechanism is shown in the figure. Figure 2 shown.
[0070] In order to further improve the model's ability to learn and extract features, the ECA attention mechanism is introduced into the hybrid feedforward neural network. In this way, the advantages of the ECA attention mechanism can be used in combination with the original self-attention mechanism to more efficiently extract key features from the patient's facial images and improve the accuracy and efficiency of semantic segmentation. This not only helps to improve the performance of the model, but also can reduce the model training time and computational complexity to a certain extent, making the model more efficient and practical.
[0071] The embodiment of the present application designs a feature extraction unit including a self-attention network and a hybrid feedforward neural network based on the ECA attention mechanism, selects an overlapping block merging unit and multiple feature extraction units to construct a Transformer module, and uses the optimized Transformer module to establish a semantic segmentation model. The semantic segmentation model can be combined with the self-attention mechanism and the ECA attention mechanism to more efficiently extract global features and detail features from the patient's facial images, which is beneficial to improving the efficiency of patient tongue image segmentation.
[0072] In an optional embodiment, some of the multiple multilayer perceptrons are target multilayer perceptrons, and the target multilayer perceptron includes a multi-level feature aggregation network.
[0073] As an example, semantic segmentation models usually adopt an encoding-decoding structure. When processing an image, this structure first extracts features from the image through an encoder, and then upsamples it through a decoder to restore the original size of the image and perform pixel-level classification. However, in this structure, the size of the feature map generated by the decoder is usually much smaller than the size of the original input image. This means that during the upsampling process, the model needs to restore the coarser feature map to the same size as the original image. In this process, it is easy to lose detail information, resulting in a decrease in the accuracy of the predicted features.
[0074] Considering that the multi-layer perceptron may suffer from loss of segmentation accuracy during upsampling, a multi-level feature aggregation (MLA) scheme can be used to avoid problems such as loss of segmentation accuracy during upsampling.
[0075] A multi-level feature aggregation upsampling scheme is adopted, that is, feature maps of different dimensions are gradually upsampled to the required size using convolutional neural networks and bilinear interpolation upsampling methods. The feature map is magnified by 2 in each iteration, and finally different features are fused. MLA itself does not involve complex neural network calculations, and will not ultimately lead to a sudden increase in the number of model parameters.
[0076] According to actual application requirements, some multilayer perceptrons are selected from multiple multilayer perceptrons, and these multilayer perceptrons are replaced with target multilayer perceptrons including a multi-level feature aggregation network, so as to establish a semantic segmentation model.
[0077] In practical applications, a multi-level feature aggregation network can be added to some of the multi-layer perceptrons in the previous cascade.
[0078] The embodiment of the present application can effectively avoid problems such as loss of segmentation accuracy during upsampling by adding a multi-level feature aggregation network to some multi-layer perceptrons in the decoder, thereby more accurately segmenting the patient's tongue image.
[0079] In an optional embodiment, the semantic segmentation model is established based on the Segformer model.
[0080] As an example, according to actual application requirements, the original Segformer model is improved in the above manner to establish a semantic segmentation model.
[0081] It is understandable that the established semantic segmentation model is based on the original Segformer model by adding parallel network layers, ECA attention mechanism and multi-level feature aggregation network.
[0082] For example, the original Segformer model structure diagram is as follows Figure 3 As shown in the figure, the structural diagram of the established semantic segmentation model is as follows Figure 4 Based on the original Segformer model, the four cascaded Transformer modules in the original Segformer model (i.e. Figure 3 The 4 Transformer Blocks in the example are improved to 4 parallel network layers in cascade (i.e. Figure 4 Each parallel network layer includes a residual network module (i.e. Figure 4 ResNet Layer in ) and the optimized Transformer module (i.e. Figure 4 The Transformer Block in the Transformer module includes an overlapping block merging unit (i.e. Figure 4 Overlap Patch Merging in ) and N feature extraction units, each feature extraction unit includes a self-attention network (i.e. Figure 4 Eifficient Self-Attn in ) and a hybrid feedforward neural network based on the ECA attention mechanism (i.e. Figure 4 Mix-FFN ECA in ), and a multi-layer perceptron cascaded in the original Segformer model (i.e. Figure 3 The MLP layer in is improved to a target multi-layer perceptron, that is, MLP and MLA are combined with convolutional neural network and bilinear interpolation to perform multi-level feature aggregation upsampling.
[0083] The working principle of the semantic segmentation model is as follows: directly connect the feature maps of multiple different dimensions in the middle layer of the encoder to the decoder, change the dimensions of multiple feature maps to the same size, and then use bilinear interpolation upsampling to enlarge the original image size by one-fourth after passing the multi-level features of the same dimension through the feedforward neural network layer. The four layers of features are spliced together to achieve feature fusion, and the number of channels is transformed through another network to make it equal to the number of categories that need to be predicted in the semantic segmentation task, which is convenient for subsequent prediction of masks.
[0084] The embodiment of the present application establishes a semantic segmentation model by improving the Segformer model, which can quickly establish a semantic segmentation model while ensuring the performance of the model, thereby helping to improve the efficiency of patient tongue image segmentation.
[0085] In an optional embodiment, before the patient's facial image is input into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model, it also includes: capturing a region of interest in the patient's facial image as a target patient's facial image; and updating the patient's facial image to the target patient's facial image.
[0086] As an example, considering that in actual applications, the acquired patient facial images may contain areas irrelevant to tongue segmentation, such as the patient's hair, upper half of the face or neck, such patient facial images can be cropped, and images of areas of interest such as the lower half of the face or mouth in the patient's facial images can be captured for subsequent tongue segmentation.
[0087] After acquiring the patient's facial image, the region of interest in the patient's facial image is captured as the target patient's facial image, and the patient's facial image is updated to the target patient's facial image, so that the updated patient's facial image is input into a pre-established semantic segmentation model, thereby obtaining the patient's tongue image output by the semantic segmentation model.
[0088] The embodiment of the present application captures the region of interest in the patient's facial image as the target patient's facial image, updates the patient's facial image to the target patient's facial image, removes the region image in the patient's facial image that is not related to tongue segmentation, and captures the region image of interest in the patient's facial image for subsequent tongue segmentation, thereby effectively reducing the processing pressure of the semantic segmentation model and facilitating improving the efficiency of patient tongue image segmentation.
[0089] In an optional embodiment, before inputting the patient's facial image into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model, it also includes: performing image preprocessing on the patient's facial image; wherein the image preprocessing includes image filtering and / or image enhancement.
[0090] As an example, considering that in actual applications, the quality of the acquired patient facial images may be poor, such patient facial images may be subjected to image preprocessing such as image filtering and image enhancement, and the preprocessed patient facial images may be used for subsequent tongue segmentation.
[0091] The embodiment of the present application can improve the image quality of the patient's facial image by first performing one or more image preprocessing such as image filtering and image enhancement on the patient's facial image, and then inputting the preprocessed patient's facial image into a pre-established semantic style model, thereby ensuring that the semantic segmentation model accurately segments the patient's tongue image, which is beneficial to improving the efficiency of patient tongue image segmentation.
[0092] The tongue image segmentation method based on the semantic segmentation model provided in the embodiment of the present application can improve the model for the problem that the Segformer model has poor segmentation effect on the traditional Chinese medicine tongue image, and solve the following technical problems:
[0093] 1. By adding parallel network layers to the Segformer model, high-dimensional detail features can be integrated while obtaining global features.
[0094] 2. In order to make the Transformer architecture more effective in extracting semantic and feature information, the Transformer architecture is combined with a convolutional neural network and the ECA attention mechanism is added to achieve more refined feature weight allocation and optimize the representation ability of the model.
[0095] 3. In order to reduce the loss of accuracy in the upsampling process, a structure combining progressive upsampling and convolutional neural network is adopted to alleviate the loss of segmentation accuracy in the upsampling process.
[0096] Please see Figure 5 , Figure 5 A structural schematic diagram of a tongue image segmentation device based on a semantic segmentation model is provided for the second embodiment of the present application. The second embodiment of the present application provides a tongue image segmentation device based on a semantic segmentation model, comprising: an image acquisition module 201, for acquiring a patient's facial image; an image segmentation module 202, for inputting the patient's facial image into a pre-established semantic segmentation model to obtain a patient's tongue image output by the semantic segmentation model; wherein the semantic segmentation model comprises an encoder and a decoder, the encoder comprises a plurality of cascaded parallel network layers, and the decoder comprises a plurality of cascaded multi-layer perceptrons.
[0097] In an optional embodiment, each of the multiple parallel network layers includes a residual network module and a Transformer module.
[0098] In an optional embodiment, the Transformer module includes an overlapping block merging unit and multiple feature extraction units, each of the multiple feature extraction units includes a self-attention network and a hybrid feedforward neural network based on the ECA attention mechanism.
[0099] In an optional embodiment, some of the multiple multilayer perceptrons are target multilayer perceptrons, and the target multilayer perceptron includes a multi-level feature aggregation network.
[0100] In an optional embodiment, the semantic segmentation model is established based on the Segformer model.
[0101] In an optional embodiment, before the patient's facial image is input into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model, it also includes: capturing a region of interest in the patient's facial image as a target patient's facial image; and updating the patient's facial image to the target patient's facial image.
[0102] In an optional embodiment, before inputting the patient's facial image into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model, it also includes: performing image preprocessing on the patient's facial image; wherein the image preprocessing includes image filtering and / or image enhancement.
[0103] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0104] Please see Figure 6 , Figure 6 A structural diagram of an electronic device is provided for the third embodiment of the present application. The third embodiment of the present application provides an electronic device 30, including a processor 301, a memory 302, and a computer program stored in the memory 302 and configured to be executed by the processor 301; when the processor 301 executes the computer program, the method described in the first embodiment of the present application is implemented, and the same beneficial effects can be achieved.
[0105] The processor 301 may implement the method described in the first embodiment of the present application when reading the computer program from the memory 302 via the bus 303 and executing the computer program.
[0106] Processor 301 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 301 can be a microprocessor.
[0107] The memory 302 may be used to store instructions executed by the processor 301 or data related to the execution of instructions. These instructions and / or data may include codes for implementing some functions or all functions of one or more modules described in the embodiments of the present application. The processor 301 of this embodiment may be used to execute instructions in the memory 302 to implement the method described in the first embodiment of the present application. The memory 302 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memory known to those skilled in the art.
[0108] The fourth embodiment of the present application provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in the first embodiment of the present application, and can achieve the same beneficial effects as the method.
[0109] In summary, the embodiments of the present application provide a tongue image segmentation method, device, equipment and medium based on a semantic segmentation model, wherein the tongue image segmentation method based on the semantic segmentation model comprises: obtaining a patient's facial image; inputting the patient's facial image into a pre-established semantic segmentation model to obtain a patient's tongue image output by the semantic segmentation model; wherein the semantic segmentation model comprises an encoder and a decoder, the encoder comprises a plurality of cascaded parallel network layers, and the decoder comprises a plurality of cascaded multilayer perceptrons. The embodiments of the present application establish a semantic segmentation model by designing an encoder comprising a plurality of cascaded parallel network layers and a decoder comprising a plurality of cascaded multilayer perceptrons, and adopting the semantic segmentation model to segment the patient's tongue image from the patient's facial image, and can utilize the parallel computing power of a plurality of parallel network layers and the nonlinear learning ability of a plurality of multilayer perceptrons to accurately segment the patient's tongue image from the patient's facial image.
[0110] The above embodiments are only described as examples. The various embodiments of the present application can be combined or nested with each other, and the present application is not limited to this.
[0111] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0112] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0113] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0114] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0115] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0116] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
Claims
1. A tongue image segmentation method based on a semantic segmentation model, characterized in that: include: Acquire a facial image of the patient; The patient's facial image is input into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model; wherein the semantic segmentation model includes an encoder and a decoder, the encoder includes a plurality of cascaded parallel network layers, and the decoder includes a plurality of cascaded multi-layer perceptrons.
2. The method according to claim 1, characterized in that Each of the multiple parallel network layers includes a residual network module and a Transformer module.
3. The method according to claim 2, characterized in that The Transformer module includes an overlapping block merging unit and multiple feature extraction units, each of which includes a self-attention network and a hybrid feedforward neural network based on an ECA attention mechanism.
4. The method according to claim 1, characterized in that: Some of the multiple multilayer perceptrons are target multilayer perceptrons, and the target multilayer perceptron includes a multi-level feature aggregation network.
5. The method according to claim 1, characterized in that The semantic segmentation model is established based on the Segformer model.
6. The method according to claim 1, characterized in that Before inputting the patient's facial image into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model, the method further includes: intercepting a region of interest in the patient's facial image as a target patient's facial image; The patient facial image is updated to the target patient facial image.
7. The method according to any one of claims 1 to 6, characterized in that: Before inputting the patient's facial image into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model, the method further includes: Perform image preprocessing on the patient's facial image; wherein the image preprocessing includes image filtering and / or image enhancement.
8. A tongue image segmentation device based on a semantic segmentation model, characterized in that: include: An image acquisition module, used for acquiring a facial image of a patient; An image segmentation module is used to input the patient's facial image into a pre-established semantic segmentation model to obtain the patient's tongue image output by the semantic segmentation model; wherein the semantic segmentation model includes an encoder and a decoder, the encoder includes a plurality of cascaded parallel network layers, and the decoder includes a plurality of cascaded multi-layer perceptrons.
9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.