Tongue picture segmentation method and device, medium and equipment

By introducing local spatiotemporal convolution and global timing attention mechanisms in dual U-net networks, the problem of insufficient dynamic change capture in tongue image segmentation is solved, and more efficient and accurate tongue image segmentation is achieved, which is suitable for real-time analysis of dynamic medical tongue images.

CN120235840AActive Publication Date: 2025-07-01BEIJING UNIV OF CHINESE MEDICINE

Patent Information

Application Number
CN202510348786.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-01
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The prior art cannot effectively capture the dynamic changes of tongue images in the time series, resulting in insufficient segmentation accuracy and traditional methods are inefficient when integrating multi-frame information.

Method used

A local spatiotemporal convolution module and a global timing attention mechanism are introduced in the dual U-net network to build an improved dual U-net network, and the tongue image segmentation is performed by collecting tongue image data sets and combining spatiotemporal convolution and timing attention mechanisms.

Benefits of technology

It realizes accurate capture of dynamic changes of tongue objects in continuous frames, improves the accuracy and efficiency of segmentation, and is suitable for medical scenarios with dynamic monitoring and real-time segmentation, providing more accurate diagnostic basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235840A_ABST
    Figure CN120235840A_ABST
Patent Text Reader

Abstract

The invention discloses a tongue picture segmentation method and device, a medium and equipment, and relates to the technical field of tongue picture segmentation. A local space-time convolution module is added to an encoder of one U-net branch of the double U-net network, and a global time sequence attention mechanism module is added to a decoder of the U-net branch, so that the double U-net network is improved. The improvement aims to solve the technical problem that the segmentation accuracy is insufficient due to the fact that tongue picture biological structure changes in a time sequence cannot be correctly expressed in the prior art. A local space-time convolution mechanism and a global time sequence attention mechanism are respectively introduced into an encoder structure and a decoder structure of one branch of the double-Unet network so as to capture more accurate dynamic changes of tongue pictures in continuous frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tongue image segmentation, and particularly relates to a tongue image segmentation method, device, medium and equipment. Background Art

[0002] Tongue diagnosis is a core method in traditional Chinese medicine. Practice has proved that the state of the tongue can directly reflect a person's health condition. In order to automatically identify potential diseases of a person through computer vision technology, it is very important to accurately segment the tongue part from tongue images.

[0003] Currently, medical image segmentation mostly relies on single-frame tongue images or basic sequence processing techniques. However, the dynamic biological structure features are difficult to accurately capture in single-frame tongue images, and conventional methods are inefficient and limited in accuracy when integrating multi-frame information. Although some systems attempt to integrate multi-frame information through simple frame stacking or averaging methods, this method often fails to correctly represent the changes in the biological structure of tongue images in the time series, resulting in insufficient segmentation accuracy. Summary of the Invention

[0004] Based on this, in order to solve the technical problems in the prior art, the present invention provides a tongue image segmentation method, device, medium and equipment.

[0005] The present invention provides a tongue image segmentation method, including: Adding a local spatio-temporal convolution module to the encoder of one U-net branch in the dual U-net network, and adding a global temporal attention mechanism module to the decoder of this U-net branch to obtain an improved U-net branch; using the unimproved U-net branch in the dual U-net network as the first branch and the improved U-net branch in the dual U-net network as the second branch to construct an improved dual U-net network including the first branch, the second branch and a segmentation layer.

[0006] Collecting a tongue image dataset to train the improved dual U-net network to obtain a tongue image segmentation model; including: Collecting a tongue image dataset, where the tongue image dataset includes consecutive-frame tongue images and the true segmentation labels corresponding to each frame of tongue image; stacking the tongue images of each frame belonging to the same sample to form a multi-dimensional tensor; Initializing the weight coefficients of the improved dual U-net network; Inputting the multi-dimensional tensor into the first branch of the improved dual U-net network to output initial tongue spatial features; inputting the multi-dimensional tensor and the initial tongue spatial features into the second branch of the improved dual U-net network to output tongue temporal global decoding features; inputting the tongue temporal global decoding features into the segmentation layer to output the predicted segmentation labels corresponding to each frame of tongue image; Calculate the loss value of the true segmentation label and the predicted segmentation label through the loss function; iteratively improve the weight coefficients in the dual U-net network according to the loss value, and re-segment the new samples according to the iterated weight coefficients; the loss function is: Wherein, is the total loss function; is the spatial loss function corresponding to the first branch, n is the number of frames of tongue image data to be segmented, is the predicted segmentation label of the t-th frame of tongue image, is the true segmentation label of the t-th frame of tongue image; is the adjustment parameter; is the temporal loss function corresponding to the second branch, is the predicted segmentation label of the (n + 1)-th frame of tongue image predicted by the second branch, is the true segmentation label of the (n + 1)-th frame of tongue image; Repeat the iterative steps until the loss function converges or reaches the preset number of iterations to obtain the trained improved dual U-net network.

[0007] Input the tongue image data to be segmented into the tongue image segmentation model, extract the spatial features of each frame of tongue image in the tongue image data to be segmented through the first branch, obtain the initial tongue image spatial features containing the morphological and size information of the tongue physiological structure of each frame of tongue image, and input the initial tongue image spatial features into the second branch; through the local spatio-temporal convolution module of the second branch encoder, perform local spatio-temporal convolution operations on the initial tongue image spatial features: Wherein, is the (n + 1)-th frame of tongue image in the tongue image data to be segmented; STCM is the spatio-temporal convolution operation, is the output of the spatio-temporal convolution operation; Conv3D is the 3D local convolution operation; is the output of the local spatio-temporal convolution operation; this operation locally encodes the initial tongue image spatial features along the time dimension to obtain the tongue image temporal local encoding features.

[0008] Decode the tongue image temporal local encoding features through the global temporal attention mechanism module of the second branch decoder: Wherein, SelfAttentionis a convolutional long short-term neural network or Transformer the attention operation in is the output of the attention operation, is the input of the attention operation; this operation captures the global temporal correlation of the tongue image spatial features in the time dimension to obtain the global decoded tongue image temporal features; the global decoded tongue image temporal features are segmented through a segmentation layer to obtain the segmentation result of the tongue image data to be segmented.

[0009] The present invention provides a tongue image segmentation device, including: A model construction module, which is used to add a local spatio-temporal convolution module to the encoder of one of the U-net branches in the double U-net network and add a global temporal attention mechanism module to the decoder of this U-net branch to obtain an improved U-net branch; using the unimproved U-net branch in the double U-net network as the first branch and the improved U-net branch in the double U-net network as the second branch, construct an improved double U-net network including the first branch, the second branch and a segmentation layer; A model training module, which is used to collect a tongue image dataset to train the improved double U-net network to obtain a tongue image segmentation model; A tongue image segmentation module, which is used to input the tongue image data to be segmented into the tongue image segmentation model, extract the spatial features of each frame of tongue image in the tongue image data to be segmented through the first branch to obtain the initial tongue image spatial features containing the physiological structure, shape and size information of the tongue image, and input the initial tongue image spatial features into the second branch; the local spatio-temporal convolution module of the second branch encoder performs local spatio-temporal convolution operations on the initial tongue image spatial features to locally encode the initial tongue image spatial features along the time dimension to obtain the tongue image temporal local encoded features; the global temporal attention mechanism module of the second branch decoder decodes the tongue image temporal local encoded features to capture the global temporal correlation of the tongue image spatial features in the time dimension to obtain the tongue image temporal global decoded features; the tongue image temporal global decoded features are segmented through a segmentation layer to obtain the segmentation result of the tongue image data to be segmented.

[0010] The present invention provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned tongue image segmentation method is implemented.

[0011] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned tongue image segmentation method is implemented.

[0012] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects: In the tongue image segmentation method provided by the present invention, local spatio-temporal convolution and global temporal attention mechanism are respectively introduced into the encoder and decoder structures of one branch of the dual Unet network to capture the more accurate dynamic changes of the tongue image in consecutive frames. Specifically, the role of the local spatio-temporal convolution module is to extract spatio-temporal features of local regions at the lower levels of the network, so as to capture the dynamic changes of the tongue image in consecutive frames. The global temporal attention mechanism, on the other hand, plays a role in the high-level network. It can focus on the temporal dependencies between frames in the whole sequence and strengthen this dependency. Through this design of spatio-temporal convolution, the Unet network can directly process the input of consecutive multi-frame images and extract joint spatio-temporal features. This not only enables the model to effectively complete the segmentation task of the current frame, but also can optimize the segmentation result of the current frame with the help of the segmentation information of historical frames, so as to achieve a more coherent and smooth segmentation effect in the time series. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0014] Figure 1 It is a schematic flow chart of a tongue image segmentation method provided by the present invention; Figure 2 It is a schematic diagram of a dual Unet network structure provided by the present invention; Figure 3 It is a schematic diagram of a computer device for implementing the tongue image segmentation method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0016] The prior art faces multiple challenges in processing dynamic medical tongue images. First, single-frame tongue image processing cannot effectively capture the changes in continuous physiological activities, such as the respiratory cycle or gastrointestinal peristalsis. Second, although some systems attempt to integrate multi-frame information through simple frame stacking or averaging methods, this method often fails to correctly represent the changes in biological structures in the time series, resulting in insufficient segmentation accuracy. To solve this problem, the present invention introduces an attention mechanism operator model, whose excellent temporal information processing ability enables the system to more accurately represent and segment the complex changes in dynamic tongue images. In addition, traditional multi-frame fusion techniques are usually complex to implement in existing systems, require high computing resources, are difficult to respond in real time, and lack interpretability for the processing results of dynamic tongue images. The present invention provides a segmentation system designed specifically for dynamic medical tongue images, which can effectively process and segment continuous tongue image data through a dual Unet structure and multi-frame fusion technology. This system is particularly suitable for medical scenarios that require dynamic monitoring and real-time segmentation, such as tongue lesion detection and dynamic change analysis, etc., and can provide doctors with more accurate diagnostic basis, enhancing the speed and quality of clinical decision-making. The present invention solves the problem of insufficient accuracy in processing dynamic tongue images by combining a dual Unet structure, multi-frame fusion technology, and an attention mechanism operator network.

[0017] Embodiment 1 Figure 1 shows the flow of the tongue image segmentation method of this embodiment. Specifically, in combination with Figure 1 This method will be described in detail, which specifically includes the following steps:

[0018] S1: Add a local spatio-temporal convolution module to the encoder of one of the U-net branches in the dual U-net network, and add a global temporal attention mechanism module to the decoder of this U-net branch to obtain an improved U-net branch; use the unimproved U-net branch in the dual U-net network as the first branch, and use the improved U-net branch in the dual U-net network as the second branch to construct an improved dual U-net network including the first branch, the second branch, and a segmentation layer.

[0019] As Figure 2As shown in the figure, a dual-stream structure is designed: the spatial stream, that is, the first branch is responsible for processing the traditional Unet spatial segmentation task. The temporal stream: that is, the second branch captures temporal dependencies through spatio-temporal convolution and temporal attention mechanisms. When starting to work, first, a dual-Unet network combined with an attention mechanism operator model processes tongue images from various medical imaging devices. These devices may include, but are not limited to, ultrasound, MRI, or CT scanners, which provide a series of dynamic or static medical tongue images. Before the tongue images are fed into the system for deep learning processing, they will undergo a series of preprocessing steps. These steps include color correction and noise removal, aiming to improve the overall quality of the tongue images and ensure the accuracy of subsequent analysis. By introducing an attention mechanism operator model in the first Unet module of the dual-Unet network, the system can be more efficient and accurate when processing temporal data. Next, the processed frames are fed into the second Unet module, which is specifically designed to process multi-frame data and learn and extract the dynamic relationships and change patterns between consecutive frames through temporal analysis techniques. In this process, the design of the attention mechanism operator allows the system to not only identify static features but also capture subtle features such as the dynamic changes of tongue images under different lighting conditions. Finally, the data fusing spatial and temporal features is subjected to final disease segmentation by a segmenter, providing accurate diagnostic results.

[0020] S101: The first branch.

[0021] The first Unet module is specifically designed to process single-frame tongue images and extract detailed spatial features in the tongue images. This step is crucial because it can accurately identify and locate key medical objects in the tongue images, such as tumors or other pathological structures. Through this efficient feature extraction technique, the system can quickly identify key medical objects in the tongue images, providing a necessary basis for subsequent multi-frame fusion analysis. This Unet module is specifically designed to process the spatial dimension of tongue images and uses its deep convolutional network architecture to extract key spatial features, such as the shape and size information of various physiological structures in medical tongue images. The output of this stage is a highly refined feature map, providing a basis for capturing detailed physiological information within each frame.

[0022] The task of the first Unet module is to process these preprocessed single-frame tongue images. It uses the multi-layer convolutional neural network structure in the deep learning network to analyze each frame of tongue image and identify key medical tongue objects, such as tumors or other lesions, from it. This process includes using complex tongue image recognition algorithms to accurately locate these objects and quickly generate information about their positions and basic attributes. The output of this stage is very crucial because it provides the necessary input data for the next step of tongue image segmentation. Then, these preliminarily identified tongue image regions are passed to the second Unet module of the system.

[0023] S102: Second branch.

[0024] The second Unet module processes multi-frame tongue image data and uses its multi-frame fusion technology and attention mechanism operator structure to analyze the dynamic changes in the tongue image sequence. The design of this module enables the system to effectively capture and analyze the changes in the time series, such as the growth or progression of lesions in tissues, which is crucial for early diagnosis and treatment planning. The attention mechanism operator not only enhances the system's ability to capture temporal information but also improves the accuracy of segmentation, enabling the system to output more reliable diagnostic results.

[0025] The second Unet module further processes these identified tongue image regions. It focuses on tongue image segmentation, using advanced tongue image processing techniques to meticulously depict the precise edges and internal details of each identified object. This includes applying multi-level tongue image segmentation algorithms to ensure that the boundaries of each lesion or tumor are accurately segmented and mapped, thereby providing an accurate tongue image basis for in-depth medical analysis and diagnosis. This segmentation result is very important because it provides detailed tongue image features and morphological information, which are the basis for the subsequent data analysis module to judge the pathological state. These processed single-frame tongue image features are fed into the second Unet module, which is similar in architecture to the first one but is specifically optimized in function to process multi-frame data in the sequence. It learns the dynamic changes from one frame to the next through advanced temporal analysis techniques, such as recurrent neural networks or attention mechanism operator networks, and effectively extracts the changes in tongue image features across multiple frames. This fusion method allows the system to not only identify static tongue image features but also capture dynamic changes, such as the subtle changes in tongue images under different conditions. Features in the time dimension are extracted and synthesized from consecutive frames, and then a comprehensive feature set containing temporal dynamic information is generated, providing more comprehensive data support for the final tongue image segmentation. These comprehensive features that integrate spatial and temporal dimensions are fed into the segmenter, and disease segmentation is performed according to the pre-trained model to output accurate diagnostic results. Local Spatiotemporal Convolution and Global Temporal Attention Mechanism are introduced into the encoder and decoder structures of Unet respectively. Local Spatiotemporal Convolution is used to capture local spatiotemporal features at a low level, while Global Temporal Attention Mechanism is used to capture global temporal correlations at a high level. To enable Unet to process multiple time frames simultaneously, a Spatiotemporal Convolution Module (STCM) is introduced to perform spatiotemporal convolution operations in the encoder part of Unet. This module not only performs convolution in the spatial dimension but also extracts features in the time dimension. Specifically, in each encoder layer of Unet, a local spatiotemporal convolution module is inserted, and 3D convolution operations are used to extract local spatiotemporal features. For the input feature map , the local spatiotemporal convolution can be expressed as:

[0026] The input image sequence is , and these frames can be stacked into a tensor with the shape (batch_size, n_frames, height, width, channels). Through the Spatiotemporal Convolution Module (STCM), 3D convolution can be performed to capture features in both time and space:

[0027] ; ; ; Among them, is the output of the l -th layer of local spatio-temporal convolution operation, is a 3D local convolution operation, is the input feature of the l -th layer encoder, is the output of the spatio-temporal convolution operation, are n + 1 tongue image frames in the tongue image data to be segmented; is the spatio-temporal convolution operation.

[0028] Global Temporal Attention Mechanism: In the decoder part of Unet, use the Transformer or ConvLSTM self-attention mechanism to capture global temporal dependencies, perform further temporal modeling on spatio-temporal features, and predict the segmentation result of the next frame , and the formula is: ; ; Among them, is the convolutional long short-term memory network, is the segmentation head. For the input feature map of the decoder, the global temporal attention mechanism can be expressed as:

[0029] ; Among them, is the global feature map after the temporal attention mechanism, SelfAttention is the convolutional long short-term neural network or Transformer the attention operation in.

[0030] S2: Collect the tongue image dataset to train the improved double U-net network to obtain the tongue image segmentation model.

[0031] Collect the tongue image dataset. The tongue image dataset includes consecutive frame tongue images and the corresponding real segmentation labels for each frame of tongue image; Stack the tongue image frames belonging to the same sample to form a multi-dimensional tensor with the shape of (batch_size, n_frames, height, width, channels).

[0032] Initialize the weight coefficients of the improved dual U-net network. Input the multi-dimensional tensor into the first branch of the improved dual U-net network, and perform 3D convolution through the spatio-temporal convolution module (STCM) to capture the features in time and space. The formula is expressed as:

[0033] ; where, is the spatio-temporal feature map. The spatio-temporal features are input into the subsequent part of the Unet for segmentation to obtain the segmentation results of each frame .

[0034] Another temporal branch further performs temporal modeling on the spatio-temporal features through the Transformer or ConvLSTM module to predict the segmentation results of the next frame , and the formula is: ; .

[0035] Spatial loss: Calculate the loss between the segmentation result of the current frame and the ground truth label : ; where, is the spatial loss function corresponding to the first branch, n is the number of frames of the tongue image data to be segmented, is the predicted segmentation label of the tongue image in the t-th frame, is the ground truth segmentation label of the tongue image in the t-th frame.

[0036] Temporal loss: Calculate the loss between the predicted segmentation result of the next frame and the ground truth label : ; where, is the temporal loss function corresponding to the second branch, is the predicted segmentation label of the tongue image in the (n + 1)-th frame predicted by the second branch, is the ground truth segmentation label of the tongue image in the (n + 1)-th frame.

[0037] The final loss is: ; where, is the weight parameter used to adjust the weights of the spatial loss and the temporal loss.

[0038] Through spatio-temporal convolution, Unet can directly process the input of multiple frames and extract spatio-temporal joint features. Through the multi-task learning framework, the model can not only handle the segmentation task of the current frame, but also predict the segmentation result of the next frame, thus improving the temporal consistency.

[0039] S3: Input the tongue image data to be segmented into the tongue image segmentation model. Through the first branch, extract the spatial features of each frame of tongue image in the tongue image data to be segmented, and obtain the initial tongue image spatial features containing the morphological and size information of the tongue physiological structure for each frame of tongue image, and input the initial tongue image spatial features into the second branch; Through the local spatio-temporal convolution module of the second branch encoder, perform local spatio-temporal convolution operations on the initial tongue image spatial features to locally encode the initial tongue image spatial features along the time dimension, and obtain the tongue image temporal local encoding features; Through the global temporal attention mechanism module of the second branch decoder, decode the tongue image temporal local encoding features to capture the global temporal correlation of the tongue image spatial features in the time dimension, and obtain the tongue image temporal global decoding features; Segment the tongue image temporal global decoding features through the segmentation layer to obtain the segmentation result of the tongue image data to be segmented.

[0040] Furthermore, to improve the operability of the system and the transparency of the diagnostic process, the present invention provides a user-friendly interface and detailed diagnostic process visualization tools. This enables medical professionals to easily operate the system and be able to deeply understand the decision-making logic of AI. Through these tools, medical professionals can intuitively see every step from tongue image acquisition to the final diagnosis, so as to better evaluate the output of the system and make clinical decisions accordingly. This system design based on dual Unet and multi-frame fusion makes full use of the latest advances in deep learning and tongue image processing technologies, providing an innovative and effective method for the segmentation and analysis of dynamic medical tongue images. Through this system, the automation and intelligence level of medical tongue image analysis have been significantly improved, greatly improving the speed and accuracy of diagnosis, thus optimizing the clinical workflow and improving the treatment effect of patients. The segmented tongue image data is then transmitted to the data analysis module. This module combines the patient's historical medical records and the tongue image data obtained by real-time processing through the Unet module. Here, deep learning algorithms are used to comprehensively consider the detailed features of the tongue image and the patient's clinical history, and conduct a comprehensive analysis of each identified and segmented object to infer their possible pathological features and conditions. This analysis process adopts the latest medical tongue image analysis technologies, including but not limited to AI-based tongue image recognition and pathological prediction models, which can extract key biomarkers and disease indicators from complex medical tongue images.

[0041] Finally, all these analyzed data and results are integrated by the system to form a detailed diagnostic report. This report not only records in detail every step and every result of tongue image analysis, but also includes specific diagnostic suggestions generated based on these data and their scientific basis. The data and suggestions in the report are based on the comprehensive judgment of deep learning algorithms and attention mechanism operator models, aiming to provide doctors with more accurate diagnostic information and treatment suggestions. Through this comprehensive and highly automated processing flow, the system not only greatly improves the efficiency and accuracy of medical diagnosis, but also improves the quality and response speed of medical services through its advanced tongue image processing and analysis techniques. In addition, the high degree of automation of this process reduces the workload of doctors, enabling them to focus more on patient care rather than heavy data processing work. This conversion process from raw medical images to detailed diagnostic reports not only demonstrates the powerful functions of modern medical technology, but also serves as a cutting-edge example of the application of deep learning and artificial intelligence technologies in the medical field.

[0042] The present invention combines a dual Unet structure, multi-frame fusion technology, and an attention mechanism operator network to solve the problem of insufficient accuracy in traditional medical tongue image segmentation systems when dealing with dynamic tongue images. The superiority of this technical solution lies in its ability to comprehensively utilize the spatial and temporal information of tongue images, significantly improving the sensitivity and precision of monitoring and segmentation of dynamic physiological processes. In addition, by combining the temporal information processing ability of the attention mechanism operator, the system can significantly improve the segmentation accuracy when processing dynamic tongue image data, making it suitable for real-time analysis and processing of complex medical scenarios.

[0043] In addition to the current embodiments, the future development of this technical solution may include further integrating additional relevant health data (such as patient symptom records) with tongue image analysis results to provide more comprehensive diagnostic support. At the same time, with the development of machine learning technologies, more advanced deep learning models can be considered in the future to further improve the segmentation performance. In addition, the system's algorithms can also be adjusted to adapt to tongue image data generated by different types of medical tongue image devices, enhancing its applicability and flexibility.

[0044] The above is the tongue image segmentation method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding tongue image segmentation device as shown in Figure 3 and includes: A model construction module, which is used to add a local spatio-temporal convolution module to the encoder of one of the U-net branches in the dual U-net network, and add a global temporal attention mechanism module to the decoder of this U-net branch to obtain an improved U-net branch; taking the unimproved U-net branch in the dual U-net network as the first branch and the improved U-net branch in the dual U-net network as the second branch, construct an improved dual U-net network including the first branch, the second branch and a segmentation layer.

[0045] A model training module, which is used to collect a tongue image dataset to train the improved dual U-net network and obtain a tongue image segmentation model.

[0046] A tongue image segmentation module, which is used to input the tongue image data to be segmented into the tongue image segmentation model, extract spatial features of each frame of tongue image in the tongue image data to be segmented through the first branch, obtain initial tongue image spatial features containing the morphological and size information of the tongue physiological structure for each frame of tongue image, and input the initial tongue image spatial features into the second branch; through the local spatio-temporal convolution module of the second branch encoder, perform local spatio-temporal convolution operations on the initial tongue image spatial features to locally encode the initial tongue image spatial features along the time dimension and obtain tongue image temporal local encoding features; through the global temporal attention mechanism module of the second branch decoder, decode the tongue image temporal local encoding features to capture the global temporal correlation of the tongue image spatial features in the time dimension and obtain tongue image temporal global decoding features; through the segmentation layer, segment the tongue image temporal global decoding features to obtain the segmentation result of the tongue image data to be segmented.

[0047] For the specific limitations of the tongue image segmentation device, reference can be made to the limitations of the tongue image segmentation method in the above text, which will not be elaborated here. Each module in the above tongue image segmentation device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of the processor, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0048] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above Figure 1 provided tongue image segmentation method.

[0049] The present invention also provides the structure of a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory. Of course, there may also be other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 provided tongue image segmentation method.

[0050] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0051] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.

Claims

1. A tongue image segmentation method, characterized in that: include: By adding a local spatiotemporal convolution module to the encoder of one of the U-net branches in the dual U-net network, and adding a global temporal attention mechanism module to the decoder of the U-net branch, an improved U-net branch is obtained; Taking the unimproved U-net branch in the dual U-net network as the first branch and the improved U-net branch in the dual U-net network as the second branch, constructing an improved dual U-net network including the first branch, the second branch and the segmentation layer; Collect tongue image data sets to train the improved dual U-net network and obtain the tongue image segmentation model; The tongue image data to be segmented is input into the tongue image segmentation model, and the spatial features of each frame of the tongue image data to be segmented are extracted through the first branch to obtain the initial tongue image spatial features of each frame of the tongue image containing the physiological structure, morphology and size information of the tongue image, and the initial tongue image spatial features are input into the second branch; the local spatiotemporal convolution module of the second branch encoder performs a local spatiotemporal convolution operation on the initial tongue image spatial features to locally encode the initial tongue image spatial features along the time dimension to obtain the tongue image temporal local encoding features; the global temporal attention mechanism module of the second branch decoder decodes the tongue image temporal local encoding features to capture the global temporal correlation of the tongue image spatial features in the time dimension to obtain the tongue image temporal global decoding features; The tongue image temporal global decoding features are segmented through the segmentation layer to obtain the segmentation results of the tongue image data to be segmented.

2. The tongue image segmentation method according to claim 1, characterized in that: The local spatiotemporal convolution module of the second branch encoder performs a local spatiotemporal convolution operation on the initial tongue image spatial features, specifically comprising: in, is the n+1 frame of tongue image data to be segmented; STCM is the spatiotemporal convolution operation, is the output of the spatiotemporal convolution operation; Conv3D It is a 3D local convolution operation; is the output of the local spatiotemporal convolution operation.

3. The tongue image segmentation method according to claim 2, characterized in that: The global temporal attention mechanism module of the second branch decoder is used to decode the tongue image temporal local encoding features, specifically including: in, SelfAttention It is a convolutional long short-term neural network or Transformer The attention operation in is the output of the attention operation, is the input of the attention operation.

4. The tongue image segmentation method according to claim 1, characterized in that: The collecting of tongue image data set to train the improved dual U-net network specifically includes: A tongue image dataset is collected, which includes continuous-frame tongue image and true segmentation labels corresponding to each frame of the tongue image; each frame of the tongue image belonging to the same sample is stacked to form a multidimensional tensor; Initialize the weight coefficient of the improved dual U-net network; Input the multidimensional tensor into the first branch of the improved dual U-net network to output the initial tongue image spatial features; input the multidimensional tensor and the initial tongue image spatial features into the second branch of the improved dual U-net network to output the tongue image temporal global decoding features; input the tongue image temporal global decoding features into the segmentation layer to output the predicted segmentation label corresponding to each frame of the tongue image; The loss function is used to calculate the loss value of the true segmentation label and the predicted segmentation label; the weight coefficient in the dual U-net network is iteratively improved according to the loss value, and the new sample is re-segmented according to the iterative weight coefficient; Repeat the iterative steps until the loss function converges or reaches the preset number of iterations to obtain the trained improved dual U-net network.

5. The tongue image segmentation method according to claim 4, characterized in that: The loss function specifically includes: in, is the total loss function; is the spatial loss function corresponding to the first branch, n is the number of tongue image data frames to be segmented, is the predicted segmentation label of the tongue image in frame t, is the true segmentation label of the tongue image in the tth frame; is the adjustment parameter; is the time loss function corresponding to the second branch, is the predicted segmentation label of the tongue image of the n+1th frame predicted by the second branch, is the true segmentation label of the tongue image in the n+1th frame.

6. A tongue image segmentation device, characterized in that: include: A model building module for obtaining an improved U-net branch by adding a local spatiotemporal convolution module to the encoder of one of the U-net branches in the dual U-net network and adding a global temporal attention mechanism module to the decoder of the U-net branch; Taking the unimproved U-net branch in the dual U-net network as the first branch and the improved U-net branch in the dual U-net network as the second branch, constructing an improved dual U-net network including the first branch, the second branch and the segmentation layer; The model training module is used to collect tongue image data sets to train the improved dual U-net network and obtain the tongue image segmentation model; The tongue image segmentation module is used to input the tongue image data to be segmented into the tongue image segmentation model, extract the spatial features of each frame of the tongue image in the tongue image data to be segmented through the first branch, obtain the initial tongue image spatial features of each frame of the tongue image containing the physiological structure, morphology and size information of the tongue image, and input the initial tongue image spatial features into the second branch; perform local spatiotemporal convolution operation on the initial tongue image spatial features through the local spatiotemporal convolution module of the second branch encoder to locally encode the initial tongue image spatial features along the time dimension to obtain the tongue image temporal local encoding features; decode the tongue image temporal local encoding features through the global temporal attention mechanism module of the second branch decoder to capture the global temporal correlation of the tongue image spatial features in the time dimension to obtain the tongue image temporal global decoding features; The tongue image temporal global decoding features are segmented through the segmentation layer to obtain the segmentation results of the tongue image data to be segmented.

7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

8. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Road extraction method and system based on high-order spatial information global automatic perception

    CN110751111A

  • Coronary artery sequence blood vessel segmentation method based on space-time discriminative feature learning

    CN112150476A

  • Double-flow U-Net image tampering detection network system and image tampering detection method thereof

    CN114998261A

  • Tongue image segmentation method based on improved Unet network, electronic equipment and storage medium

    CN118608533A

  • Heart mitral valve medical image segmentation model, training method, segmentation method and equipment

    CN119648719A

Cited By

  • Tongue picture image enhancement system

    CN120655527A