A method, system, and storage medium for segmenting tongue images.
By constructing a three-dimensional tongue model, calculating the degree of concavity and convexity, and setting tongue image contour constraints, the error problem of tongue image segmentation methods in complex backgrounds was solved, achieving high-precision tongue image segmentation and defect identification, and improving diagnostic accuracy.
Patent Information
- Application Number
- CN202510319679.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Existing tongue segmentation methods perform poorly in complex backgrounds or when the tongue image is unclear. They lack detailed analysis of the tongue's morphology and are difficult to automatically correct errors generated by segmentation images.
By acquiring tongue images from multiple angles, a three-dimensional tongue model is constructed. The tongue edge features are extracted for upper tongue image model segmentation. The degree of concavity and convexity is calculated and the tongue body contraction and expansion deformation curve is constructed. The tongue image contour constraints are set for region segmentation and defect recognition, thereby enhancing defect representation.
It improves the accuracy of tongue image segmentation and defect recognition capabilities, ensuring that the segmentation results match the actual structure of the tongue, enabling timely detection of abnormal areas and improving diagnostic efficiency.
Smart Images

Figure CN120235896B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image segmentation technology, and in particular to a method, system and storage medium for segmenting tongue images. Background Technology
[0002] Early tongue image segmentation methods primarily relied on edge detection, thresholding, and morphological operations from image processing. Edge detection algorithms such as Canny edge detection and the Sobel operator were widely used for edge extraction in tongue images, while thresholding distinguished the tongue from the background by classifying image grayscale values. However, these traditional methods required high image quality and performed poorly in complex backgrounds or when the tongue image was unclear. With the rise of deep learning, especially the application of convolutional neural networks (CNNs), tongue image segmentation methods have made significant progress. CNN-based segmentation methods can be trained on large amounts of labeled data to automatically learn tongue features, overcoming the limitations of traditional methods in processing complex images. However, current technologies often lack detailed analysis of tongue morphology, making it difficult to automatically correct errors generated during image segmentation. Summary of the Invention
[0003] Therefore, it is necessary to provide a method, system, and storage medium for segmenting tongue images to solve at least one of the aforementioned technical problems.
[0004] To achieve the above objective, a method for segmenting a tongue image is provided, the method comprising the following steps:
[0005] Step S1: Obtain a set of tongue images from various perspectives; construct a three-dimensional tongue image model based on the global perspective of the tongue image set, and extract the tongue edge features of the three-dimensional tongue image model to perform upper tongue image model segmentation to obtain the upper tongue image three-dimensional model.
[0006] Step S2: Extract the original pixel information of the three-dimensional model of the upper tongue image, and perform regional image segmentation on the three-dimensional model of the upper tongue image to generate the upper tongue image region image; calculate the concavity and convexity of the upper tongue image region image, and construct the tongue body contraction and expansion deformation curve; perform morphological region segmentation on the upper tongue image region image based on the tongue body contraction and expansion deformation curve to generate the tongue image morphological region segmentation image.
[0007] Step S3: Set tongue image contour constraints; perform tongue surface region discrimination on the tongue image morphology region segmentation image according to the tongue image contour constraints. When the discrimination result is false, perform region recalibration on the tongue image morphology region segmentation image until the discrimination result is true.
[0008] Step S4: Perform tongue defect identification on the segmented tongue image morphology region and enhance the defect representation of the identified tongue defects to obtain an enhanced tongue image segmentation image.
[0009] This invention acquires tongue images from multiple angles and constructs a 3D model based on a global perspective, accurately capturing the tongue's three-dimensional structure and edge features. This provides a solid foundation for subsequent morphological analysis and defect detection. The segmented 3D model of the upper tongue helps clarify the key areas for analysis, improving the accuracy of subsequent processing. The extracted raw pixel information provides rich data support for subsequent analysis. Region image segmentation helps extract meaningful areas from the tongue image. The calculation of concavity and convexity further refines the morphological features of the tongue. The construction of the tongue's contraction and expansion deformation curve can reveal the deformation of the tongue under different shapes, contributing to a deeper understanding of the tongue's physical state. The constraint of the tongue contour helps ensure that the segmented image conforms to... To avoid unreasonable segmentation or misjudgment, the actual tongue structure is used. By performing contour discrimination and recalibration on the segmented image, the segmentation results can be continuously optimized, ultimately ensuring high accuracy of image segmentation. Defect identification is the core of this process, which can accurately locate abnormal areas in the tongue image. By enhancing the representation of defects, the visualization effect of defective areas in the image can be improved, thereby improving the diagnostic efficiency of doctors or related technicians. The enhanced tongue image segmentation image can provide clearer and more detailed information, helping to better assess the health status of the tongue. Therefore, this invention effectively improves the segmentation accuracy and defect identification capability of tongue images through multi-view 3D modeling, refined region segmentation, contour constraint and defect enhancement, thus improving the accuracy of tongue image segmentation.
[0010] Preferably, the construction of the 3D tongue image model based on the global perspective of the tongue image set includes:
[0011] Extract the photographic angles of the tongue image set, and perform image viewpoint overlap discrimination and labeling on the tongue image set based on the photographic angles of the tongue image set to obtain the tongue overlapping viewpoint image;
[0012] By stitching together overlapping images of the tongue, a tongue image with a global perspective is generated.
[0013] Calculate the depth information of each pixel in the tongue image from a global perspective, and generate a tongue depth map by calculating parallax.
[0014] The pixel spatial coordinates of the tongue depth map are confirmed, and the three-dimensional mesh reconstruction of the tongue image from the global perspective is performed using the pixel spatial coordinates to generate an initial three-dimensional tongue image model.
[0015] The initial 3D tongue image model is rendered in 3D to obtain a 3D tongue image model.
[0016] This invention extracts the photographic angles of tongue images, ensuring that the perspective information of each image is fully utilized and avoiding duplicate or unnecessary perspective data. The perspective overlap discrimination and marking process can effectively filter out suitable images, thereby reducing the impact of inconsistent or duplicate information on subsequent model construction and improving model accuracy. By stitching multiple tongue images together to generate a global perspective image, the final image can contain more comprehensive tongue information. Through this stitching technology, a more complete tongue image view can be obtained, thus providing sufficient visual data support for subsequent 3D reconstruction. Depth information is the foundation of 3D reconstruction. By calculating the depth information of each pixel, the relative spatial position of each part in the image can be accurately determined. The generated depth map provides important spatial data for subsequent 3D modeling and mesh reconstruction, ensuring that the model can realistically reflect the 3D structure of the tongue. By determining the spatial coordinates of pixels, each image pixel can be accurately mapped into 3D space, ensuring the accuracy of the mesh reconstruction process. By using these coordinates to reconstruct a 3D mesh from a global perspective of the tongue image, a preliminary 3D tongue image model can be generated. This model includes the geometric structure of the tongue and serves as the basis for further analysis and processing. Through 3D rendering technology, the preliminary 3D model can be transformed into a more visually realistic 3D tongue image model. This step not only makes the model more visually appealing but also improves the interpretability of the image.
[0017] Preferably, the step of extracting the original pixel information of the region of the three-dimensional model of the upper tongue and performing region image segmentation on the three-dimensional model of the upper tongue includes:
[0018] Extract the original pixel information of the three-dimensional model of the upper tongue;
[0019] Two-dimensional planar projection is performed on the three-dimensional model of the upper tongue to obtain two-dimensional images of the upper tongue from various perspectives;
[0020] The original pixel information is used to perform initial region segmentation on the two-dimensional image of the upper tongue from various perspectives to generate the initial segmentation region of the upper tongue.
[0021] The mean value of the feature components is calculated for the initial segmented region of the upper tongue image from various perspectives to obtain the mean value of the feature components;
[0022] The size of the segmentation grid in the initial segmentation region of the upper tongue image is adjusted by using the mean of the feature components to generate an image of the upper tongue region.
[0023] This invention provides fundamental data support for subsequent analysis and segmentation by extracting raw pixel information. This raw pixel information contains all the details of the tongue image, ensuring that subsequent processing preserves the true structure and features of the tongue to the greatest extent possible, avoiding information loss. Two-dimensional planar projection converts the information of the three-dimensional model into easily processed two-dimensional images. This step helps simplify the three-dimensional data into two-dimensional images from multiple different perspectives, facilitating segmentation and analysis. The multi-angle projected images make the segmentation results more comprehensive, allowing for analysis of the tongue image from multiple viewpoints. Initial segmentation of the tongue image from each viewpoint using the raw pixel information can separate important regions of the tongue image (such as the tongue surface and tip) from the overall image, laying the foundation for more detailed processing later. Initial segmentation provides a coarse division of regions, facilitating optimization and fine-tuning in subsequent steps. Calculating the mean of feature components in the initial segmented regions from various perspectives helps analyze the overall characteristics of each region, such as color, texture, and shape. This calculation provides crucial data support for subsequent fine-tuning, making the region division more accurate and consistent with the actual morphological characteristics of the tongue. Adjusting the region grid size by using the mean of feature components allows for adjustments to the level of detail in each region based on its actual characteristics. This step ensures the flexibility of image segmentation, enabling smaller grid divisions for regions with rich details or prominent features, thereby improving segmentation accuracy. It helps avoid errors caused by coarse segmentation, especially in regions with significant variations in tongue morphology.
[0024] Preferably, the calculation of the concavity and convexity of the upper tongue image region and the construction of the tongue body contraction and expansion deformation curve include:
[0025] For each pixel within the upper tongue region image, obtain its depth value Z. i ;
[0026] The average depth value Z of the region is obtained by averaging the depth values of all pixels in the upper tongue region image. avg ;
[0027] The maximum depth value Z of the upper tongue region image was obtained by performing depth extremum calculation. max and minimum depth value Z min ;
[0028] The degree of concavity / convexity in the upper tongue region image is calculated, and the formula for calculating the degree of concavity / convexity is as follows:
[0029]
[0030] In the formula, C represents the concavity / convexity of the upper tongue image region, N represents the number of pixels in the selected region of the image, and Z represents the depth of the upper tongue image region. i Z is the depth value of the i-th pixel. avgZ is the average depth value of all pixels within the selected region. max Z represents the maximum depth value of all pixels within the selected region. min The minimum depth value for all pixels within the selected area;
[0031] The tongue body contraction and expansion deformation curves are constructed by using the concavity and convexity of the upper tongue image region. The tongue body contraction and expansion deformation curves include the tongue body contraction curve and the tongue body expansion curve.
[0032] This invention reflects the relative position of each pixel in three-dimensional space through its depth value. By obtaining the depth value of each pixel, the geometric morphology of the tongue can be accurately analyzed, providing precise spatial information for subsequent calculation of the degree of concavity and convexity. By calculating the average depth value of a region, the overall depth level of that region can be obtained. The average depth value provides a benchmark for subsequent calculation of the degree of concavity and convexity, making it possible to compare the depth changes of different pixels. By calculating the maximum and minimum depth values of a region, the depth range within the region can be clearly defined. These extreme values are crucial for describing the "undulation" characteristics of the tongue and help reveal the structural variability of the tongue. Through the formula for calculating the degree of concavity and convexity, the degree of concavity and convexity C is quantified as the squared average of the depth differences within the region. This index provides a numerical measure of the degree of "undulation" of the tongue and can quantitatively describe the physical state of the tongue region, such as the health status of the tongue or a marker of certain diseases. A higher degree of concavity and convexity indicates that there are greater morphological changes in the tongue, while a lower degree of concavity and convexity indicates that the tongue surface is relatively smooth.
[0033] Preferably, the morphological region segmentation of the upper tongue image based on the tongue's contraction and expansion deformation curve includes:
[0034] Based on the tongue's expansion and contraction deformation curve, the depth change of each pixel in the upper tongue image region relative to its neighborhood is calculated to obtain depth change data. The formula for calculating the depth change is as follows:
[0035]
[0036] In the formula, Z(x,y) is the depth value of the point (x,y). Z is the gradient operator, where Z is the surface deformation, x is the x-coordinate of the pixel, and y is the y-coordinate of the pixel.
[0037] By fitting a high-order polynomial to describe the local deformation of depth variation data, a nonlinear deformation modeling function is generated, where the formula for describing local deformation is shown below:
[0038] Z(x,y,t)=f(x,y)+g(t)+∈;
[0039] In the formula, (xz,y,t) represents the nonlinear deformation modeling function at point (x,y) at time t, f(x,y) represents the local deformation at point (x,y), x is the x-coordinate of the pixel, y is the y-coordinate of the pixel, and ∈ represents the error term;
[0040] The contraction and expansion process of the tongue is simulated in the upper tongue region image based on the nonlinear deformation modeling function, generating a local region deformation fitting function. The formula for simulating the contraction and expansion process of the tongue is shown below:
[0041] Z(x,y,t)=Z0(x,y)·(1+δ(t));
[0042] In the formula, Z(x,y,t) is the deformation fitting function of the local region, Z0 is the depth value in the initial state, δ(t) is the deformation coefficient at time t, x is the horizontal coordinate of the pixel, and y is the vertical coordinate of the pixel.
[0043] The deformation process of the upper tongue region image at different time points is calculated based on the local region deformation fitting function, and the upper tongue region image is segmented into morphological regions through the deformation process to generate a tongue morphological region segmentation image, which includes convex region segmentation, concave region segmentation and flat region segmentation.
[0044] This invention reveals local deformation information of the tongue image region by calculating the depth change of each pixel relative to its neighborhood. The depth change data helps describe subtle changes on the tongue surface, thereby identifying potential protrusions and depressions, providing a foundation for subsequent morphological segmentation. By fitting with a high-order polynomial, the local deformation of the tongue image region can be described more accurately. This not only captures the nonlinear characteristics of tongue surface deformation but also effectively reduces error terms, ensuring the accuracy of the deformation model and thus improving the quality of the segmentation process. Simulating the contraction and expansion process of the tongue helps to fully understand the dynamic changes of the tongue image in different states. By introducing the deformation coefficient δ(t), the changes in tongue shape can be finely adjusted, thereby accurately identifying the morphological features of the tongue image at different time points. By analyzing the deformation process of the tongue image region, different morphological regions can be accurately segmented.
[0045] Preferably, step S3 includes the following steps:
[0046] Step S31: Set tongue image contour constraints;
[0047] Step S32: Based on the tongue image contour constraints, perform tongue surface region discrimination on the tongue image morphology region segmentation image. If the discrimination result is true, the tongue image morphology region segmentation image is still output.
[0048] Step S33: When the discrimination result is false, the tongue image morphology region segmentation image is recalibrated until the discrimination result is true.
[0049] This invention provides a clear reference standard for subsequent image discrimination and adjustment by setting tongue image contour constraints. Contour constraints ensure that the image segmentation results conform to the physiological structure of the tongue image, which helps to remove noise or erroneous regions that do not conform to the actual tongue shape and improves the realism of the segmented image. By judging whether the segmented image of the tongue image morphology region conforms to the tongue image contour constraints, the correctness of the segmentation results can be verified in real time. If the judgment result is true, it means that the segmented image meets the expectations and no further adjustment is needed, saving computational resources. When the judgment result is true, step S32 directly outputs the segmented image, avoiding repeated processing and thus improving the overall processing efficiency. When the judgment result is false, the segmented image of the tongue image morphology region is adjusted by region recalibration to make it closer to the actual tongue image contour. This mechanism ensures that the algorithm can self-correct when faced with unsatisfactory initial segmentation results and finally produce more accurate segmentation results. Through repeated region recalibration, the accuracy of tongue image segmentation can be optimized and errors eliminated, especially in the edges or special morphology regions of the tongue image. Recalibration ensures that the final tongue image segmentation result conforms to the real tongue image features and avoids blurred or discontinuous edges.
[0050] Preferably, step S33 includes the following steps:
[0051] Step S331: When the discrimination result is false, the tongue image morphology region segmentation image is assigned a region identifier to obtain the region image identifier.
[0052] Step S332: Perform label consistency check on the region image identifier, and locate the abnormal region of the region image identifier according to the tongue surface region discrimination result to obtain the abnormal region of the tongue image morphology segmentation image.
[0053] Step S333: Based on the preset minimum area threshold, noise regions are removed from abnormal regions, and custom morphological description labels are applied to the tongue image morphology region segmentation image. Then, return to step S32 until the judgment result is true.
[0054] This invention assigns a unique identifier to each region, enabling clear differentiation of the boundaries and morphological features of each tongue image region. This facilitates precise localization and analysis in subsequent processing, reducing interference between different regions. The assignment of region identifiers provides a foundation for subsequent label consistency checks and abnormal region localization, ensuring the smooth operation of the entire process. Label consistency checks ensure that the labels of different regions do not overlap or conflict during image segmentation, thereby avoiding erroneous region fusion or misjudgment. Based on the tongue surface region discrimination results, abnormal regions in the image (such as regions with segmentation errors or that do not conform to physiological structures) can be quickly located. The goal of locating abnormal regions is to ensure that the final segmentation result conforms to the actual morphology of the tongue image, rather than inconsistent regions caused by algorithm or image errors. By using a preset minimum area threshold, noisy regions in the image are removed, thereby eliminating small regions that do not belong to the tongue image morphology or are caused by image noise, blurring, etc. This reduces the influence of irrelevant factors on the tongue image analysis results, improving the reliability and accuracy of the final image. By using custom labels to describe the morphological features of the tongue image, more detailed labels that conform to the actual physiological morphology can be provided for each region. This provides higher quality input data for subsequent analysis, especially in clinical applications, which can better assist doctors in making accurate judgments.
[0055] Preferably, step S4 includes the following steps:
[0056] Step S41: Perform tongue defect recognition on the segmented tongue image morphology region to obtain the tongue defect recognition result;
[0057] Step S42: Based on the tongue image defect recognition results, perform defect region image detail enhancement on the tongue image morphology region segmentation image to obtain an enhanced tongue image segmentation image.
[0058] This invention, through tongue image defect recognition, can accurately identify defects (such as cracks, color changes, or irregular shapes) in tongue images. Defect recognition can help the system detect potential tongue diseases or lesions at an early stage, improving the overall diagnostic accuracy. The precise identification of defects through algorithms effectively avoids oversights or misjudgments during manual examination. Detail enhancement makes the details of defective areas in the tongue image clearer, especially for small or low-contrast defects. Detail enhancement can effectively highlight their morphological features, ensuring that analysts can better observe these subtle changes. By enhancing the details of defective areas, the contrast and clarity of the image are improved, which helps to more accurately define the location and shape of defects and reduce the interference of blurred areas on subsequent analysis. The enhanced tongue image segmentation provides doctors with more practically valuable data.
[0059] This specification provides a tongue image segmentation system for performing the tongue image segmentation method described above. The tongue image segmentation system includes:
[0060] The upper tongue image segmentation module is used to acquire a set of tongue images from various perspectives; a three-dimensional tongue image model is constructed based on the global perspective of the tongue image set, and the tongue edge features of the three-dimensional tongue image model are extracted to perform upper tongue image model segmentation on the three-dimensional tongue image model to obtain the upper tongue image three-dimensional model.
[0061] The tongue image morphology segmentation module is used to extract the original pixel information of the upper tongue image 3D model, and perform regional image segmentation on the upper tongue image 3D model to generate upper tongue image region images; calculate the concavity and convexity of the upper tongue image region images, and construct the tongue body contraction and expansion deformation curve; and perform morphological region segmentation on the upper tongue image region images based on the tongue body contraction and expansion deformation curve to generate tongue image morphology region segmentation images.
[0062] The region discrimination module is used to set tongue image contour constraints; based on the tongue image contour constraints, it performs tongue surface region discrimination on the tongue image morphology region segmentation image. When the discrimination result is false, the tongue image morphology region segmentation image is recalibrated until the discrimination result is true.
[0063] The defect recognition module is used to identify tongue defects in the segmented image of the tongue image morphology region, and to enhance the appearance of the identified tongue defects to obtain an enhanced segmented image of the tongue image.
[0064] The present invention also provides a storage medium storing a computer program, which, when executed, implements the above-described method for segmenting tongue images.
[0065] The beneficial effects of this invention lie in its ability to construct a complete three-dimensional tongue model by acquiring tongue images from multiple perspectives and combining them with three-dimensional reconstruction technology. This model provides a precise geometric basis for subsequent analysis, ensuring higher accuracy. Extracting tongue edge features helps to accurately define the boundaries of the tongue image, providing a reliable reference for the segmentation process and avoiding unnecessary errors. Through upper tongue image model segmentation, tongue features can be accurately extracted from the complete three-dimensional image model, thereby generating an upper tongue image three-dimensional model, laying the foundation for further tongue image analysis. By segmenting the regions of the upper tongue image three-dimensional model, different regions of the tongue image (such as the tongue tip and tongue body) can be clearly distinguished, providing a more refined spatial division for subsequent analysis. By calculating the concavity and convexity of the tongue image region, the morphological characteristics of the tongue image can be effectively identified, further revealing subtle changes in the health status of the tongue. The deformation curve of the tongue body is the basis for dynamic analysis of the tongue image, reflecting the morphological changes of the tongue image in different states, which is of great significance for diagnosing tongue diseases. By segmenting the morphological regions through the tongue body contraction and expansion deformation curve, a more detailed division of the tongue image morphological regions can be obtained, providing clinical... This invention provides more accurate data support by setting contour constraints to effectively discriminate regions in tongue image morphology segmentation images, avoiding interference from irrelevant regions. Based on tongue contour discrimination, it effectively avoids interference from erroneous regions and ensures the accuracy and consistency of the final segmented image through region recalibration. Continuous recalibration ensures that the segmented tongue image conforms to the actual tongue shape, improving analysis accuracy and avoiding misjudgments. By identifying defects in the segmented tongue image, it can promptly detect abnormal or lesion areas (such as cracks and ulcers), providing important evidence for early disease diagnosis. Enhancement of defective regions highlights the details of defects, improving their visualization and accuracy. Defect enhancement makes defective regions more obvious, helping doctors or systems accurately identify tongue lesions and providing more reliable tongue image data support. Therefore, this invention effectively improves the segmentation accuracy and defect recognition capability of tongue images through multi-view 3D modeling, refined region segmentation, contour constraints, and defect enhancement, thus enhancing the precision of tongue image segmentation. Attached Figure Description
[0066] Figure 1 This is a flowchart illustrating the steps of a tongue image segmentation method.
[0067] Figure 2 for Figure 1 A detailed flowchart illustrating the implementation steps of step S3.
[0068] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0069] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0070] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0071] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0072] To achieve the above objectives, please refer to Figures 1 to 2 A method for segmenting a tongue image, the method comprising the following steps:
[0073] Step S1: Obtain a set of tongue images from various perspectives; construct a three-dimensional tongue image model based on the global perspective of the tongue image set, and extract the tongue edge features of the three-dimensional tongue image model to perform upper tongue image model segmentation to obtain the upper tongue image three-dimensional model.
[0074] Step S2: Extract the original pixel information of the three-dimensional model of the upper tongue image, and perform regional image segmentation on the three-dimensional model of the upper tongue image to generate the upper tongue image region image; calculate the concavity and convexity of the upper tongue image region image, and construct the tongue body contraction and expansion deformation curve; perform morphological region segmentation on the upper tongue image region image based on the tongue body contraction and expansion deformation curve to generate the tongue image morphological region segmentation image.
[0075] Step S3: Set tongue image contour constraints; perform tongue surface region discrimination on the tongue image morphology region segmentation image according to the tongue image contour constraints. When the discrimination result is false, perform region recalibration on the tongue image morphology region segmentation image until the discrimination result is true.
[0076] Step S4: Perform tongue defect identification on the segmented tongue image morphology region and enhance the defect representation of the identified tongue defects to obtain an enhanced tongue image segmentation image.
[0077] This invention acquires tongue images from multiple angles and constructs a 3D model based on a global perspective, accurately capturing the tongue's three-dimensional structure and edge features. This provides a solid foundation for subsequent morphological analysis and defect detection. The segmented 3D model of the upper tongue helps clarify the key areas for analysis, improving the accuracy of subsequent processing. The extracted raw pixel information provides rich data support for subsequent analysis. Region image segmentation helps extract meaningful areas from the tongue image. The calculation of concavity and convexity further refines the morphological features of the tongue. The construction of the tongue's contraction and expansion deformation curve can reveal the deformation of the tongue under different shapes, contributing to a deeper understanding of the tongue's physical state. The constraint of the tongue contour helps ensure that the segmented image conforms to... To avoid unreasonable segmentation or misjudgment, the actual tongue structure is used. By performing contour discrimination and recalibration on the segmented image, the segmentation results can be continuously optimized, ultimately ensuring high accuracy of image segmentation. Defect identification is the core of this process, which can accurately locate abnormal areas in the tongue image. By enhancing the representation of defects, the visualization effect of defective areas in the image can be improved, thereby improving the diagnostic efficiency of doctors or related technicians. The enhanced tongue image segmentation image can provide clearer and more detailed information, helping to better assess the health status of the tongue. Therefore, this invention effectively improves the segmentation accuracy and defect identification capability of tongue images through multi-view 3D modeling, refined region segmentation, contour constraint and defect enhancement, thus improving the accuracy of tongue image segmentation.
[0078] In this embodiment of the invention, reference is made to Figure 1 The diagram shown is a flowchart illustrating the steps of a tongue image segmentation method according to the present invention. In this example, the tongue image segmentation method includes the following steps:
[0079] Step S1: Obtain a set of tongue images from various perspectives; construct a three-dimensional tongue image model based on the global perspective of the tongue image set, and extract the tongue edge features of the three-dimensional tongue image model to perform upper tongue image model segmentation to obtain the upper tongue image three-dimensional model.
[0080] In this embodiment of the invention, a structured light 3D scanner (e.g., David SLS-3, Artec Eva, etc.) is used. These devices can quickly generate high-precision tongue surface image data. Each scan of the tongue is performed from one perspective, ensuring coverage of all areas of the tongue (including the top, bottom, sides, tip, and root). At least 8-12 perspectives are required (with a rotation angle of 30°-45° between each perspective). Structured light scanners typically provide an accuracy of 0.1mm to 0.5mm, ensuring sufficiently detailed tongue surface information is obtained. Using the multiple perspective data obtained from the structured light scan, multi-view fusion technology (such as a registration algorithm based on image feature matching) is used to register the scan data from different perspectives. The ICP (Iterative Closest Point) algorithm is used to register multiple scanning perspectives, ensuring that the 3D point clouds from different perspectives can be accurately aligned. A point cloud is generated by combining the data from multiple perspectives, achieving a point cloud density of 10,000 points per square centimeter. The PoissonSurface algorithm is then used to... The Reconstruction algorithm reconstructs the surface of the tongue from point cloud data, generating a complete 3D mesh model. An edge detection algorithm is applied to the 3D tongue model, using the Sobel operator to extract the edges of the tongue in 3D space. Tongue edge features include the tongue's contour, with a focus on extracting the tip, root, and lateral edges. A GraphCut algorithm is used for tongue image segmentation. This algorithm uses energy minimization to segment the entire tongue image into different regions (such as the tip, back, and base of the tongue). By selecting an appropriate segmentation threshold, the segmentation results are ensured to be accurate, with an error controlled within 2 mm. To avoid missegmentation, local optimization is performed based on the anatomical features of the tongue (such as the location of the tip and root) to ensure accurate extraction of the upper tongue image region. Filtering algorithms (such as Gaussian filtering) are used to remove surface noise and smooth the model surface, ensuring the smoothness of the upper tongue image 3D model. The final output upper tongue image 3D model is in STL format for further analysis, processing, or application.
[0081] Step S2: Extract the original pixel information of the three-dimensional model of the upper tongue image, and perform regional image segmentation on the three-dimensional model of the upper tongue image to generate the upper tongue image region image; calculate the concavity and convexity of the upper tongue image region image, and construct the tongue body contraction and expansion deformation curve; perform morphological region segmentation on the upper tongue image region image based on the tongue body contraction and expansion deformation curve to generate the tongue image morphological region segmentation image.
[0082] In this embodiment of the invention, the three-dimensional model of the upper tongue image (usually in STL, PLY, or OBJ format) is converted into a two-dimensional projected image. This process can be accomplished by projecting the three-dimensional model onto multiple viewpoints to generate two-dimensional images, ensuring coverage of different parts of the tongue image. The image resolution should be set to at least 3000x3000 pixels to ensure fine capture of image details. Three-dimensional image rendering techniques, such as OpenGL or PBR (Physically-Based Rendering) rendering models, are used to render the three-dimensional data into two-dimensional images. Specifically, projection methods can be used to generate two-dimensional images from different angles. A deep convolutional neural network (CNN) is used for tongue image region segmentation. Specifically, the U-Net model can be selected, which is a network architecture widely used in medical image segmentation, particularly suitable for processing detailed three-dimensional data. A certain number of labeled tongue image region images (including the tip, back, and root of the tongue) are prepared for training the segmentation model. During network training, an appropriate loss function, such as the Dice coefficient, should be selected to ensure segmentation accuracy. The learning rate is set to 0.001, and the batch size is... The size is 16, and the Adam optimizer is used to optimize model performance. The output of the segmentation model is a labeled image that can effectively segment the tongue image region into different parts (such as the tongue tip, tongue surface, and tongue root). Based on the geometric features of the upper tongue image region, curvature analysis is used to calculate the degree of concavity and convexity. Specifically, the local curvature of the tongue image can be calculated by solving the second derivative of the tongue image surface. Concavity and convexity can be analyzed by Gaussian curvature and mean curvature. By using 3D mesh processing tools such as MeshLab or CGAL library, the curvature of each point on the surface is calculated to obtain the distribution map of concavity and convexity. The tongue contraction and expansion deformation curve is a curve describing the deformation of each region of the tongue image during the process of tongue morphological change (from contraction to opening). Based on the curvature data of the tongue image and the region segmentation results, the degree of concavity and convexity is calculated. To calculate tongue deformation, assuming the tongue's normal shape is a contracted state, a curve is constructed by comparing the geometric deformation under different conditions (such as when the tongue is open and closed). The degree of deformation can be quantified by comparing the differences between different tongue images, and a deformation curve is established. The deformation amount can be calculated by analyzing the curvature changes of each region. Based on the calculation results of the contraction and expansion deformation curve, morphological region segmentation of the tongue image can be performed using a threshold-based segmentation method. The specific steps are as follows: Based on the results of the deformation curve, a threshold range is set (e.g., curvature greater than a certain value indicates that the region is a convex region, and curvature less than a certain value indicates that the region is a concave region). The morphological region is segmented using a region growing algorithm or threshold segmentation to obtain different tongue image morphological regions (such as convex, concave, flat, etc.).
[0083] Step S3: Set tongue image contour constraints; perform tongue surface region discrimination on the tongue image morphology region segmentation image according to the tongue image contour constraints. When the discrimination result is false, perform region recalibration on the tongue image morphology region segmentation image until the discrimination result is true.
[0084] In this embodiment of the invention, the contour constraint of the tongue image refers to the geometric features of the tongue edges in the tongue image, such as the shape of the tongue tip and tongue margins, which must conform to specific standards or rules. These standards are usually based on anatomical tongue morphological features, including: the symmetry between the tongue tip and the tongue root, the tongue contour must conform to certain curve features, and the smoothness and continuity of the tongue margins. An edge detection algorithm (such as Canny edge detection) is used to extract the tongue image contour. The purpose of edge detection is to find the obvious edges of the tongue image and establish the tongue image contour lines. When using the Canny algorithm, a low threshold of 50 and a high threshold of 150 are set to ensure effective extraction of the main edges of the tongue image. The process involves setting geometric conditions, such as: the widest part of the tongue is located in the middle of the tongue and is symmetrically distributed; the angle between the tip and root of the tongue varies within a certain range (e.g., 20° to 40°); and the length of the tongue edge does not exceed a preset maximum value (e.g., 300mm). The contour extraction results are compared with the segmented image to check whether the tongue image morphology region conforms to the set contour features. If the tongue contour lines are irregular, missing, or do not conform to the preset morphological constraints, the result is judged as "false." If the boundary of the segmented tongue image region is discontinuous or shows severe distortion, it is also judged as "false." A morphological reconstruction algorithm is used to check the integrity and edge continuity of the segmented tongue image region. The system employs geometric shape matching algorithms, such as least squares fitting, to fit the contour of the tongue and determine if it conforms to normal geometric constraints. If the result does not meet the requirements, region recalibration is performed. When the result is false, it indicates that the contour of the current segmented image is inconsistent or abnormal, and the image needs to be recalibrated to meet the set contour constraints. The original segmentation result is corrected using a region growing algorithm. Based on the set contour constraints, the algorithm starts growing from the center region of the tongue image and gradually expands to the edge of the tongue until the contour constraints are met. Geometric deformation correction is performed on the tongue image region by adjusting parameters such as the angle and length of the tongue edge. The deviation of the contour constraint is minimized using an optimization algorithm (such as gradient descent) until the constraint conditions are met. The recalibrated region is then re-labeled and the region label is updated to ensure that all constraints are met. After each adjustment of the tongue morphology region segmentation image, the contour constraint needs to be judged again until the judgment result is true. Once the contour of the tongue morphology region segmentation image meets all the set geometric constraints, the judgment result is true and the process ends. To prevent excessive iterations, a maximum number of iterations is set (e.g., 5 times). If the contour constraint is not met after exceeding this number, a warning should be output to indicate the possible error.
[0085] Step S4: Perform tongue defect identification on the segmented tongue image morphology region and enhance the defect representation of the identified tongue defects to obtain an enhanced tongue image segmentation image.
[0086] In this embodiment of the invention, Gaussian blur is used to reduce background noise in the image, ensuring the accuracy of defect detection. The image is converted to grayscale to simplify computation. A trained CNN model is used to extract the features of tongue defects. A deep learning framework (such as TensorFlow or PyTorch) is used for defect recognition. CNN can effectively learn complex features in the image, including cracks, surface irregularities, color changes, etc., and can train the model to recognize different defect types, such as cracks, dents, color anomalies, etc., and output the location and category of the defect area. Local feature descriptors (such as SIFT, SURF) or image gradient methods (such as the Sobel operator) are used to enhance edge information and detect subtle defect areas in the tongue image. Combined with a region growing algorithm, the region is expanded from the initially detected defect points to accurately extract the defect area. Local contrast enhancement is performed on the identified defect area to make the defect area more clearly visible. Commonly used contrast enhancement techniques include: Histogram equalization is applied to defect areas to improve local image contrast. Adaptive enhancement is performed on defect areas to avoid over-enhancing other areas, focusing on highlighting defect areas. Specific color mapping is used to highlight defective parts, for example, red or yellow is used to mark cracks or fissures to make them more visually prominent. Pseudo-color processing is used to convert the parts representing defects in the grayscale image to different colors through mapping algorithms to enhance the visual effect of defects. Morphological dilation is applied to defect areas to make the boundaries of defects clearer, facilitating subsequent defect diagnosis. If the defect area is small, noise can be removed by erosion to retain important defect details. Local region contrast enhancement algorithms (such as Laplacian-based algorithms) are used to improve the display effect of defects. The identified defect area and the enhanced area are synthesized into the original tongue image to obtain the final enhanced image. The defect area is recalibrated, and the identified defect area is merged with the segmented tongue image area to ensure the coherence and accuracy of the final image.
[0087] Preferably, the construction of the 3D tongue image model based on the global perspective of the tongue image set includes:
[0088] Extract the photographic angles of the tongue image set, and perform image viewpoint overlap discrimination and labeling on the tongue image set based on the photographic angles of the tongue image set to obtain the tongue overlapping viewpoint image;
[0089] By stitching together overlapping images of the tongue, a tongue image with a global perspective is generated.
[0090] Calculate the depth information of each pixel in the tongue image from a global perspective, and generate a tongue depth map by calculating parallax.
[0091] The pixel spatial coordinates of the tongue depth map are confirmed, and the three-dimensional mesh reconstruction of the tongue image from the global perspective is performed using the pixel spatial coordinates to generate an initial three-dimensional tongue image model.
[0092] The initial 3D tongue image model is rendered in 3D to obtain a 3D tongue image model.
[0093] In this embodiment of the invention, for each image, the camera's angle information is extracted using EXIF metadata, typically including the camera's orientation, pitch angle, and rotation angle. If EXIF data is missing, the image shooting angle can be manually calibrated or marked using a positioning device (such as a gyroscope). An angle difference threshold of 5 degrees is set, meaning that when the difference in shooting angles between two images is less than 5 degrees, the two images are considered to have overlapping perspectives. SIFT (Scale Invariant Feature Transform) is used to extract feature points, with nfeatures = 500, extracting 500 feature points for each image. Then, FLANN (Fast Nearest Neighbor Search) is used to... Feature matching is performed with a matching threshold of 0.7. If the feature point matching degree of two images is greater than this value, they are considered to have viewpoint overlap. The RANSAC algorithm is used to remove mismatched points, with 1000 iterations and a tolerance of 5 pixels. The Laplacian pyramid stitching method is used, with 5 layers, each layer having half the resolution of the previous layer. The stitching error is set to be less than 1 pixel. Image fusion is performed using a weighted average method, i.e., in the overlapping area, a pixel-weighted average is applied. Semi-Global image fusion is also used. The Matching (SGM) algorithm is used, with numDisparities = 64 (disparity range of 64 pixels), blockSize = 5 (each matching block is 5x5 pixels), baseline distance (B) set to 0.1 meters, and camera focal length (f) set to 35mm. The depth value of each pixel in the depth map is calculated using the formula Z = f·B / d. The depth map accuracy is set to 1mm to ensure high detail in the generated depth map. The camera intrinsic parameter matrix K is set to... The focal length is 800 pixels, the principal point is located at the center of the image, the depth value of each pixel is accurate to 1 mm, Poisson reconstruction is used, the point cloud accuracy is set to 1 mm, and the default tree depth (e.g., 10) is used for mesh construction. The Laplace smoothing algorithm is used, with a smoothing factor of 0.1 to reduce noise on the mesh surface. Delaunay triangulation is used for initial meshing with an accuracy of 0.2 mm. The Phong lighting model is used, with ambient light intensity of 0.1, diffuse light intensity of 0.7, specular light intensity of 0.2, and light source position of (10,10,10). When using texture mapping, the texture resolution is set to 2048x2048 pixels to ensure clear details, generating a 3D tongue image model.
[0094] Preferably, the step of extracting the original pixel information of the region of the three-dimensional model of the upper tongue and performing region image segmentation on the three-dimensional model of the upper tongue includes:
[0095] Extract the original pixel information of the three-dimensional model of the upper tongue;
[0096] Two-dimensional planar projection is performed on the three-dimensional model of the upper tongue to obtain two-dimensional images of the upper tongue from various perspectives;
[0097] The original pixel information is used to perform initial region segmentation on the two-dimensional image of the upper tongue from various perspectives to generate the initial segmentation region of the upper tongue.
[0098] The mean value of the feature components is calculated for the initial segmented region of the upper tongue image from various perspectives to obtain the mean value of the feature components;
[0099] The size of the segmentation grid in the initial segmentation region of the upper tongue image is adjusted by using the mean of the feature components to generate an image of the upper tongue region.
[0100] In this embodiment of the invention, the resolution of the three-dimensional model of the upper tongue image is set to 1mm, which is the smallest unit size of each three-dimensional grid. The RGB color space is selected, and the color information (R, G, B) of each pixel in the image is extracted as the original pixel information. The color value of each point is extracted from the point cloud data of the upper tongue image three-dimensional model, and converted using an intrinsic parameter matrix and depth information to generate the corresponding pixel information. An orthographic projection method is selected, meaning that when the three-dimensional model is projected onto a two-dimensional plane, perspective transformation is not considered to ensure the accurate projection position of each point. Different projection angles, such as 10 degrees, 45 degrees, and 90 degrees, are set to obtain two-dimensional images from different angles. The output 2D image is set to a resolution of 1024x1024 pixels. K-means clustering is used with 3 clusters. Pixels are initially classified using the RGB color space. The maximum number of iterations is set to 100, and the error threshold is 1e-4. The initial segmentation region size is determined by the image size. Assuming the 2D image resolution is 1024x1024, the image is divided into 8x8 blocks, each 128x128 pixels in size. Euclidean distance is used to measure the color difference between pixels. Color features (RGB mean), texture features (such as Local Binary Pattern (LBP)), and shape features are selected. Features (such as region area and contour) are used as feature components. Statistical analysis is performed on pixels within each segmented region to calculate the mean color value (average value of each color channel) within the region, and to extract the mean values of texture and shape. Each segmented region is 128x128 pixels in size. The mean feature value is calculated within each region. Based on the calculated mean feature values (color, texture, shape, etc.), a grid size adjustment standard is set. For example, the grid size can be reduced for regions with high feature similarity, while the grid size can be increased for regions with large feature differences. The initial grid size is 128x128 pixels, and the adjusted grid size range is 64x64 pixels. The grid is set to 256x256 pixels. Euclidean distance is used to measure the feature differences between different regions. A threshold of 0.1 is set. Regions below this value are merged, and the grid size is increased. For similar regions, the grid size is adjusted by merging regions to finally generate a complete region image. When merging regions, a merging threshold of 0.1 is set (merging is performed when the similarity measure is less than this value). The final grid size after merging is 256x256 pixels to ensure that the features of each segmented region have strong discriminative power. The final generated region image is smoothed using Gaussian smoothing (standard deviation set to 3) to remove noise.
[0101] Preferably, the calculation of the concavity and convexity of the upper tongue image region and the construction of the tongue body contraction and expansion deformation curve include:
[0102] For each pixel within the upper tongue region image, obtain its depth value Z. i ;
[0103] The average depth value Z of the region is obtained by averaging the depth values of all pixels in the upper tongue region image. avg ;
[0104] The maximum depth value Z of the upper tongue region image was obtained by performing depth extremum calculation. max and minimum depth value Z min ;
[0105] The degree of concavity / convexity in the upper tongue region image is calculated, and the formula for calculating the degree of concavity / convexity is as follows:
[0106]
[0107] In the formula, C represents the concavity / convexity of the upper tongue image region, N represents the number of pixels in the selected region of the image, and Z represents the depth of the upper tongue image region. i Z is the depth value of the i-th pixel. avg Z is the average depth value of all pixels within the selected region. max Z represents the maximum depth value of all pixels within the selected region. min The minimum depth value for all pixels within the selected area;
[0108] The tongue body contraction and expansion deformation curves are constructed by using the concavity and convexity of the upper tongue image region. The tongue body contraction and expansion deformation curves include the tongue body contraction curve and the tongue body expansion curve.
[0109] In this embodiment of the invention, the depth value Z i It is extracted from the depth information of each pixel in the upper tongue image region, usually obtained from a depth map. The depth map can be generated by parallax calculation (using multi-view image pairs) or by obtaining depth information based on laser scanning. Assuming the resolution of the region image is 1024x1024, each pixel contains a depth value Z. i The depth value of each pixel is stored as a matrix, the size of which is consistent with the image resolution. The average depth value of all pixels is calculated, and the maximum and minimum depth values in the region image are obtained. These values are extracted from the depth data of all pixels in the region. The deviation of each pixel from the average depth is calculated using the formula for calculating the degree of concavity and convexity, then normalized to a proportion within the depth range, and finally squared and averaged. The formula for calculating the degree of concavity and convexity is shown below: In the formula, C represents the concavity / convexity of the upper tongue image region, N represents the number of pixels in the selected region of the image, and Z represents the depth of the upper tongue image region. i Z is the depth value of the i-th pixel. avg Z is the average depth value of all pixels within the selected region. max Z represents the maximum depth value of all pixels within the selected region. minThe minimum depth value for all pixels within the selected area; by modeling the contraction state of the tongue, the curve of its concavity and convexity changes with time or posture is calculated. Typically, the contraction curve can be established using continuous posture data (e.g., image sequences from different angles) or by directly applying pressure and the deformation data of the tongue. Similarly, the change in the expansion state of the tongue is also represented by the change in concavity and convexity. By collecting image data under different expansion states and calculating the concavity and convexity under each state, an expansion curve can be generated.
[0110] Preferably, the morphological region segmentation of the upper tongue image based on the tongue's contraction and expansion deformation curve includes:
[0111] Based on the tongue's expansion and contraction deformation curve, the depth change of each pixel in the upper tongue image region relative to its neighborhood is calculated to obtain depth change data. The formula for calculating the depth change is as follows:
[0112]
[0113] In the formula, Z(x,y) is the depth value of the point (x,y). Z is the gradient operator, where Z is the surface deformation, x is the x-coordinate of the pixel, and y is the y-coordinate of the pixel.
[0114] By fitting a high-order polynomial to describe the local deformation of depth variation data, a nonlinear deformation modeling function is generated, where the formula for describing local deformation is shown below:
[0115] Z(x,y,t)=f(x,y)+g(t)+∈;
[0116] In the formula, Z(x,y,t) represents the nonlinear deformation modeling function at point (x,y) at time t, f(x,y) represents the local deformation at point (x,y), x is the x-coordinate of the pixel, y is the y-coordinate of the pixel, and ∈ represents the error term;
[0117] The contraction and expansion process of the tongue is simulated in the upper tongue region image based on the nonlinear deformation modeling function, generating a local region deformation fitting function. The formula for simulating the contraction and expansion process of the tongue is shown below:
[0118] Z(x,y,t)=Z0(x,y)·(1+δ(t));
[0119] In the formula, Z(x,y,t) is the deformation fitting function of the local region, Z0 is the depth value in the initial state, δ(t) is the deformation coefficient at time t, x is the horizontal coordinate of the pixel, and y is the vertical coordinate of the pixel.
[0120] The deformation process of the upper tongue region image at different time points is calculated based on the local region deformation fitting function, and the upper tongue region image is segmented into morphological regions through the deformation process to generate a tongue morphological region segmentation image, which includes convex region segmentation, concave region segmentation and flat region segmentation.
[0121] In this embodiment of the invention, given the depth value Z(x,y) of each pixel in the upper tongue region image, gradient calculation is required for this region image. Taking each pixel (x,y) as the center and considering its surrounding 3×3 neighborhood, the depth change calculation formula is as follows: In the formula, Z(x,y) is the depth value of the point (x,y). Here, Z represents the gradient operator, Z is the surface deformation, x is the x-coordinate of a pixel, and y is the y-coordinate of a pixel. The gradient of the image is calculated using common edge detection methods such as the Sobel or Prewitt operators. The gradient function in scipy.ndimage or OpenCV is used to calculate the depth change of each pixel relative to its neighborhood. A high-order polynomial model is fitted to describe the local deformation of the depth change. It is assumed that a quadratic or cubic polynomial is chosen for fitting to capture nonlinear deformation. The formula for describing the local deformation is: Z(x,y,t)=f(x,y) The formula is: Z(x,y,t) + g(t) + ∈; where Z(x,y,t) represents the nonlinear deformation modeling function at point (x,y) at time t, f(x,y) represents the local deformation at point (x,y), x is the x-coordinate of the pixel, y is the y-coordinate of the pixel, and ∈ represents the error term. The depth variation data is fitted using least squares or polynomial regression to obtain f(x,y) and g(t). By fitting the time variable t, the deformation coefficient g(t) is obtained. Time series analysis or curve fitting techniques (such as exponential decay function or sine wave fitting) can be used to capture the time-varying deformation. Deformation, based on the local region deformation fitting function formula, simulates the contraction and expansion process of the tongue by selecting different times t: Z(x,y,t)=Z0(x,y)·(1+δ(t)); where Z(x,y,t) is the local region deformation fitting function, Z0 is the depth value in the initial state, δ(t) is the deformation coefficient at time t, x is the horizontal coordinate of the pixel, and y is the vertical coordinate of the pixel; by calculating the depth change Z(x,y,t) at different time points t, the deformation process of the region is captured. Based on the concavity and convexity of the tongue, depth changes, etc., a segmentation threshold is set to ensure... Different region types are defined: when Z(x,y,t) is higher than a certain threshold, the region is considered a convex region; when Z(x,y,t) is lower than a certain threshold, the region is considered a concave region; and when Z(x,y,t) is within a certain range, the region is considered a flat region. Based on the deformed depth data, by setting appropriate thresholds, the image is divided into different morphological regions. The Otsu method or K-means clustering can be used for threshold selection. According to the depth value and deformation process of different regions, the image is divided into three regions: convex, concave, and flat, generating a segmented tongue image.
[0122] As an example of the present invention, reference is made to Figure 2 As shown, step S3 in this example includes:
[0123] Step S31: Set tongue image contour constraints;
[0124] Step S32: Based on the tongue image contour constraints, perform tongue surface region discrimination on the tongue image morphology region segmentation image. If the discrimination result is true, the tongue image morphology region segmentation image is still output.
[0125] Step S33: When the discrimination result is false, the tongue image morphology region segmentation image is recalibrated until the discrimination result is true.
[0126] In this embodiment of the invention, edge detection is performed on the tongue image. Commonly used algorithms include Canny edge detection or the Sobel operator to extract the tongue contour from the image. The edge image is then processed, such as closing operations (filling small holes) and dilation operations (expanding edges), to ensure that the extracted tongue contour is as complete as possible. The morphological information of the tongue contour is used to set the boundary conditions for region segmentation. The constraints can be defined by specifying a maximum contour region, a minimum contour area, etc., to ensure that only regions related to the tongue are considered. The accuracy of the constraints can be further improved by contour approximation algorithms. The set contour constraints are used to judge the morphological region segmentation image. If the currently segmented region meets the contour constraints (for example, the segmented region is completely within the tongue contour and the region area meets the requirements), the judgment result is true. Morphological detection or connected component analysis is used to determine the final result. The system determines whether the tongue image region meets the contour constraints. If the determination result is true (i.e., the region meets the tongue contour requirements), the current tongue image morphology region segmentation image is output. If the determination result is false, the system continues to step S33 to perform region recalibration. An image recalibration algorithm, such as region growing or image restoration, is used to adjust the tongue image morphology region segmentation image and correct the current region segmentation result. Starting from the known tongue image contour region, the region is gradually expanded according to the similarity of pixels (such as depth value, color value, etc.). Dilation or erosion operations are used to expand or prune the region to adapt to the contour requirements. After each recalibration, the tongue image contour constraint determination is performed again until a result that is determined to be true is obtained. The maximum number of iterations can be set during recalibration to prevent entering an infinite loop. If the final determination is true, the recalibrated tongue image morphology region segmentation image is output.
[0127] Preferably, step S33 includes the following steps:
[0128] Step S331: When the discrimination result is false, the tongue image morphology region segmentation image is assigned a region identifier to obtain the region image identifier.
[0129] Step S332: Perform label consistency check on the region image identifier, and locate the abnormal region of the region image identifier according to the tongue surface region discrimination result to obtain the abnormal region of the tongue image morphology segmentation image.
[0130] Step S333: Based on the preset minimum area threshold, noise regions are removed from abnormal regions, and custom morphological description labels are applied to the tongue image morphology region segmentation image. Then, return to step S32 until the judgment result is true.
[0131] In this embodiment of the invention, a connected component labeling algorithm is used to identify each independent region in the morphological region segmentation image. Each independent region is assigned a unique identifier for subsequent operations. The `cv2.connectedComponents` function provided by OpenCV is used for connected component analysis to obtain the identifier for each region. This identifier is assigned to each region to ensure that each segmented region in the image has a unique identifier, facilitating subsequent operations. The region image identifiers are checked to ensure the integrity and consistency of each region label. If some region identifiers are not correctly assigned or the identifier assignments are discontinuous, adjustments are made. The size of the connected components is checked to ensure that no identifier is assigned to regions that are too small or too large. Based on the discrimination results of the tongue surface region (such as contour constraints), it is checked whether the current identifier region meets the expected tongue image morphology. If some regions deviate from the preset tongue surface morphology range... These regions will be considered anomalous regions. Contour matching or morphological detection can be used to confirm whether the regions conform to the conventional shape of the tongue. Anomalous regions are marked to prepare for subsequent processing (such as noise removal and region recalibration). Based on a preset minimum area threshold, anomalous regions are removed, eliminating regions that are too small or do not conform to the shape of the tongue. Morphological operations (such as erosion and dilation) are performed on the removed regions to optimize their shape. The regions after noise removal are relabeled to generate custom morphological description labels. This can be based on region shape descriptors (such as the roundness and aspect ratio of the region). The labels of each region are redefined through morphological features to make them more consistent with the shape of the tongue. Based on the custom labels and the denoised segmented image, the process returns to step S32 to perform tongue surface region discrimination. If the discrimination result is true, the subsequent steps continue; otherwise, the iteration continues.
[0132] Preferably, step S4 includes the following steps:
[0133] Step S41: Perform tongue defect recognition on the segmented tongue image morphology region to obtain the tongue defect recognition result;
[0134] Step S42: Based on the tongue image defect recognition results, perform defect region image detail enhancement on the tongue image morphology region segmentation image to obtain an enhanced tongue image segmentation image.
[0135] In this embodiment of the invention, edge detection algorithms (such as Canny and Sobel algorithms) are used to perform preliminary processing on the segmented image of the tongue image morphology region to identify defective regions in the image. These edge features can reveal cracks or damage areas. Morphological operations (erosion, dilation, etc.) are used to further enhance the features of the defective regions and increase their prominence for clearer detection. Connected component analysis is used to analyze the regions and mark the defective regions. Based on the shape, area, aspect ratio, and other features of the regions, it is further determined whether they are defective regions. According to the defect type in the tongue image (such as cracks, depressions, damage, etc.), machine learning or deep learning algorithms (such as SVM and CNN) are used to classify the image defects and automatically label the defective regions. The detected defective regions are classified to obtain specific defect types and regions. The defect recognition results are visualized, and the outlines or regions of the defective regions are marked to obtain the tongue image defect recognition results. The results include information such as the location, shape, and type of the defective regions. Based on the defect recognition results, local image enhancement is performed on the defective regions, such as using histogram equalization or contrast enhancement, to enhance the defects in the image. To enhance the detail of the defective area, local brightness adjustment is used to improve its contrast, making it more distinct from the surrounding normal area. Edge sharpening techniques (such as Laplacian operator, Sobel operator, etc.) are applied to the defective area to make the defect edges clearer and more prominent, thereby enhancing the visual effect of the defective area. Sharpening filters are used to process the defective area to highlight edge features. High-pass filtering is used to enhance the high-frequency components of the tongue image, making the image details richer, especially in the defective area. Frequency domain filtering can be combined to enhance the defective area in the high-frequency part, ensuring that the details of the defect are significantly improved. The processed defective area is then combined with the original image to obtain the enhanced tongue image segmentation image. The enhanced image should clearly display the defective area and provide a clearer basis for subsequent defect analysis.
[0136] This specification provides a tongue image segmentation system for performing the tongue image segmentation method described above. The tongue image segmentation system includes:
[0137] The upper tongue image segmentation module is used to acquire a set of tongue images from various perspectives; a three-dimensional tongue image model is constructed based on the global perspective of the tongue image set, and the tongue edge features of the three-dimensional tongue image model are extracted to perform upper tongue image model segmentation on the three-dimensional tongue image model to obtain the upper tongue image three-dimensional model.
[0138] The tongue image morphology segmentation module is used to extract the original pixel information of the upper tongue image 3D model, and perform regional image segmentation on the upper tongue image 3D model to generate upper tongue image region images; calculate the concavity and convexity of the upper tongue image region images, and construct the tongue body contraction and expansion deformation curve; and perform morphological region segmentation on the upper tongue image region images based on the tongue body contraction and expansion deformation curve to generate tongue image morphology region segmentation images.
[0139] The region discrimination module is used to set tongue image contour constraints; based on the tongue image contour constraints, it performs tongue surface region discrimination on the tongue image morphology region segmentation image. When the discrimination result is false, the tongue image morphology region segmentation image is recalibrated until the discrimination result is true.
[0140] The defect recognition module is used to identify tongue defects in the segmented image of the tongue image morphology region, and to enhance the appearance of the identified tongue defects to obtain an enhanced segmented image of the tongue image.
[0141] The present invention also provides a storage medium storing a computer program, which, when executed, implements the above-described method for segmenting tongue images.
[0142] The beneficial effects of this invention lie in its ability to construct a complete three-dimensional tongue model by acquiring tongue images from multiple perspectives and combining them with three-dimensional reconstruction technology. This model provides a precise geometric basis for subsequent analysis, ensuring higher accuracy. Extracting tongue edge features helps to accurately define the boundaries of the tongue image, providing a reliable reference for the segmentation process and avoiding unnecessary errors. Through upper tongue image model segmentation, tongue features can be accurately extracted from the complete three-dimensional image model, thereby generating an upper tongue image three-dimensional model, laying the foundation for further tongue image analysis. By segmenting the regions of the upper tongue image three-dimensional model, different regions of the tongue image (such as the tongue tip and tongue body) can be clearly distinguished, providing a more refined spatial division for subsequent analysis. By calculating the concavity and convexity of the tongue image region, the morphological characteristics of the tongue image can be effectively identified, further revealing subtle changes in the health status of the tongue. The deformation curve of the tongue body is the basis for dynamic analysis of the tongue image, reflecting the morphological changes of the tongue image in different states, which is of great significance for diagnosing tongue diseases. By segmenting the morphological regions through the tongue body contraction and expansion deformation curve, a more detailed division of the tongue image morphological regions can be obtained, providing clinical... This invention provides more accurate data support by setting contour constraints to effectively discriminate regions in tongue image morphology segmentation images, avoiding interference from irrelevant regions. Based on tongue contour discrimination, it effectively avoids interference from erroneous regions and ensures the accuracy and consistency of the final segmented image through region recalibration. Continuous recalibration ensures that the segmented tongue image conforms to the actual tongue shape, improving analysis accuracy and avoiding misjudgments. By identifying defects in the segmented tongue image, it can promptly detect abnormal or lesion areas (such as cracks and ulcers), providing important evidence for early disease diagnosis. Enhancement of defective regions highlights the details of defects, improving their visualization and accuracy. Defect enhancement makes defective regions more obvious, helping doctors or systems accurately identify tongue lesions and providing more reliable tongue image data support. Therefore, this invention effectively improves the segmentation accuracy and defect recognition capability of tongue images through multi-view 3D modeling, refined region segmentation, contour constraints, and defect enhancement, thus enhancing the precision of tongue image segmentation.
[0143] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0144] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for segmenting a tongue image, characterized in that, Includes the following steps: Step S1: Obtain a set of tongue images from various perspectives; construct a three-dimensional tongue image model based on the global perspective of the tongue image set, and extract the tongue edge features of the three-dimensional tongue image model to perform upper tongue image model segmentation to obtain the upper tongue image three-dimensional model. Step S2: Extract the original pixel information of the three-dimensional model of the upper tongue image, and perform regional image segmentation on the three-dimensional model of the upper tongue image to generate the upper tongue image region image; calculate the concavity and convexity of the upper tongue image region image, and construct the tongue body contraction and expansion deformation curve; perform morphological region segmentation on the upper tongue image region image based on the tongue body contraction and expansion deformation curve to generate the tongue image morphological region segmentation image. Step S3: Set tongue image contour constraints; perform tongue surface region discrimination on the tongue image morphology region segmentation image according to the tongue image contour constraints. When the discrimination result is false, perform region recalibration on the tongue image morphology region segmentation image until the discrimination result is true. Step S4: Perform tongue defect identification on the segmented tongue image morphology region and enhance the defect representation of the identified tongue defects to obtain an enhanced tongue image segmentation image.
2. The tongue image segmentation method according to claim 1, characterized in that, The construction of a 3D tongue image model based on a global perspective of a tongue image set includes: Extract the photographic angles of the tongue image set, and perform image viewpoint overlap discrimination and labeling on the tongue image set based on the photographic angles of the tongue image set to obtain the tongue overlapping viewpoint image; By stitching together overlapping images of the tongue, a tongue image with a global perspective is generated. Calculate the depth information of each pixel in the tongue image from a global perspective, and generate a tongue depth map by calculating parallax. The pixel spatial coordinates of the tongue depth map are confirmed, and the three-dimensional mesh reconstruction of the tongue image from the global perspective is performed using the pixel spatial coordinates to generate an initial three-dimensional tongue image model. The initial 3D tongue image model is rendered in 3D to obtain a 3D tongue image model.
3. The method for segmenting tongue images according to claim 1, characterized in that, The process of extracting the original pixel information of the three-dimensional model of the upper tongue and performing region image segmentation on the three-dimensional model of the upper tongue includes: Extract the original pixel information of the three-dimensional model of the upper tongue; Two-dimensional planar projection is performed on the three-dimensional model of the upper tongue to obtain two-dimensional images of the upper tongue from various perspectives; The original pixel information is used to perform initial region segmentation on the two-dimensional image of the upper tongue from various perspectives to generate the initial segmentation region of the upper tongue. The mean value of the feature components is calculated for the initial segmented region of the upper tongue image from various perspectives to obtain the mean value of the feature components; The size of the segmentation grid in the initial segmentation region of the upper tongue image is adjusted by using the mean of the feature components to generate an image of the upper tongue region.
4. The tongue image segmentation method according to claim 1, characterized in that, The calculation of the concavity and convexity of the upper tongue image region and the construction of the tongue body contraction and expansion deformation curve include: For each pixel in the upper tongue region image, obtain its depth value. ; The average depth value of the region is obtained by averaging the depth values of all pixels in the upper tongue region image. ; The maximum depth value of the region was obtained by calculating the depth extremum of the upper tongue image. and minimum depth value ; The degree of concavity / convexity in the upper tongue region image is calculated, and the formula for calculating the degree of concavity / convexity is as follows: ; In the formula, The degree of concavity / convexity of the upper tongue image region. The number of pixels in the selected region of the image. For the first The depth value of each pixel. This is the average depth value of all pixels within the selected area. The maximum depth value of all pixels within the selected area. The minimum depth value for all pixels within the selected area; The tongue body contraction and expansion deformation curves are constructed by using the concavity and convexity of the upper tongue image region. The tongue body contraction and expansion deformation curves include the tongue body contraction curve and the tongue body expansion curve.
5. The method for segmenting tongue images according to claim 1, characterized in that, The morphological region segmentation of the upper tongue image based on the tongue's expansion and contraction deformation curve includes: Based on the tongue's expansion and contraction deformation curve, the depth change of each pixel in the upper tongue image region relative to its neighborhood is calculated to obtain depth change data. The formula for calculating the depth change is as follows: ; In the formula, For point The depth value, For gradient operators, For surface deformation, The x-coordinate of the pixel is The ordinate of the pixel; By fitting a high-order polynomial to describe the local deformation of depth variation data, a nonlinear deformation modeling function is generated, where the formula for describing local deformation is shown below: ; In the formula, Indicated as in Time point The nonlinear deformation modeling function at the location, Represented as at point Local deformation at the location, The x-coordinate of the pixel is The ordinate of the pixel is Represented as an error term, The deformation coefficient; The contraction and expansion process of the tongue is simulated in the upper tongue region image based on the nonlinear deformation modeling function, generating a local region deformation fitting function. The formula for simulating the contraction and expansion process of the tongue is shown below: ; In the formula, In order to be in Time point The local deformation fitting function at the location. This represents the depth value in the initial state. for The deformation coefficient at time t. The x-coordinate of the pixel is The ordinate of the pixel; The deformation process of the upper tongue region image at different time points is calculated based on the local region deformation fitting function, and the upper tongue region image is segmented into morphological regions through the deformation process to generate a tongue morphological region segmentation image, which includes convex region segmentation, concave region segmentation and flat region segmentation.
6. The method for segmenting a tongue image according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Set tongue image contour constraints; Step S32: Based on the tongue image contour constraints, perform tongue surface region discrimination on the tongue image morphology region segmentation image. If the discrimination result is true, the tongue image morphology region segmentation image is still output. Step S33: When the discrimination result is false, the tongue image morphology region segmentation image is recalibrated until the discrimination result is true.
7. The method for segmenting tongue images according to claim 6, characterized in that, Step S33 includes the following steps: Step S331: When the discrimination result is false, the tongue image morphology region segmentation image is assigned a region identifier to obtain the region image identifier. Step S332: Perform label consistency check on the region image identifier, and locate the abnormal region of the region image identifier according to the tongue surface region discrimination result to obtain the abnormal region of the tongue image morphology segmentation image. Step S333: Based on the preset minimum area threshold, noise regions are removed from abnormal regions, and custom morphological description labels are applied to the tongue image morphology region segmentation image. Then, return to step S32 until the judgment result is true.
8. The method for segmenting a tongue image according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Perform tongue defect recognition on the segmented tongue image morphology region to obtain the tongue defect recognition result; Step S42: Based on the tongue image defect recognition results, perform defect region image detail enhancement on the tongue image morphology region segmentation image to obtain an enhanced tongue image segmentation image.
9. A segmentation system for tongue image, characterized in that, For performing the tongue image segmentation method as described in claim 1, the tongue image segmentation system comprises: The upper tongue image segmentation module is used to acquire a set of tongue images from various perspectives; a three-dimensional tongue image model is constructed based on the global perspective of the tongue image set, and the tongue edge features of the three-dimensional tongue image model are extracted to perform upper tongue image model segmentation on the three-dimensional tongue image model to obtain the upper tongue image three-dimensional model. The tongue image morphology segmentation module is used to extract the original pixel information of the upper tongue image 3D model, and perform regional image segmentation on the upper tongue image 3D model to generate upper tongue image region images; calculate the concavity and convexity of the upper tongue image region images, and construct the tongue body contraction and expansion deformation curve; and perform morphological region segmentation on the upper tongue image region images based on the tongue body contraction and expansion deformation curve to generate tongue image morphology region segmentation images. The region discrimination module is used to set tongue image contour constraints; based on the tongue image contour constraints, the tongue image morphology region segmentation image is used to discriminate the tongue surface region. When the discrimination result is false, the tongue image morphology region segmentation image is recalibrated until the discrimination result is true. The defect recognition module is used to identify tongue defects in the segmented image of the tongue image morphology region, and to enhance the appearance of the identified tongue defects to obtain an enhanced segmented image of the tongue image.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed, it implements the tongue image segmentation method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Artificial intelligence-based tongue picture recognition method and system
CN113139971A
Tongue picture image segmentation method and device and medium
CN113781488A