A tongue image semantic reconstruction method fusing priori knowledge graph and image segmentation
By constructing a prior knowledge graph and a multi-task segmentation network, the problem of unifying image segmentation and TCM semantic links in intelligent tongue image analysis was solved, achieving stable segmentation and consistency judgment of tongue images, and improving the interpretability and clinical auxiliary value of tongue image analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINAN XUNWANG INTERNET TECHNOLOGY CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-23
AI Technical Summary
In existing intelligent tongue image analysis technology, the image segmentation link and the traditional Chinese medicine semantic link lack a unified connection, which leads to a disconnect between the combination of site features and the diagnostic basis. Furthermore, the model training relies on pixel supervision and lacks knowledge constraints, resulting in insufficient stability of feature extraction, low consistency of semantic judgment, and insufficient interpretability and clinical auxiliary reference value of the output results.
By constructing a prior knowledge graph, including a morphological feature layer, a microstructure layer, a spatial location layer, and a logical rule layer, and training it in conjunction with a multi-task segmentation network, a unified organization and knowledge constraint of tongue images are achieved. This enables tongue body segmentation, location segmentation, and multi-feature extraction. Furthermore, joint reasoning is performed based on logical rules to generate a tongue image semantic map and a comprehensive diagnostic report.
It improves the stability of tongue image segmentation results and the consistency of semantic judgment, enhances the interpretability and clinical reference value of the results, and realizes unified modeling and collaborative reasoning of tongue imagery and TCM interpretation rules.
Smart Images

Figure CN122265200A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent tongue diagnosis technology, and in particular to a method for semantic reconstruction of tongue images that integrates prior knowledge graphs and image segmentation. Background Technology
[0002] With the development of digital TCM and medical image analysis technology, intelligent tongue image analysis usually adopts a technical approach of standardized acquisition, tongue segmentation, color and texture extraction, and classification. Some solutions use deep learning to complete the recognition of signs such as tongue color and coating, while others combine knowledge graphs or rule bases to express TCM diagnostic knowledge to assist in the interpretation of results and the generation of reports.
[0003] In existing related technologies, image segmentation links and TCM semantic links are mostly segmented processing. On the one hand, the segmentation sites, microscopic signs and pathogenesis rules lack a unified correlation, which can easily lead to a disconnect between site feature combinations and diagnostic basis. On the other hand, model training usually relies mainly on pixel supervision and lacks knowledge constraints, resulting in insufficient stability of feature extraction and low consistency of semantic judgment under complex tongue conditions. Furthermore, there is still room for improvement in the interpretability and clinical auxiliary reference value of the output results. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a tongue image semantic reconstruction method that integrates prior knowledge graphs and image segmentation to solve the problem that it is difficult to unify the modeling and collaborative reasoning of tongue image segmentation results, site semantics and traditional Chinese medicine interpretation rules.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a tongue image semantic reconstruction method that integrates prior knowledge graphs and image segmentation, comprising, The original tongue image is acquired, and color correction, brightness normalization, and scale normalization are performed on the original tongue image to obtain a preprocessed tongue image. By organizing the tongue shape, color, coating, papillae, saliva, cracks, segmentation sites, and logical rules in the theory of tongue diagnosis in traditional Chinese medicine, a priori knowledge graph is obtained. Obtain labeled tongue images and train a multi-task segmentation network using a prior knowledge graph to obtain a knowledge-constrained multi-task segmentation network. The preprocessed tongue image is input into a knowledge-constrained multi-task segmentation network, and tongue body segmentation, site segmentation, and multi-feature extraction are performed based on the segmentation sites in the prior knowledge graph to obtain tongue body segmentation results, site segmentation results, and feature distribution results. The site segmentation results and feature distribution results are combined, and joint reasoning is performed based on the logical rules in the prior knowledge graph to obtain the semantic judgment result; The results of tongue segmentation, site segmentation, feature distribution, and semantic determination are presented in a correlated manner to obtain a tongue image semantic map and a comprehensive diagnostic report.
[0007] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in this invention, the acquisition of the original tongue image specifically includes: Standardized light source acquisition is used to acquire tongue images output by the tongue image acquisition device to obtain acquired tongue images; The tongue region is located in the acquired tongue image to obtain the localized tongue image; The tongue image is cropped to obtain the original tongue image.
[0008] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in this invention, the construction of the prior knowledge graph specifically includes: Extract the morphological features corresponding to tongue shape, tongue color, and tongue coating to obtain the morphological feature layer; Microscopic features corresponding to lingual papillae and saliva were extracted to obtain the microstructure layer; Extract the spatial positional relationships corresponding to the segmentation sites to obtain the spatial site layer; Extract the combination relationships of tongue color, coating, tongue shape, tongue papillae, saliva, and cracks to obtain a logical rule layer; By associating the morphological feature layer, microstructure layer, spatial location layer, and logical rule layer, a prior knowledge graph is obtained.
[0009] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation described in this invention, the training of the multi-task segmentation network specifically includes: Annotated tongue images containing tongue outline, location regions, tongue color grade, coating type, tongue papilla location, and saliva grade are organized into a sample set. The logical rules in the prior knowledge graph are converted into computable constraints to obtain a set of knowledge constraints. The labeled sample set is input into the multi-task segmentation network to obtain the training feature results; Based on the training feature results and the labeled sample set, a segmentation loss is constructed to obtain the pixel constraint results; The consistency of the site feature combination corresponding to the training feature result is compared with the knowledge constraint itemset to obtain the semantic constraint result. By jointly updating the pixel constraint results and semantic constraint results, a knowledge-constrained multi-task segmentation network is obtained.
[0010] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in this invention, the generation of the tongue segmentation results, site segmentation results, and feature distribution results specifically includes: The preprocessed tongue image is input into an encoder-decoder structure, and a knowledge-constrained multi-task segmentation network with spatial attention and channel attention mechanisms is introduced at the decoder end to obtain a multi-task feature map. Tongue segmentation is performed based on the multi-task feature map to obtain the tongue segmentation result; Region segmentation is performed based on segmentation sites in the multi-task feature map and prior knowledge graph to obtain the site segmentation results; Simultaneous parsing is performed on the multi-task feature map and the site segmentation results to obtain the feature distribution results.
[0011] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation described in this invention, the step of simultaneously parsing the multi-task feature map and the site segmentation results specifically includes: Morphological and color analyses were performed on the corresponding sites of the lingual papillae in the site segmentation results to obtain the distribution results of the lingual papillae. Texture analysis was performed on the corresponding sites of tongue coating in the site segmentation results, and color analysis was performed on the corresponding sites of tongue color in the site segmentation results to obtain the distribution results of tongue coating and tongue color. Brightness analysis and region connectivity analysis were performed on the tongue region corresponding to the tongue segmentation results to obtain the body fluid distribution results. Contour analysis is performed on the tongue contour corresponding to the tongue segmentation result, and texture analysis and region connectivity analysis are performed on the crack features in the multi-task feature map to obtain the tongue state result and crack result; The results of tongue papillae distribution, tongue coating distribution, tongue color distribution, saliva distribution, tongue shape, and cracks are combined to obtain the characteristic distribution results.
[0012] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation described in this invention, the step of obtaining the semantic determination result specifically includes: The segmentation sites corresponding to the site segmentation results are combined with the tongue color distribution results, tongue coating distribution results, tongue papilla distribution results, saliva distribution results, tongue shape results, and crack results in the feature distribution results to obtain a structured feature set; Candidate rules are extracted from logical rules based on the segmentation points to obtain a candidate rule set; The structured feature set is matched with the candidate rule set to obtain structured diagnostic suggestions; By associating structured diagnostic prompts with corresponding segmentation sites and features, semantic determination results are obtained.
[0013] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation described in this invention, wherein obtaining structured diagnostic prompts specifically includes, The candidate rule set is filtered to obtain the matching rule set; Consistency checks are performed on the structured feature set and the matching rule set to obtain structured diagnostic prompts.
[0014] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation described in this invention, the method for obtaining a tongue image semantic map and a comprehensive diagnostic report specifically includes: The tongue segmentation results are overlaid with the site segmentation results to obtain the site annotation map; By hierarchically associating the site annotation map with the feature distribution results, a tongue image semantic map is obtained. The tongue image semantic map and semantic judgment results are combined to obtain a comprehensive diagnostic report.
[0015] As a preferred embodiment of the tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation described in this invention, the merging of the tongue image semantic map and semantic determination results specifically includes: The results of tongue color distribution, tongue coating distribution, tongue papillae distribution, and saliva distribution are correlated with the site annotation map to obtain the site feature correspondence results; The site feature correspondence results and semantic determination results are combined into a graphic and textual representation to obtain the report content set; The report content set is combined with the original tongue image to obtain a comprehensive diagnostic report.
[0016] The beneficial effects of this invention are as follows: By performing color correction, brightness normalization, and scale normalization on the original tongue image, the impact of differences in acquisition conditions on subsequent analysis is reduced; by constructing a prior knowledge graph containing morphological feature layers, microstructure layers, spatial location layers, and logical rule layers, a unified organization of tongue imagery, segmentation sites, and TCM interpretation rules is achieved; by using a knowledge-constrained multi-task segmentation network in conjunction with pixel constraints and semantic constraints, the stability of tongue segmentation, site segmentation, and multi-feature extraction is improved; furthermore, by combining structured feature combinations with logical rule matching to complete semantic judgment, and presenting it in association with a tongue image semantic map and a comprehensive diagnostic report, the consistency of judgment results, the integrity of the basis links, and the interpretability of the output results are improved. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart for a tongue image semantic reconstruction method that integrates prior knowledge graphs and image segmentation.
[0019] Figure 2 This is a flowchart of tongue body segmentation and feature analysis.
[0020] Figure 3 This is a flowchart of feature combination and semantic judgment reasoning.
[0021] Figure 4 A flowchart for generating a comprehensive diagnostic report. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a tongue image semantic reconstruction method that integrates prior knowledge graph and image segmentation, including the following steps: S1. Obtain the original tongue image and perform color correction, brightness normalization, and scale normalization on the original tongue image to obtain a preprocessed tongue image.
[0026] S1.1. Standardize the light source acquisition of the tongue image output by the tongue image acquisition device. Specifically, acquire the tongue image output by the tongue image acquisition device under constant lighting conditions, and check the clarity, exposure status and integrity of the tongue image output by the tongue image acquisition device. Retain the tongue image with complete tongue outline, normal exposure and no obvious obstruction to obtain the acquired tongue image.
[0027] The tongue region is located in the acquired tongue image. Specifically, based on the color and brightness differences between the tongue region and the background region in the acquired tongue image, foreground separation is performed to obtain a foreground separated image. Connected component analysis is performed on the foreground separated image to remove small discrete regions and non-tongue interference regions, retaining the largest closed region as the tongue candidate region. The outer boundary of the tongue candidate region is extracted, and boundary continuity trimming is performed on the outer boundary to determine the range of the tongue region, thus obtaining the located tongue image. Tongue region localization can reduce the interference of the background region, lip region, and teeth region on subsequent image cropping.
[0028] The tongue image is cropped by determining the cropping boundary based on the range of the tongue region in the tongue image, and cropping the tongue image along the cropping boundary to retain the complete tongue region and remove the background region and non-tongue interference region outside the tongue region to obtain the original tongue image.
[0029] S1.2. Color correction is performed on the original tongue image. Specifically, the original tongue image is converted to a color space, white balance correction is performed on the original tongue image, and color compensation is performed on the comprehensive color deviation in the original tongue image to bring the color distribution of the original tongue image back to a unified color reference, thus obtaining a color-corrected image. Color correction can reduce the impact of color deviation on the recognition of tongue color and tongue coating color.
[0030] Brightness normalization is performed on the color-corrected image. Specifically, brightness equalization is applied to the brightness channel of the color-corrected image, adjusting the brightness of both high-brightness and low-brightness areas to achieve a more uniform brightness distribution across the tongue region, resulting in a brightness-normalized image. Brightness normalization reduces the impact of localized reflections, shadows, and uneven brightness on the display of tongue texture and tongue boundaries.
[0031] The brightness-normalized image is scaled by scaling it proportionally according to the uniform image size requirements. While keeping the length-to-width ratio of the tongue unchanged, the tongue region is adjusted to a uniform size, and the boundary region formed after scaling is padded to obtain the preprocessed tongue image.
[0032] S2. Organize the tongue shape, color, coating, papillae, saliva, cracks, segmentation sites, and logical rules in the theory of tongue diagnosis in traditional Chinese medicine to obtain a priori knowledge graph.
[0033] S2.1. Classify and organize the tongue shape, tongue color, and tongue coating in the theory of tongue diagnosis in Traditional Chinese Medicine (TCM). Specifically, collect the content on tongue shape, tongue color, and tongue coating identification from TCM tongue diagnosis literature, tongue diagnosis textbooks, and clinical tongue diagnosis records, remove duplicate expressions with different names but the same meaning, and obtain a set of tongue shape entries, a set of tongue color entries, and a set of tongue coating entries; merge the set of tongue shape entries, the set of tongue color entries, and the set of tongue coating entries according to the overall shape, overall color, and surface coating that can be directly observed by the naked eye, where the overall shape corresponds to the tongue shape, the overall color corresponds to the tongue color, and the surface coating corresponds to the tongue coating, to obtain a set of macroscopic observation items; record each item in the set of macroscopic observation items according to the characteristic name, identification meaning, and category, to obtain the morphological feature layer.
[0034] It should be noted that macroscopic observation items refer to items that can be directly observed and judged at the overall level of the tongue image. In this step, these specifically include tongue shape, tongue color, and tongue coating. Among them, tongue shape is used to characterize the overall shape of the tongue, tongue color is used to characterize the overall color of the tongue, and tongue coating is used to characterize the surface coating of the tongue.
[0035] It should be noted that by uniformly organizing the content on tongue shape, tongue color, and tongue coating in TCM tongue diagnosis literature, tongue diagnosis textbooks, and clinical tongue diagnosis records, and merging them according to overall shape, overall color, and surface coating, a morphological feature layer with clear boundaries and unified names can be formed. This reduces semantic confusion caused by homonyms and synonyms in the subsequent knowledge organization process, and improves the semantic consistency and callability of the prior knowledge graph.
[0036] The feature items in the morphological feature layer are further expanded to microscopic features. Specifically, using the tongue shape, tongue color and tongue coating items in the morphological feature layer as the basis for upper-level observation, we collect tongue papilla items and body fluid items related to microscopic changes of the tongue surface from TCM tongue diagnosis literature, tongue diagnosis textbooks and clinical tongue diagnosis records, remove duplicate expressions with different names but the same meaning, and obtain tongue papilla item set and body fluid item set. Among them, the tongue papillae item is used to characterize changes in the microstructure of the tongue surface, such as red dots, swelling, and atrophy, while the saliva item is used to characterize changes in the moistness of the tongue surface, such as moist, dry, slippery, slightly dry, and mixed with saliva. The tongue papillae item set and the saliva item set are classified and organized according to the observation objects, where the observation objects are changes in the microstructure of the tongue surface and changes in the moistness of the tongue surface. Changes in the microstructure of the tongue surface correspond to the tongue papillae item, and changes in the moistness of the tongue surface correspond to the saliva item, resulting in a microscopic observation item set. Each item in the microscopic observation item set is recorded according to its feature name, feature meaning, and category, and a hierarchical correspondence is established with the corresponding item in the morphological feature layer to obtain the microstructure layer.
[0037] It should be noted that by further organizing the tongue papillae and saliva on the basis of the morphological feature layer and forming a microstructure layer, the macroscopic observation results and microscopic observation results can be kept in the same semantic link. This makes it easier for subsequent image segmentation results to correspond to both macroscopic and microscopic features, and improves the ability of the prior knowledge graph to constrain fine-grained image features.
[0038] S2.2. Spatial localization and organization of feature items in the morphological feature layer and microstructure layer are performed. Specifically, the tongue shape, tongue color, and tongue coating items in the morphological feature layer, and the tongue papilla and saliva items in the microstructure layer are used as feature items to be localized. The corresponding relationship of segmentation sites in the TCM tongue diagnosis theory is extracted. Among them, the segmentation site refers to the observation position on the tongue surface with fixed spatial significance. The corresponding relationship of segmentation site refers to the correspondence between the tongue surface positions such as the front, middle, root, and sides of the tongue and the TCM internal organs. The feature items to be localized are located one by one according to the tongue surface position, position boundary, and corresponding internal organs to obtain the spatial site layer.
[0039] The site items and aforementioned feature items in the spatial site layer are organized into rules, specifically as follows: using the segmentation sites in the spatial site layer as the site basis, tongue shape, tongue color, and tongue coating items in the morphological feature layer, and tongue papilla, saliva, and fissure items in the microstructure layer are retrieved to obtain a site feature correspondence set; each correspondence in the site feature correspondence set is checked item by item, and feature items that appear simultaneously at the same segmentation site or adjacent segmentation sites are counted to obtain a co-occurrence relationship set; the segmentation sites, tongue shape, tongue color, tongue coating, tongue papilla, saliva, and fissure items in the co-occurrence relationship set are combined and recorded according to site conditions, feature conditions, and discrimination conclusions to obtain a combination relationship set; the accompanying features and non-primary features in the combination relationship set are screened, and the correspondences used to limit the primary discrimination conclusion or to exclude other pathogenesis conclusions are recorded as an exclusion relationship set; the co-occurrence relationship set, combination relationship set, and exclusion relationship set are recorded as items according to rule number, rule name, site conditions, feature conditions, and discrimination conclusions to obtain a logical rule layer.
[0040] Among them, the aforementioned characteristic items refer to the tongue shape, tongue color, and tongue coating in the morphological characteristic layer, as well as the tongue papillae, saliva, and cracks in the microstructure layer.
[0041] It should be noted that by organizing the co-occurrence, combination, and exclusion relationships among various features into a logical rule layer based on the spatial location layer, the prior knowledge graph can not only have spatial positioning capabilities but also feature combination discrimination capabilities, which facilitates subsequent joint reasoning based on the location segmentation results and feature distribution results.
[0042] S2.3. Perform inter-layer associations on the morphological feature layer, microstructure layer, spatial location layer, and logical rule layer. Specifically, the tongue shape, tongue color, and tongue coating items in the morphological feature layer are treated as first-class graph entities; the tongue papillae, saliva, and crack items in the microstructure layer are treated as second-class graph entities; the segmentation site items in the spatial location layer are treated as third-class graph entities; and the combination relationships in the logical rule layer are treated as graph relations. Connect the first-class graph entities, second-class graph entities, third-class graph entities, and graph relations item by item, and check for duplicate items, conflicting items, and missing connection items. Retain connection results with consistent names, consistent positions, and consistent relationships to obtain the prior knowledge graph.
[0043] It should be noted that by unifying and associating the morphological feature layer, microstructure layer, spatial location layer, and logical rule layer to obtain a prior knowledge graph, tongue image features, spatial locations, and discrimination rules can be organized into a unified knowledge structure that can be directly invoked. This facilitates the direct use of knowledge in subsequent knowledge-constrained multi-task segmentation network training, location segmentation, and semantic judgment processes, thereby improving the completeness and interpretability of the tongue image semantic reconstruction process.
[0044] S3.1. Organize the labeled tongue images into samples, specifically: collect labeled tongue images and perform integrity checks, consistency checks, and name unification on the tongue outline, location region, tongue color grade, coating type, tongue papilla position, and saliva grade corresponding to the labeled tongue images; remove missing annotations, annotations with unclosed boundaries, and conflicting annotations; organize the retained labeled tongue images according to image files, annotation masks, and label entries to obtain a labeled sample set.
[0045] Among them, the labeled tongue image refers to a tongue image that includes information such as the outline of the tongue body, location area, tongue color grade, coating type, location of tongue papillae, and saliva grade.
[0046] It should be noted that by uniformly checking and organizing the labeled tongue images to obtain the labeled sample set, the impact of missing and conflicting labels on the stability of network training can be reduced, and the learning consistency of the multi-task segmentation network on tongue contour, location region and multi-feature labels can be improved.
[0047] The logical rules in the prior knowledge graph are converted into computable constraints. Specifically, the rule entries in the logical rule layer are read, and the segmentation sites, feature conditions, and discrimination conclusions in the rule entries are split. The segmentation sites are converted into site constraint entries, the feature conditions are converted into label constraint entries corresponding to the tongue color grade, tongue coating type, tongue papilla position, and saliva grade, and the discrimination conclusions are converted into consistency judgment entries. The site constraint entries, label constraint entries, and consistency judgment entries are then organized one by one to obtain the knowledge constraint item set.
[0048] S3.2. Input the labeled sample set into the multi-task segmentation network. Specifically, the labeled tongue images in the labeled sample set are input into a multi-task segmentation network with an encoder-decoder structure and spatial attention and channel attention mechanisms set at the decoder end. The encoder end extracts multi-scale image features, the decoder end restores the pixel spatial distribution, and simultaneously outputs training feature results corresponding to the tongue contour, location region, tongue color grade, tongue coating type, tongue papilla position, and saliva grade. When inputting the labeled sample set into the multi-task segmentation network, the labeled sample set is input in batches according to the batch size, and the learning rate, number of iterations, and parameter update method are set to obtain the parameter set for the training process.
[0049] It should be noted that by using an encoder-decoder structure for the input of the labeled sample set and setting spatial attention and channel attention mechanisms on the decoder end of the multi-task segmentation network, tongue segmentation, site segmentation and multi-feature extraction tasks can be learned simultaneously during the same training process, thereby improving the response capability of key segmentation sites and multi-feature regions.
[0050] The segmentation loss is constructed based on the training feature results and the labeled sample set. Specifically, the training feature results are aligned item by item with the corresponding labels in the labeled sample set. Dice loss and cross-entropy loss are calculated for the tongue contour correspondence results and the site region correspondence results. Pixel-level classification error is calculated for the tongue color grade correspondence results, the tongue coating type correspondence results, the tongue papilla position correspondence results, and the saliva grade correspondence results. The Dice loss, cross-entropy loss, and pixel-level classification error are assigned loss weight coefficients based on the degree of influence of the tongue contour correspondence results, the site region correspondence results, the tongue color grade correspondence results, the tongue coating type correspondence results, the tongue papilla position correspondence results, and the saliva grade correspondence results on the training target, the magnitude of each loss, and the category distribution in the labeled sample set. The losses are then weighted and summed according to the loss weight coefficients to obtain the segmentation loss and pixel constraint results.
[0051] It should be noted that by aligning the training feature results and the labeled sample set item by item and forming pixel constraint results, it is possible to ensure that the output of the multi-task segmentation network is consistent with the labeled information in terms of spatial location and category discrimination, thereby improving pixel-level classification accuracy.
[0052] S3.3. Perform a consistency comparison between the site feature combinations corresponding to the training feature results and the knowledge constraint itemset. Specifically, extract the tongue color grade, tongue coating type, tongue papilla position, and saliva grade corresponding to each segmentation site based on the site region correspondence results in the training feature results to form site feature combinations; match the site feature combinations with the site constraint items, label constraint items, and consistency judgment items in the knowledge constraint itemset one by one, and record inconsistencies and penalty items for site feature combinations that do not meet the knowledge constraint itemset to obtain semantic constraint results; set semantic constraint weight coefficients for inconsistencies and penalty items based on the importance of the segmentation site, the influence of feature conditions on the discrimination conclusion, and the degree of damage to the consistency of the knowledge graph.
[0053] Inconsistencies located at key segmentation sites and involving core feature conditions correspond to higher semantic constraint weight coefficients, while inconsistencies located at non-key segmentation sites or involving only accompanying features correspond to lower semantic constraint weight coefficients.
[0054] The penalty term corresponding to each inconsistency is multiplied by the corresponding semantic constraint weight coefficient to obtain the weighted penalty value of each inconsistency term; the weighted penalty values of each inconsistency term are accumulated or averaged to obtain the weight value of the semantic constraint result; and the semantic constraint result is formed based on the weight value of the semantic constraint result.
[0055] It should be noted that by comparing the consistency of the site feature combination with the knowledge constraint itemset to obtain the semantic constraint results, the deviation of the multi-task segmentation network output from the prior knowledge graph can be limited, thereby enhancing the semantic rationality and interpretability of the multi-task segmentation network output.
[0056] The pixel constraint results and semantic constraint results are jointly updated. Specifically, the pixel constraint results and semantic constraint results are jointly weighted to obtain the joint constraint results. Based on the parameter set of the training process, the loss weight coefficients and the semantic constraint weight coefficients, backpropagation and parameter iterative update are performed on the multi-task segmentation network. The generation of training feature results, segmentation loss calculation and consistency comparison are repeatedly performed until the pixel constraint results and semantic constraint results simultaneously stabilize, thus obtaining the knowledge-constrained multi-task segmentation network.
[0057] It should be noted that by jointly updating the pixel constraint results and semantic constraint results to obtain the knowledge-constrained multi-task segmentation network, the multi-task segmentation network can simultaneously satisfy pixel-level segmentation accuracy and prior knowledge graph consistency, thereby improving the overall stability and semantic reliability of tongue segmentation, site segmentation and multi-feature extraction.
[0058] S4. Input the preprocessed tongue image into the knowledge-constrained multi-task segmentation network, and perform tongue body segmentation, site segmentation and multi-feature extraction based on the segmentation sites in the prior knowledge graph to obtain the tongue body segmentation result, site segmentation result and feature distribution result.
[0059] S4.1. The preprocessed tongue image is input into a knowledge-constrained multi-task segmentation network with an encoder-decoder structure and spatial and channel attention mechanisms set at the decoder end. Specifically, the preprocessed tongue image is input into the encoder of the knowledge-constrained multi-task segmentation network, and layer-by-layer convolution and downsampling are performed on the preprocessed tongue image to extract multi-scale image features corresponding to the tongue body boundary, tongue surface texture, color distribution, and local structure. The multi-scale image features are input into the decoder of the knowledge-constrained multi-task segmentation network, and layer-by-layer upsampling and feature fusion are performed on the multi-scale image features. The spatial response of the corresponding areas of the tongue tip, tongue middle, tongue root, and sides of the tongue is enhanced through the spatial attention mechanism, and the response of the corresponding feature channels of tongue color, tongue coating, tongue papillae, saliva, and cracks is enhanced through the channel attention mechanism. The enhanced features are then mapped at the pixel level to obtain a multi-task feature map.
[0060] Tongue segmentation is performed based on the multi-task feature map. Specifically, the pixel responses of the corresponding tongue region and background region in the multi-task feature map are read, and pixel-level classification of the tongue region and background region is performed on each pixel in the multi-task feature map to form an initial tongue region. Hole filling and isolated small region removal are performed on the initial tongue region to remove holes inside the tongue region and scattered points outside the tongue region. Boundary smoothing is performed on the boundary of the region after hole filling and isolated small region removal to obtain the tongue segmentation result.
[0061] S4.2. Perform region segmentation based on the segmentation sites in the multi-task feature map and prior knowledge graph. Specifically, the effective area of the tongue surface is defined by the tongue body segmentation result, and the correspondence of segmentation sites in the prior knowledge graph is read. The direction of the tongue tip is determined according to the front position of the tongue body segmentation result, the direction of the tongue root is determined according to the rear position of the tongue body segmentation result, and the directions of the sides of the tongue are determined according to the left and right boundaries of the tongue body segmentation result. The effective area of the tongue surface is spatially divided according to the length direction of the tongue body and the left and right boundary positions, dividing the effective area of the tongue surface into the anterior tongue region, the middle tongue region, the tongue root region, and the sides of the tongue region. The anterior tongue region is further divided into the anterior 1 / 3 sides of the tongue region and the tip edge region, and each region is mapped one-to-one with the segmentation sites in the prior knowledge graph to obtain the site segmentation result.
[0062] It should be noted that by combining the segmentation site correspondence in the prior knowledge graph with the tongue segmentation results to perform regional site segmentation, subsequent multi-feature extraction can directly correspond to the anterior 1 / 3 of the tongue, the tip and side of the tongue, the middle of the tongue, the root of the tongue, and the sides of the tongue, thereby enhancing the consistency between image region division and the semantics of TCM sites.
[0063] S4.3. Perform simultaneous parsing on the multi-task feature map and the site segmentation results. Specifically: using the anterior 1 / 3 of the tongue in the site segmentation results as the corresponding sites for the tongue papillae, perform morphological and color analysis on the anterior 1 / 3 of the tongue to identify the number, density, and color distribution of the areas corresponding to the red dots, thus obtaining the tongue papillae distribution results; using the anterior 1 / 3 of the tongue in the site segmentation results as the corresponding sites for the tongue coating, perform texture analysis on the anterior 1 / 3 of the tongue to extract the complexity and density of the tongue coating texture; perform color analysis on the tongue color corresponding sites in the site segmentation results to extract the hue distribution and color grade, thus obtaining the tongue coating distribution results. The results of tongue color distribution are analyzed. Using the tongue region corresponding to the tongue segmentation results as the saliva analysis region, brightness analysis and region connectivity analysis are performed on the tongue region corresponding to the tongue segmentation results to identify reflective areas and saliva line distribution, thus obtaining the saliva distribution results. Using the tongue segmentation results as the basis for tongue state analysis, contour analysis is performed on the tongue segmentation results to extract tongue length, width, edge morphology, and overall contour changes, thus obtaining the tongue state results. Using the crack response region in the multi-task feature map as the basis for crack analysis, texture analysis and region connectivity analysis are performed on the crack features in the multi-task feature map to extract crack direction, length, and connectivity range, thus obtaining the crack results.
[0064] It should be noted that the results of tongue papilla distribution are used to characterize changes in the microstructure of the tongue surface, such as red dots; the results of tongue coating distribution are used to characterize the thickness, density, and texture of the tongue coating; the results of tongue color distribution are used to characterize the color distribution of the tongue body; the results of saliva distribution are used to characterize the moistness of the tongue surface; the results of tongue shape are used to characterize the overall morphology of the tongue body; and the results of cracks are used to characterize the distribution of cracks on the tongue surface.
[0065] By performing simultaneous parsing on multi-task feature maps and site segmentation results, parallel extraction of tongue papillae, tongue coating, tongue color, saliva, tongue shape, and cracks can be completed within the same process. This allows macroscopic and microscopic features to be uniformly parsed under the same spatial site system, improving the hierarchy and semantic correspondence of feature extraction.
[0066] The results of tongue papillae distribution, tongue coating distribution, tongue color distribution, saliva distribution, tongue shape, and cracks were compiled. Specifically, the results of tongue papillae distribution, tongue coating distribution, tongue color distribution, saliva distribution, tongue shape, and cracks were matched item by item according to the correspondence of segmentation sites. The feature results corresponding to the anterior 1 / 3 lateral regions of the tongue, the tip and edge regions of the tongue, the middle region of the tongue, the root region of the tongue, and the lateral regions of the tongue were uniformly recorded into the same feature entry. After the feature entries were completed, the format was standardized and the results were summarized to obtain the feature distribution results.
[0067] S5. Combine the site segmentation results and feature distribution results, and perform joint reasoning based on the logical rules in the prior knowledge graph to obtain the semantic judgment result.
[0068] S5.1. Combine the segmentation sites corresponding to the site segmentation results with the tongue color distribution results, tongue coating distribution results, tongue papilla distribution results, saliva distribution results, tongue shape results, and crack results in the feature distribution results. Specifically, read the tongue anterior 1 / 3 bilateral regions, tongue tip and side regions, tongue middle region, tongue root region, and tongue bilateral regions in the site segmentation results, and retrieve the tongue color distribution results, tongue coating distribution results, tongue papilla distribution results, saliva distribution results, tongue shape results, and crack results corresponding to each segmentation site. Record each segmentation site and its corresponding feature results according to the site name, feature name, feature value, and site attribution relationship. After completing the corresponding records, organize all the results to obtain a structured feature set.
[0069] Candidate rules are extracted from the logical rule layer based on the segmentation sites. Specifically, the segmentation sites in the structured feature set are read, and the rule entries corresponding to the segmentation sites are retrieved in the logical rule layer. The retrieved rule entries are initially screened according to the site conditions, and the rule entries that are consistent with the segmentation sites in the structured feature set are retained. The retained rule entries are summarized to obtain the candidate rule set.
[0070] S5.2. Rule filtering is performed on the candidate rule set, specifically as follows: read the site conditions, feature conditions, and discrimination conclusions in the candidate rule set, and compare the consistency of the site conditions with the segmentation sites in the structured feature set; based on the consistency of the site conditions, compare the feature conditions in the candidate rule set with the tongue color distribution results, tongue coating distribution results, tongue papilla distribution results, saliva distribution results, tongue shape results, and crack results in the structured feature set item by item, retain the rule entries with complete feature conditions and consistent with the structured feature set, and remove the rule entries with inconsistent site conditions, missing feature conditions, or conflicting feature conditions; organize the retained rule entries to obtain the matching rule set.
[0071] S5.3. Perform consistency verification between the structured feature set and the matching rule set. Specifically, verify the correspondence between the features corresponding to each segmentation site in the structured feature set and the site conditions and feature conditions in the matching rule set one by one, retain the rule entries that simultaneously satisfy both site conditions and feature conditions, and obtain a set of valid rule entries; read the discrimination conclusion field in the set of valid rule entries, extract the discrimination conclusion corresponding to each segmentation site, and obtain a set of discrimination conclusions; where the discrimination conclusion refers to the rule output conclusion recorded in the logical rule layer for feature combinations that satisfy both site conditions and feature conditions; organize the set of discrimination conclusions with the corresponding segmentation sites, corresponding feature conditions, and corresponding exclusion items one by one to obtain structured diagnostic prompts.
[0072] The structured diagnostic prompts are associated with corresponding segmentation sites and features. Specifically, the following steps are taken: The discriminant conclusion, site conditions, feature conditions, and exclusion items in the structured diagnostic prompts are read. The discriminant conclusion is used as the judgment result. The site conditions, feature conditions, and exclusion items are matched item by item with the corresponding segmentation sites, tongue color distribution results, tongue coating distribution results, tongue papillae distribution results, saliva distribution results, tongue shape results, and crack results in the structured feature set to obtain a detailed explanation of the basis. The judgment result and the detailed explanation of the basis are then correlated and organized to obtain a TCM pathogenesis prompt. Here, the judgment result refers to the rule-based judgment result directly formed by the discriminant conclusion; the detailed explanation of the basis refers to the record of the judgment basis formed by the site conditions, feature conditions, exclusion items, and corresponding features in the structured feature set; and the TCM pathogenesis prompt refers to the pathogenesis prompt content formed by the combination of the judgment result and the detailed explanation of the basis. The judgment result, the detailed explanation of the basis, and the TCM pathogenesis prompt are then uniformly correlated and recorded to obtain a semantic judgment result.
[0073] It should be noted that the semantic judgment result refers to the judgment result formed after establishing a correspondence between the rule judgment content in the structured diagnostic prompt and the corresponding site and corresponding feature in the structured feature set.
[0074] S6. The tongue segmentation results, site segmentation results, feature distribution results, and semantic judgment results are presented together to obtain a tongue image semantic map and a comprehensive diagnostic report.
[0075] S6.1. Overlay the tongue body segmentation results with the site segmentation results. Specifically, read the outer boundary of the tongue body segmentation results and the boundaries of the anterior 1 / 3 lateral regions, tip region, middle region, root region, and lateral regions of the tongue in the site segmentation results; use the outer boundary of the tongue body as the outer boundary and the boundaries of each segmented site as the inner boundary line; perform coordinate alignment and boundary overlay on the outer boundary of the tongue body and the boundaries of each segmented site; label the site name and site range of each segmented site after boundary overlay to obtain the site labeling map.
[0076] It should be noted that by aligning the tongue body segmentation results with the site segmentation results and overlaying the boundaries, the spatial position of each segmentation site can be directly presented within the tongue body region. This facilitates the subsequent hierarchical correspondence of feature distribution results according to the site range, thereby improving the spatial clarity and site identification of the tongue image visualization results.
[0077] S6.2. Perform hierarchical association between the site annotation map and the feature distribution results. Specifically, read the boundaries of each segmentation site in the site annotation map, and retrieve the tongue color distribution result, tongue coating distribution result, tongue papilla distribution result, saliva distribution result, tongue shape result, and crack result from the feature distribution results item by item; perform hierarchical mapping of the tongue color distribution result, tongue coating distribution result, tongue papilla distribution result, and saliva distribution result according to the segmentation site range in the site annotation map; record the tongue shape result according to the overall outline position of the tongue; record the crack result according to the crack location area and the corresponding segmentation site; perform unified association on the results of each layer after completing the hierarchical mapping and hierarchical recording to obtain the tongue image semantic map.
[0078] S6.3. Merge the tongue image semantic map and semantic judgment results, specifically: read the tongue color distribution results, tongue coating distribution results, tongue papilla distribution results, and saliva distribution results in the tongue image semantic map, and match the tongue color distribution results, tongue coating distribution results, tongue papilla distribution results, and saliva distribution results with the corresponding segmentation sites in the site annotation map item by item to obtain the site feature correspondence results; read the judgment results, detailed explanations of the basis, and TCM pathogenesis hints in the semantic judgment results, and merge the site feature correspondence results with the judgment results, detailed explanations of the basis, and TCM pathogenesis hints in a graphical and textual manner to obtain the report content set.
[0079] S6.4. Combine the report content set with the original tongue image, specifically: read the original tongue image and use it as the original image page; read the report content set and arrange the site feature correspondence results, judgment results, detailed explanations of the basis, and TCM pathogenesis suggestions in the report content set in the order of image results first and text results second; combine the results of the original image page and the report content set after the unified arrangement to obtain a comprehensive diagnostic report.
[0080] In summary, this invention reduces the impact of differences in acquisition conditions on subsequent analysis by performing color correction, brightness normalization, and scale normalization on the original tongue image; it achieves a unified organization of tongue imagery, segmentation sites, and TCM interpretation rules by constructing a prior knowledge graph containing morphological feature layers, microstructure layers, spatial site layers, and logical rule layers; it improves the stability of tongue segmentation, site segmentation, and multi-feature extraction by using a knowledge-constrained multi-task segmentation network that combines pixel constraints and semantic constraints; and it further combines structured feature combinations with logical rule matching to complete semantic judgment, presenting it in conjunction with a tongue image semantic map and a comprehensive diagnostic report, thereby improving the consistency of judgment results, the integrity of the basis links, and the interpretability of the output results.
[0081] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for semantic reconstruction of tongue image by integrating prior knowledge graph and image segmentation, characterized in that: include, The original tongue image is acquired, and color correction, brightness normalization, and scale normalization are performed on the original tongue image to obtain a preprocessed tongue image. By organizing the tongue shape, color, coating, papillae, saliva, cracks, segmentation sites, and logical rules in the theory of tongue diagnosis in traditional Chinese medicine, a priori knowledge graph is obtained. Obtain labeled tongue images and train a multi-task segmentation network using a prior knowledge graph to obtain a knowledge-constrained multi-task segmentation network. The preprocessed tongue image is input into a knowledge-constrained multi-task segmentation network, and tongue body segmentation, site segmentation, and multi-feature extraction are performed based on the segmentation sites in the prior knowledge graph to obtain tongue body segmentation results, site segmentation results, and feature distribution results. The site segmentation results and feature distribution results are combined, and joint reasoning is performed based on the logical rules in the prior knowledge graph to obtain the semantic judgment result; The results of tongue segmentation, site segmentation, feature distribution, and semantic determination are presented in a correlated manner to obtain a tongue image semantic map and a comprehensive diagnostic report.
2. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 1, characterized in that: The acquisition of the original tongue image specifically includes, Standardized light source acquisition is used to acquire tongue images output by the tongue image acquisition device to obtain acquired tongue images; The tongue region is located in the acquired tongue image to obtain the localized tongue image; The tongue image is cropped to obtain the original tongue image.
3. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 1, characterized in that: The construction of the prior knowledge graph specifically includes, Extract the morphological features corresponding to tongue shape, tongue color, and tongue coating to obtain the morphological feature layer; Microscopic features corresponding to lingual papillae and saliva were extracted to obtain the microstructure layer; Extract the spatial positional relationships corresponding to the segmentation sites to obtain the spatial site layer; Extract the combination relationships of tongue color, coating, tongue shape, tongue papillae, saliva, and cracks to obtain a logical rule layer; By associating the morphological feature layer, microstructure layer, spatial location layer, and logical rule layer, a prior knowledge graph is obtained.
4. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 3, characterized in that: The training of the multi-task segmentation network specifically includes, Annotated tongue images containing tongue outline, location regions, tongue color grade, coating type, tongue papilla location, and saliva grade are organized into a sample set. The logical rules in the prior knowledge graph are converted into computable constraints to obtain a set of knowledge constraints. The labeled sample set is input into the multi-task segmentation network to obtain the training feature results; Based on the training feature results and the labeled sample set, a segmentation loss is constructed to obtain the pixel constraint results; The consistency of the site feature combination corresponding to the training feature result is compared with the knowledge constraint itemset to obtain the semantic constraint result. By jointly updating the pixel constraint results and semantic constraint results, a knowledge-constrained multi-task segmentation network is obtained.
5. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 4, characterized in that: The generation of the tongue segmentation results, site segmentation results, and feature distribution results specifically includes, The preprocessed tongue image is input into an encoder-decoder structure, and a knowledge-constrained multi-task segmentation network with spatial attention and channel attention mechanisms is introduced at the decoder end to obtain a multi-task feature map. Tongue segmentation is performed based on the multi-task feature map to obtain the tongue segmentation result; Region segmentation is performed based on segmentation sites in the multi-task feature map and prior knowledge graph to obtain the site segmentation results; Simultaneous parsing is performed on the multi-task feature map and the site segmentation results to obtain the feature distribution results.
6. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 5, characterized in that: The simultaneous parsing of the multi-task feature map and the site segmentation results specifically includes, Morphological and color analyses were performed on the corresponding sites of the lingual papillae in the site segmentation results to obtain the distribution results of the lingual papillae. Texture analysis was performed on the corresponding sites of tongue coating in the site segmentation results, and color analysis was performed on the corresponding sites of tongue color in the site segmentation results to obtain the distribution results of tongue coating and tongue color. Brightness analysis and region connectivity analysis were performed on the tongue region corresponding to the tongue segmentation results to obtain the body fluid distribution results. Contour analysis is performed on the tongue contour corresponding to the tongue segmentation result, and texture analysis and region connectivity analysis are performed on the crack features in the multi-task feature map to obtain the tongue state result and crack result; The results of tongue papillae distribution, tongue coating distribution, tongue color distribution, saliva distribution, tongue shape, and cracks are combined to obtain the characteristic distribution results.
7. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 6, characterized in that: The obtained semantic determination result specifically includes, The segmentation sites corresponding to the site segmentation results are combined with the tongue color distribution results, tongue coating distribution results, tongue papilla distribution results, saliva distribution results, tongue shape results, and crack results in the feature distribution results to obtain a structured feature set; Candidate rules are extracted from logical rules based on the segmentation points to obtain a candidate rule set; The structured feature set is matched with the candidate rule set to obtain structured diagnostic suggestions; By associating structured diagnostic prompts with corresponding segmentation sites and features, semantic determination results are obtained.
8. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 7, characterized in that: The obtained structured diagnostic prompts specifically include, The candidate rule set is filtered to obtain the matching rule set; Consistency checks are performed on the structured feature set and the matching rule set to obtain structured diagnostic prompts.
9. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 8, characterized in that: The obtained tongue image semantic map and comprehensive diagnostic report specifically include, The tongue segmentation results are overlaid with the site segmentation results to obtain the site annotation map; By hierarchically associating the site annotation map with the feature distribution results, a tongue image semantic map is obtained. The tongue image semantic map and semantic judgment results are combined to obtain a comprehensive diagnostic report.
10. The tongue image semantic reconstruction method integrating prior knowledge graph and image segmentation as described in claim 9, characterized in that: The merging of the tongue image semantic map and semantic determination results specifically includes, The results of tongue color distribution, tongue coating distribution, tongue papillae distribution, and saliva distribution are correlated with the site annotation map to obtain the site feature correspondence results; The site feature correspondence results and semantic determination results are combined into a graphic and textual representation to obtain the report content set; The report content set is combined with the original tongue image to obtain a comprehensive diagnostic report.