A traditional chinese medicine inspection method and system based on ar glasses
By synchronously collecting tongue image and physiological sign data through AR glasses, generating privacy authorization tokens, obtaining desensitized medical history data, and performing dynamic light and shadow compensation and multimodal dialectical reasoning, the problems of data security and diagnostic accuracy in intelligent TCM observation diagnosis are solved, realizing an efficient and safe TCM observation diagnosis method.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-07-10
AI Technical Summary
Existing intelligent TCM diagnostic techniques suffer from insufficient standardization in data collection process control, weak data privacy and security, significant interference from ambient lighting, susceptibility to tongue feature recognition, low diagnostic accuracy, reliance on single physical signs, and lack of longitudinal medical history reference.
By synchronously collecting dynamic video streams of tongue images, ambient lighting parameters, and physiological signs data through AR glasses, a privacy authorization token is generated. Desensitized medical history data is obtained through a privacy computing channel, and dynamic light and shadow compensation and temporal feature extraction are performed. Combined with a knowledge graph of traditional Chinese medicine syndromes, multimodal data-based syndrome differentiation and reasoning are carried out to achieve AR visualization annotation.
It standardized the data collection process for diagnosis and treatment, ensured data privacy and security, improved the accuracy of tongue feature recognition, enhanced the scientific nature of dialectical reasoning and the intuitiveness of result presentation, and improved the efficiency of diagnosis and treatment and clinical acceptance.
Smart Images

Figure CN122369890A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of AR glasses technology, and in particular to a traditional Chinese medicine diagnostic method and system based on AR glasses. Background Technology
[0002] In traditional Chinese medicine (TCM), visual diagnosis is a core component of clinical practice. Identifying physical signs such as tongue appearance and facial complexion, along with syndrome differentiation, are key aspects of visual diagnosis. With the rapid development of augmented reality (AR), artificial intelligence, and medical big data technologies, intelligent TCM visual diagnosis techniques are gradually emerging. These techniques utilize AR glasses combined with intelligent algorithms to collect TCM physical signs, automate syndrome differentiation, and visualize treatment results. This has become an important direction for promoting the modernization, standardization, and convenience of TCM diagnosis and treatment, and is widely applied in outpatient clinics and primary healthcare settings.
[0003] In existing technologies, intelligent TCM diagnostic techniques have basic functions such as physical sign collection, tongue image recognition, and preliminary syndrome differentiation. They can use AR devices to collect tongue images, rely on artificial intelligence algorithms to identify static tongue features and make simple syndrome judgments, and also support basic retrieval of electronic medical records and display of text-based diagnostic results. However, existing technologies lack standardized management of the data collection process and have weak authorization and security guarantees for data retrieval. Secondly, in the process of acquiring cross-institutional medical history data, data availability and privacy security cannot be simultaneously guaranteed, resulting in a high risk of data privacy leakage. Furthermore, tongue feature recognition is easily affected by ambient lighting factors and mostly focuses on static image analysis, failing to accurately capture the temporal dynamic changes in tongue texture and characteristics. Moreover, the diagnostic reasoning process often relies on single physical sign data, leading to low diagnostic accuracy.
[0004] Therefore, it is necessary to improve one or more of the problems existing in the above-mentioned related technical solutions.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this disclosure is to provide a TCM diagnostic method and system based on AR glasses, thereby overcoming, to at least to some extent, one or more problems caused by the limitations and defects of related technologies.
[0007] Firstly, this application provides a traditional Chinese medicine diagnostic method based on AR glasses, including: After obtaining the patient's authorization, the system simultaneously collects dynamic video streams of the patient's tongue image, ambient lighting parameters, and physiological signs data to obtain all raw diagnostic data and generate a bound privacy authorization token. Based on the privacy authorization token, an encrypted medical history query is initiated to the medical institution through the privacy computing channel to obtain the de-identified structured medical history data; Based on the ambient lighting parameters, the dynamic video stream of the tongue image is subjected to dynamic lighting compensation and standardization processing, and the temporal dynamic change features of tongue color and texture are extracted through a residual 3D convolutional network to obtain the dynamic feature vector of the tongue image. The dynamic feature vector of the tongue image, the physiological signs data, and the structured medical history data are aligned in time sequence. Spatial correlation reasoning and temporal evolution deduction are performed through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph. The syndrome probability is calculated and the diagnostic basis is marked. The diagnostic results and abnormal signs are displayed through AR semantic annotation.
[0008] Secondly, this application provides a TCM diagnostic system based on AR glasses, used to perform the aforementioned TCM diagnostic method based on AR glasses, including: The privacy authorization module is used to synchronously collect dynamic video streams of the patient's tongue image, ambient light parameters, and physiological signs data after the patient authorizes it, to obtain all the original diagnostic and treatment data and generate a bound privacy authorization token. The medical history acquisition module is used to initiate an encrypted medical history query to the medical institution through a privacy computing channel based on the privacy authorization token, and obtain the de-identified structured medical history data; The tongue image analysis module is used to perform dynamic light and shadow compensation and standardization processing on the dynamic video stream of the tongue image based on the ambient lighting parameters, and to extract the temporal dynamic change features of tongue color and texture through a residual 3D convolutional network to obtain the dynamic feature vector of the tongue image. The syndrome differentiation and annotation module is used to align the dynamic feature vector of the tongue image, the physiological signs data and the structured medical history data in time sequence, and to perform spatial correlation reasoning and temporal evolution deduction through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph, calculate the syndrome probability and annotate the syndrome differentiation basis, and display the syndrome differentiation results and abnormal signs through AR semantic annotation.
[0009] The technical solution provided in this application may include the following beneficial effects: This application presents a TCM diagnostic method and system based on AR glasses. Upon patient authorization, it can simultaneously collect dynamic video streams of the tongue image, ambient lighting parameters, and physiological signs, generating a bound privacy authorization token. This standardizes the collection process of original diagnostic data and ensures the security of authorized data retrieval. Furthermore, based on the privacy authorization token, it initiates encrypted queries via a privacy computing channel to obtain desensitized structured medical history data, achieving secure acquisition and privacy protection of cross-institutional medical history data. Simultaneously, it performs standardized processing of the dynamic video stream of the tongue image with light and shadow compensation and extracts temporal dynamic features. These features are then integrated with multimodal data for dialectical reasoning and AR visualization annotation, improving the accuracy of feature recognition, the scientific nature of dialectical reasoning, and the intuitiveness of result presentation in TCM diagnostic methods.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0012] Figure 1 A flowchart illustrating a traditional Chinese medicine diagnostic method based on AR glasses in an exemplary embodiment of this disclosure is shown. Figure 2 A detailed flowchart of step S100 of the traditional Chinese medicine diagnostic method based on AR glasses in an exemplary embodiment of this disclosure is shown. Figure 3 A detailed flowchart of step S200 of the traditional Chinese medicine diagnostic method based on AR glasses in an exemplary embodiment of this disclosure is shown. Figure 4 A detailed flowchart of step S300 of the traditional Chinese medicine diagnostic method based on AR glasses in an exemplary embodiment of this disclosure is shown. Figure 5 A detailed flowchart of step S400 of the traditional Chinese medicine observation diagnosis method based on AR glasses in an exemplary embodiment of this disclosure is shown. Figure 6 This diagram illustrates the structure of a traditional Chinese medicine diagnostic system based on AR glasses, as shown in an exemplary embodiment of this disclosure. Detailed Implementation
[0013] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0014] This example implementation first provides a traditional Chinese medicine diagnostic method based on AR glasses. This method can be applied to a terminal device, such as a mobile terminal like a mobile phone, desktop computer, personal digital assistant, laptop, tablet, or smartwatch. (Reference) Figure 1 As shown, the method may include the following steps: Step S100: After obtaining the patient's authorization, the system simultaneously collects the patient's tongue image dynamic video stream, ambient light parameters, and physiological signs data to obtain the full amount of original diagnostic and treatment data and generate a bound privacy authorization token.
[0015] Step S200: Based on the privacy authorization token, initiate an encrypted medical history query to the medical institution through the privacy computing channel to obtain the de-identified structured medical history data.
[0016] Step S300: Based on the ambient lighting parameters, perform dynamic light and shadow compensation standardization processing on the tongue image dynamic video stream, and extract the temporal dynamic change features of tongue color and texture through a residual 3D convolutional network to obtain the tongue image dynamic feature vector.
[0017] Step S400: The tongue dynamic feature vector, the physiological signs data and the structured medical history data are aligned in time sequence, and spatial correlation reasoning and temporal evolution inference are performed through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph. The syndrome probability is calculated and the diagnostic basis is marked. The diagnostic results and abnormal signs are displayed through AR semantic annotation.
[0018] The above method addresses the pain points of traditional Chinese medicine (TCM) observation diagnosis, such as static analysis, fragmented diagnostic logic, insufficient data privacy and security, poor scenario adaptability, and weak interpretability of diagnostic results. It achieves intelligent TCM observation diagnosis that is safe, accurate, and interpretable throughout the entire process: A dual-modal authorization mechanism with triple verification of gestures, voice, and voiceprints ensures the compliance and security of patient data collection and medical history retrieval. Simultaneously collected dynamic video streams of tongue images, ambient lighting parameters, and physiological signs provide a multi-dimensional, time-consistent raw data foundation for subsequent dynamic diagnosis, avoiding data leakage and authorization disputes from the source. Relying on a privacy computing channel, it achieves usability without visibility of cross-institutional medical history data, ensuring that original sensitive data does not leave the medical institution throughout the process. While ensuring data privacy and compliance, it obtains desensitized structured medical history data, providing complete support for diagnosis based on the course of the disease and past treatment information, solving the problem of isolated medical history information and lack of longitudinal disease course reference in traditional observation diagnosis. Dynamic light and shadow compensation eliminates the inconsistencies in... To mitigate the interference of ambient lighting on tongue image recognition, this method combines residual 3D convolutional networks to extract dynamic features of the tongue image, overcoming the limitations of traditional static tongue diagnosis which relies solely on single-frame images. It increases the accuracy of identifying key features such as tongue moisture and dryness from 82% to 91%, significantly improving the precision and robustness of tongue feature extraction. This approach is suitable for various diagnostic scenarios, including clinics, wards, and outdoor settings. Furthermore, by incorporating a spatiotemporal graph convolutional network with a TCM syndrome knowledge graph, it integrates tongue image, physiological signs, and medical history data to achieve dynamic syndrome differentiation and reasoning. This simulates the spatial correlation logic of TCM syndrome differentiation based on symptoms, and enhances the continuity and individualized adaptability of syndrome differentiation through temporal evolution modeling and medical history correction. Simultaneously, it generates interpretable diagnostic evidence with classic TCM texts, and combines AR semantic annotation to intuitively present the diagnostic results and abnormal signs. This reduces the doctor's diagnostic steps from 7 to 3, significantly improving diagnostic efficiency and clinical acceptance. The overall method balances data security, diagnostic accuracy, scenario adaptability, and clinical practicality.
[0019] Below, we will refer to Figures 2 to 5 The steps of the method described above in this example embodiment will be explained in more detail.
[0020] In step S100, after obtaining the patient's authorization, the system simultaneously collects the patient's tongue image dynamic video stream, ambient light parameters, and physiological signs data to obtain the full amount of original diagnostic and treatment data and generate a bound privacy authorization token.
[0021] It should be noted that a dual-modal biometric authentication method combining gestures and voice is used to complete patient authorization. The synchronously collected dynamic video stream of the tongue image is used to capture dynamic vital signs such as the tongue's moisture and color changes. Ambient lighting parameters are used to eliminate the interference of light on vital sign recognition. Physiological vital sign data is heart rate variability data calculated from neck PPG signals, which is used to match syndrome-related vital signs such as pulse rate in traditional Chinese medicine. All collected raw diagnostic and treatment data are stored in the edge computing unit built into the AR glasses, using a localized processing mode, without transmitting sensitive data such as the patient's face and tongue image to the cloud, which complies with medical data security standards. The generated privacy authorization token is a blockchain encrypted token with a timestamp and digital signature. It is only bound to the current medical session and has a short validity period, which can prevent the theft and reuse of the token. Blockchain storage can also realize traceability of the diagnosis and treatment process.
[0022] In one embodiment, such as Figure 2 As shown, step S100 may include the following sub-steps.
[0023] In step S110, the doctor initiates visual diagnosis and medical history retrieval commands via voice or gesture through AR glasses, and guides the patient to complete the specified gesture or authorized voice expression.
[0024] It should be noted that doctors can preset voice commands through AR glasses, such as requesting a patient's medical history or initiating a visual examination, or preset gestures, such as clenching a fist or swiping, to initiate commands. After the system triggers the authorization request, it can issue authorization guidance prompts to the patient through bone conduction headphones, clearly requiring the patient to look at the camera, say the specified authorization voice, and make the specified gesture. The authorization prompts are simultaneously presented in a visual interface in the patient's AR glasses field of view, ensuring that the patient clearly understands the authorization operation process and improving the convenience and accuracy of the authorization operation.
[0025] In step S120, the motion trajectory of the key points of the patient's 3D skeleton is collected in parallel by the depth sensor of the AR glasses, and the voiceprint features are extracted after the patient's authorized speech is collected by the microphone array, thus completing the collection of triple biometric features of gesture, voice and voiceprint.
[0026] It should be noted that the depth sensor of the AR glasses can use a ToF sensor or a fisheye camera to capture depth and RGB images of the patient's hand area. Gesture recognition uses a lightweight 3D convolutional neural network or graph convolutional network to identify the 3D spatial position and movement trajectory of 21 key skeletal points of the hand in real time. Authorized gestures are preset standardized actions such as drawing circles and opening and closing fists. The microphone array can use beamforming technology to filter out ambient noise in the examination room and accurately focus on the patient's voice. Speech recognition converts the audio stream into authorized text, and voiceprint feature extraction is used to confirm that the voice comes from the patient, avoiding unauthorized operations such as recording playback and others answering on behalf of the patient. Gesture, speech, and voiceprint acquisition are executed in parallel and synchronously to ensure consistency of bimodal data time.
[0027] In step S130, based on the triple biometrics, the confidence level of gesture matching, the matching degree of voice content, and the uniqueness of voiceprint are weighted and judged. After the judgment is passed, the patient authorization is completed, and a privacy authorization token containing the query scope, validity period and patient ID is generated.
[0028] It should be noted that a lightweight multimodal fusion decision engine can be built in, designed with reference to the concept of multimodal data fusion engines. This engine is used for the weighted comprehensive calculation of gesture recognition confidence, voice content matching confidence, and voiceprint uniqueness confidence. Among them, the gesture matching confidence threshold is >95%, and the voice content and voiceprint matching confidence threshold is >90%. Authorization is considered successful only if all three criteria are met. If any modality fails, the failure type is clearly indicated and a retry is guided. The generated privacy authorization token includes the query scope, validity period, and patient ID. The patient ID can be an anonymized ID generated from the patient's unique identifier using a hash algorithm to protect the patient's identity privacy.
[0029] Furthermore, step S130 includes: In step S131, the motion trajectory of the key points of the 3D skeleton of the hand, the authorized voice text and voiceprint features are matched with the locally pre-stored authorized templates, and the corresponding confidence values are output.
[0030] It should be noted that the locally pre-stored authorized template is stored in the edge computing unit of the AR glasses. It can include the standard gesture trajectory, authorized voice text, and voiceprint feature benchmark data pre-recorded by the patient. The matching process is executed locally throughout and does not transmit biometric data to the outside world. The 3D skeletal trajectory of the hand, authorized voice text, and voiceprint features are compared with the template independently and in parallel, and the confidence value in the range of 0-1 is output. The closer the value is to 1, the higher the degree of matching.
[0031] In step S132, the three types of confidence are weighted according to preset rules, and the validity of the authorization is determined based on the weighted result.
[0032] It should be noted that the preset weighting rule can give higher weight to gesture features than to voice and voiceprint features. After weighted calculation, a comprehensive authorization score is obtained. If the score meets the standard, the authorization is valid; otherwise, it is invalid. This rule can be set in combination with the security requirements of TCM diagnosis and treatment scenarios. When authorization fails, the failure mode can be accurately located, eliminating the need for the patient to repeat all operations and simplifying the retry process.
[0033] In step S133, after determining that the authorization is valid, a privacy authorization token bound to the current session is generated based on the scope of medical history query, the effective duration of the session, and the patient ID.
[0034] It should be noted that the validity period of the privacy authorization token can be preset according to the medical treatment scenario, such as 5 minutes for regular outpatient visits and 24 hours for long-term medical treatment. The token is uniquely bound to the current session and will automatically expire when the session ends or expires. After the token is generated, it is immediately recorded on the blockchain hash for evidence storage. The evidence storage information includes the authorization time, token identifier, and query scope, which can serve as a valid certificate of authorization compliance in medical disputes and enhance the legal effect of the medical treatment process.
[0035] In step S140, a continuous video stream of the patient's tongue is acquired at a frame rate of 30fps, and parameters such as color temperature, illuminance, and main light source direction are acquired through an ambient light sensor, while physiological sign data are acquired through a physiological sign sensor.
[0036] It should be noted that the dynamic video stream of the tongue image is acquired at a frame rate of 30fps, with a continuous acquisition time of ≥10 seconds, to ensure complete capture of dynamic features such as changes in tongue color and tongue coating moisture. The ambient light sensor simultaneously acquires color temperature, illuminance, and main light source direction parameters, which serve as the core input for subsequent dynamic light and shadow compensation. The physiological sign sensor acquires neck PPG signals, and the edge computing unit calculates heart rate variability data in real time, corresponding to pulse characteristics such as pulse rate and pulse slowness in traditional Chinese medicine. All data are acquired synchronously to ensure the temporal consistency of multimodal data.
[0037] In step S200, based on the privacy authorization token, an encrypted medical history query is initiated to the medical institution through the privacy computing channel to obtain the de-identified structured medical history data.
[0038] It should be noted that cross-institutional medical history queries are achieved based on federated learning, secure multi-party computation, and trusted execution environment technologies. Following the principle of data remaining stationary while the model moves and queries remaining stationary while data moves, the patient's original medical history data is stored entirely within the medical institution's internal network, and only anonymized structured data is returned, achieving data usability without visibility. At the same time, it provides longitudinal data such as past diagnoses, medication history, and disease progression for traditional Chinese medicine diagnosis, solving the problems of isolated medical history information and lack of continuity in diagnosis.
[0039] In one embodiment, such as Figure 3 As shown, step S200 may include the following sub-steps.
[0040] In step S210, an encrypted medical history query request containing the privacy authorization token, query semantics and patient ID is constructed locally and sent to the relay unit through a TLS / SSL encrypted channel.
[0041] It should be noted that the encrypted medical history query request is built locally by the edge computing unit of the AR glasses. It does not contain sensitive information such as the patient's name or ID number, but only carries a privacy authorization token, the patient ID, and standardized query semantics. The query semantics are structured statements, which can accurately target specific symptoms and medical histories. The request is transmitted to the privacy computing scheduling center through a TLS / SSL encrypted channel. This center only schedules the request and does not store the patient's original data.
[0042] In step S220, the encrypted medical history query request is parsed in the relay unit, the encrypted query task is distributed to the medical institution with relevant data, and the query is executed on the privacy computing node deployed in the internal network of the medical institution to obtain the query result; the privacy computing node runs in a trusted execution environment.
[0043] It should be noted that the relay unit can serve as a privacy computing scheduling center, performing only compliance parsing of encrypted requests without accessing sensitive patient medical history data. After parsing, it distributes encrypted query tasks to partner medical institutions. Each medical institution deploys privacy computing nodes internally, which run in an Intel SGX or ARM TrustZone trusted execution environment. All query operations are executed in a memory-isolated black box, preventing external parties from spying on or stealing data.
[0044] Furthermore, step S220 includes: In step S221, the relay unit decrypts and parses the received encrypted medical history query request to extract the authorization token, query semantics, and patient anonymous ID information.
[0045] It should be noted that the relay unit decryption and parsing only extracts necessary information such as the validity of the authorization token, query semantics, and patient ID, without parsing patient privacy information. At the same time, it verifies the validity of the token and the digital signature. If the token expires or the signature is invalid, the query will be rejected directly to ensure the compliance of the query.
[0046] In step S222, the encrypted query task is distributed to the privacy computing node of the corresponding institution by matching the query semantics with the medical institution holding the corresponding medical history data.
[0047] It should be noted that the privacy computing scheduling center has a built-in data node matching library of cooperating medical institutions. Based on the disease type and time range of the query semantics, it automatically matches the medical institutions that store the corresponding medical history, accurately distributes the encrypted query task to the target node, avoids invalid queries, and improves query efficiency.
[0048] In step S223, within the trusted execution environment of the privacy computing node, a retrieval is performed in the database of the corresponding institution based on the parsed query conditions to obtain the original medical history query results.
[0049] It should be noted that the trusted execution environment is a hardware-level secure isolation environment with encryption, isolation, and anti-tampering features. The privacy computing node only searches the local database according to the query conditions, and only retrieves the patient's relevant medical history data within the query scope. It does not obtain irrelevant medical data, thus ensuring the accuracy and security of the query.
[0050] In step S230, the query result is replaced, generalized, and subjected to differential privacy desensitization processing, and then encrypted by the privacy computing node before being returned to the relay unit.
[0051] It should be noted that the query results must undergo three layers of desensitization before leaving the medical institution: replacement processing converts the specific drug name to the ATC classification code, generalization processing converts the precise date to the year / quarter, and differential privacy processing adds a small amount of noise to prevent reverse identification; after desensitization, the data is encrypted again by the privacy computing node to ensure transmission security.
[0052] In step S240, fragmented data from multiple institutions are aggregated into an encrypted complete dataset through secure multi-party computation and returned to the local machine.
[0053] It should be noted that the privacy computing scheduling center uses secure multi-party computation technology to aggregate fragmented medical history data into a complete encrypted dataset without decrypting the anonymized and encrypted data of each institution. The aggregation process does not expose any data, thus ensuring the privacy and security of multi-institutional data fusion.
[0054] In step S250, the encrypted aggregated data is decrypted locally to obtain desensitized structured medical history data containing previous diagnoses, medications, and disease duration.
[0055] It should be noted that the decryption of the encrypted dataset is completed locally on the edge computing unit of the AR glasses. After decryption, the data includes previous diagnoses, medications, disease course, and physical constitution identification results. All sensitive identity information is removed, and the data can be directly input into the dialectical reasoning engine to support individualized dialectical modeling.
[0056] In step S300, dynamic light and shadow compensation and standardization processing is performed on the tongue image dynamic video stream based on the ambient lighting parameters, and the temporal dynamic change features of tongue color and texture are extracted through a residual 3D convolutional network to obtain the tongue image dynamic feature vector.
[0057] It should be noted that dynamic light and shadow compensation is used to address the tongue image recognition bias caused by lighting environment, and residual 3D convolutional network is used to address the shortcomings of traditional visual diagnosis, which only processes single static images and ignores the dynamic changes of tongue texture. The residual 3D convolutional network combines the spatial extraction capability of 2D CNN with the temporal modeling capability of 3D CNN to accurately capture the temporal changes of tongue texture color and texture, improving the accuracy of tongue texture moisture and dryness recognition from 82% to 91%.
[0058] In one embodiment, such as Figure 4 As shown, step S300 may include the following sub-steps.
[0059] In step S310, based on the ambient lighting parameters, the tongue image dynamic video stream is decoupled from the ambient light through physical rendering to remove ambient light interference, and then re-rendered using preset light source parameters to generate a standardized tongue image sequence.
[0060] It should be noted that dynamic lighting compensation is based on PBR physical rendering technology. The system has a built-in skin bidirectional reflectance distribution function model to simulate the interaction between light and tongue tissue. Through lighting decoupling and standard light source re-rendering, the tongue image under full scene lighting is converted into a standard clinic image, so that the facial color and tongue image recognition error is less than 5%, and the system's environmental adaptability is improved.
[0061] Furthermore, step S310 includes: In step S311, based on the ambient lighting parameters, the tongue dynamic video stream is decoupled into an intrinsic reflectivity component and an ambient lighting component through physical rendering; the ambient lighting parameters include color temperature, illuminance, and main light source direction.
[0062] It should be noted that the physical rendering model decouples the tongue image video frame into an intrinsic reflectivity component and an ambient lighting component. The intrinsic reflectivity is the true color and texture of the tongue, while the ambient lighting component is the interference such as color temperature, shadows, and highlights. Accurate acquisition of ambient lighting parameters is the core prerequisite for decoupling.
[0063] In step S312, the ambient light component in the tongue image dynamic video stream is removed, while the inherent reflectivity component is retained.
[0064] It should be noted that removing the ambient light component can completely eliminate interference such as color temperature deviation, shadow occlusion, and high light reflection, retaining only the true inherent reflectivity information of the tongue, thus eliminating the influence of light on tongue image feature recognition from the root.
[0065] In step S313, the inherent reflectivity component is re-rendered using preset light source parameters to generate a standardized tongue image sequence.
[0066] It should be noted that the preset light source can be the D65 standard daylight specifically used for TCM tongue diagnosis, which is a front diffuse light with a color temperature of 6500K. The inherent reflectivity component is re-rendered using this light source to generate a tongue image sequence with a unified lighting standard, ensuring the consistency of feature extraction.
[0067] In step S320, the standardized tongue image sequence is input into a residual 3D convolutional network to simultaneously extract the spatial texture features of the tongue body and the temporal variation features of the RGB channels of the tongue.
[0068] It should be noted that the input to the residual 3D convolutional network can be 16 consecutive frames of RGB normalized tongue images with a resolution of 224×224. The network consists of stacked 3D residual modules, each containing two 3×3×3 convolutional layers, batch normalization, ReLU activation, and identity skip connections. Skip connections solve the gradient vanishing problem, allowing the network to accurately learn subtle residual changes in tongue color and simultaneously extract spatial texture and RGB temporal variation features.
[0069] In step S330, the spatial texture features of the tongue body and the temporal change features of the RGB channels of the tongue are subjected to 3D pooling and global average pooling to obtain a dynamic feature vector of the tongue image containing information on the static color, dynamic change rate, and texture evolution of the tongue.
[0070] It should be noted that the network uses 2×2×2 3D max pooling to reduce the spatiotemporal dimension and extract high-level features, and global average pooling to generate a fixed-length one-dimensional tongue image feature vector. The vector contains the static color of the tongue, the RGB change rate, and the texture evolution features. Among them, the ΔR / Δt red change rate can directly indicate key diagnostic features such as the aggravation of the thermal image of the tongue.
[0071] In step S400, the tongue dynamic feature vector, the physiological signs data and the structured medical history data are time-series aligned, and spatial correlation reasoning and temporal evolution deduction are performed through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph. The syndrome probability is calculated and the diagnostic basis is marked. The diagnostic results and abnormal signs are displayed through AR semantic annotation.
[0072] It should be noted that by integrating multimodal data of tongue appearance, physical signs, and medical history, and incorporating a TCM syndrome knowledge graph that incorporates the "Treatise on Febrile Diseases," "Treatise on Differentiation of Febrile Diseases," and "TCM Diagnostics," dynamic reasoning based on symptoms, temporal evolution, and medical history correction is achieved through spatiotemporal graph convolutional networks. This generates interpretable diagnostic results with classic texts, which are presented intuitively through AR annotations. This approach aligns with the logic of integrating the four diagnostic methods in TCM, and solves the problems of fragmented diagnostic logic and lack of theoretical support.
[0073] In one embodiment, such as Figure 5 As shown, step S400 may include the following sub-steps.
[0074] In step S410, the tongue image dynamic feature vector, the physiological sign data, and the desensitized structured medical history data are time-series aligned to obtain the syndrome differentiation input feature set.
[0075] It should be noted that temporal alignment uses the collection timestamp as a unified benchmark to accurately match the temporal sequence of tongue appearance, physical signs, and medical history data, eliminating inference bias caused by asynchronous collection. The diagnostic input feature set generated after alignment is the unified data foundation for inference in spatiotemporal graph convolutional networks.
[0076] In step S420, the diagnostic input feature set is mapped to the TCM syndrome knowledge graph, tongue appearance quantitative features and physical signs quantitative features are assigned to symptom nodes, and past medical history, medication history and disease progression attributes are associated with syndrome nodes to construct a dynamic reasoning graph including real-time features and medical history attributes.
[0077] It should be noted that the TCM syndrome knowledge graph includes four types of nodes: symptoms, syndromes, treatment principles, and prescriptions / acupoints, as well as six types of relation edges: has_symptom, indicates, contrasts, followed_by, treats, and composed_of. It embeds the Wei Qi Ying Xue and San Jiao syndrome differentiation theories. After the feature set is mapped to the graph, symptom nodes are assigned quantitative features, and syndrome nodes are associated with medical history attributes. The constructed dynamic reasoning graph realizes individualized syndrome differentiation, which is consistent with the patient's disease course.
[0078] In step S430, spatial graph convolution is performed on the symptom nodes and syndrome nodes in the dynamic reasoning graph, and the quantified features of tongue appearance, quantified features of physical signs, and medical history association attributes are aggregated along the has_symptom and indicate relationship edges between nodes to obtain the initial syndrome probability.
[0079] It should be noted that spatial graph convolution can be performed on the dynamic inference graph at each time step. Symptom node features are aggregated along the has_symptom and indicate edges to syndrome nodes, simulating the TCM thinking of inferring syndromes from symptoms. The aggregated information includes tongue appearance, physical signs, and medical history features, and outputs the uncorrected initial syndrome probability.
[0080] In step S440, the continuous temporal feature sequences of the same symptom node and the same syndrome node in the dynamic reasoning graph are subjected to temporal convolution to obtain the syndrome temporal evolution coefficients.
[0081] It should be noted that temporal convolution can be a one-dimensional temporal convolution, which models the continuous temporal features of the same node after spatial convolution, simulates the evolution of syndromes over time, and outputs the temporal evolution coefficients of syndromes, which can characterize the trend of syndrome aggravation, relief, and stabilization, and are the core parameters of dynamic syndrome differentiation.
[0082] In step S450, the weights of historical disease courses consistent with the current syndrome are extracted from the desensitized structured medical history data, and the initial syndrome probability is corrected by combining the syndrome time-series evolution coefficient.
[0083] It should be noted that by combining medical history and course data to correct the probability of syndromes, the bias of single-time physical sign identification can be eliminated. It is especially suitable for the course analysis of chronic and recurrent diseases, improving the continuity and accuracy of syndrome differentiation. The correction process is in line with the course and outcome of TCM diseases.
[0084] Furthermore, step S450 includes: In step S451, the historical disease records corresponding to the current syndrome are matched from the desensitized structured medical history data, and the corresponding historical disease weights are extracted.
[0085] It should be noted that when matching the current syndrome with the historical course records from the desensitized medical history, the weight of the historical course is determined by the matching degree between the current syndrome and the historical syndrome and the consistency of the course trend. The higher the matching degree and the trend consistency, the greater the weight value, and the weight is a quantitative value of 0-1.
[0086] In step S452, the initial syndrome probability, the syndrome time-series evolution coefficient, and the historical disease course weight are weighted and calculated.
[0087] It should be noted that weighted calculation can use multiplication operations to multiply the initial syndrome probability, the time-series evolution coefficient, and the weight of the historical course of the disease, and integrate the three bases of real-time signs, time-series changes, and past medical history to improve the reliability of the diagnosis results.
[0088] In step S453, the initial syndrome probability is calibrated based on the weighted calculation result to obtain the corrected syndrome probability.
[0089] It should be noted that the probability calibration adjusts the initial probability according to the weighted results. If the evolution of the syndrome is consistent with the trend of the medical history, the probability is increased; if they are inconsistent, the probability is decreased. After calibration, the final syndrome differentiation probability is obtained by integrating multi-dimensional information.
[0090] In step S460, the syndrome with the highest probability of the initial syndrome after correction is selected as the final diagnosis result, and the contribution of core tongue appearance features, physiological signs and desensitization history is traced back by attention weight to generate the basis for diagnosis.
[0091] It should be noted that spatiotemporal graph convolutional networks can calculate the contribution of each feature through the attention mechanism, trace back the path of tongue appearance, physical signs, and medical history with high contribution, and generate structured diagnostic evidence containing contribution ratios and classic TCM provisions, thereby improving the interpretability and clinical credibility of AI diagnosis.
[0092] In step S470, the final diagnostic results, diagnostic basis, and abnormal physical signs are superimposed and rendered onto the AR glasses' field of view through AR semantic annotation.
[0093] It should be noted that AR semantic annotation can use SLAM technology to achieve 3D spatial registration of abnormal signs. Tongue teeth marks, cyanosis, etc. are labeled in three levels: mild (green), moderate (yellow), and severe (red). The annotation is calibrated in real time through head / eye tracking and attached to the corresponding sign position. The diagnosis results and evidence are rendered synchronously, reducing the doctor's operation steps from 7 steps to 3 steps and shortening the patient's symptom input time by 40%, which greatly improves the diagnostic efficiency.
[0094] Furthermore, such as Figure 6 As shown, this application also provides a TCM diagnostic system based on AR glasses, used to perform the aforementioned TCM diagnostic method based on AR glasses, including: The privacy authorization module is used to synchronously collect dynamic video streams of the patient's tongue image, ambient light parameters, and physiological signs data after the patient authorizes it, to obtain all the original diagnostic and treatment data and generate a bound privacy authorization token. The medical history acquisition module is used to initiate an encrypted medical history query to the medical institution through a privacy computing channel based on the privacy authorization token, and obtain the de-identified structured medical history data; The tongue image analysis module is used to perform dynamic light and shadow compensation and standardization processing on the dynamic video stream of the tongue image based on the ambient lighting parameters, and to extract the temporal dynamic change features of tongue color and texture through a residual 3D convolutional network to obtain the dynamic feature vector of the tongue image. The syndrome differentiation and annotation module is used to align the dynamic feature vector of the tongue image, the physiological signs data and the structured medical history data in time sequence, and to perform spatial correlation reasoning and temporal evolution deduction through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph, calculate the syndrome probability and annotate the syndrome differentiation basis, and display the syndrome differentiation results and abnormal signs through AR semantic annotation.
[0095] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. A traditional Chinese medicine diagnostic method based on AR glasses, characterized in that, include: After obtaining the patient's authorization, the system simultaneously collects dynamic video streams of the patient's tongue image, ambient lighting parameters, and physiological signs data to obtain all raw diagnostic data and generate a bound privacy authorization token. Based on the privacy authorization token, an encrypted medical history query is initiated to the medical institution through the privacy computing channel to obtain the de-identified structured medical history data; Based on the ambient lighting parameters, the dynamic video stream of the tongue image is subjected to dynamic lighting compensation and standardization processing, and the temporal dynamic change features of tongue color and texture are extracted through a residual 3D convolutional network to obtain the dynamic feature vector of the tongue image. The dynamic feature vector of the tongue image, the physiological signs data, and the structured medical history data are aligned in time sequence. Spatial correlation reasoning and temporal evolution deduction are performed through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph. The syndrome probability is calculated and the diagnostic basis is marked. The diagnostic results and abnormal signs are displayed through AR semantic annotation.
2. The traditional Chinese medicine diagnostic method based on AR glasses according to claim 1, characterized in that, The steps described above, after patient authorization, to simultaneously collect dynamic video streams of the patient's tongue image, ambient lighting parameters, and physiological signs data to obtain complete raw diagnostic data and generate a bound privacy authorization token, include: Doctors can initiate visual diagnosis and medical history retrieval commands via voice or gesture through AR glasses, and guide patients to complete specified gestures or authorize voice statements. The AR glasses use a depth sensor to collect the motion trajectory of key points of the 3D skeleton of the patient's hand in parallel, and use a microphone array to collect the patient's authorized speech and extract voiceprint features to complete the collection of triple biometric features of gesture, voice and voiceprint. Based on the aforementioned triple biometrics, the confidence level of gesture matching, the matching degree of voice content, and the uniqueness of voiceprint are weighted and judged. After the judgment is passed, the patient authorization is completed, and a privacy authorization token containing the query scope, validity period and patient ID is generated. The system acquires a continuous video stream of the patient's tongue at a frame rate of 30fps, and collects parameters such as color temperature, illuminance, and main light source direction through an ambient light sensor, and collects physiological sign data through a physiological sign sensor.
3. The traditional Chinese medicine diagnostic method based on AR glasses according to claim 2, characterized in that, The step of weighting and determining the confidence level of gesture matching, the matching degree of voice content, and the uniqueness of voiceprint based on the three biometric features, and completing the patient authorization after the determination is successful, and generating a privacy authorization token containing the query scope, validity period and patient ID, includes: The motion trajectory of the key points of the 3D skeleton of the hand, the authorized speech text and voiceprint features are matched with the local pre-stored authorization templates respectively, and the corresponding confidence values are output. The three types of confidence levels are weighted according to preset rules, and the validity of the authorization is determined based on the weighted result. After determining that the authorization is valid, a privacy authorization token bound to the current session is generated based on the scope of medical history query, the validity duration of the session, and the patient ID.
4. The traditional Chinese medicine diagnostic method based on AR glasses according to claim 1, characterized in that, The step of initiating an encrypted medical history query to a medical institution through a privacy computing channel based on the privacy authorization token to obtain anonymized structured medical history data includes: An encrypted medical history query request containing the privacy authorization token, query semantics, and patient ID is constructed locally and sent to the relay unit via a TLS / SSL encrypted channel; The relay unit parses the encrypted medical history query request, distributes the encrypted query task to medical institutions with relevant data, and executes the query on a privacy computing node deployed in the medical institution's internal network to obtain the query result; the privacy computing node runs in a trusted execution environment. The query results are replaced, generalized, and subjected to differential privacy desensitization processing, and then encrypted through the privacy computing node before being returned to the relay unit; Fragmented data from multiple institutions is aggregated into an encrypted, complete dataset through secure multi-party computation and returned locally; The encrypted aggregated data is decrypted locally to obtain desensitized structured medical history data containing previous diagnoses, medications, and disease duration.
5. The TCM diagnostic method based on AR glasses according to claim 4, characterized in that, The relay unit parses the encrypted medical history query request, distributes the encrypted query task to medical institutions with relevant data, and executes the query on a privacy computing node deployed in the medical institution's internal network to obtain the query result. The steps of running the privacy computing node in a trusted execution environment also include: The relay unit decrypts and parses the received encrypted medical history query request, extracting the authorization token, query semantics, and patient anonymous ID information; Based on the semantic matching of the query, the medical institutions holding the corresponding medical history data are matched, and the encrypted query task is distributed to the privacy computing nodes of the corresponding institutions. Within the trusted execution environment of the privacy computing node, a retrieval is performed in the database of the corresponding institution based on the parsed query conditions to obtain the original medical history query results.
6. The traditional Chinese medicine diagnostic method based on AR glasses according to claim 1, characterized in that, The steps of performing dynamic lighting compensation and standardization processing on the tongue image dynamic video stream based on the ambient lighting parameters, and extracting the temporal dynamic change features of tongue color and texture through a residual 3D convolutional network to obtain the tongue image dynamic feature vector include: Based on the ambient lighting parameters, the dynamic video stream of the tongue image is decoupled from the ambient light through physical rendering to remove ambient light interference, and then re-rendered using preset light source parameters to generate a standardized tongue image sequence. The standardized tongue image sequence is input into a residual 3D convolutional network to simultaneously extract the spatial texture features of the tongue body and the temporal variation features of the RGB channels of the tongue. The spatial texture features of the tongue body and the temporal variation features of the RGB channels of the tongue are processed by 3D pooling and global average pooling to obtain a dynamic feature vector of the tongue image containing information on the static color, dynamic change rate, and texture evolution of the tongue.
7. The traditional Chinese medicine diagnostic method based on AR glasses according to claim 6, characterized in that, The steps of decoupling the dynamic video stream of the tongue image based on the ambient lighting parameters through physical rendering to remove ambient light interference, and re-rendering using preset light source parameters to generate a standardized tongue image sequence include: Based on the ambient lighting parameters, the tongue dynamic video stream is decoupled into an intrinsic reflectivity component and an ambient lighting component through physical rendering; the ambient lighting parameters include color temperature, illuminance, and main light source direction. Remove the ambient light component from the tongue image dynamic video stream, and retain the intrinsic reflectivity component; The inherent reflectivity component is re-rendered using preset light source parameters to generate a standardized tongue image sequence.
8. The traditional Chinese medicine diagnostic method based on AR glasses according to claim 1, characterized in that, The steps of aligning the tongue dynamic feature vector, the physiological sign data, and the structured medical history data in a temporal sequence, performing spatial association reasoning and temporal evolution deduction through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph, calculating syndrome probabilities and labeling diagnostic criteria, and displaying diagnostic results and abnormal signs through AR semantic annotation include: The tongue image dynamic feature vector, the physiological sign data, and the desensitized structured medical history data are time-series aligned to obtain the syndrome differentiation input feature set; The diagnostic input feature set is mapped to the TCM syndrome knowledge graph, and the symptom nodes are assigned quantified tongue appearance features and quantified physical signs features. The syndrome nodes are associated with past medical history, medication history, and disease progression attributes, and a dynamic reasoning graph including real-time features and medical history attributes is constructed. Spatial graph convolution is performed on the symptom nodes and syndrome nodes in the dynamic reasoning graph. The quantified features of tongue appearance, the quantified features of physical signs, and the medical history association attributes are aggregated along the has_symptom and indicate relationship edges between nodes to obtain the initial syndrome probability. Temporal convolution is performed on the continuous temporal feature sequences of the same symptom node and the same syndrome node in the dynamic reasoning graph to obtain the syndrome temporal evolution coefficients; Extract the weights of historical disease courses that are consistent with the current syndrome from the desensitized structured medical history data, and combine them with the syndrome time-series evolution coefficient to correct the initial syndrome probability; The syndrome with the highest probability of the initial syndrome after correction is selected as the final diagnosis result, and the contribution of core tongue appearance features, physiological signs and desensitization history is traced back by attention weight to generate the basis for diagnosis. Through AR semantic annotation, the final diagnostic results, diagnostic basis, and abnormal physical signs are overlaid and rendered onto the AR glasses' field of view.
9. The traditional Chinese medicine diagnostic method based on AR glasses according to claim 8, characterized in that, The step of extracting the weights of historical disease courses consistent with the current syndrome from the desensitized structured medical history data, and correcting the initial syndrome probability by combining the syndrome time-series evolution coefficient, includes: Match the historical disease records corresponding to the current syndrome from the desensitized structured medical history data, and extract the corresponding historical disease weights; The initial syndrome probability, the syndrome time-series evolution coefficient, and the historical disease course weight are weighted and calculated. The initial syndrome probability is calibrated based on the weighted calculation results to obtain the corrected syndrome probability.
10. A traditional Chinese medicine diagnostic system based on AR glasses, characterized in that, For performing the traditional Chinese medicine diagnostic method based on AR glasses as described in any one of claims 1-9, comprising: The privacy authorization module is used to synchronously collect dynamic video streams of the patient's tongue image, ambient light parameters, and physiological signs data after the patient authorizes it, to obtain all the original diagnostic and treatment data and generate a bound privacy authorization token. The medical history acquisition module is used to initiate an encrypted medical history query to the medical institution through a privacy computing channel based on the privacy authorization token, and obtain the de-identified structured medical history data; The tongue image analysis module is used to perform dynamic light and shadow compensation and standardization processing on the dynamic video stream of the tongue image based on the ambient lighting parameters, and to extract the temporal dynamic change features of tongue color and texture through a residual 3D convolutional network to obtain the dynamic feature vector of the tongue image. The syndrome differentiation and annotation module is used to align the dynamic feature vector of the tongue image, the physiological signs data and the structured medical history data in time sequence, and to perform spatial correlation reasoning and temporal evolution deduction through a spatiotemporal graph convolutional network equipped with a TCM syndrome knowledge graph, calculate the syndrome probability and annotate the syndrome differentiation basis, and display the syndrome differentiation results and abnormal signs through AR semantic annotation.