Multi-part image acquisition system and method for inspection diagnosis of traditional Chinese medicine, medium, program product and terminal
By integrating multimodal sensors and adaptive light sources into traditional Chinese medicine diagnostic equipment, and combining them with lightweight models for attitude and image quality assessment, the instability and uniformity of image acquisition in traditional Chinese medicine diagnostic equipment are solved. This enables automated, high-quality image acquisition of multiple body parts, making it suitable for standardized applications in various scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-12
- Publication Date
- 2026-04-10
AI Technical Summary
Existing TCM diagnostic equipment suffers from problems such as large fluctuations in color temperature and illumination during image acquisition, limited acquisition sites, lack of guaranteed image quality, reliance on manual guidance for operation, and lack of unified standards, which affect diagnostic accuracy and user experience.
Employing a multimodal sensor and adaptive light source component within a closed imaging cavity, combined with a lightweight tongue posture detection model and a site-adaptive image quality assessment model, it achieves automated acquisition of the face, tongue surface, and tongue base. Through multimodal in-situ detection and a closed-loop feedback mechanism, it ensures image quality and posture compliance.
It achieves fully automated, high-quality image acquisition of the face, tongue surface, and tongue base, improving image visibility and accuracy. It is suitable for hospital self-service terminals, community health stations, and home health devices, meeting the requirements of standardization and traceability.
Smart Images

Figure CN121817812A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent control technology for traditional Chinese medicine inspection by observation, in particular to a multi-site image acquisition system, method, medium, program product and terminal for traditional Chinese medicine inspection by observation. Background Art
[0002] Traditional Chinese medicine "inspection by observation", as the top of the four diagnostic methods (inspection by observation, auscultation and olfaction, interrogation, and palpation), has an irreplaceable position in clinical syndrome differentiation. Among them, facial image, tongue surface and sublingual region are the three core observation objects: the facial image can reflect the prosperity and decline of qi and blood and the state of zang-fu organs. For example, sallow complexion often indicates spleen deficiency; the tongue surface is used to judge key signs such as cold and heat, deficiency and excess, the thickness of the tongue coating, and cracks and tooth marks; while the sublingual region (i.e., the sublingual collaterals) is an important indicator of blood stasis syndrome and heart and liver diseases. However, due to its hidden position and the need for patients to actively turn the tongue for cooperation, there has long been a lack of standardized and automated acquisition means. In recent years, with the development of artificial intelligence technology, AI-assisted traditional Chinese medicine syndrome differentiation has gradually become a research hotspot, and many enterprises have launched "traditional Chinese medicine four diagnostic instruments" to automatically collect tongue image or facial image through a camera. However, existing devices still have significant technical defects in actual applications.
[0003] (1) Most devices rely on ambient light or only equipped with simple LED fill light, resulting in large fluctuations in color temperature and illuminance, which极易造成舌色失真,例如将正常的“淡红舌”误判为“绛舌”。
[0004] (2) The acquisition sites are generally single, and most devices only support tongue surface or facial image shooting, and it is almost impossible to achieve standardized acquisition of the sublingual region, severely restricting the recognition ability of complex syndrome types such as blood stasis syndrome.
[0005] (3) There is no effective guarantee mechanism for image quality. The system usually lacks functions such as in-place detection, clarity evaluation and reflection suppression, resulting in a large number of invalid or low-quality images being sent to the AI model, affecting the diagnostic accuracy.
[0006] (4) The operation process highly depends on manual guidance. For example, medical staff need to prompt users on-site to "stick out the tongue" and "raise the head", which is difficult to be applied to self-service terminals or home health scenarios.
[0007] (5) Due to the lack of a unified acquisition standard, the image results obtained at different times and by different devices are not comparable, hindering multi-center clinical research and the generalization ability of the AI model.
[0008] It should be noted that there is an unclear expression in the translation of item (11). You may need to check and correct it according to the specific meaning.Currently, the State Administration of Traditional Chinese Medicine is vigorously promoting the standardization of TCM diagnostic and treatment equipment, explicitly requiring that the image acquisition process be repeatable and traceable. Meanwhile, the National Medical Products Administration (NMPA) emphasizes the controllability of input data quality in the review of AI medical devices, with relevant standards such as the YY / T 1833 series specifying requirements for the stability and consistency of image acquisition. On the application side, hospitals and community health stations urgently need "unattended, one-click" terminals for the four diagnostic methods (inspection, diagnosis, and treatment) to reduce labor costs and improve service efficiency. At the same time, the demand for family-based TCM health management is rapidly increasing, but the inconsistent quality of user-taken images in open environments seriously affects user trust and clinical value. Summary of the Invention
[0009] In view of the shortcomings of the prior art, the present invention provides a multi-site image acquisition system, method, medium, program product and terminal for traditional Chinese medicine visual diagnosis, which is used to solve at least one of the technical problems in the prior art.
[0010] To achieve the above and other related objectives, the first aspect of this application provides a multi-site image acquisition system for traditional Chinese medicine visual diagnosis, comprising: an imaging cavity and a control module; wherein, the imaging cavity is equipped with a light source component, an image acquisition component, and a sensor component, for acquiring images of the patient's face, tongue surface, and tongue base; the control module comprises: a multimodal perception unit, for receiving multimodal sensor data acquired by the sensor component, and triggering the image acquisition component to acquire images of the patient's current posture in real time based on the multimodal sensor data, and inputting the patient's current posture image into a trained tongue posture light detection model for posture compliance assessment, so as to obtain the posture assessment result of the patient's current posture image; The adaptive light source adjustment unit is used to adaptively adjust the light source component based on multimodal sensor data according to different patient areas to be sampled. The image quality assessment unit is used to trigger the image acquisition component to acquire images of the patient's areas to be sampled in real time according to the posture assessment results after the light source component is adaptively adjusted, obtain images of the areas to be sampled, and input the images of the areas to be sampled into the adaptive image quality assessment model for quality assessment to obtain quality assessment results. Based on the quality assessment results, it is determined whether the images of the areas to be sampled need to be re-acquired, until the quality assessment results of the acquired face images, tongue surface images, and tongue underside images all meet the preset quality assessment requirements, and then output standardized TCM inspection image results.
[0011] In some embodiments of the first aspect of this application, the multimodal sensing unit is further configured to determine whether the patient's current posture is acceptable based on the posture assessment result. If the posture is unacceptable, the unit generates corresponding augmented reality guidance information based on the posture assessment result so that the patient can adjust their posture in real time.
[0012] In some embodiments of the first aspect of this application, the sensor assembly includes: an ambient light sensor, a color temperature sensor, and an infrared sensor.
[0013] In some embodiments of the first aspect of this application, the image acquisition component includes: a low-resolution camera and a high-resolution camera.
[0014] In some embodiments of the first aspect of this application, the control module further includes: a data storage unit; the data storage unit is used to store data during the operation of the image acquisition system.
[0015] In some embodiments of the first aspect of this application, the system further includes: a human-computer interaction module; the human-computer interaction module is used to display augmented reality guidance information to guide the patient to adjust their posture.
[0016] To achieve the above and other related objectives, a second aspect of this application provides a method for acquiring multi-site images for traditional Chinese medicine (TCM) visual diagnosis, applied to the multi-site image acquisition system for TCM visual diagnosis as described above. The method includes: receiving multimodal sensor data acquired by a sensor component; triggering an image acquisition component in real-time to acquire an image of the patient's current posture based on the multimodal sensor data; inputting the patient's current posture image into a trained tongue posture light detection model for posture compliance assessment to obtain a posture assessment result for the patient's current posture image; adaptively adjusting a light source component based on the multimodal sensor data according to the different patient sites to be acquired; after the light source component is adaptively adjusted, triggering the image acquisition component in real-time to acquire images of the patient's sites based on the posture assessment result to obtain images of the sites to be acquired; inputting the images of the sites to be acquired into a site-adaptive image quality assessment model for quality assessment to obtain a quality assessment result; determining whether to re-acquire images of the sites to be acquired based on the quality assessment result; and outputting standardized TCM visual diagnosis image results when the quality assessment results of the acquired face image, tongue surface image, and tongue underside image all meet preset quality assessment requirements.
[0017] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-site image acquisition method for traditional Chinese medicine visual diagnosis.
[0018] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product, which includes computer program code that, when executed on a computer, enables the computer to implement the multi-site image acquisition method for traditional Chinese medicine visual diagnosis.
[0019] To achieve the above and other related objectives, a fifth aspect of this application provides an electronic terminal, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the multi-site image acquisition method for traditional Chinese medicine visual diagnosis.
[0020] As described above, the multi-site image acquisition system, method, medium, program product, and terminal for traditional Chinese medicine visual diagnosis provided in this application have the following beneficial effects:
[0021] (1) This application realizes fully automatic closed acquisition of three key parts of TCM visual diagnosis: face, tongue surface and tongue base, filling the technical gap of TCM visual diagnosis image acquisition methods that have long lacked standardized acquisition methods.
[0022] (2) This application integrates an ambient light sensor and a color temperature sensor in a closed imaging cavity to construct a closed-loop feedback mechanism, precisely control the light source components, and ensure that the color temperature is stable within the medically effective range.
[0023] (3) This application proposes a multimodal in-situ detection mechanism that integrates infrared proximity sensing and a lightweight tongue posture detection model to accurately determine whether each part is exposed in compliance with regulations and to guide posture compliance, thereby significantly improving the visibility of key features such as sublingual veins in the image.
[0024] (4) This application constructs a site-adaptive image quality assessment model, optimizes the evaluation weights of clarity, color and texture for different sites of TCM physical signs, and improves the accuracy of image quality assessment.
[0025] (5) This application achieves automatic acquisition of high-quality images of multiple parts without human intervention by using a closed imaging cavity structure, multimodal positioning detection and closed-loop guidance mechanism, from user guidance, light adjustment, quality assessment to automatic shooting. It is applicable to various scenarios such as hospital self-service terminals, community health stations and home health equipment. Attached Figure Description
[0026] Figure 1 The diagram shown is a schematic representation of a multi-site image acquisition system for traditional Chinese medicine visual diagnosis, as described in one embodiment of this application.
[0027] Figure 2 The diagram shown is a flowchart illustrating a method for acquiring images of multiple body parts in traditional Chinese medicine visual diagnosis, as described in one embodiment of this application.
[0028] Figure 3 The diagram shown is a structural schematic of an electronic terminal according to an embodiment of this application. Detailed Implementation
[0029] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0030] The multi-site image acquisition system for TCM visual diagnosis provided in this application can serve as the core imaging module of a TCM four diagnostic instruments, or it can be independently integrated into remote consultation terminals, community health kiosks, and home health management devices. It supports a three-tiered service network of hospitals, communities, and families, and promotes the transformation of TCM visual diagnosis from "experience-based subjective judgment" to "standardized data-driven" approaches. It has a clear industrialization path and social value.
[0031] To facilitate understanding of the embodiments of this application, firstly, in conjunction with Figure 1 Detailed explanation. Figure 1 A schematic diagram of a multi-site image acquisition system 100 for traditional Chinese medicine visual diagnosis is shown in an embodiment of the present invention. The system 100 includes: an imaging cavity 110 and a control module 120; the imaging cavity 110 and the control module 120 are communicatively connected.
[0032] Among them, such as Figure 1 As shown, the imaging cavity 110 is equipped with a light source assembly 1101, an image acquisition assembly 1102, and a sensor assembly 1103, which are used to acquire images of the patient's face, tongue surface, and tongue base.
[0033] In this embodiment, the imaging cavity 110 is a relatively enclosed space designed to isolate it from external ambient light interference. Furthermore, the inner wall of the imaging cavity 110 is made of a matte black material, which effectively reduces specular reflection.
[0034] In this embodiment, the light source assembly 1101 includes: a main ring LED group composed of LED arrays with different color temperatures, and a dedicated LED bead group. The color temperature range of the LED array is 3000K to 6500K. An optical diffuser is also installed in front of all the LED arrays. The optical diffuser can convert point light sources into uniform surface light sources, thereby ensuring uniform and soft illumination. The optical diffuser is made of matte PMMA material with high light transmittance or matte PC material with high light transmittance. The dedicated LED bead group is used to illuminate the underside of the tongue. An optical diffuser is installed in front of the dedicated LED bead group to avoid forming glaring light spots.
[0035] Specifically, multiple LED arrays with different color temperatures are set around the light source component inside the shooting cavity; a set of LED beads is set at the bottom inside the shooting cavity, with its optical axis tilted upward, for example, at an angle of 30° to 45° with the horizontal plane. The purpose of the tilt is to enhance the shadow contrast of the grooves (such as the sublingual veins) in the tongue area.
[0036] In this embodiment, the image acquisition component 1102 includes a low-resolution camera and a high-resolution camera. The low-resolution camera has a resolution of 320×240; the high-resolution camera is a main camera with at least 8 megapixels and supports RAW format output. When using the high-resolution camera, the fixed white balance gain calibrated at the camera's factory settings is used during image acquisition, and automatic white balance is disabled to avoid color cast.
[0037] In this embodiment, the sensor assembly 1103 includes: an ambient light sensor, a color temperature sensor, and an infrared sensor.
[0038] like Figure 1 As shown, the control module 120 includes: a multimodal sensing unit 1201, an adaptive light source adjustment unit 1202, and an image quality assessment unit 1203.
[0039] The multimodal sensing unit 1201 is used to receive multimodal sensor data collected by the sensor component, and trigger the image acquisition component to acquire the patient's current posture image in real time based on the multimodal sensor data. The patient's current posture image is then input into the trained tongue posture light detection model for posture compliance assessment to obtain the posture assessment result of the patient's current posture image.
[0040] It should be noted that the system uses data collected by the infrared sensor in the sensor assembly to determine whether the user is close to the image acquisition assembly. If the user is confirmed to be close, the low-resolution camera in the image acquisition assembly is triggered to capture an image of the patient's current posture. This image is then input into a trained tongue posture detection model to determine whether the patient is in position and whether the current posture is acceptable. This enables the determination of "target presence" and "posture compliance" under conditions of low resolution, complex posture, and partial occlusion.
[0041] In this embodiment, the multimodal sensor data includes, but is not limited to: light intensity collected by the ambient light sensor, LED color temperature collected by the color temperature sensor, and data on whether the user is approaching detected by the infrared sensor.
[0042] The tongue pose detection model in this embodiment is an improvement on the ultra-lightweight object detection model FastestDet, specifically designed for traditional Chinese medicine diagnostic scenarios. The structure of the tongue pose detection model includes a backbone network and an attention layer.
[0043] The backbone network is constructed by fusing ShuffleNetV2 network with depthwise separable convolutions; the attention layer uses the SE attention module, and the fully connected layer in the SE attention module is replaced with a 1×1 convolutional layer; the output layer of the tongue pose detection model is used to output regular bounding boxes (BBoxes), key point heatmaps, and pose compliance confidence scores, and the output regular bounding boxes (BBoxes), key point heatmaps, and pose compliance confidence scores are used as the pose evaluation results of the input image of the model.
[0044] It should be noted that the backbone network adopts a ShuffleNetV2 network fusion structure with depthwise separable convolutions to form a lightweight detection model. Furthermore, attention layers are introduced in the last two layers of the tongue pose detection model. These attention layers use the SE attention module, and the fully connected layers in the SE attention module are replaced with 1×1 convolutional layers to enhance the response in key regions. In this embodiment, the total number of parameters for the tongue pose detection model is only 1.2MB (after INT8 quantization). Experiments have shown that the inference speed on the Snapdragon 665 platform reaches 14ms / frame, significantly outperforming models such as YOLOv7-tiny (>50ms / frame), thus meeting the real-time interaction requirements of self-service terminals.
[0045] In this embodiment, the conventional bounding box is the target recognition region (such as a face, tongue surface, or tongue underside). The key point heatmap includes the coordinates of anatomical key points. For example, the tip and root of the tongue are output as anatomical key points in the tongue surface image; the origins of the left and right sublingual veins are output as anatomical key points in the tongue underside image; and the nose tip and chin are output as anatomical key points in the face image. A tongue posture light inspection model is used to determine whether the tongue is straight and centered and whether the tongue is fully exposed by turning it over to judge whether the patient's current posture is acceptable.
[0046] The confidence level for posture compliance is obtained as follows: after identifying key anatomical points such as the tip and root of the tongue from the image, the geometric relationship between each key anatomical point is calculated and compared with preset judgment conditions. A value is assigned based on the comparison result to obtain the confidence level for posture compliance. The range of the confidence level for posture compliance is pose_valid ∈ [0, 1]. The preset judgment conditions are set according to the requirements of the "Technical Specifications for Traditional Chinese Medicine Inspection (Trial Implementation)" to determine whether the shooting conditions are met (such as the tongue being straight, centered, unobstructed, and fully extended). For example, the preset judgment conditions for the tongue image are: the absolute value of the tilt angle of the line connecting the tip and root of the tongue is less than 15°, and the horizontal coordinate of the tip of the tongue is within ±20% of the center area of the image. If the calculated absolute value of the tilt angle of the line connecting the tip and root of the tongue is less than 15°, and the horizontal coordinate of the tip of the tongue is within ±20% of the center area of the tongue image, then pose_valid ≥ 0.8; otherwise, pose_valid is less than 0.5.
[0047] The multimodal perception unit 1201 is also used to determine whether the patient's current posture is acceptable based on the posture assessment results. If it is not acceptable, it generates corresponding augmented reality guidance information based on the posture assessment results so that the patient can adjust their posture in real time.
[0048] It should be noted that after obtaining the posture assessment results, the results are compared with a preset posture assessment threshold to determine whether the patient's current posture is acceptable. The specific comparison process is as follows: if the confidence level of posture compliance is greater than or equal to the preset posture assessment threshold, the patient's current posture is considered acceptable; if the confidence level is less than the preset posture assessment threshold, the patient's current posture is considered unacceptable. Corresponding augmented reality guidance information is then generated, and the patient is guided to adjust their posture in real time according to this information until the current posture is detected as acceptable before proceeding with subsequent image acquisition and analysis.
[0049] For example, in a tongue image, if the output `pose_valid` ≥ 0.7, meaning the confidence level for posture compliance is greater than or equal to the preset posture assessment threshold, it is judged as "positioning qualified," indicating that the patient's current posture is acceptable. Otherwise, it indicates that the patient's current posture is abnormal. When the patient's current posture is abnormal, it will affect subsequent image acquisition and analysis; therefore, patient posture adjustment is necessary. Abnormal patient posture may be caused by the following reasons: excessive tilt angle (tongue up or down), tongue tip deviated to the left / right (head rotation or tongue deviation), etc. Based on the differences between the geometric relationships between various anatomical key points and the preset judgment conditions, the offset of the anatomical key points is calculated, and augmented reality (AR) guidance information is further generated, such as: "Please move your tongue to the left," "Please look up at the green box," or "Please gently flip the tongue tip with a tongue depressor."
[0050] The adaptive light source adjustment unit 1202 is used to adaptively adjust the light source components based on multimodal sensor data according to the different areas of the patient to be sampled.
[0051] It should be noted that during the acquisition of images of the face, tongue surface, and under the tongue within the imaging cavity, different LED lights and other operations are adjusted at different acquisition stages, meaning that the light source conditions differ for image acquisition of different areas. For example, when acquiring face and tongue surface images, uniform surface lighting is required for both the face and tongue areas to create shadowless, soft, and uniform frontal illumination, avoiding highlights and shadows on the face and meeting requirements for skin tone and tongue texture reproduction. However, when acquiring under the tongue images, oblique side lighting is used to enhance the contrast of the sublingual veins. Oblique light creates a transition of light and dark at microstructures such as the sublingual venous groove and papillae, significantly improving texture contrast and making the veins in the acquired images easier to identify, thus meeting the requirement of "observing veins requires oblique lighting" in traditional Chinese medicine.
[0052] Once the multimodal sensing unit 1201 determines that the patient's current posture is acceptable, it receives multimodal sensor data from the sensor component 1103, namely, real-time light source data from the light source component collected by the ambient light sensor and color temperature sensor. This data is then compared with the light source conditions of the patient's current sampling area. If there is a difference between the real-time light source data and the light source conditions of the current sampling area, the illumination parameters of the light source component are adjusted until the real-time light source data of the light source component is detected to match the light source conditions of the patient's current sampling area, thus completing the adaptive adjustment of the light source component.
[0053] In this embodiment, when the patient's area to be photographed is the face or tongue, the specific process of adaptively adjusting the light source component is as follows: based on the light intensity collected by the ambient light sensor and the LED color temperature collected by the color temperature sensor, the drive current is adjusted through PWM or analog dimming (DIM pin) to precisely control the overall light intensity (e.g., 500 lux) and color temperature of the light source component. This creates shadowless, soft, and uniform frontal illumination, avoiding facial highlights and shadows, and meeting the requirements for reproducing facial skin tone and tongue texture.
[0054] In this embodiment, when the patient's sampling site is the under-tongue area, the specific process of adaptively adjusting the light source component is as follows: when the patient's sampling site is the under-tongue area, the light intensity of the main ring LED group is turned off or significantly reduced, and then the dedicated LED bead group is turned on. The driving current of the dedicated LED bead group is finely adjusted according to the light intensity collected by the ambient light sensor to ensure that the light intensity is within the preset light range (such as 300–600 lux).
[0055] The image quality assessment unit 1203 is used to trigger the image acquisition component to acquire images of the patient's body part in real time according to the posture assessment result after the light source component is adaptively adjusted, obtain the image of the body part to be acquired, input the image of the body part to be acquired into the body part adaptive image quality assessment model for quality assessment, obtain the quality assessment result, determine whether the body part to be acquired should be re-acquired according to the quality assessment result, and output standardized TCM inspection image results when the quality assessment results of the acquired face image, tongue surface image, and tongue underside image all meet the preset quality assessment requirements.
[0056] It should be noted that when the posture assessment result indicates that the patient's current posture is acceptable, and after adaptively adjusting the light source component according to the patient's area to be captured, the high-resolution camera in the image acquisition component is triggered to start image acquisition. During the sequential acquisition of face, tongue, and under-tongue images, quality assessment is performed on the images acquired at different stages. The quality assessment methods and requirements for images of different areas are different. Different quality assessment methods are used for face, tongue, and under-tongue images. Based on preset quality assessment requirements, corresponding assessment networks are constructed and integrated to form an area-adaptive image quality assessment model. When the patient's area to be acquired is captured in real time, it is input into the area-adaptive image quality assessment model, and the corresponding network is found to output the quality assessment result of the area to be acquired.
[0057] Furthermore, the site-adaptive image quality assessment model optimizes the evaluation weights of clarity, color, and texture for different sites according to TCM physical signs. Facial images are used for TCM facial diagnosis, and key visual features include overall skin tone, luster, and moisture. Therefore, color fidelity has the highest weight, brightness uniformity has a lower weight, and clarity has the lowest weight. For example, a sallow, pale, or reddish complexion corresponds to different syndromes such as spleen deficiency, blood deficiency, and excess heat; color deviation can lead to misdiagnosis. If the site-adaptive image quality assessment model's quality assessment result for the current site does not meet the preset quality assessment requirements, image acquisition is repeated until a site image that meets the preset quality assessment requirements is obtained. Then, the next site image is acquired, and the same site-adaptive image quality assessment model is used to evaluate and judge the quality of the next acquired image. This process is repeated for facial images, tongue surface images, and tongue underside images, resulting in the final standardized TCM visual diagnosis image.
[0058] It is important to emphasize that by working collaboratively with the tongue posture detection model and the location-adaptive image quality assessment model, the tongue posture detection model determines "whether it is permissible to take a picture," while the location-adaptive image quality assessment model judges "whether the picture is good." The tongue posture detection model optimizes shooting conditions, avoids background interference, and further improves image assessment accuracy, resulting in high-quality data under correct location, good lighting, and clear imaging conditions. Therefore, the final output image data possesses characteristics such as clear location, compliant posture, consistent lighting, and meeting quality standards. It can be directly used as standardized input for downstream models such as constitution identification and syndrome classification for face analysis, tongue surface analysis, and tongue underside analysis, avoiding model performance degradation due to data noise. Experiments have shown that this system improves the classification accuracy of face images, tongue surface images, and tongue underside images acquired in subsequent analysis processes by 5.8–9.3 percentage points.
[0059] In one embodiment of this application, the control module further includes a data storage unit; the data storage unit is used to store data during the operation of the image acquisition system. It should be noted that the system in this embodiment records data such as the light source parameters, ambient illuminance, posture evaluation results, and quality evaluation results of the light source component through the data storage unit. This meets the explicit requirements of the NMPA's "Key Points for the Review of Artificial Intelligence Medical Devices" regarding "controllable input data and traceable processing," and lays the foundation for subsequent registration and application.
[0060] In one embodiment of this application, the system further includes a human-computer interaction module; the human-computer interaction module is used to display augmented reality guidance information to guide the patient to adjust their posture. The system guides the user to their designated position via a display screen or voice prompts, such as guiding the user to complete the three steps of "looking at the face → extending the tongue → turning the tongue" in sequence, achieving a fully automated process.
[0061] The specific workflow of the multi-site image acquisition system for traditional Chinese medicine visual diagnosis provided in this application is as follows:
[0062] When the infrared sensor in the sensor assembly detects the patient's approach, the low-resolution camera in the image acquisition assembly is activated, and the initialized tongue posture detection model is loaded for real-time inference. The output of the tongue posture detection model includes: target category (such as face, tongue surface, tongue floor) and its bounding box; anatomical key points (such as the position of the tongue tip); and posture compliance confidence score. If the posture compliance confidence score pose_valid ≥ 0.7, it is judged as "position qualified"; otherwise, augmented reality guidance information is generated based on the offset of the anatomical key points, such as: "Please move your tongue to the left", "Please look up and face the green box", or "Please gently flip the tip of your tongue with the tongue depressor".
[0063] Once the patient is properly positioned, the system reads data from the ALS and color temperature sensors built into the imaging cavity, dynamically adjusts the light source components, and activates the site-adaptive image quality assessment model. After imaging, the system uses the site-adaptive image quality assessment model to determine if the image quality meets the preset quality assessment requirements. The image, along with structured metadata such as the acquisition site, LED color temperature, illuminance estimate, sharpness score, positioning score, and timestamp, is saved. The system automatically switches to the next acquisition site until the face, tongue surface, and tongue base are all acquired, ultimately generating standardized TCM visual diagnosis images.
[0064] To facilitate understanding of the tongue pose light detection model of this application, the following specific embodiments are provided for illustration. The training and deployment process of the tongue pose light detection model is as follows:
[0065] Step 1: Data Acquisition and Preprocessing. The acquisition device was a closed-type TCM visual diagnosis image acquisition terminal, equipped with a built-in high-definition camera (resolution 3264×2448), an adjustable color temperature LED light source array (range 4000K–6500K), and an ambient light sensor. A total of 850 subjects were recruited, and three types of images were acquired under standard lighting conditions: frontal facial photographs, naturally extended tongue (tongue surface), and tongue depressor-assisted tongue retraction (tongue floor), resulting in a total of 2550 raw images. The acquired raw images underwent preprocessing as follows: all raw images were RAW domain white balance corrected, uniformly cropped, and downsampled to 640×480 resolution. The preprocessed image data was used as the standard annotation baseline data.
[0066] Step 2: Perform multi-granularity annotation on the preprocessed image data. Each image was independently annotated by two TCM physicians with more than 5 years of clinical experience, using a customized annotation tool, and included the following three layers of information:
[0067] Bounding box (BBox): Labels the target region, with categories including face, tongue_surface, and tongue_underside;
[0068] Key anatomical points: (1) Tongue surface image: mark the tip and root of the tongue; (2) Tongue floor image: mark the origin of the left and right sublingual veins; (3) Face image: mark the nose and chin.
[0069] The pose compliance label (pose_valid) determines whether the shooting conditions are met according to the "Technical Specifications for Traditional Chinese Medicine Inspection (Trial)" (such as tongue straight, centered, unobstructed, and tongue fully turned out). The value is a Boolean value of true or false.
[0070] If the Kappa consistency coefficient of the annotation results of two physicians reaches 0.87, the discrepancy image samples will be arbitrated by a third chief physician to determine the final label, and the annotated image data will be used as training data.
[0071] Step 3: Model Training Phase. The tongue pose detection model is trained using labeled image data.
[0072] Step 4: Model Deployment and Operation. After training, the tongue pose detection model is exported in ONNX (Open Neural Network Exchange) format and further compressed. The final size of the tongue pose detection model is 1.18MB. On an Android 10 device equipped with a Qualcomm Snapdragon 665 chip, the average inference latency is 13.4ms / frame, meeting the requirements for real-time interaction (≥30FPS).
[0073] It is important to emphasize that the multi-site image acquisition system for traditional Chinese medicine (TCM) visual diagnosis provided in this application offers a closed, multi-site, adaptive, and fully automated TCM visual diagnosis image acquisition system, ensuring that the acquired images meet clinical and algorithmic requirements in four aspects: medical validity, color accuracy, texture clarity, and AI usability. It has the following beneficial effects:
[0074] (1) This application realizes fully automatic closed acquisition of three key parts of TCM visual diagnosis: face, tongue surface and tongue base, filling the technical gap of TCM visual diagnosis image acquisition methods that have long lacked standardized acquisition methods.
[0075] (2) This application integrates an ambient light sensor and a color temperature sensor in a closed imaging cavity to construct a closed-loop feedback mechanism, precisely control the light source components, and ensure that the color temperature is stable within the medically effective range.
[0076] (3) This application proposes a multimodal in-situ detection mechanism that integrates infrared proximity sensing and a lightweight tongue posture detection model to accurately determine whether each part is exposed in compliance with regulations and to guide posture compliance, thereby significantly improving the visibility of key features such as sublingual veins in the image.
[0077] (4) This application constructs a site-adaptive image quality assessment model, optimizes the evaluation weights of clarity, color and texture for different sites of TCM physical signs, and improves the accuracy of image quality assessment.
[0078] (5) This application achieves automatic acquisition of high-quality images of multiple parts without human intervention by using a closed imaging cavity structure, multimodal positioning detection and closed-loop guidance mechanism, from user guidance, light adjustment, quality assessment to automatic shooting. It is applicable to various scenarios such as hospital self-service terminals, community health stations and home health equipment.
[0079] In the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect, without limiting their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" do not necessarily imply that they are different.
[0080] It should be noted that, in the embodiments of this application, the words "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0081] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0082] Figure 2 The schematic block diagram of the multi-site image acquisition method for traditional Chinese medicine visual diagnosis provided in this application embodiment is applied to the multi-site image acquisition system for traditional Chinese medicine visual diagnosis. The method includes:
[0083] Step S21: Receive multimodal sensor data collected by the sensor component, and trigger the image acquisition component to collect the patient's current posture image in real time based on the multimodal sensor data. Input the patient's current posture image into the trained tongue posture light detection model to perform posture compliance assessment, so as to obtain the posture assessment result of the patient's current posture image.
[0084] Step S22: Adaptively adjust the light source component based on multimodal sensor data according to the different patient sites to be sampled;
[0085] Step S23: After the light source component is adaptively adjusted, the image acquisition component is triggered in real time to acquire images of the patient's area to be acquired based on the posture assessment results. The images of the areas to be acquired are then input into the adaptive image quality assessment model to perform quality assessment and obtain the quality assessment results. Based on the quality assessment results, it is determined whether the images of the areas to be acquired should be re-acquired. The standardized TCM visual diagnosis image results are output when the quality assessment results of the acquired face image, tongue surface image, and tongue underside image all meet the preset quality assessment requirements.
[0086] It should be understood that the specific process of each module performing the above-mentioned steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.
[0087] It should also be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0088] Figure 3 This is a schematic block diagram of the electronic terminal provided in an embodiment of this application. Figure 3 As shown, the electronic terminal includes at least one processor 301, a memory 302, at least one network interface 303, and a user interface 305. The various components in the device are coupled together via a bus system 304. It is understood that the bus system 304 is used to implement communication between these components. In addition to a data bus, the bus system 304 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 3 The general will label all buses as bus systems.
[0089] The user interface 305 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0090] It is understood that memory 302 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0091] In this embodiment of the invention, the memory 302 is used to store various types of data to support the operation of the electronic terminal 300. Examples of this data include: any executable program for operation on the electronic terminal 300, such as the operating system 3021 and application programs 3022; the operating system 3021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 3022 may contain various applications, such as media players, browsers, etc., for implementing various application services. The method for multi-site image acquisition for traditional Chinese medicine visual diagnosis provided in this embodiment of the invention can be included in the application program 3022.
[0092] The methods disclosed in the above embodiments of the present invention can be applied to processor 301, or implemented by processor 301. Processor 301 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 301 or by instructions in the form of software. The processor 301 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 301 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 301 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0093] In an exemplary embodiment, the electronic terminal 300 may be used to execute the aforementioned method by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs).
[0094] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute the multi-site image acquisition method for traditional Chinese medicine visual diagnosis according to any of the embodiments shown.
[0095] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to execute the multi-site image acquisition method for traditional Chinese medicine visual diagnosis according to any of the embodiments shown.
[0096] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0097] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0098] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0100] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0101] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0102] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs, DVDs), or semiconductor media (e.g., solid-state disks, SSDs, etc.).
[0103] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0105] In summary, the multi-site image acquisition system, method, medium, program product, and terminal for traditional Chinese medicine visual diagnosis provided in this application include: a shooting cavity and a control module; wherein, the shooting cavity is equipped with a light source component, an image acquisition component, and a sensor component, used to acquire images of the patient's face, tongue surface, and tongue base; the control module includes: a multimodal perception unit, used to receive multimodal sensor data acquired by the sensor component, and to trigger the image acquisition component to acquire the patient's current posture image in real time based on the multimodal sensor data, and input the patient's current posture image into a trained tongue posture light detection model for posture compliance assessment, thereby obtaining the posture assessment result of the patient's current posture image; The adaptive light source adjustment unit is used to adaptively adjust the light source component based on multimodal sensor data according to different patient areas to be sampled. The image quality assessment unit is used to trigger the image acquisition component to acquire images of the patient's areas to be sampled in real time according to the posture assessment results after the light source component is adaptively adjusted, obtain images of the areas to be sampled, and input the images of the areas to be sampled into the area adaptive image quality assessment model for quality assessment to obtain quality assessment results. Based on the quality assessment results, it is determined whether the images of the areas to be sampled need to be re-acquired, until the quality assessment results of the acquired face images, tongue surface images, and tongue underside images all meet the preset quality assessment requirements, and then output standardized TCM inspection image results.
[0106] This application achieves fully automated closed-loop image acquisition of three key areas in Traditional Chinese Medicine (TCM) visual diagnosis: face, tongue surface, and tongue base, filling a long-standing technological gap in the lack of standardized acquisition methods for TCM visual diagnosis images. This application integrates an ambient light sensor and a color temperature sensor within the closed imaging cavity, constructing a closed-loop feedback mechanism to precisely control the light source components and ensure the color temperature remains stable within the medically effective range. This application proposes a multimodal in-situ detection mechanism, integrating infrared proximity sensing and a lightweight tongue posture detection model to accurately determine whether each area is compliantly exposed and provide posture guidance, significantly improving the visibility of key features such as sublingual veins in the image. This application constructs a site-adaptive image quality assessment model, optimizing the evaluation weights of clarity, color, and texture for different TCM features in various areas, improving the accuracy of image quality assessment. Through the closed imaging cavity structure, multimodal in-situ detection, and closed-loop guidance mechanism, this application achieves high-quality, multi-site image acquisition without human intervention, from user guidance, lighting adjustment, quality assessment to automatic shooting. It is applicable to various scenarios such as hospital self-service terminals, community health stations, and home health devices. Therefore, this application effectively overcomes the various shortcomings of existing technologies and has high industrial application value.
[0107] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A multi-site image acquisition system for traditional Chinese medicine visual diagnosis, characterized in that, include: Imaging cavity and control module; The imaging cavity is equipped with a light source assembly, an image acquisition assembly, and a sensor assembly, used to acquire images of the patient's face, tongue surface, and tongue base. The control module includes: The multimodal sensing unit is used to receive multimodal sensor data collected by the sensor component, and trigger the image acquisition component to acquire the patient's current posture image in real time based on the multimodal sensor data. The patient's current posture image is then input into the trained tongue posture light detection model for posture compliance assessment to obtain the posture assessment result of the patient's current posture image. An adaptive light source adjustment unit is used to adaptively adjust the light source components based on multimodal sensor data according to different patient sites to be sampled. The image quality assessment unit is used to trigger the image acquisition component to acquire images of the patient's body parts in real time based on the posture assessment results after the light source component has been adaptively adjusted. The image of the body parts to be acquired is then input into the body part adaptive image quality assessment model for quality assessment to obtain the quality assessment results. Based on the quality assessment results, it is determined whether the image of the body parts to be acquired should be re-acquired. The standardized TCM inspection image results are output when the quality assessment results of the acquired face image, tongue image, and tongue underside image all meet the preset quality assessment requirements.
2. The multi-site image acquisition system for traditional Chinese medicine visual diagnosis according to claim 1, characterized in that, The multimodal perception unit is also used to determine whether the patient's current posture is acceptable based on the posture assessment results. If it is not acceptable, it generates corresponding augmented reality guidance information based on the posture assessment results so that the patient can adjust their posture in real time.
3. The multi-site image acquisition system for traditional Chinese medicine visual diagnosis according to claim 1, characterized in that, The sensor assembly includes: an ambient light sensor, a color temperature sensor, and an infrared sensor.
4. The multi-site image acquisition system for traditional Chinese medicine visual diagnosis according to claim 1, characterized in that, The image acquisition components include: a low-resolution camera and a high-resolution camera.
5. The multi-site image acquisition system for traditional Chinese medicine visual diagnosis according to claim 1, characterized in that, The control module further includes a data storage unit; the data storage unit is used to store data during the operation of the image acquisition system.
6. The multi-site image acquisition system for traditional Chinese medicine visual diagnosis according to claim 1, characterized in that, The system also includes a human-computer interaction module; the human-computer interaction module is used to display augmented reality guidance information to guide the patient to adjust their posture.
7. A method for acquiring images of multiple body parts for traditional Chinese medicine visual diagnosis, characterized in that, The method, applied to the multi-site image acquisition system for traditional Chinese medicine visual diagnosis as described in any one of claims 1 to 6, comprises: The system receives multimodal sensor data collected by the sensor component and triggers the image acquisition component to acquire the patient's current posture image in real time based on the multimodal sensor data. The patient's current posture image is then input into the trained tongue posture light detection model to perform posture compliance assessment in order to obtain the posture assessment result of the patient's current posture image. The light source components are adaptively adjusted based on multimodal sensor data, depending on the patient's site of data collection. After the light source component is adaptively adjusted, the image acquisition component is triggered in real time to acquire images of the patient's area to be acquired based on the posture assessment results. The images of the areas to be acquired are then input into the adaptive image quality assessment model for quality assessment to obtain the quality assessment results. Based on the quality assessment results, it is determined whether the images of the areas to be acquired need to be re-acquired. Standardized TCM visual diagnosis image results are output when the quality assessment results of the acquired face image, tongue image, and tongue underside image all meet the preset quality assessment requirements.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-site image acquisition method for traditional Chinese medicine visual diagnosis as described in claim 7.
9. A computer program product, characterized in that, The computer program product includes computer program code, which, when run on a computer, enables the computer to implement the multi-site image acquisition method for traditional Chinese medicine visual diagnosis as described in claim 7.
10. An electronic terminal, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the multi-site image acquisition method for traditional Chinese medicine visual diagnosis as described in claim 7.