A System and Method for Determining Traditional Chinese Medicine Pulse Diagnosis Based on Facial Video Using Multi-Wavelength Light Sources

By using a multi-wavelength light source fusion method, combining visible light and near-infrared light sources, and utilizing a deep neural network model to extract and filter facial video signals, the problem of low accuracy in TCM pulse diagnosis in existing technologies has been solved, achieving a higher accuracy rate in pulse diagnosis.

CN119279529BActive Publication Date: 2026-01-06CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411440026.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2026-01-06
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

In existing technologies, the signal acquisition light source is limited to the visible light band, ignoring the ability of near-infrared light to acquire skin signals. This results in low accuracy in TCM pulse diagnosis, and the chaotic nature of the signal acquisition light source makes it prone to interference, increasing the difficulty of signal analysis.

Method used

A multi-wavelength light source fusion method is adopted, combining visible light and near-infrared light sources. The pulse wave in the facial video is extracted through a deep neural network model, and the signal is preprocessed, filtered and fused. Multi-task learning is used to identify the eight elements of pulse and upgrade the model.

Benefits of technology

It improved the accuracy of TCM pulse diagnosis by 6.9%, improved signal noise issues, and enhanced the extraction of deep-level facial pulse information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119279529B_ABST
    Figure CN119279529B_ABST
Patent Text Reader

Abstract

The present application relates to the fields of medical treatment and information technology, in particular to a system and method for determining traditional Chinese pulse conditions based on multi-wavelength light sources. A video acquisition module acquires facial videos under different wavelengths of light sources, a video preprocessing module crops the videos into images of fixed frame numbers and sizes, a waveform extraction module extracts coarse signals of the facial regions of interest by combining deep neural networks, a waveform fusion unit obtains fusion representation information after filtering by using multiple filters, a pulse condition recognition module determines eight elements of pulse conditions by combining deep neural networks, and the eight elements are predicted simultaneously. The model upgrade module is used to retrain the misjudged data, so as to realize the model closed loop. The present application can obtain filtered multi-wavelength pulse waves, improve the pertinence of facial pulse information extraction, further extract facial deep pulse information, and improve the accuracy of pulse condition diagnosis by combining deep neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of medical and information technology, in particular to a system and method for determining traditional Chinese medicine pulse conditions based on multi-wavelength light source fusion in facial video. BACKGROUND

[0002] Pulse diagnosis is one of the characteristic diagnosis methods in the four diagnostic methods of traditional Chinese medicine, and has important clinical guiding significance. In recent years, with the continuous breakthroughs in modern science and technology, the objectification of pulse diagnosis has made rapid progress.

[0003] Imaging photoplethysmography, namely IPPG, is a transcutaneous light signal acquisition method that can be used to detect various indicators related to arterial pulsation. The IPPG method records the periodic changes in reflected light intensity through a detector, captures the subtle changes in the irradiated part caused by pulse pulsation, and thus realizes pulse diagnosis; the process uses imaging equipment as the detector and records the signal in the form of video, effectively avoiding physical connection with the patient, and has the characteristics of non-contact, easy-to-implement platform, convenient operation, and low cost.

[0004] Among the many IPPG detection sites, the face is highly similar to the "cunkou pulse" of traditional Chinese medicine pulse diagnosis in terms of traditional Chinese medicine theory and holographic restoration, providing the optimal acquisition site for non-contact pulse diagnosis.

[0005] The key technology of the facial video detection method based on IPPG for traditional Chinese medicine pulse is to extract the pulse wave from the facial video, so as to determine the traditional Chinese medicine pulse; the specific method is as follows: according to the interaction between light and human skin and the photoelectric effect such as skin penetration and reflection, the periodic changes of reflected light caused by blood vessel pulsation are captured, and the pulse wave is obtained through noise filtering and other methods. The main steps of the current method include:

[0006] (1) Natural light / conventional light is used to irradiate the face for facial video acquisition;

[0007] (2) After image processing, the red, green, and blue channel signals are extracted, and only the green channel is extracted for signal processing to obtain the pulse wave;

[0008] The current technical solution has the following problems and defects:

[0009] (1) The video acquisition light source is limited to the visible light band, ignoring the ability of near-infrared light to obtain skin signals, which reduces the accuracy of determining traditional Chinese medicine pulse;

[0010] (2) The signal acquisition light source is messy and has complex photoelectric interaction with the skin, which easily produces large interference, increases the difficulty of subsequent signal analysis, and reduces the accuracy of determining traditional Chinese medicine pulse. SUMMARY

[0011] The present application aims to improve the prior art signal acquisition light source is limited to natural light, resulting in insufficient extraction of deep facial pulse information and high signal noise, and the existing low accuracy of traditional Chinese medicine pulse discrimination, and provides a system and method for discriminating traditional Chinese medicine pulse based on multi-wavelength light source fusion.

[0012] To achieve the above-mentioned purpose of the application, the present application provides a method for discriminating traditional Chinese medicine pulse based on multi-wavelength light source fusion of facial video, comprising the following steps:

[0013] S1: The video acquisition module acquires the facial video under the irradiation of different wavelength light sources, which is used for discriminating traditional Chinese medicine pulse;

[0014] S2: The video preprocessing module pre-processes the acquired facial video to generate images with fixed frame number and fixed size;

[0015] S3: The coarse signal extraction unit of the waveform extraction module detects the facial region of interest in the image using a deep neural network model, and extracts coarse signals of different wavelengths through spatial pixel averaging;

[0016] S4: The filtering unit of the waveform extraction module filters the coarse signals of different wavelengths to obtain pure signals under different wavelengths;

[0017] S5: The waveform fusion unit constructs a signal fusion network by processing the pure signals under different wavelengths under different wavelengths to obtain fusion representation information;

[0018] S6: The pulse recognition module performs pulse recognition according to the fusion representation information, combines the deep neural network model to obtain the recognition result of the eight elements of pulse, and combines the multi-task learning method to predict the eight elements of pulse;

[0019] S7: The model upgrading module sends the pulse data with incorrect discrimination into the deep neural network model for retraining.

[0020] Preferably, in step S1, the video acquisition module acquires the facial video under the irradiation of different wavelength light sources, which is used for discriminating traditional Chinese medicine pulse, comprising a multi-wavelength light acquisition unit and a video shooting unit, the multi-wavelength light acquisition unit comprising a visible light and near-infrared light light source acquisition platform, respectively emitting light source signals of corresponding wavelengths, the video shooting unit being a darkroom operation, using a high-definition black and white camera to acquire facial video under the irradiation of different wavelength light sources, and the subject sits or stands in front of the camera, closes his eyes, and avoids wearing glasses, masks and other facial obstructions.

[0021] The video preprocessing module in step S2 uses an audio and video codec to process the face video respectively, and cuts the face video into image frames with fixed frame number and fixed size as the input of the waveform extraction module.

[0022] Preferably, the waveform extraction module comprises a coarse signal extraction unit and a filtering unit. The coarse signal extraction unit detects a face region of interest using a deep neural network, and obtains coarse signals of different wavelengths through spatial pixel averaging. The filtering unit uses multiple filters for signal filtering. In step S3, the coarse signal extraction unit detects a face region of interest using a deep neural network, and obtains coarse signals C(n) of different wavelengths through spatial pixel averaging, as follows:

[0023]

[0024] wherein M(x, y, n) is the pixel intensity of pixel point coordinate (x, y) in the nth picture, X skin is the face region of interest in the nth picture, and K is the number of pixels detected in X skin .

[0025] Preferably, in step S4, the filtering unit of the waveform extraction module filters the coarse signals to obtain pure signals at different wavelengths. The filtering unit uses Bland-Altman analysis, ensemble empirical mode decomposition, and Savitzky-Golay filter to filter the multi-wavelength coarse signals and obtain pure signals at different wavelengths. The Bland-Altman analysis filters abrupt noise, the Savitzky-Golay filter filters high-frequency noise, and the ensemble empirical mode decomposition filters low-frequency noise to obtain the pure signals at different wavelengths.

[0026] Preferably, in step S5, the waveform fusion unit processes the pure signals at different wavelengths to construct a signal fusion network. The transformer and its variants are used to cooperatively represent and encode the pure signals at different wavelengths. Cross-attention mechanism is used for pairwise interaction training. The obtained interaction fusion features are fused again to automatically obtain the fusion representation information of the pulse wave corresponding to the multi-wavelength video.

[0027] Preferably, in step S6, the pulse pattern recognition module obtains the recognition results of the eight elements of pulse pattern based on the extracted features through a deep neural network. Multi-task learning is applied to predict the eight elements of pulse pattern. The eight elements of pulse pattern are predicted through shared bottom features, which are the features common to the eight elements of pulse pattern.

[0028] Step S7 carries out model training, the pulse condition result of each patient is expert certified and marked, and the expert certification result is taken as the correct label of the case, when the model prediction result is different from it, the result is identified as model determination error data, input into the model upgrading module, retrained in the model, and the attention of the model to difficult to classify samples is improved, so as to realize closed loop optimization of the model.

[0029] The application provides a system for judging traditional Chinese pulse conditions based on multi-wavelength light source fusion. Figure 1 As shown in the figure, the system comprises a video acquisition module, a video preprocessing module, a waveform extraction module, a waveform fusion unit, a pulse condition recognition module, a model upgrading module, information is input into the video acquisition module, the multi-wavelength light acquisition unit builds a light source acquisition platform comprising visible light and near-infrared light, the light source is composed of LED circuits of different wavelengths, which respectively emit light source signals of corresponding wavelengths, the video shooting unit is operated in a dark room, and a high-definition black-and-white camera is used to acquire facial videos under irradiation of light sources of different wavelengths, when acquiring, the subject sits or stands in front of the camera, closes eyes and avoids wearing eyes, masks and other facial obstructions, the acquired multi-wavelength facial video signals are input into the video preprocessing module, the video preprocessing module inputs the processed picture frames into the waveform extraction module, the coarse signal extraction unit in the waveform extraction module extracts coarse signals of different wavelengths in the picture frames and sends them into the filtering unit, the filtering unit inputs the extracted pure signals under different wavelengths into the waveform fusion unit, the waveform fusion unit inputs the fusion representation information of pulse waves corresponding to the processed multi-wavelength videos into the pulse condition recognition module, the pulse condition recognition module inputs the judged result into the model upgrading module, through the improvement of the application, compared with the existing method for detecting traditional Chinese pulse conditions based on facial video of IPPG, the diagnostic accuracy rate of eight elements of pulse conditions is improved by about 6.9%.

[0030] Compared with the prior art, the application has the advantages of

[0031] 1. The application provides a system and method for judging traditional Chinese pulse conditions based on multi-wavelength light sources, acquires facial videos under irradiation of visible light and near-infrared light to judge pulse conditions, processes and fuses signals of different wavelengths, obtains fusion representation information of pulse waves corresponding to multi-wavelength videos, further extracts facial deep pulse wave information, and compared with the existing method for detecting traditional Chinese pulse conditions based on facial video of IPPG, the diagnostic accuracy rate of eight elements of pulse conditions is improved by 6.9%.

[0032] 2. The application provides a system and method for facial video discrimination of traditional Chinese pulse conditions based on a multi-wavelength light source, which extracts and filters multi-wavelength signals to obtain pure signals at different wavelengths, thereby improving the problem of high signal noise and reducing pulse diagnosis, and improving the diagnosis accuracy of the eight elements of pulse by 6.9% compared with the existing method of detecting traditional Chinese pulse conditions based on IPPG. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 The structure block diagram of the system and method for facial video discrimination of traditional Chinese pulse conditions based on multi-light source fusion.

[0034] Figure 2 The principle diagram of the waveform fusion unit.

[0035] Figure 3 The schematic diagram of the video shooting module. DETAILED DESCRIPTION

[0036] The application will be further described in detail below in combination with test examples and specific embodiments. However, it should not be understood that the above-mentioned subject matter of the application is limited to the following examples only, and any technology realized based on the content of the application falls within the scope of the application.

[0037] Example 1

[0038] The application provides a system and method for facial video discrimination of traditional Chinese pulse conditions based on a multi-wavelength light source, as shown in the figure, after information input, the information is sequentially sent to the video acquisition module, the video preprocessing module, the waveform extraction module, the waveform fusion unit, the pulse recognition module and the model upgrading module. Figure 1

[0039] Firstly, the video acquisition module. The doctor logs in the system, inputs the subject information (name, gender, age, ID number, etc.), and starts the traditional Chinese pulse recognition. The shooting environment in the multi-wavelength light acquisition unit of the video acquisition module is a dark room, the subject takes a sitting or standing position, keeps the face still, closes the eyes, and places the black and white high-definition camera in front of the human face at the same level. The multi-wavelength lamp is placed above the camera, and the video shooting unit shoots the facial video at each wavelength for about 40s.

[0040] After the video acquisition is completed, the video preprocessing module is entered. The video coding and decoding tool FFmpeg is used to read different wavelength video streams and perform frame division, and the video frame is adjusted to a fixed number of frames and a fixed size of image frame. In this embodiment, the frame number is 256 frames, and the size is 256x256 pixels.

[0041] ​The acquired image frames are sent to the waveform extraction module. The coarse signal extraction unit uses a deep neural network to detect image frames, detects regions of interest on the face, and obtains a coarse signal through spatial pixel averaging. This coarse signal is then sent to the filtering unit for signal filtering using various filters. The deep neural network determines the largest skin area on the face as the region of interest and extracts it using a semantic segmentation model based on the Diffusion deep neural network. The filtering unit uses Bland-Altman analysis, ensemble empirical mode decomposition algorithm, and Savitzky-Golay filter to denoise high-frequency, low-frequency, and abrupt noise.

[0042] The clean signals at different wavelengths extracted by the waveform extraction module are sent to the waveform fusion unit, where they are collaboratively represented and encoded using transformers and their variants, such as... Figure 2 As shown, a cross-attention mechanism is used to train the pure signals at different wavelengths in pairs to construct a signal fusion network, extract the fused pulse wave features, and obtain the fused representation information of the pulse waves corresponding to the multi-wavelength video.

[0043] The pulse identification module uses a deep neural network to derive the identification results of eight pulse elements based on the fused representation information. Multi-task learning is applied to predict these eight elements: pulse location, pulse length, pulse width, pulse rate, tension, fluency, strength, and evenness. This prediction is achieved by sharing underlying features. In this embodiment, the eight pulse elements are divided into eight sub-tasks, each of which is further refined into a binary or tri-classification task depending on the outcome indicator.

[0044] The model upgrade module classifies and identifies the pulse of the subject. If the system identification result is inconsistent with the expert diagnosis result, the data is sent into the model for retraining using the OHME training method, thereby achieving closed-loop optimization of the model.

[0045] Example 2

[0046] The video acquisition module of the present invention includes the multi-wavelength light acquisition unit and the video shooting unit.

[0047] The multi-wavelength optical acquisition unit mainly uses LED light sources and cameras of different wavelengths to emit light source signals of corresponding wavelengths. In this embodiment, three LED light sources with wavelengths of 530nm, 650nm, and 940nm are selected, with an optical output power >20mW and an electrical power ≥20W. This embodiment selects the Medvision MV-GE134 industrial camera as the video acquisition device, with an MV-LD-4-4M-G lens, a focal length of 6mm, and an f / 1.4 aperture.

[0048] like Figure 3As shown, the video shooting unit shoots in a dark room environment. During the acquisition process, the subject sits quietly facing the camera and about 80 cm away horizontally, with eyes closed, avoiding wearing masks, glasses or other face coverings, and avoiding hair covering the face. Multi-wavelength LED lights are placed above the camera, and the LED lights are turned on to shoot facial videos for about 40 seconds at each wavelength.

[0049] After video recording, the video is preprocessed, waveform extracted, and fused before being sent to the pulse identification module for diagnosis of the eight elements of pulse. The study found that compared with the existing method of detecting TCM pulse based on facial video using IPPG, the accuracy of diagnosis of the eight elements of pulse is improved by 6.9%.

[0050] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying traditional Chinese pulse conditions in a facial video based on multi-wavelength light source fusion, characterized in that, It comprises the following steps: S1: a video acquisition module collects face videos under illumination of light sources of different wavelengths, for discriminating traditional Chinese medicine pulse conditions; the video acquisition module comprises a multi-wavelength light acquisition unit and a video shooting unit; The wavelengths of the multi-wavelength light acquisition unit include 530 nm, 650 nm and 940 nm, the light output power is > 20 mw, and the electric power is ≥ 20 w; the video shooting unit is darkroom operation, and a high-definition black-and-white camera is used to collect face videos under illumination of light sources of different wavelengths; during the collection process, the subject sits still in front of the camera and is horizontally 80 cm away from the camera, closes his eyes, avoids wearing face coverings, and avoids hair covering the face, and shoots face videos for 40 s under each wavelength; S2: a video preprocessing module pre-processes the collected face videos to generate images with fixed frame numbers and fixed sizes; the fixed frame number is 256 frames, and the fixed size is 256x256 pixels; S3: a coarse signal extraction unit of a waveform extraction module detects a face region of interest in the image by using a deep neural network model, and extracts coarse signals of different wavelengths through spatial pixel averaging; The deep neural network determines the maximum skin region of the face as the face region of interest, and uses a semantic segmentation model based on Diffusion deep neural network for extraction; The coarse signals of different wavelengths C(n) are obtained through spatial pixel averaging, and the method is as follows: Wherein, M(x, y, n) is the pixel intensity of the pixel point with coordinate (x, y) in the nth frame of picture, X skin is the face region of interest in the nth frame of picture, K is the number of pixels detected in X skin ​ S4: a filtering unit of the waveform extraction module filters the coarse signals of different wavelengths to obtain pure signals under different wavelengths; the filtering unit uses Bland-Altman analysis method, ensemble empirical mode decomposition method and Savitzky-Golay filter to remove high-frequency and mutation noise; S5: a waveform fusion unit processes the pure signals under different wavelengths to construct a signal fusion network, and obtains fusion representation information; the signal fusion network adopts a cross-attention mechanism for pairwise interaction training, and the obtained interaction fusion features are fused again to automatically obtain fusion representation information of pulse waves corresponding to the multi-wavelength video; S6: a pulse condition recognition module recognizes pulse conditions according to the fusion representation information, combines a deep neural network model to obtain recognition results of eight elements of pulse conditions, and combines a multi-task learning method to predict the eight elements of pulse conditions; in step S6, the pulse condition recognition module obtains recognition results of eight elements of pulse conditions through a deep neural network based on the extracted features, applies multi-task learning to predict the eight elements of pulse conditions, and predicts the eight elements of pulse conditions through shared bottom features; the eight elements of pulse conditions are divided into eight sub-tasks, and each sub-task is refined into a binary classification or ternary classification task according to an outcome index; S7: a model upgrading module sends pulse condition data with incorrect discrimination into a deep neural network model for retraining; the model upgrading module classifies and recognizes pulse conditions of the subject, and if the recognition result is inconsistent with the expert diagnosis result, the data is sent into the model for retraining again using an OHME training method, so as to realize closed-loop optimization of the model.

2. The method of claim 1, wherein the method is characterized by, The multi-wavelength light acquisition unit in step S1 includes a visible light and near-infrared light light source acquisition platform that respectively emits light source signals of corresponding wavelengths. 3.The method of claim 1, wherein the method is characterized by, In step S4, the Bland-Altman analysis method filters mutation noise, the Savitzky-Golay filter filters high-frequency noise, and the ensemble empirical mode decomposition method filters low-frequency noise.

4. The method of claim 1, wherein the method is characterized by, The waveform fusion unit in step S5 uses a transformer and its variants to cooperatively characterize and encode signals of different wavelengths.

5. A system for identifying traditional Chinese pulse conditions in a face video based on multi-wavelength light source fusion, characterized in that, The system comprises a video acquisition module, a video preprocessing module, a waveform extraction module, a waveform fusion unit, a pulse condition identification module, and a model upgrading module. The video acquisition module comprises a multi-wavelength light acquisition unit and a video shooting unit. The multi-wavelength light acquisition unit uses multi-wavelength light sources to irradiate the human face. The video shooting unit acquires facial videos under irradiation of different wavelengths of light sources as the input of the video preprocessing module. The wavelengths of the multi-wavelength light acquisition unit include 530 nm, 650 nm, and 940 nm, and the light output power is > 20 mw, and the electric power is ≥ 20 w. The video shooting unit is operated in a dark room, and a high-definition black-and-white camera is used to acquire facial videos under irradiation of different wavelengths of light sources. During the acquisition process, the subject sits still in front of the camera and is horizontally 80 cm away from the camera, closes his eyes, avoids wearing face coverings, and avoids having hair covering the face. The facial videos under each wavelength are shot for about 40 s. The video preprocessing module uses an audio and video codec to process the facial videos acquired by the video acquisition module. The videos are cropped into image frames of a fixed frame number and a fixed size as the input of the waveform extraction module. The fixed frame number is 256 frames, and the fixed size is 256 x 256 pixels. The waveform extraction module comprises a coarse signal extraction unit and a filtering unit. The coarse signal extraction unit uses a deep neural network to detect a facial region of interest, spatially averages the image frames output by the video preprocessing module, and obtains coarse signals of different wavelengths. The deep neural network determines the maximum skin region of the face as the facial region of interest and extracts it using a semantic segmentation model based on a Diffusion deep neural network. The coarse signals C(n) of different wavelengths are obtained by spatially averaging the image frames. Wherein, M(x, y, n) is the pixel intensity of the pixel point coordinate (x, y) in the nth frame of picture, X skin is the face region of interest in the nth frame of picture, K is the number of pixels detected in X skin The filter unit uses a plurality of filters to filter the coarse signal to obtain a pure signal at different wavelengths as the input of the waveform fusion unit, and the filter unit uses Bland-Altman analysis method, ensemble empirical mode decomposition method and Savitzky-Golay filter to remove high-frequency and mutation noise. The waveform fusion unit uses a transformer and its variants to cooperatively characterize and encode. The cross-attention mechanism is used to train and construct a signal fusion network by interacting with each other, to obtain fusion characterization information of the pulse wave corresponding to the multi-wavelength video as the input of the pulse condition identification module. The signal fusion network is trained and constructed by interacting with each other using the cross-attention mechanism, and the obtained interaction fusion features are fused again to automatically obtain the fusion characterization information of the pulse wave corresponding to the multi-wavelength video. The pulse identification module obtains an identification result of eight elements of pulse based on the fusion feature information through a deep neural network, applies multi-task learning to share bottom features, and simultaneously performs eight-element prediction, and takes the identification result and the prediction result as input of the model upgrading module; the eight elements of pulse are divided into eight sub-tasks, and each sub-task is refined into a two-classification or three-classification task according to an outcome index; The model upgrading module sends model determination error data into model retraining, improves the attention of the model to difficult-to-classify samples, and thus realizes closed-loop optimization of the model; the model upgrading module performs pulse classification and identification on a subject, and if the identification result is inconsistent with expert diagnosis result, the data is sent into the model for retraining in an OHME training mode.

6. The system for identifying the pulse condition of traditional Chinese medicine in facial video based on multi-wavelength light source fusion according to claim 5, characterized in that, The waveform fusion unit includes a transformer and its variants and a cross-attention mechanism, inputs the pure signals under different wavelengths, and outputs fusion feature information of pulse waves corresponding to multi-wavelength videos.

7. The system for identifying the pulse condition of traditional Chinese medicine in facial video based on multi-wavelength light source fusion according to claim 5, characterized in that, The pulse identification module includes a deep neural network and multi-task learning, inputs the fusion feature information, and outputs a discrimination result of pulse and a prediction of eight elements of pulse.

Citation Information

Patent Citations

  • Multi-spectrum facial physiological signal acquisition system and method

    CN111419201A

  • Traditional Chinese medicine pulse judgment system, method and device based on face video

    CN116109818A