Automatic system for traditional chinese medicine diagnosis by sound based on intelligent voice technology
By designing a TCM physiological model and using speech signal analysis technology, the shortcomings in the application of low-frequency speech components in TCM diagnosis have been addressed, thereby improving the personalization and accuracy of TCM diagnosis and providing individual health monitoring and intervention plans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 芦煜
- Filing Date
- 2020-07-24
- Publication Date
- 2026-04-17
AI Technical Summary
Modern medicine has not fully explored the diagnostic value of low-frequency components in speech signals in health diagnosis, especially in traditional Chinese medicine diagnosis, where there is a lack of effective technical means to analyze and utilize this information.
A TCM physiological model of speech signals was designed. Specific low-frequency signals in human speech samples were obtained through frequency domain extension. Combined with TCM theory, the physical attributes of the organs in speech were determined and data output was achieved. This included automatic speech recognition, feature factor generation and analysis, auxiliary information support, and a TCM comprehensive identification module, which were integrated to form a TCM auscultation and diagnosis system.
It can support the TCM diagnostic process, provide individual organ function assessment and health status analysis, improve the accuracy and comprehensiveness of diagnosis, and provide personalized health monitoring and intervention recommendations.
Smart Images

Figure CN112002342B_ABST
Abstract
Description
Technical Field
[0001] This patent covers the following fields: Traditional Chinese Medicine, automated design of TCM diagnostic methods, signal analysis and information integration, and intelligent recognition. Background Technology
[0002] The resonant frequency of atoms is approximately 10¹⁵ Hz, while at the molecular level it decreases to approximately 10⁹ Hz. As volume increases, the frequency of physical vibration decreases. Body tissues, organs, and cells all maintain certain vibrational frequencies, with the frequency range of internal organs primarily in the infrasound range (1-20 Hz). Although the vocal cords are the primary vibrational site in human speech, it is actually a process of the entire body resonating according to its physical properties. This is why human speech possesses distinctive characteristics, allowing for individual identification. However, the health implications of this resonant characteristic have not yet been fully explored by modern medicine.
[0003] The modern medical applications of speech information are primarily within the audible range (20-8000Hz), and it is currently capable of being used for the functional evaluation of the speech organs. For example, it has been developed for the diagnostic assessment of vocal cord hypoplasia in children, postoperative rehabilitation evaluation after tracheotomy, and pre- and post-treatment sound and auditory perception analysis in patients with speech disorders. Using artificial intelligence, pattern recognition technology can distinguish emotions and psychological states, supporting the diagnosis of mental illnesses such as depression, social phobia, anxiety disorders, and bipolar disorder, and predicting whether a sample exhibits manic behavior or suicidal tendencies. However, analysis in the low-frequency range is limited. Firstly, the vibrational energy in this frequency range is low, making detection difficult. Secondly, the strong penetrating power of infrasound makes it difficult to identify which infrasound components are carried in speech. Furthermore, modern medicine has not yet explored the diagnostic value of this information, thus leading to its unintentional neglect by modern science.
[0004] Traditional Chinese medicine (TCM) itself does not have the concept of "resonance," but the low-frequency characteristics experienced during TCM diagnosis are expressed in many diagnostic methods and described as a subjective experience that can be intuitively perceived. Take pulse diagnosis as an example: during pulse diagnosis, one often experiences a strange "feeling," like a higher-frequency vibration superimposed on the pulse and heart rate. This vibration falls within the low-frequency range and can be detected by a finger pressure sensor. By analyzing the characteristics of this fluctuation, i.e., the frequency distribution, the function of the internal organs can be identified.
[0005] Human speech, heartbeat, and respiration all originate from the human body, thus carrying information about the body during these processes. If pulse diagnosis information possesses low-frequency attributes that support diagnosis, then the low-frequency components of speech signals can similarly support diagnosis. Therefore, by designing speech features in this frequency band, we can support the TCM diagnostic process from another perspective and discover the body's frequency attributes. This information can be passed down as diagnostic experience. Mining health-related features from this information can support digital diagnosis. Summary of the Invention
[0006] This invention patent is titled "Traditional Chinese Medicine Physiological Model of Speech Signals and Auscultation and Diagnosis System," abbreviated as Auscultation and Diagnosis (or Voicediagnosis in English), and the design version is V1.0. The system's design principle is based on the physical resonance effect of the human body. It uses the original speech signal as the input segment and, through frequency domain expansion (1-20000Hz), acquires specific low-frequency signals from human speech samples to determine the physical attributes of the internal organs and output data. The system consists of a speech signal input terminal, a speech recognition (VAD) module, a speech feature factor generation and analysis module, an auxiliary information support module, and a comprehensive TCM identification module. The device can derive diagnostic conclusions based on speech related to the Eight Principles, the Three Jiaos, Qi, Blood, Body Fluids, and the internal organs.
[0007] The overall workflow is as follows: (1) Input voice and personal information. The voice is converted into a raw audio signal through the information input terminal. The input information is converted into corresponding indicator data of TCM theory through the auxiliary information support module. (2) The raw audio signal is distinguished from the noise segment by the speech automatic recognition VAD module. The speech feature factor generation and analysis module analyzes the speech time domain, frequency domain and composite domain features to form speech feature factors. (3) The speech feature factors are converted into TCM diagnostic related feature indicators by the TCM comprehensive identification module to form quantitative diagnostic data including the five tones and five elements, viscera, triple burner, eight principles, qi, blood and body fluid differentiation. After information integration, a TCM auscultation diagnosis comprehensive diagnostic conclusion is generated. (4) The TCM comprehensive identification module calls personal information and compares it with the voice conclusion to form a reliability report. A matching intervention plan is formed in the high reliability conclusion.
[0008] 1. Design of a wideband speech model method
[0009] The audio signal input hardware of the auscultation system is designed with a wide frequency domain. The corresponding physiological pronunciation model is an innovative resonant speech model, which includes four components: excitation model U(z), vocal tract model G(z), radiation model R(z), and resonance model R(z). The low-frequency (1-20Hz) component of the audio superposition is attributed to the resonance model R(z).
[0010] Traditional speech models are cognitions achieved through physiological analysis: speech signals are decomposed into three components: an excitation model, a vocal tract model, and a radiation model.
[0011] S(z)=U(z)·G(z)·R(z)
[0012] Excitation model U(z): When producing a voiced sound, the continuous opening and closing of the vocal cords generates intermittent pulse waves. This pulse wave is similar to a slanted triangular pulse train. When producing an unvoiced sound, it can be equivalent to random white noise.
[0013] The vocal tract model G(z): There are currently two viewpoints on the mathematical model of the vocal tract. One is to regard the vocal tract as a system formed by multiple tubes with different cross-sectional areas connected in series, namely the "vocal tube model". The other is to regard the vocal tract as a resonant cavity, namely the "resonant peak model".
[0014] Radiation model R(z): The radiation model characterizes the radiation effect of the mouth and lips and the diffraction effect of the round head.
[0015] By extending the understanding of human speech function and pronunciation information through Traditional Chinese Medicine (TCM) principles, this system's design innovatively proposes a speech resonance model based on TCM concepts. This model updates the traditional digital model of speech signals to four components: an excitation model, a vocal tract model, a radiation model, and a resonance cavity model based on abdominal resonance. The resonance cavity is divided into audible and infrasound bands according to its frequency range. The formula is as follows:
[0016] S(z)=U(z)·G(z)·R(z)·A(z)
[0017] The first three terms are the same as in the previous formula. The fourth term, the abdominal resonance model A(z), represents the vibrational enhancement of the relevant frequencies of different organs and their harmonic effects under the action of a speech excitation source. Using this innovative model, infrasound signals can be extracted from speech signals to complete the information extraction of the visceral resonance component.
[0018] 2. Speech Feature Factor Conversion
[0019] After acquiring the speech signal, the speech components are identified by the Automatic Voice Analyzer (VAD) module. Further, the speech feature factor generation and analysis module extracts diagnostically valuable speech factors, categorized into overall information features and local features. The overall analysis module acquires the following information: Mel-frequency spectral feature map, Mel-frequency cepstral coefficients (MFCC), dominant frequency (DF), average frequency, perceptual spectral center (PS), power spectral characteristics (FBEs), and signal-to-noise ratio (SNR) at 100.5Hz, 215.7Hz, 345.5Hz, 492.0Hz, 655.3Hz, 839.6Hz, 1048.9Hz, 1285.5Hz, and 1551.4Hz. The speech energy curves are segmented at endpoint frequencies of 1850.8Hz, 2190.0Hz, 2571.1Hz, 3000.3Hz, 3483.9Hz, 4030.4Hz, 4645.9Hz, 5338.9Hz, 6122Hz, 7005.5Hz, and 8000.0Hz; the highest five formants, envelope characteristics, and frequency domain characteristics of speech energy temporal fluctuations are analyzed; the inter-frame frequency distribution and frame shift temporal fluctuation characteristics of the speech signal are calculated synchronously.
[0020] The local feature analysis module employs VAD technology to accurately annotate speech intervals, enabling single-character feature analysis. The acquired information includes: duration of single-character phonemes in Chinese character pronunciation, pitch, timbre, and intensity; zero-crossing rate (ZCR) of the pronunciation segment; power spectral density, flatness, slope, and kurtosis characteristics of the spectral curve; mean, median, standard deviation, discrete intervals, inter-character spacing, fluctuation trends, and minimum and maximum values of single-character components in the time-domain curve; and the acquisition of a speech power model and rhythm model for the speaker's Chinese character pronunciation, completing speech rate and intonation complexity analysis. Based on time-frequency cross-linking analysis technology, it completes the classification of emotions and moods, as well as tone and voice recognition, along with environmental noise component and low-frequency feature analysis, obtaining a comprehensive and multi-perspective dataset of key speech feature factors.
[0021] 3. Traditional Chinese Medicine Diagnostic Methods
[0022] The speech factor indicators output by the speech feature factor generation and analysis module can be intelligently diagnosed by the TCM comprehensive identification module. The TCM comprehensive identification module incorporates the TCM theory of five tones and five elements, viscera and bowels, triple burner, eight principles, qi, blood and body fluids, and can realize speech diagnosis based on the special characteristics of speech signals. The methods are as follows: (1) Triple burner diagnosis method: by analyzing the speech patterns of laryngeal, chest and abdominal speech, the main speech methods used by the sample in normal speech are evaluated and used to judge the distribution of triple burner qi in the speech process; (2) Five tone viscera and bowels diagnosis method: by using the five tone features of jiao, zhi, gong, shang and yu within the audible threshold, the individual's five tones are identified, and the individual bias based on the five tones is completed to form the five tone viscera and bowels diagnosis system; (3) Qi, blood and body fluids diagnosis method: the 'noise' component in speech is separated and evaluated. (3) Using the body fluid status, combined with the volume of speech, frequency formant and frequency domain distribution, the sample’s qi and blood deficiency and excess and body fluid distribution status are judged; (4) Using the eight principles of dialectics, the stress position of single word pronunciation is analyzed to obtain the “cold and hot exterior and interior” characteristics of the sample, the five-tone intonation is analyzed to judge the “rising and falling” characteristics, the speech energy distribution is analyzed to judge the “deficiency and excess of middle qi” of the sample; (5) Extract the pronunciation speed of Chinese characters to obtain the “rapid and slow” characteristics of emotions; (6) Calculate the complexity of speech temporal frequency to analyze the richness of the expression of emotions in the sample language.
[0023] 4. Information support for the auxiliary information support module
[0024] To enhance the system's functionality, the device incorporates a wide range of information understanding methods, integrated into an auxiliary information support module. This module integrates relatively less common methods from Traditional Chinese Medicine, such as the Five Elements and Six Qi theory and the calculation of Heavenly Stems and Earthly Branches, as well as disease prediction and information analysis methods for emotional and health inquiries. It also includes voice diagnosis to analyze the accuracy of predictions.
[0025] Taking the principle of Heavenly Stems and Earthly Branches calculation as an example, the Heavenly Stems and Earthly Branches module of the auxiliary information support module of the sound diagnosis system utilizes the traditional Chinese medicine principle of Heavenly Stems and Earthly Branches for timekeeping. It converts the birth date into Heavenly Stems and Earthly Branches factors and uses the principles of interpretation and understanding of these factors in traditional Chinese medicine to summarize their ascending / descending, cold / heat, and yin / yang attributes. This yields the ascending / descending attributes, yin / yang balance, and cold / heat distribution of the sample based on the characteristics of Heavenly Stems and Earthly Branches. Finally, it derives the cold / heat factors of the five internal organs.
[0026] The method is as follows: At the user input terminal, users enter their birthdate information (10 digits), including 4 for the year, 2 for the month, 2 for the day, and 2 for the time. Using the traditional Chinese medicine (TCM) sexagenary cycle (stem-branch) system, the time is converted into 8 symbols. Utilizing the TCM interpretations of these symbols, they are broken down into attributes of cold / heat, yin / yang, five tones and five elements, entry / exit, and rise / fall. The descriptions of these attributes are based on... Figure 3 As shown, the quantization calculation method is as follows:
[0027] In the lifting attribute, lifting force (F) udThis mainly attempts to correspond to the rise and fall of individual Qi, and its literal meaning expresses the factors representing upward and downward trends in the attributes of Heavenly Stems and Earthly Branches. The calculation method for its comprehensive effect is as follows:
[0028]
[0029] n represents the total number of sexagenary cycle symbols, θ i This indicates the sexagenary cycle corresponding to each serial number. Figure 3 The position of θ in i Angle. The same applies below.
[0030] In the entry and exit attributes, the entry and exit force (F oi This primarily attempts to correspond to the opening and closing of an individual's spiritual consciousness; its literal meaning expresses the factors with discrete and aggregate attributes in the Heavenly Stems and Earthly Branches, and the calculation method for their combined effect is as follows:
[0031]
[0032] In the attributes of cold and heat, the cold and heat force (F) ch This primarily attempts to correspond to individual hot and cold elements. Its literal meaning expresses the factors with hot and cold attributes within the Heavenly Stems and Earthly Branches, and its total value is calculated as follows:
[0033]
[0034] In the Yin-Yang attribute, Yin-Yang power (F) yy This study primarily attempts to correlate the overall functional hyperactivity of an individual with the richness of their spiritual expression. It expresses factors with both Yang and Yin attributes within the Heavenly Stems and Earthly Branches (stems and branches) attributes; those in odd-numbered positions are considered Yang, and those in even-numbered positions are considered Yin. The quantitative calculation method is as follows:
[0035]
[0036] The first term is the sum of the cold and heat values of all Yang components (pre-designed), and the second term is the sum of the cold and heat values of all Yin components.
[0037] These attributes are used as attribute features derived from the sample's birth date.
[0038] During the large-sample data calculation process, by comparing the differences in this feature, it was found that the distribution of low-frequency characteristics expressed in the speech of different people showed a certain correlation with the three calculation conclusions mentioned above. Although the correlation was not high and the values varied greatly between individuals, it suggests that individuals born at different times may have unique influences in their corresponding speech information.
[0039] This design can support the discovery and exploration of relevant Traditional Chinese Medicine (TCM) principles and patterns. Firstly, supported by this feature and combined with TCM theory, the relationship between an individual's vocal "constitution" characteristics and birth time can be explored. This design uses the physical concept of resonant fragility to express the organ resonance characteristics of speech. The stronger the resonance, the more stimulation the area receives, indicating that its condition is more readily perceived and functionally manifested, and its "yang" attribute is more likely to be inferred. If the proportion of the organ system corresponding to the infrasound of speech is found to be consistent with the yin-yang attribute distribution ratio calculated by the Heavenly Stems and Earthly Branches, and this consistency maintains a certain high probability, then the organ in the sample can be considered equivalent to the "function" in the body's constitution-function relationship. That is, it has a synergistic relationship with multiple physical attributes, and the bias of this part will be more significantly expressed.
[0040] This correspondence can be expanded quantitatively by further broadening the data scope. By combining its metabolic rate, it can be used to assess cold and heat according to the Heavenly Stems and Earthly Branches. By combining the emotional attributes corresponding to speech analysis, it can express the coming and going of the spirit. This corresponds to the matching of the coming and going of the spirit attributes of the Heavenly Stems and Earthly Branches. However, regarding the infrasound component, its positive expression in the speech pronunciation process can not only use traditional Chinese medicine theory to deduce the corresponding organ characteristics, but also express its correspondence.
[0041] The auxiliary information support module within the auditory diagnosis system includes an optimized questionnaire for assessing emotional characteristics. Its design references authoritative psychological assessment questionnaires, including the 92-item MBTI emotional characteristic assessment questionnaire, depression questionnaires, and anxiety assessment questionnaires. Content with TCM diagnostic value was extracted and integrated into a new questionnaire assessment system. The questionnaire interpretation process employed a depersonalization design to prevent morally constrained responses, refined repetitive questions, and re-summarized the TCM emotional characteristics corresponding to each question. Alongside the original individual assessment conclusions, it includes interpretations of the five-element TCM emotional characteristics. The final optimized system comprises 31 questions, capable of outputting quantifiable five-element emotional attributes. Correlation comparisons are performed with the five-element emotional conclusions output from the audio samples to calibrate the accuracy of emotional prediction.
[0042] The auxiliary information support module of the voice diagnosis system incorporates standardized methods for expressing TCM consultation information. It establishes a self-assessment-based TCM diagnostic approach using 30 optimized TCM consultation questionnaires. Within this module, quantitative information output for Five Elements Differentiation, Eight Principles Differentiation, and Qi-Blood-Body Fluid Differentiation can be achieved through consultation. The consultation questions are designed for precise questioning, quantitative answers, and recording the time spent on each answer, evaluating the reliability of the answer, and obtaining conclusions on the Eight Principles, Qi-Blood-Body Fluid, and Five Elements of the sample's health characteristics. The accuracy of the diagnostic conclusions obtained from voice analysis can be compared and judged.
[0043] 5. Comprehensive Traditional Chinese Medicine Identification Module
[0044] The TCM comprehensive identification module is designed in a distributed manner, realizing multiple TCM diagnostic theories and data analysis strategies for parallel processing and internal competition to improve data accuracy. By integrating voice samples with the transformed data obtained from personal basic information input, the TCM comprehensive identification module can expand its information integration function. With the support of independent emotional and health consultation data conclusions, it can achieve an accurate assessment of the recognition accuracy and prediction quality of voice output information.
[0045] This module has an external interface, allowing direct incorporation of embedded ECG data obtained from external measurements into the data structure, and extraction and merging of relevant data from the ECG signals for analysis. The signal analysis process for merging and analysis is as follows: Speech is converted into raw audio digital signals via the information input terminal. Low-frequency features of 1-20Hz are extracted from the signal to obtain heart rate data and the coefficient of variation of heart rate, estimating the range of heart rate variation. Through the frequency doubling effect, a high-risk frequency band for resonance in the low-frequency range is formed. Then, the vibration characteristics, amplitude, and fluctuation correlation of speech in this frequency band are evaluated to assess the cardiac resonance participation. Similarly, the analysis further utilizes the resonance features of internal organs.
[0046] Due to the dynamic nature of the heart, its intrinsic frequency is difficult to measure practically. However, its intrinsic resonant frequency is generally considered to be related to the heart's pumping rhythm. Assuming a heart rate (HR) of 70 beats / min, and the SDNN index in the coefficient of variation (HRV) calculation (i.e., the average of the standard deviation of the RR interval every 5 minutes (288 values) in a 24-hour ECG recording) is 50 ms, this indicates that the cardiac range is mostly between 63.5 and 76.5 beats / min, or 1.06-1.275 Hz. Then, the frequency range of twice this range is calculated, and by multiplying it by a prediction coefficient μ, the influence of this frequency on the heart's original intrinsic frequency is obtained. In this design, μ = 0.5t, where t is the octave. A higher octave corresponds to a smaller direct impact on resonance, but also a larger range of influence. Within the original intrinsic frequency range, the resonance characteristics decrease quadratically with distance from the center point. This forms an estimate of the comprehensive effect of resonance on the internal organs.
[0047] When this frequency is detected in speech, it can be considered the sum of all organs in the body whose frequencies are consistent with the heart's intrinsic frequency. In Traditional Chinese Medicine (TCM) thinking, structures with physical properties equivalent to the heart can also support the heart's pumping function using their own physical properties, and are thus a component of the functional heart. Therefore, we name this the cardiac-dominated resonance system, equivalent to the heart's resonance supporting speech frequencies.
[0048] The human body has vertical resonant peaks in the 3-8Hz frequency band, corresponding to internal organs such as the liver and spleen. The natural frequencies of most internal organs are between 3-18Hz; around 2Hz, the body can be considered as a whole. Extending this method to individual organs, based on their fundamental resonant frequencies, we can analyze and distribute the resonant frequency response of speech to each organ. Using the classification method derived from this model, we design a calculation method to decompose the infrasound component in a speech segment into the intensity of resonance of each organ.
[0049] When the infrasound component of speech is closely related to a particular organ, we can equivalently assume that speech plays a greater role in that organ, and the speech expression process has a greater impact on that organ. This physical characteristic can be further derived from the TCM theory of auditory diagnosis—specifically, the theory of organ function—to express the corresponding organ characteristics. The associated content is predicted to be related to emotions. Furthermore, a design utilizes the low-frequency resonance characteristics of speech to identify organs, enabling an analysis of individual organ imbalances based on the Five Elements theory. When a weak or excessively large low-frequency resonance range is found, the health of that organ can be closely monitored, and its physiological characteristics can be deduced based on the TCM Five Elements theory.
[0050] This external integration capability provides a richer source of information for the diagnosis of viscera through voice, enhances the comprehensiveness and stability of TCM auscultation diagnosis information, and enables personalized health monitoring services by embedding it into mobile terminals.
[0051] 6. Conclusion Output:
[0052] The data output of the voice diagnosis system is divided into two parts. The first part is the basic physical attribute characteristics, which are the direct feature components of speech. These serve as the basic features for backend data analysis and can be accessed and viewed, but are not displayed on the front end. The second part is the TCM syndrome differentiation results, which are displayed as five-tone-five-element data based on speech analysis; articulation location-triple jiao syndrome differentiation data; infrasound-organ fragility data; and intonation-organ syndrome differentiation data. After weighted integration, the final main syndrome identification conclusion is obtained.
[0053] 7. How to use
[0054] This patented auditory diagnosis system has two usage methods: a standardized diagnostic method and a free analysis method. The former involves pronouncing words according to a specified method, completing data collection and analysis, and applying it to basic sound feature analysis to obtain basic steady-state feature analysis without emotional interference, including objective emotional characteristics, personality tone, five-element imbalance, eight principles, and syndrome-related health status such as qi, blood, and body fluids. The latter involves any form of speech signal expressed in various natural contexts in a real-world environment, used to evaluate transient emotions and subjective syndrome characteristics. Based on the combination of steady-state and transient features formed by these two modes, the system can track and monitor the speaker's emotional and health status in real time.
[0055] In addition, the extended active intervention method developed in this patent includes the following:
[0056] (1) Low-frequency feedback intervention
[0057] The excessively low infrasound component of speech was discovered, and an active infrasound emission device was used to amplify this frequency and use it as an output source to feed back to the human body. Under the resonant intervention of this frequency, the duration and intensity of the corresponding organ resonance were increased, and the organ response rate was improved in terms of physical vibration characteristics, ultimately achieving targeted organ conditioning based on infrasound low frequency.
[0058] (2) intonation training
[0059] Preliminary studies have found higher vocal frequencies in patients with depression and other mental health conditions, suggesting that pitch-down vocal training may antagonize depressive tendencies. Empirical observations show that in verbal communication, a higher pitch (frequency) relative to an individual's fundamental tone often represents a sense of respect and deference. When individuals are speaking to superiors, elders, teachers, or people they admire, they often use a higher pitch, with the vocal power source shifting upwards, possibly through chest or throat articulation. Conversely, when speaking to subordinates, younger people, or students, the pitch tends to be lower and more subdued, with the vocal power source shifting downwards, tending towards diaphragmatic articulation. The higher pitch in the speech of patients with depression expresses self-deprecation, an internal tension and upward-stretching state unrelated to the recipient, which may be due to habitual subjective will, external factors, or induced by internal emotional abnormalities. However, such vocal pitch can also reflect and reinforce depressive tendencies. By training individuals with depression or those at high risk of depression to speak with a lower intonation, it is possible to counteract their habitual rising intonation during information expression. This negative feedback from their own speech patterns can help balance their emotional biases and improve their depressive state. Therefore, through speech analysis, individuals can receive continuous prompts and train their lower intonation during daily speech, potentially reducing factors that trigger depression.
[0060] The system combines the subjective features of voice samples with objective information to form a voice diagnosis conclusion supported by dual information, which further enhances the comprehensiveness and stability of TCM voice diagnosis. Based on the conclusions obtained by this model, it is possible to form immediate conditioning feedback in the form of diet therapy, health preservation methods, voice, sound frequency, and music, so as to realize immediate non-drug conditioning, frequency intervention, and energy guidance for health.
[0061] The summary of the application methods of this auscultation-based diagnosis system reveals that it not only enables intelligent prediction based on voice but also provides a self-training evaluation tool for health monitoring. In terms of data output strategy, it adopts a comprehensive reference data output scheme, satisfying users' application of existing knowledge. This feature also allows the auscultation-based diagnosis system to serve as a tool for quantitative research in traditional Chinese medicine theory, providing researchers and scientific workers with a massive and accurate data source, facilitating the discovery of more correspondences and regularities in future experiments. The following specific implementation examples illustrate several application scenarios and demonstrate its functionality. Attached Figure Description
[0062] Figure 1 Overall signal analysis process
[0063] In the diagram: 1. Human-Machine Interface (HMI) information input interface, i.e., voice signal input terminal; 2. Automatic Voice Recognition (VAD) module; 3. Voice feature factor generation and analysis module; 4. Auxiliary information support module; 5. Traditional Chinese Medicine comprehensive identification module. A diagnostic conclusion report is generated after these five steps are completed.
[0064] Figure 2 Voice-based ECG comprehensive assessment and TCM syndrome differentiation plan
[0065] In this diagram, S1 represents the electrocardiogram (ECG) signal, and S2 represents the speech signal. RH indicates the obtained resonance frequency range; R-Vs represents the amplitude characteristics and energy distribution corresponding to the infrasonic frequency range of the speech. TCM represents the Traditional Chinese Medicine (TCM) knowledge database, and Co represents the conclusion. Through relevant interpretation methods, the diagnostic characteristics of the heart based on resonance effects can be obtained. When the speech resonance reaches a certain level within the heart's resonance frequency range and exceeds that of other organs, this person can be considered to have the element of fire in their Five Elements theory. If their birth date and the Heavenly Stem and Earthly Branch of their birth also have these characteristics, they can be identified as a person with the fire element.
[0066] Figure 3 Based on the theory of viscera and bowels in Traditional Chinese Medicine, the five viscera plus the three jiaos are used for dialectical differentiation of speech resonance characteristics.
[0067] In this model, Li, H, S, Lu, K, and A represent the liver, heart, spleen, lung, and kidney systems according to Traditional Chinese Medicine (TCM), respectively. The entire abdominal cavity and trunk are considered as a whole, representing the Sanjiao (Triple Burner), all participating in the TCM Five Elements model. The intrinsic frequency ranges of each component, obtained through relevant calculations, are combined with the frequency distribution in speech to determine the distribution characteristics of the infrasound components in various parts of the body. Using the Hidden Markov Model (HMM), the optimal organs can be identified, and the correspondence between infrasound and organ systems can be obtained.
[0068] Figure 4 A quantitative model of the ascending and descending, hot and cold, entering and exiting, and yin and yang of the Heavenly Stems and Earthly Branches.
[0069] Among them, the angle θ is used in the calculation of the ascending and descending, hot and cold, entering and exiting, and yin and yang quantitative models of the Heavenly Stems and Earthly Branches in the main text of the instruction manual. Detailed Implementation
[0070] Example 1. Depression identification.
[0071] Thirty clinical samples were collected from individuals with depression (DP), 30 from individuals with kidney disease (UD), and 60 from a control group without a clinically diagnosed disease (CM). A total of 92 key speech feature factors were compared and statistically analyzed among the groups, revealing the following:
[0072] There were 7 highly significant differences (***P<0.0001), 10 significant differences (**P<0.001), and 22 significant differences (*P<0.05) between the UD and DP groups. The difference rate reached 42.39%.
[0073] Nine indicators showed highly significant differences between the UD group and the CM group*** (P < 0.0001), six indicators showed significant differences** (P < 0.001), and eleven indicators showed significant differences* (P < 0.05), with a difference rate of 28.26%.
[0074] There was one highly significant difference*** (P < 0.0001), two significant differences** (P < 0.001), and four significant differences* (P < 0.05) between the DP group and the CM group, with a difference rate of 7.6%.
[0075] The data shows that the device has multiple identifiable indicators for depression with obvious emotional symptoms, indicating that it has a certain ability to diagnose the disease.
[0076] Meanwhile, the trends derived from the indicators, such as the Five Elements classification and the differentiation of syndromes based on the internal organs, are consistent with the disease characteristics discovered by TCM physicians in their diagnosis and identification.
[0077] Taking the Qi mechanism of the Triple Burner as an example, the statistical values are shown in the table below:
[0078] Table 1. Comparison of numerical differences in trifocal weight between groups
[0079]
[0080] Note: (1) In pairwise comparisons, the LSD comparison method is used for groups with homogeneous variances, and the Dummett T3 comparison method is used for groups with unequal variances; (2) The trifocal data are calculated based on the original data and are not normalized, which results in heterogeneity between groups in the mean and standard deviation of the trifocal data, and there is no comparability between the data in the groups.
[0081] As shown in the table above, the inter-group differences in Qi mechanism data of the Middle Jiao (Triple Burner) were the highest among the three groups. The differences between the DP group and the UD group, and between the DP group and the CM group, were all extremely significant (**P < 0.000), indicating that the Qi mechanism of the Middle Jiao in depression was significantly lower than that in the kidney disease group and the normal group. The differences between the other groups were not significant. The range of emotional fluctuations in depression is wider than in normal individuals, which is attributed in Traditional Chinese Medicine to an imbalance of ascending and descending Qi due to insufficient Middle Jiao Qi. The numerical characteristics expressed by the speech data are consistent with the symptom characteristics of depression.
[0082] Example 2. Implementation plan for the correlation between emotion and organ fragility
[0083] Samples were included in the RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song) database. The original sample consisted of 24 actors, each recording speech information with the same semantic meaning for one of eight emotions: 01. neutral, 02. calm, 03. happy, 04. sad, 05. angry, 06. fearful, 07. disgusted, and 08. surprised. The samples were grouped according to this emotion classification, and relevant features of their speech signals were extracted using a voice diagnosis system. The audio energy in the 1-20Hz frequency band was divided into 10 equal parts and accumulated. The statistical distribution of the proportion of samples with the highest infrasound data in the 1-20Hz frequency band during the expression of different emotions was compared and analyzed, resulting in the data shown in Table 1 below.
[0084] Table 2. Distribution of infrasound peak frequency in speech samples from different emotion groups.
[0085]
[0086] Note: Bold text indicates the peak number of people in the infrasound frequency range of the sample speech. Since the speech amplitude increases significantly after 18-20Hz, only speech data within 18Hz are observed.
[0087] The table below summarizes the intrinsic frequencies of various organs in human physiology and their corresponding emotional relationships:
[0088] Table 3. Relationship between the five internal organs, emotions, and infrasound trends in emotional speech in Traditional Chinese Medicine.
[0089]
[0090] Application testing demonstrates that the device, in analyzing the infrasound frequency band of speech, can identify infrasound differences between different emotions by observing changes in the amplitude of infrasound components, while maintaining consistency with the principles of organ resonance in Traditional Chinese Medicine (TCM). This demonstrates the effectiveness of TCM theory and yields relevant indicative data and measurement methods.
[0091] Example 3. Comparison of differential speech features between smokers and non-smokers
[0092] By comparing a large sample of random speech, the speech characteristics of smokers and non-smokers were analyzed, revealing differences in 22 indicators, as shown in Table 3 below:
[0093] Table 4. Statistical differences in speech features between smoking and non-smoking samples in the male sample.
[0094]
[0095] Analysis of the data in the table above shows that segments 4, 12, and 20 of the MEL spectrum correspond to the frequency bands of 345.5Hz-492.0Hz, 2190.0Hz-2571.1Hz, and 7005.5Hz-8000.0Hz, respectively. Significant inter-group differences were observed in the speech data within these segments, specifically, the volume of non-smoking samples in these frequency bands was higher than that of smoking samples. Correspondingly, regarding the ratio of infrasound volume in the 1-20Hz infrasound band to the overall speech volume, the amplitude of infrasound volume in male smoking samples was found to be higher than that in male non-smoking samples.
[0096] In the peripheral speech modulation frequency data, the non-smoking samples were higher than the smoking samples, indicating that the non-smoking samples had a wider range of speech rhythm variations, i.e., a fuller intonation. In the male speech intensity index, the smoking samples were higher than the non-smoking samples. These two indicators suggest that smokers have a stronger speech "consciousness," i.e., they are more excited and have richer intonations. This suggests that the auditory diagnosis system can measure the neural excitation state produced by tobacco in smokers and express it through speech.
[0097] The standard deviation of the MAG Level 1 extrema represents the number of fundamental frequency amplitude fluctuation extrema in the spectrum generated by the Fourier transform of each frame, and the change in the number of these extrema with frame shift of the speech signal. This indicator shows significantly less variation in non-smoking samples than in smoking samples. The MAG Level 1 extrema, often simply referred to as ripple, expresses the basic fluctuation characteristics of the frequency distribution in the speech signal. The frame with the highest ripple rate represents the frame with the highest amount of noise in the speech, and its position within the speech segment is the location of the maximum ripple rate. This position is earlier in non-smoking samples, closer to the midpoint, while it is significantly later in smoking samples, indicating that the noise appears later in the smoking samples.
[0098] The above embodiments demonstrate that the auditory diagnosis system has the ability to identify the impact of different lifestyles and habits on the speech articulation system in its indicator design, and indirectly analyzes its syndrome characteristics and sympathetic excitation state by quantitatively expressing the resulting speech features.
[0099] The technical solutions of this invention are understood. However, these embodiments are merely illustrative examples and should not be construed as limiting the specific implementation of this invention to these embodiments. For those skilled in the art, several simple deductions and modifications can be made without departing from the concept of this invention, and all such modifications should be considered within the scope of protection of this invention.
Claims
1. An automated TCM auscultation and diagnosis system supported by intelligent voice technology, characterized in that, The system includes: a voice signal input terminal, a voice recognition (VAD) module, a voice feature factor generation and analysis module, an auxiliary information support module, and a comprehensive TCM identification module; the system derives diagnostic conclusions based on voice, including the Eight Principles, the Three Jiaos, Qi, Blood, Body Fluids, and Zang-Fu organs. The voice signal input terminal is a wideband voice recording device; the voice signal analysis adopts a wideband resonance model, which includes four components: excitation model U(z), vocal tract model G(z), radiation model R(z), and resonance model R(z); wherein the resonance model R(z) is the 1-20Hz low frequency component outside the audible threshold contained in the audio. The overall workflow is as follows: Input voice and personal information. The voice is converted into raw audio signal through the information input terminal. The entered information is converted into corresponding indicator data of traditional Chinese medicine theory through the auxiliary information support module. The original audio signal is distinguished from speech and noise segments by the speech recognition (VAD) module, and then the speech feature factor generation and analysis module analyzes the speech time domain, frequency domain and composite domain features to form speech feature factors. The speech feature factors are transformed into TCM diagnostic feature indicators by the TCM comprehensive identification module, forming quantitative diagnostic data including the five tones and five elements, viscera, triple burner, eight principles, and qi, blood and body fluid differentiation. After information integration, a comprehensive TCM diagnosis conclusion based on sound hearing is generated. The TCM comprehensive identification module compares personal information with voice conclusions to generate a reliability report, and formulates corresponding intervention plans based on high reliability conclusions. The speech feature factor analysis module incorporates the following TCM theory-based speech dialectics methods: The Triple Burner Differentiation Method analyzes the speech patterns of glottal, chest, and abdominal articulation to evaluate the main articulation methods used by a sample during normal speech, and is used to determine the distribution of the Triple Burner Qi in the speech process. The Five-Tone Zang-Fu Differentiation Method utilizes the characteristics of the five tones (Jiao, Zhi, Gong, Shang, Yu) within the audible threshold range to identify an individual's five tones, complete the individual bias based on the five tones, and form the Five-Tone Zang-Fu Differentiation System. The Qi, Blood and Body Fluid Differentiation Method separates the "noise" component in speech, assesses the body's body fluid status, and combines speech volume, frequency formants, and frequency domain distribution to determine the deficiency or excess of Qi and Blood and the distribution of body fluids in the sample. The Eight Principles of Dialectics analyze the stress position of single-character pronunciation, obtain the "cold and hot, exterior and interior" characteristics of the sample, analyze the five tones, determine the "rising and falling" characteristics, analyze the distribution of speech energy, and judge the "deficiency and excess of the middle qi" of the sample. Extract the pronunciation speed of Chinese characters to obtain the "rapid" or "slow" characteristics of emotions; The complexity of speech temporal frequency calculations is used to analyze the expressive richness of sentiment in sample language. Among them, the physical resonance effect of human organs is determined based on low-frequency characteristics. By extracting low-frequency characteristics of 1-20Hz from the signal and comparing them with the intrinsic frequencies of the organs, the proportion of speech resonance of the organs is determined, and the physical fragility of related organs is indirectly inferred, thus expanding the scope of TCM diagnostic information and the perspective of review. The auxiliary information support module includes a time calculation section, which has the following functions: Using the sexagenary cycle timekeeping principle, the birth date is converted into sexagenary cycle factors; Using the principles of Traditional Chinese Medicine, the ascending and descending attributes, yin-yang balance attributes, exterior-interior distribution, and cold-heat distribution attributes of the eight stem-branch factors in the sample were obtained. By incorporating information from the Five Elements and Six Qi theory of Traditional Chinese Medicine, the system automatically generates Qi data at two times: the time of birth and the time of speech. This includes the main Qi, the guest Qi, the main Qi, the guest Qi, and the TCM predictions regarding the excess or deficiency of each factor. These conclusions are corroborated by the identification conclusions obtained from speech analysis.
2. The automated TCM auscultation and diagnosis system supported by intelligent voice technology according to claim 1, characterized in that, The auxiliary information support module also includes an emotional calculation part, which is designed with an emotional characteristic assessment questionnaire. The questionnaire has 31 questions and has been designed to be depersonalized. The conclusion part designs and develops a traditional Chinese medicine five-line emotion discrimination method to form a quantitative conclusion of the five-line emotion, which is compared with the emotional conclusion of the voice sample output and serves to calibrate the accuracy of voice emotion prediction.
3. The automated TCM auscultation and diagnosis system supported by intelligent voice technology according to claim 1, characterized in that, The auxiliary information support module also includes a consultation information collection section, which incorporates standardized TCM consultation information to form 30 questions. By optimizing the TCM diagnostic connotations represented by each question, the TCM theories of Five Elements Differentiation, Eight Principles Differentiation, and Qi, Blood and Body Fluid Differentiation are quantitatively embedded to form a quantitative TCM self-diagnosis consultation conclusion. Based on the voice analysis conclusion, the system synchronously records the time used for each question and evaluates the reliability of the answer.
4. The automated TCM auscultation and diagnosis system supported by intelligent voice technology according to claim 1, characterized in that, The system has two usage methods: standardized diagnostic method and free analysis. The former involves pronouncing words according to a specified pronunciation method, completing data collection and analysis, and applying it to basic sound feature analysis to obtain basic steady-state feature analysis without emotional interference, including objective emotional characteristics, personality tone, five-element bias, eight principles, and syndrome-related health status of qi, blood, and body fluids. The latter involves any form of speech signal expressed in various natural contexts in a real-world environment to evaluate transient emotions and subjective syndrome characteristics. Based on the two modes, steady-state and transient features are combined to track and monitor the emotional and health status of the speaker in real time.
5. The automated TCM auscultation and diagnosis system supported by intelligent voice technology according to claim 1, characterized in that, The TCM comprehensive identification module is designed in a distributed manner, with multiple TCM diagnostic theories and data analysis strategies processed in parallel and internal competition completed to improve data accuracy. The device information integrates voice samples and personal basic information, or embeds electrocardiogram data externally to integrate electrocardiogram signals and improve the comprehensiveness and stability of TCM auscultation diagnosis information. It can also realize personalized health monitoring services by embedding it into mobile terminals.
Citation Information
Patent Citations
Human health status detection system based on meridian point measurement
CN101703397A
Chinese medicine sound diagnosis acquisition and analysis system
CN102342858A
Medical diagnosis system based on artificial intelligence and diagnosis method
CN109285605A
Method for interactive inquiry of traditional Chinese medicine syndrome factor inquiry voice robot
CN109820478A