Voice voiceprint characteristic value clustering method and system for representing human health state
By using cluster analysis of voiceprint feature values, a correlation model between voiceprint features and physiological systems/diseases is established, which solves the problem that existing technologies have not deeply explored the correlation between voiceprint features and human physiological systems, and realizes non-invasive health status assessment and disease auxiliary diagnosis.
Patent Information
- Application Number
- CN202512054870.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies have failed to effectively utilize voiceprint features to establish a correlation with human physiological systems or disease characteristics, lacking a systematic mapping relationship, and relying on invasive detection or wearable devices, resulting in complex operation and high costs.
By acquiring voiceprint data from different populations, preprocessing and feature extraction are performed to construct a baseline model for healthy individuals. Using difference comparison and cluster analysis, physiological voiceprint block maps are drawn, forming explicit and stable feature data clusters to achieve non-invasive health status assessment and disease-aided diagnosis.
It enables health status assessment without contact with the human body or collection of biological samples, reducing the threshold and discomfort of testing, improving the objectivity and reliability of assessment results, and providing a new dimension of technical support for disease auxiliary diagnosis.
Smart Images

Figure CN121528253A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer application technology, and in particular to a method and system for clustering voiceprint feature values to characterize human health status. Background Technology
[0002] In the fields of medical health monitoring and disease-aided diagnosis, traditional technologies mainly rely on invasive tests (such as blood analysis and imaging examinations) or wearable devices to collect physiological indicators (such as heart rate and body temperature), which have limitations such as complex operation, high cost, or the need for continuous device wearing. Although existing voiceprint recognition technology is widely used in fields such as identity authentication and voice interaction, it only stays at the level of modeling the identity features of voice signals, and no technology has yet established a correlation between voiceprint features and human physiological function or disease characteristics.
[0003] Therefore, a method is urgently needed to solve at least one of the above problems. Summary of the Invention
[0004] This application provides a speech signature feature clustering method and system for characterizing human health status. It aims to address the problem that although there are attempts in the prior art to analyze emotions or psychological states through speech signals, such methods only involve the paralinguistic features of speech (such as intonation and speech rate), and do not deeply explore the intrinsic relationship between voiceprint features (such as Mel frequency cepstral coefficients, fundamental frequency, etc.) and human physiological systems (such as respiratory system and nervous system). Furthermore, it lacks a systematic solution for constructing a mapping relationship between voiceprint features and physiological / disease features through cluster analysis.
[0005] In a first aspect, embodiments of this application provide a method for clustering voiceprint feature values to characterize human health status, the method comprising: Acquire voiceprint data from different population groups, including healthy individuals, specific population groups, or populations with diseases; preprocess the acquired voiceprint data; extract voiceprint feature values from the preprocessed voiceprint data to characterize the voiceprint. A baseline model of healthy individuals is constructed using the voiceprint feature values of healthy individuals; the voiceprint feature values of individuals with specific systems or diseases are compared with the baseline model of healthy individuals, and feature data related to human physiological systems or diseases are marked based on the comparison results. Physiological voiceprint block maps are drawn based on labeled feature data, and cluster analysis is performed on the feature data in the physiological voiceprint block maps to form explicit and stable feature data groups that are associated with human physiological systems or diseases.
[0006] In some embodiments, after generating a set of explicit and stable feature data associated with human physiological systems or diseases, the method further includes: validating the generated feature data set, and using the validated feature data set for non-invasive, quantitative characterization of human health status, as well as for disease-aided diagnosis and health monitoring.
[0007] In some embodiments, the verification of the formed feature data group includes: verifying the feature data group using reserved test voiceprint data, comparing the clustering results of the test data with known human health status or disease diagnosis results, calculating the accuracy, recall and specificity of the clustering results, and determining that the feature data group has passed verification if the accuracy, recall and specificity all reach a preset threshold; otherwise, readjusting the parameters or steps of the clustering analysis until verification is passed.
[0008] In some embodiments, the preprocessing of the acquired voiceprint data includes: removing silent segments and non-speech signals from the voiceprint data using a voice activity detection algorithm; performing noise reduction processing on the voiceprint data using a bandpass filter to remove environmental noise and high-frequency and low-frequency interference signals; and performing volume normalization processing on the noise-reduced voiceprint data to ensure that the volume amplitude of different voiceprint data is within the same numerical range.
[0009] In some embodiments, the step of extracting voiceprint feature values from preprocessed speech voiceprint data includes: segmenting the preprocessed speech voiceprint data into multiple short-time speech frames, extracting time-domain features and frequency-domain features for each short-time speech frame, wherein the time-domain features include short-time energy, short-time average zero-crossing rate and fundamental frequency, and the frequency-domain features include Mel frequency cepstral coefficients, spectral centroid and spectral roll-off point, and combining the extracted time-domain features and frequency-domain features to form a voiceprint feature vector.
[0010] In some embodiments, the construction of a benchmark model for healthy individuals using voiceprint feature values includes: performing statistical analysis on the voiceprint feature values of healthy individuals, calculating the mean, variance, and probability distribution of each voiceprint feature value, establishing a multidimensional Gaussian mixture model of voiceprint features of healthy individuals based on the statistical analysis results, and using the Gaussian mixture model as a benchmark model for healthy individuals to characterize the normal distribution range of voiceprint features of healthy individuals.
[0011] In some embodiments, the step of comparing the voiceprint feature values of a specific system or disease population with a baseline model of a healthy population, and marking feature data related to the human physiological system or disease based on the comparison results, includes: calculating the degree of difference between the voiceprint feature values of a specific system or disease population and the corresponding voiceprint feature values in the baseline model of a healthy population, wherein the degree of difference is the ratio of the absolute difference of the voiceprint feature values to the standard deviation of the voiceprint feature values of the healthy population, and when the degree of difference is greater than a preset difference threshold, the corresponding voiceprint feature value is marked as feature data related to the human physiological system or disease.
[0012] In some embodiments, the step of drawing a physiological voiceprint block map based on labeled feature data includes: visually mapping the labeled feature data in a multidimensional feature space according to the human physiological system or disease category associated with the labeled feature data; dividing different block areas in the visualization interface according to the physiological system or disease category; displaying the distribution of feature data related to the physiological system or disease in each block area; and forming a physiological voiceprint block map that reflects the relationship between the feature data and the physiological system or disease.
[0013] In some embodiments, the clustering analysis of feature data in the physiological voiceprint block map to form a feature data group that is explicit and stable and associated with human physiological systems or diseases includes: using a density clustering algorithm to cluster the feature data in the physiological voiceprint block map; determining cluster centers based on the distribution density of feature data in a multidimensional feature space; dividing the density-reachable feature data into the same cluster, with each cluster corresponding to a human physiological system or disease category; and iteratively optimizing the cluster centers and cluster division to ensure that the feature data within each cluster is highly similar and that the feature data between different clusters are significantly different, thus forming an explicit and stable feature data group.
[0014] Secondly, this application provides a voiceprint feature value clustering system for characterizing human health status, the system comprising: The data acquisition unit is used to acquire voiceprint data of different groups of people, including healthy people, specific groups, or people with diseases; preprocess the acquired voiceprint data; and extract voiceprint feature values from the preprocessed voiceprint data to characterize the voiceprint. The data comparison unit is used to construct a baseline model of healthy people using the voiceprint feature values of healthy people; compare the voiceprint feature values of specific systems or disease groups with the baseline model of healthy people, and mark the feature data related to human physiological systems or diseases based on the comparison results. The block drawing unit is used to draw physiological voiceprint block maps based on labeled feature data, and to perform cluster analysis on the feature data in the physiological voiceprint block maps to form explicit and stable feature data groups that are associated with human physiological systems or diseases.
[0015] This application eliminates the need for contact with the human body or collection of biological samples, enabling health status assessment through voice signals, thus lowering the threshold for detection and reducing discomfort. Through statistical modeling and cluster analysis of voiceprint feature values, health status is transformed into a quantifiable cluster of feature data, enhancing the objectivity of the assessment results. A correlation model between voiceprint features and physiological systems / diseases is established, providing a new dimension of technical support for disease-aided diagnosis. The "physiological voiceprint block map" visually presents the mapping relationship between feature data and physiological systems, facilitating medical analysis and clinical application.
[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic flowchart illustrating the steps of a speech signature feature value clustering method for characterizing human health status provided in an embodiment of this application; Figure 2 This is a schematic block diagram of a voiceprint feature value clustering system for characterizing human health status provided in an embodiment of this application; Figure 3 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0022] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0023] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0024] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] In the fields of medical health monitoring and disease-aided diagnosis, traditional technologies mainly rely on invasive tests (such as blood analysis and imaging examinations) or wearable devices to collect physiological indicators (such as heart rate and body temperature), which have limitations such as complex operation, high cost, or the need for continuous device wearing. Although existing voiceprint recognition technology is widely used in fields such as identity authentication and voice interaction, it only stays at the level of modeling the identity features of voice signals, and no technology has yet established a correlation between voiceprint features and human physiological function or disease characteristics.
[0026] Although there are attempts to analyze emotions or psychological states through speech signals in existing technologies, such methods only involve the paralinguistic features of speech (such as intonation and speech rate), and do not delve into the intrinsic relationship between voiceprint features (such as Mel frequency cepstral coefficients, fundamental frequency, etc.) and human physiological systems (such as respiratory system and nervous system), and lack a systematic solution for constructing a mapping relationship between voiceprint features and physiological / disease features through cluster analysis.
[0027] To resolve the above issues, please refer to... Figure 1This application provides a voiceprint feature value clustering method for characterizing human health status, applied to computer devices. The computer devices can be deployed on a single server or server cluster, or on handheld terminals, laptops, wearable devices, or robots. It should be noted that all information involved in the method provided in this application is extracted with the authorization of the relevant user and in accordance with relevant regulations, and will not infringe on user privacy.
[0028] The provided speech signature feature value clustering method for representing human health status includes steps S101 to S103. Details are as follows: Step S101. Obtain voiceprint data from different groups of people, including healthy people, specific groups, or people with diseases; preprocess the obtained voiceprint data; extract voiceprint feature values from the preprocessed voiceprint data to characterize the voiceprint.
[0029] Specifically, voiceprint data from different groups of people are collected using computer equipment. To address the issues of noise interference and feature diversity in voice signals, invalid signals are removed and the data format is standardized through preprocessing. Then, key parameters that can characterize voiceprint features are separated from the voice signal using feature extraction algorithms.
[0030] Computer equipment collects voice signals through microphones or voice input interfaces, covering voice samples from healthy people, people with specific systemic diseases (such as patients with respiratory diseases or neurological diseases), and sub-healthy people. The collection scenarios include quiet indoor environments or controlled noise environments to ensure data diversity.
[0031] During the data collection process, the corresponding population tags (such as health status and disease type) of the samples are recorded, and data collection is completed with user authorization, which complies with privacy protection regulations.
[0032] Preprocessing includes: Voice Activity Detection (VAD): Voice segments and non-voice segments are identified using a dual-threshold energy detection algorithm or a Hidden Markov Model (HMM), eliminating silent segments and environmental noise (such as background noise and equipment hum), retaining only valid voice data. Bandpass Filtering Noise Reduction: A 40Hz-3.4kHz bandpass filter is used to filter the voice signal, removing low-frequency mechanical vibration noise (such as air conditioner fan noise) and high-frequency electromagnetic interference, retaining the effective frequency band of human speech. Volume Normalization: The amplitude of the voice signal is linearly scaled, adjusting the volume amplitude of all samples to the standardized range of [-1, 1], eliminating the impact of differences in pronunciation intensity among different users on subsequent analysis.
[0033] The preprocessed speech signal is divided into short speech frames of 20-30 milliseconds, with a frame shift of 10 milliseconds. The Hamming window function is used to reduce inter-frame signal truncation distortion.
[0034] Temporal feature extraction reflects the intensity changes of the speech signal by calculating the short-time energy of each frame; it also calculates the short-time average zero-crossing rate to characterize the frequency and voiced / unvoiced characteristics of the speech signal; and it extracts the fundamental frequency by using the autocorrelation method or the average amplitude difference function (AMDF) to reflect the periodic characteristics of vocal cord vibration.
[0035] Frequency domain feature extraction involves transforming each frame of speech into the frequency domain using a Fast Fourier Transform (FFT), calculating the Mel-frequency cepstral coefficients (MFCCs) to capture frequency components sensitive to the human ear, and calculating the spectral centroid and spectral rolloff to characterize the central location of the spectral energy distribution and high-frequency attenuation characteristics, respectively. The time-domain and frequency-domain features are then concatenated sequentially into a multi-dimensional feature vector (e.g., 39-dimensional MFCCs + 6-dimensional time-domain features) as the final representation of the speaker's feature values.
[0036] Step S102. Construct a baseline model of healthy individuals using voiceprint feature values of healthy individuals; compare the voiceprint feature values of individuals with those of specific systems or diseases with the baseline model of healthy individuals, and mark the feature data related to human physiological systems or diseases based on the comparison results.
[0037] Specifically, based on the voiceprint characteristics of healthy individuals, a voiceprint characteristic distribution model under normal physiological conditions is established. By comparing the differences between the voiceprint characteristics of a specific population and the benchmark model, characteristic parameters related to physiological systems or diseases are identified.
[0038] The baseline model for healthy individuals was constructed by randomly selecting 80% of the samples from the voiceprint feature values of healthy individuals as the training set. A multidimensional Gaussian mixture model (GMM) was used for modeling, and the model parameters, including the mean vector, covariance matrix, and mixture component weights of each feature dimension, were estimated using the expectation-maximization (EM) algorithm.
[0039] The model outputs the probability density distribution of voiceprint features of healthy individuals and defines the normal fluctuation range of each feature dimension (such as mean ± 2 standard deviations) as a benchmark for judging abnormal features.
[0040] For voiceprint feature values of specific systems or disease populations, the difference between each feature dimension and the corresponding dimension of the health benchmark model is calculated: Difference = |Feature value - Health mean| / Health standard deviation; A difference threshold is set (e.g., 1.5 times the standard deviation). When the difference of a feature dimension exceeds the threshold, the feature is marked as an "abnormal feature" and associated with the corresponding physiological system or disease category according to the preset mapping relationship (e.g., abnormal fundamental frequency is associated with respiratory diseases, and abnormal MFCC is associated with nervous system dysfunction).
[0041] Statistical screening is performed on the marked abnormal features to remove random fluctuation features and retain stable difference features that repeatedly appear in the same population.
[0042] Step S103. Draw a physiological voiceprint block map based on the labeled feature data, and perform cluster analysis on the feature data in the physiological voiceprint block map to form a feature data group that is explicit and stable and associated with the human physiological system or disease.
[0043] Specifically, the marked abnormal features are mapped to a multi-dimensional space using visualization technology to form a "physiological voiceprint block map". The feature data are then grouped using a clustering algorithm to establish an explicit association model between voiceprint features and physiological systems / diseases.
[0044] Physiological voiceprint block mapping employs principal component analysis (PCA) or t-distributed random neighborhood embedding (t-SNE) algorithms to reduce the dimensionality of high-dimensional feature data, mapping the feature vectors to a two-dimensional or three-dimensional visualization space.
[0045] In the visualization interface, different blocks are divided according to physiological systems or disease categories (such as respiratory, nervous, and cardiovascular). Each block displays the characteristic data points of the corresponding population in the form of a scatter plot, and uses color coding to distinguish the degree of feature difference (such as red for high difference features and blue for low difference features).
[0046] The block diagram supports interactive operations. Users can zoom and pan to view the data distribution of specific feature intervals, or click on data points to view the detailed health labels of samples.
[0047] Cluster analysis employs density-based clustering algorithms (such as DBSCAN) to cluster feature data in the block diagram, setting the neighborhood radius ε and the minimum number of samples MinPts as clustering parameters. The algorithm traverses all data points, grouping points with achievable density (i.e., containing at least MinPts samples within the ε radius) into the same cluster, with each cluster corresponding to a physiological state or disease type. The effectiveness of the clustering results is evaluated by calculating the silhouette coefficient to determine intra-cluster compactness and inter-cluster separation. Through iterative adjustment of the ε and MinPts parameters, the feature data within each cluster exhibit high similarity (e.g., Euclidean distance less than 0.5 standard deviations), while different clusters show significant differences (e.g., silhouette coefficient greater than 0.7). The final feature data cluster, represented by the cluster center vector, serves as a voiceprint feature template for the physiological system or disease, used for subsequent health status matching and diagnostic assistance.
[0048] In some embodiments, after generating a set of explicit and stable feature data associated with human physiological systems or diseases, the method further includes: validating the generated feature data set, and using the validated feature data set for non-invasive, quantitative characterization of human health status, as well as for disease-aided diagnosis and health monitoring.
[0049] After forming a feature data set, the reliability of the model is ensured through a validation process, and the validated model is applied to actual health monitoring and disease auxiliary diagnosis scenarios to achieve non-invasive quantitative assessment.
[0050] During the data acquisition phase, 20% of the voiceprint data from healthy individuals and individuals with specific diseases is reserved as an independent test set to ensure no data overlap between the test set and the training set. The voiceprint feature values from the test set are input into the established feature data clustering model, and a clustering algorithm is used to output the prediction results for health status or disease category.
[0051] The prediction results are compared with known clinical diagnostic results in the test set (such as disease diagnosis reports issued by hospitals and physiological indicator test results), and the accuracy (number of correctly predicted samples / total number of samples), recall (number of correctly predicted positive samples / actual number of positive samples), and specificity (number of correctly predicted negative samples / actual number of negative samples) are calculated.
[0052] Set a preset validation threshold (e.g., accuracy ≥ 90%, recall ≥ 85%, specificity ≥ 85%). If all indicators meet the threshold, the model is considered validated. If not, readjust the clustering analysis parameters (e.g., neighborhood radius ε and minimum sample size MinPts for density clustering) or return to step S102 to adjust the difference comparison threshold until validation is successful.
[0053] The validated feature data cluster model is integrated into computer devices (such as medical diagnostic terminals and smart voice devices). By collecting users' voice signals in real time, the model is called to output quantitative results of health status or auxiliary diagnostic suggestions for diseases.
[0054] In some embodiments, the verification of the formed feature data group includes: verifying the feature data group using reserved test voiceprint data, comparing the clustering results of the test data with known human health status or disease diagnosis results, calculating the accuracy, recall and specificity of the clustering results, and determining that the feature data group has passed verification if the accuracy, recall and specificity all reach a preset threshold; otherwise, readjusting the parameters or steps of the clustering analysis until verification is passed.
[0055] Define a validation method for the feature data clusters, and quantitatively evaluate the classification performance of the model by comparing independent test data with clinical diagnostic results to ensure the reliability of the clustering results.
[0056] Test samples that are not used in model training are randomly selected from historically collected voiceprint data. Each sample is labeled with a clear health status label (such as "healthy", "respiratory system disease", "nervous system disease"). The label is confirmed by a professional physician based on clinical examination results.
[0057] The voiceprint feature values of the test samples are input into the feature data cluster model, and the clustering algorithm (such as DBSCAN) is used to determine the cluster to which each sample belongs, and map it to the corresponding physiological system or disease category.
[0058] Confusion matrix between predicted results and true labels is constructed, and the number of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) samples is counted.
[0059] Performance metrics are calculated as follows: Accuracy = (TP + TN) / (TP + TN + FP + FN); Recall = TP / (TP + FN); Specificity = TN / (TN + FP). If any metric fails to reach the preset threshold, the parameter tuning process is automatically triggered: If the cluster division is too fine (low recall), increase the neighborhood radius ε or decrease MinPts; If the clusters overlap (low specificity), decrease ε or increase MinPts; Repeat the steps until the metric meets the target.
[0060] In some embodiments, the preprocessing of the acquired voiceprint data includes: removing silent segments and non-speech signals from the voiceprint data using a voice activity detection algorithm; performing noise reduction processing on the voiceprint data using a bandpass filter to remove environmental noise and high-frequency and low-frequency interference signals; and performing volume normalization processing on the noise-reduced voiceprint data to ensure that the volume amplitude of different voiceprint data is within the same numerical range.
[0061] To address noise and invalid signals in the raw speech and voiceprint data, a preprocessing procedure is used to improve data quality and ensure the accuracy of subsequent feature extraction.
[0062] Voice Activity Detection (VAD) employs a dual-threshold energy detection algorithm: it calculates the short-time energy and zero-crossing rate of the speech signal, sets high and low energy thresholds and a zero-crossing rate threshold, and determines the signal as a speech segment when the short-time energy is higher than the high threshold or the zero-crossing rate is higher than the low threshold; otherwise, it is determined as a non-speech segment (such as silence or coughing), and non-speech segment data is discarded. Alternatively, a VAD based on a Hidden Markov Model (HMM) can be used. By training a silence model and a speech model, the input signal is sequence-labeled to segment out effective speech regions.
[0063] The bandpass filter noise reduction design uses a Butterworth bandpass filter with a frequency range of 40Hz-3400Hz to filter speech signals, attenuating low-frequency noise below 40Hz (such as mechanical vibration) and high-frequency noise above 3400Hz (such as electromagnetic interference), while preserving the main frequency components of human speech (200Hz-3000Hz).
[0064] Volume normalization calculates the root mean square amplitude (RMS) of the speech signal, linearly scaling the amplitude of each sample to a standardized range of [-1, 1]. The formula is: xnorm = (x xmin) / (xmax) xmin)×2 1; where xmin and xmax are the minimum and maximum amplitude values of the sample, ensuring that volume fluctuations caused by differences in pronunciation force among different users do not affect subsequent feature analysis.
[0065] In some embodiments, the step of extracting voiceprint feature values from preprocessed speech voiceprint data includes: segmenting the preprocessed speech voiceprint data into multiple short-time speech frames, extracting time-domain features and frequency-domain features for each short-time speech frame, wherein the time-domain features include short-time energy, short-time average zero-crossing rate and fundamental frequency, and the frequency-domain features include Mel frequency cepstral coefficients, spectral centroid and spectral roll-off point, and combining the extracted time-domain features and frequency-domain features to form a voiceprint feature vector.
[0066] By using short-time analysis and multi-dimensional feature extraction, key parameters that characterize voiceprint features are separated from the preprocessed speech signal to form a multi-dimensional feature vector.
[0067] Short-time speech frame segmentation divides the preprocessed speech signal into 25-millisecond short-time frames along the time axis, with a frame shift of 10 milliseconds (i.e., adjacent frames overlap by 15 milliseconds) to ensure the continuity of features between frames. A Hamming window function is used to weight the signal of each frame to reduce spectral leakage.
[0068] Temporal Feature Extraction: Short-Time Energy: Calculates the sum of squares of the signal in each frame, reflecting the intensity changes of the speech. The formula is: ; Where s(n) is the speech signal, w(m) is the Hamming window function, and N is the number of frame length points.
[0069] Short-time average zero-crossing rate: This is the number of times the signal waveform crosses the zero level in each frame. The formula is: ; Used to distinguish between voiced and unvoiced sounds (voiceless consonants have a high zero-crossing rate, while voiced consonants have a low zero-crossing rate). The fundamental frequency is calculated by using the autocorrelation method to determine the autocorrelation function of the speech frame, finding the delay period corresponding to the main peak as the fundamental frequency period, and the reciprocal of this period is the fundamental frequency (F0), which reflects the fundamental frequency of vocal cord vibration.
[0070] Frequency domain feature extraction: Mel frequency cepstral coefficients (MFCC) are obtained by performing an FFT transform on each frame of speech to the frequency domain and calculating the power spectrum; the power spectrum is filtered by a Mel filter bank (usually 40 filters) to simulate the characteristics of human hearing; the logarithm of the filtered energy is taken and a discrete cosine transform (DCT) is performed to extract the first 12-13 order coefficients as MFCC features.
[0071] The centroid of the spectrum reflects the brightness of the spectrum (the proportion of high-frequency components) by calculating the centroid frequency of the spectral energy distribution. The roll-off point is used to detect high-frequency attenuation in speech by finding the frequency point in the spectrum where the cumulative energy accounts for 95% of the total energy.
[0072] The time-domain features (such as 3D: short-time energy, zero-crossing rate, F0) and frequency-domain features (such as 13D MFCC + 2D spectral centroid / roll-off point) are concatenated into an 18-dimensional feature vector, which serves as the final representation of the voiceprint feature value.
[0073] In some embodiments, the construction of a benchmark model for healthy individuals using voiceprint feature values includes: performing statistical analysis on the voiceprint feature values of healthy individuals, calculating the mean, variance, and probability distribution of each voiceprint feature value, establishing a multidimensional Gaussian mixture model of voiceprint features of healthy individuals based on the statistical analysis results, and using the Gaussian mixture model as a benchmark model for healthy individuals to characterize the normal distribution range of voiceprint features of healthy individuals.
[0074] Based on the statistical distribution of voiceprint features in healthy individuals, a multidimensional Gaussian mixture model is established as a benchmark to define the fluctuation range of voiceprint features under normal physiological conditions.
[0075] Collect voiceprint feature vectors (such as the 18-dimensional features described in Example 4) from at least 1000 healthy individuals, and calculate the mean (μj) and variance (σj) of each feature dimension. 2 ) and probability density distribution, to identify the correlation between features (such as the correlation between MFCC and fundamental frequency).
[0076] Gaussian Mixture Model (GMM) modeling assumes that the voiceprint features of healthy individuals follow a multimodal distribution. It employs a GMM model containing 2-5 Gaussian components, with the following formula: ; Where M is the number of mixture components, ωi is the weight of the i-th component (∑ωi=1), μi is the mean vector, and Σi is the covariance matrix. The model parameters are iteratively optimized using the Expectation-Maximization (EM) algorithm: ωi, μi, and Σi of each component are initialized (e.g., random initialization or k-means pre-clustering); the E-step calculates the posterior probability of each sample belonging to each component; the M-step updates the parameters of each component based on the posterior probability until the log-likelihood function converges.
[0077] For each feature dimension, the normal fluctuation range is defined with the mean μj of the healthy population as the center, ±2 standard deviations (μj±2σj). Feature values exceeding this range are judged as abnormal.
[0078] In some embodiments, the step of comparing the voiceprint feature values of a specific system or disease population with a baseline model of a healthy population, and marking feature data related to the human physiological system or disease based on the comparison results, includes: calculating the degree of difference between the voiceprint feature values of a specific system or disease population and the corresponding voiceprint feature values in the baseline model of a healthy population, wherein the degree of difference is the ratio of the absolute difference of the voiceprint feature values to the standard deviation of the voiceprint feature values of the healthy population, and when the degree of difference is greater than a preset difference threshold, the corresponding voiceprint feature value is marked as feature data related to the human physiological system or disease.
[0079] By quantifying the differences between voiceprint characteristics of a specific population and health benchmarks, key features related to physiological systems or diseases can be identified, providing labeled data for subsequent cluster analysis.
[0080] For each voiceprint feature value xj in a specific system or disease population, the difference between it and the corresponding feature mean μj in the health baseline model is calculated: Diff j = |xj μj∣ / σj; where σj is the standard deviation of characteristic j of the healthy population, reflecting the fluctuation range of the characteristics of the healthy population.
[0081] Set a difference threshold T (e.g., 1.5 or 2.0, adjusted according to clinical needs). When Diff j > T, label feature j as "physiological system / disease-related feature" and associate it with a specific physiological system or disease category according to a preset mapping rule: High fundamental frequency (F0) difference: associated with respiratory system diseases (e.g., vocal cord inflammation affecting vibration frequency); High spectral roll-off point difference: associated with neurological system diseases (e.g., motor aphasia leading to high-frequency speech disorders).
[0082] For all samples from the same disease population, the frequency of each marker feature is counted. Features that occur less than 70% of the time in the same population are removed (considered as random fluctuations), and high-frequency and stable differential features are retained as valid marker data.
[0083] In some embodiments, the step of drawing a physiological voiceprint block map based on labeled feature data includes: visually mapping the labeled feature data in a multidimensional feature space according to the human physiological system or disease category associated with the labeled feature data; dividing different block areas in the visualization interface according to the physiological system or disease category; displaying the distribution of feature data related to the physiological system or disease in each block area; and forming a physiological voiceprint block map that reflects the relationship between the feature data and the physiological system or disease.
[0084] The labeled feature data is visualized and mapped to a multi-dimensional space, and displayed in sections according to physiological systems or disease categories, forming a "physiological voiceprint block map" that intuitively reflects the feature-physiological relationship.
[0085] Feature dimensionality reduction reduces labeled high-dimensional feature data (e.g., 18-dimensional) to 2-3 dimensions using principal component analysis (PCA) or t-distributed random neighborhood embedding (t-SNE) algorithms, retaining principal components (PCA) with a cumulative variance contribution rate ≥90% or optimizing the consistency of probability distribution from high-dimensional space to low-dimensional space (t-SNE).
[0086] Visualization mapping uses a two-dimensional plane (XY axis) or a three-dimensional space (XYZ axis) where each data point represents a dimensionality-reduced feature of a sample. The color of the point indicates the disease category (e.g., red = respiratory diseases, blue = nervous system diseases), and the size of the point indicates the magnitude of the difference (the larger the value, the more significant the deviation of the feature from the health baseline).
[0087] In the visualization interface, rectangular or circular blocks are drawn according to physiological systems or disease categories (such as "respiratory", "neural", "cardiovascular"). Each block displays only the feature data points associated with that category. The block boundaries are automatically generated based on the density threshold of the data distribution (such as the minimum convex hull containing 90% of the same type of data points).
[0088] It supports hovering the mouse to view the original feature values and health tags corresponding to data points; it provides zoom tools (such as scroll wheel zoom) and filtering functions (such as filtering by disease type) to facilitate users in analyzing the distribution patterns of specific category features.
[0089] In some embodiments, the clustering analysis of feature data in the physiological voiceprint block map to form a feature data group that is explicit and stable and associated with human physiological systems or diseases includes: using a density clustering algorithm to cluster the feature data in the physiological voiceprint block map; determining cluster centers based on the distribution density of feature data in a multidimensional feature space; dividing the density-reachable feature data into the same cluster, with each cluster corresponding to a human physiological system or disease category; and iteratively optimizing the cluster centers and cluster division to ensure that the feature data within each cluster is highly similar and that the feature data between different clusters are significantly different, thus forming an explicit and stable feature data group.
[0090] Density clustering algorithm is used to group the visualized feature data, and compact and distinguishable feature data clusters are formed through iterative optimization to establish an explicit association between voiceprint features and physiological systems / diseases.
[0091] Density clustering initialization: Select the DBSCAN algorithm, set the neighborhood radius ε (e.g., Euclidean distance 0.8 in the reduced-dimensional space) and the minimum number of samples MinPts (e.g., 5), traverse all data points, and label core points (containing ≥MinPts points within the radius ε), boundary points (non-core points but belonging to the neighborhood of a certain core point) and noise points (non-core points and not belonging to the neighborhood of any core point).
[0092] Cluster partitioning starts from any unvisited core point and recursively searches for all points that are density-reachable from it to form a cluster; this process is repeated until all core points are visited, and each cluster corresponds to a physiological system or disease category.
[0093] Iterative optimization calculates the silhouette coefficient for each cluster using the following formula: ; Where a(i) is the average distance from sample i to other samples in the same cluster, b(i) is the average distance from sample i to the nearest heterogeneous cluster, and s(i) takes values in the range [-1, 1]. The larger the value, the better the intra-cluster compactness and inter-cluster separation. If the global average silhouette coefficient is <0.5, ε and MinPts are automatically adjusted (e.g., ε±0.1, MinPts±1), and re-clustering is performed until the average silhouette coefficient is ≥0.7.
[0094] For the final cluster, the feature mean vector of all samples within the cluster is calculated as the cluster center to form the feature template of the physiological system or disease; noise points (profile coefficient < 0) are removed to ensure that each feature data group contains only stable features that are strongly correlated with a specific physiological state.
[0095] Please see Figure 2 As shown, Figure 2 This is a schematic diagram of the structure of a voiceprint feature value clustering system 200 representing human health status provided in this application embodiment. This voiceprint feature value clustering system 200 is used to execute the steps of the voiceprint feature value clustering method representing human health status shown in the above embodiments. The voiceprint feature value clustering system 200 can be a single server or a server cluster, or it can be a terminal, such as a handheld terminal, laptop computer, wearable device, or robot.
[0096] like Figure 2 As shown, the voiceprint feature value clustering system 200, which characterizes the health status of the human body, includes: The data acquisition unit 201 is used to acquire voiceprint data of different groups of people, including healthy people, specific groups of people or people with diseases; preprocess the acquired voiceprint data; and extract voiceprint feature values from the preprocessed voiceprint data to characterize the voiceprint. The data comparison unit 202 is used to construct a baseline model of healthy people using the voiceprint feature values of healthy people; compare the voiceprint feature values of specific systems or disease groups with the baseline model of healthy people, and mark the feature data related to human physiological systems or diseases based on the comparison results. Block drawing unit 203 is used to draw physiological voiceprint block maps based on labeled feature data, and to perform cluster analysis on the feature data in the physiological voiceprint block maps to form a group of explicit and stable feature data that are associated with human physiological systems or diseases.
[0097] In some embodiments, after generating a set of explicit and stable feature data associated with human physiological systems or diseases, the method further includes: validating the generated feature data set, and using the validated feature data set for non-invasive, quantitative characterization of human health status, as well as for disease-aided diagnosis and health monitoring.
[0098] In some embodiments, the verification of the formed feature data group includes: verifying the feature data group using reserved test voiceprint data, comparing the clustering results of the test data with known human health status or disease diagnosis results, calculating the accuracy, recall and specificity of the clustering results, and determining that the feature data group has passed verification if the accuracy, recall and specificity all reach a preset threshold; otherwise, readjusting the parameters or steps of the clustering analysis until verification is passed.
[0099] In some embodiments, the preprocessing of the acquired voiceprint data includes: removing silent segments and non-speech signals from the voiceprint data using a voice activity detection algorithm; performing noise reduction processing on the voiceprint data using a bandpass filter to remove environmental noise and high-frequency and low-frequency interference signals; and performing volume normalization processing on the noise-reduced voiceprint data to ensure that the volume amplitude of different voiceprint data is within the same numerical range.
[0100] In some embodiments, the step of extracting voiceprint feature values from preprocessed speech voiceprint data includes: segmenting the preprocessed speech voiceprint data into multiple short-time speech frames, extracting time-domain features and frequency-domain features for each short-time speech frame, wherein the time-domain features include short-time energy, short-time average zero-crossing rate and fundamental frequency, and the frequency-domain features include Mel frequency cepstral coefficients, spectral centroid and spectral roll-off point, and combining the extracted time-domain features and frequency-domain features to form a voiceprint feature vector.
[0101] In some embodiments, the construction of a benchmark model for healthy individuals using voiceprint feature values includes: performing statistical analysis on the voiceprint feature values of healthy individuals, calculating the mean, variance, and probability distribution of each voiceprint feature value, establishing a multidimensional Gaussian mixture model of voiceprint features of healthy individuals based on the statistical analysis results, and using the Gaussian mixture model as a benchmark model for healthy individuals to characterize the normal distribution range of voiceprint features of healthy individuals.
[0102] In some embodiments, the step of comparing the voiceprint feature values of a specific system or disease population with a baseline model of a healthy population, and marking feature data related to the human physiological system or disease based on the comparison results, includes: calculating the degree of difference between the voiceprint feature values of a specific system or disease population and the corresponding voiceprint feature values in the baseline model of a healthy population, wherein the degree of difference is the ratio of the absolute difference of the voiceprint feature values to the standard deviation of the voiceprint feature values of the healthy population, and when the degree of difference is greater than a preset difference threshold, the corresponding voiceprint feature value is marked as feature data related to the human physiological system or disease.
[0103] In some embodiments, the step of drawing a physiological voiceprint block map based on labeled feature data includes: visually mapping the labeled feature data in a multidimensional feature space according to the human physiological system or disease category associated with the labeled feature data; dividing different block areas in the visualization interface according to the physiological system or disease category; displaying the distribution of feature data related to the physiological system or disease in each block area; and forming a physiological voiceprint block map that reflects the relationship between the feature data and the physiological system or disease.
[0104] In some embodiments, the clustering analysis of feature data in the physiological voiceprint block map to form a feature data group that is explicit and stable and associated with human physiological systems or diseases includes: using a density clustering algorithm to cluster the feature data in the physiological voiceprint block map; determining cluster centers based on the distribution density of feature data in a multidimensional feature space; dividing the density-reachable feature data into the same cluster, with each cluster corresponding to a human physiological system or disease category; and iteratively optimizing the cluster centers and cluster division to ensure that the feature data within each cluster is highly similar and that the feature data between different clusters are significantly different, thus forming an explicit and stable feature data group.
[0105] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the speech and voiceprint feature value clustering system and each module that characterizes human health status described above can be referred to the corresponding content in the various embodiments of the speech and voiceprint feature value clustering method that characterizes human health status described above, and will not be repeated here.
[0106] The aforementioned voiceprint feature value clustering method for representing human health status can be implemented as a computer program, which can be used in various ways, such as... Figure 2 It runs on the device shown.
[0107] Please see Figure 3 , Figure 3 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application. The computer device includes a processor, a memory, and a network interface connected via a device bus, wherein the memory may include a storage medium and internal memory.
[0108] The storage medium may store operating devices and computer programs. The computer program includes program instructions that, when executed, cause the processor to perform any speech signature clustering method that characterizes human health status.
[0109] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0110] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to perform any speech signature clustering method that characterizes the health status of the human body.
[0111] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 3The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the terminal to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0112] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0113] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: Acquire voiceprint data from different population groups, including healthy individuals, specific population groups, or populations with diseases; preprocess the acquired voiceprint data; extract voiceprint feature values from the preprocessed voiceprint data to characterize the voiceprint. A baseline model of healthy individuals is constructed using the voiceprint feature values of healthy individuals; the voiceprint feature values of individuals with specific systems or diseases are compared with the baseline model of healthy individuals, and feature data related to human physiological systems or diseases are marked based on the comparison results. Physiological voiceprint block maps are drawn based on labeled feature data, and cluster analysis is performed on the feature data in the physiological voiceprint block maps to form explicit and stable feature data groups that are associated with human physiological systems or diseases.
[0114] In some embodiments, after generating a set of explicit and stable feature data associated with human physiological systems or diseases, the method further includes: validating the generated feature data set, and using the validated feature data set for non-invasive, quantitative characterization of human health status, as well as for disease-aided diagnosis and health monitoring.
[0115] In some embodiments, the verification of the formed feature data group includes: verifying the feature data group using reserved test voiceprint data, comparing the clustering results of the test data with known human health status or disease diagnosis results, calculating the accuracy, recall and specificity of the clustering results, and determining that the feature data group has passed verification if the accuracy, recall and specificity all reach a preset threshold; otherwise, readjusting the parameters or steps of the clustering analysis until verification is passed.
[0116] In some embodiments, the preprocessing of the acquired voiceprint data includes: removing silent segments and non-speech signals from the voiceprint data using a voice activity detection algorithm; performing noise reduction processing on the voiceprint data using a bandpass filter to remove environmental noise and high-frequency and low-frequency interference signals; and performing volume normalization processing on the noise-reduced voiceprint data to ensure that the volume amplitude of different voiceprint data is within the same numerical range.
[0117] In some embodiments, the step of extracting voiceprint feature values from preprocessed speech voiceprint data includes: segmenting the preprocessed speech voiceprint data into multiple short-time speech frames, extracting time-domain features and frequency-domain features for each short-time speech frame, wherein the time-domain features include short-time energy, short-time average zero-crossing rate and fundamental frequency, and the frequency-domain features include Mel frequency cepstral coefficients, spectral centroid and spectral roll-off point, and combining the extracted time-domain features and frequency-domain features to form a voiceprint feature vector.
[0118] In some embodiments, the construction of a benchmark model for healthy individuals using voiceprint feature values includes: performing statistical analysis on the voiceprint feature values of healthy individuals, calculating the mean, variance, and probability distribution of each voiceprint feature value, establishing a multidimensional Gaussian mixture model of voiceprint features of healthy individuals based on the statistical analysis results, and using the Gaussian mixture model as a benchmark model for healthy individuals to characterize the normal distribution range of voiceprint features of healthy individuals.
[0119] In some embodiments, the step of comparing the voiceprint feature values of a specific system or disease population with a baseline model of a healthy population, and marking feature data related to the human physiological system or disease based on the comparison results, includes: calculating the degree of difference between the voiceprint feature values of a specific system or disease population and the corresponding voiceprint feature values in the baseline model of a healthy population, wherein the degree of difference is the ratio of the absolute difference of the voiceprint feature values to the standard deviation of the voiceprint feature values of the healthy population, and when the degree of difference is greater than a preset difference threshold, the corresponding voiceprint feature value is marked as feature data related to the human physiological system or disease.
[0120] In some embodiments, the step of drawing a physiological voiceprint block map based on labeled feature data includes: visually mapping the labeled feature data in a multidimensional feature space according to the human physiological system or disease category associated with the labeled feature data; dividing different block areas in the visualization interface according to the physiological system or disease category; displaying the distribution of feature data related to the physiological system or disease in each block area; and forming a physiological voiceprint block map that reflects the relationship between the feature data and the physiological system or disease.
[0121] In some embodiments, the clustering analysis of feature data in the physiological voiceprint block map to form a feature data group that is explicit and stable and associated with human physiological systems or diseases includes: using a density clustering algorithm to cluster the feature data in the physiological voiceprint block map; determining cluster centers based on the distribution density of feature data in a multidimensional feature space; dividing the density-reachable feature data into the same cluster, with each cluster corresponding to a human physiological system or disease category; and iteratively optimizing the cluster centers and cluster division to ensure that the feature data within each cluster is highly similar and that the feature data between different clusters are significantly different, thus forming an explicit and stable feature data group.
[0122] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the steps of the voiceprint feature value clustering method for characterizing human health status as provided in any embodiment of this application.
[0123] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0124] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for clustering voiceprint feature values to characterize human health status, characterized in that, include: Acquire voiceprint data from different groups of people, including healthy people, specific groups, or people with diseases; The acquired voiceprint data is preprocessed; Voiceprint feature values are extracted from the preprocessed speech voiceprint data to represent the speech voiceprint; A baseline model of healthy individuals is constructed using the voiceprint feature values of healthy individuals; the voiceprint feature values of individuals with specific systems or diseases are compared with the baseline model of healthy individuals, and feature data related to human physiological systems or diseases are marked based on the comparison results. Physiological voiceprint block maps are drawn based on labeled feature data, and cluster analysis is performed on the feature data in the physiological voiceprint block maps to form explicit and stable feature data groups that are associated with human physiological systems or diseases.
2. The method according to claim 1, characterized in that, Following the group of explicit and stable feature data associated with human physiological systems or diseases, the following is also included: The resulting feature data set is validated, and the validated feature data set is used for non-invasive, quantitative characterization of human health status, as well as for disease-aided diagnosis and health monitoring.
3. The method according to claim 2, characterized in that, The verification of the formed feature data group includes: The feature data group is verified using reserved test voiceprint data. The clustering results of the test data are compared with known human health status or disease diagnosis results. The accuracy, recall and specificity of the clustering results are calculated. If the accuracy, recall and specificity all reach the preset threshold, the feature data group is determined to have passed the verification. Otherwise, the parameters or steps of the clustering analysis are readjusted until the verification is passed.
4. The method according to claim 1, characterized in that, The preprocessing of the acquired voiceprint data includes: The speech activity detection algorithm removes silent segments and non-speech signals from the speech voiceprint data. A bandpass filter is used to denoise the speech voiceprint data to remove environmental noise and high-frequency and low-frequency interference signals. The volume of the denoised speech voiceprint data is then normalized to ensure that the volume amplitude of different speech voiceprint data is within the same numerical range.
5. The method according to claim 1, characterized in that, The extraction of voiceprint feature values from the preprocessed speech voiceprint data includes: The preprocessed speech voiceprint data is segmented into multiple short-time speech frames. Time-domain features and frequency-domain features are extracted for each short-time speech frame. The time-domain features include short-time energy, short-time average zero-crossing rate, and fundamental frequency. The frequency-domain features include Mel-frequency cepstral coefficients, spectral centroid, and spectral roll-off point. The extracted time-domain features and frequency-domain features are combined to form a voiceprint feature vector.
6. The method according to claim 1, characterized in that, The construction of a baseline model for healthy individuals using voiceprint feature values includes: Statistical analysis was performed on the voiceprint feature values of healthy individuals, and the mean, variance, and probability distribution of each voiceprint feature value were calculated. Based on the statistical analysis results, a multidimensional Gaussian mixture model of the voiceprint features of healthy individuals was established. The Gaussian mixture model was used as a benchmark model for healthy individuals to characterize the normal distribution range of voiceprint features of healthy individuals.
7. The method according to claim 1, characterized in that, The process involves comparing the voiceprint feature values of a specific system or disease population with a baseline model of a healthy population, and marking feature data related to the human physiological system or disease based on the comparison results, including: The difference between the voiceprint feature values of a specific system or disease population and the corresponding voiceprint feature values in the baseline model of healthy population is calculated. The difference is the ratio of the absolute difference of the voiceprint feature values to the standard deviation of the voiceprint feature values of healthy population. When the difference is greater than a preset difference threshold, the corresponding voiceprint feature values are marked as feature data related to human physiological system or disease.
8. The method according to claim 1, characterized in that, The process of drawing physiological voiceprint block maps based on labeled feature data includes: Based on the human physiological system or disease category associated with the labeled feature data, the labeled feature data is visualized and mapped in a multi-dimensional feature space. In the visualization interface, different block areas are divided according to the physiological system or disease category. Each block area displays the distribution of feature data related to that physiological system or disease, forming a physiological voiceprint block map that reflects the relationship between feature data and physiological system or disease.
9. The method according to claim 1, characterized in that, The clustering analysis of feature data in the physiological voiceprint block map forms a group of explicit and stable feature data associated with the human physiological system or disease, including: Density clustering algorithm is used to cluster feature data in physiological voiceprint block maps. Cluster centers are determined based on the distribution density of feature data in multidimensional feature space. Feature data with achievable density are divided into the same cluster, and each cluster corresponds to a human physiological system or disease category. By iteratively optimizing the cluster centers and cluster division, the feature data within each cluster are highly similar, and the feature data between different clusters are significantly different, forming an explicit and stable feature data group.
10. A voiceprint feature value clustering system for characterizing human health status, characterized in that, The method applied to any one of claims 1-9 includes: The data acquisition unit is used to acquire voiceprint data of different groups of people, including healthy people, specific groups, or people with diseases; preprocess the acquired voiceprint data; and extract voiceprint feature values from the preprocessed voiceprint data to characterize the voiceprint. The data comparison unit is used to construct a baseline model of healthy people using the voiceprint feature values of healthy people; compare the voiceprint feature values of specific systems or disease groups with the baseline model of healthy people, and mark the feature data related to human physiological systems or diseases based on the comparison results. The block drawing unit is used to draw physiological voiceprint block maps based on labeled feature data, and to perform cluster analysis on the feature data in the physiological voiceprint block maps to form explicit and stable feature data groups that are associated with human physiological systems or diseases.
Citation Information
Patent Citations
Abnormal sound diagnostic device
CN103325387A
Blade icing identification method and system based on conditional variation auto-encoder
CN117469105A
Cognitive disease intelligent risk management method based on multi-mode voiceprint data analysis
CN120452481A
Voice voiceprint frequency division acquisition and detail extraction method and system
CN121938417A
Information processing device, program, and information processing method
WO2020084680A1