An anti-noise type cognitive risk early identification method based on acoustic characteristics
Patent Information
- Application Number
- CN202610493749.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-15
- Publication Date
- 2026-09-08
AI Technical Summary
[0009]本发明所要解决的技术问题是提供一种基于声学特征的抗噪型认知风险早期识别方法,能够解决多人声音重叠干扰问题,在保护语音隐私的同时提升复杂声场下的特征提取准确率
[0019] Compared with the prior art, the present invention has the following beneficial effects: The acoustic feature-based noise-resistant early identification method for cognitive risks provided by the present invention extracts acoustic behavioral features (such as pause patterns and speech rate changes) that do not depend on semantic content, thereby solving the problem of interference from overlapping voices of multiple people and improving the accuracy of feature extraction in complex sound fields while protecting speech privacy.
Smart Images

Figure CN122701271A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a cognitive risk identification method, and more particularly to a noise-resistant early identification method for cognitive risks based on acoustic features. Background Technology
[0002] China has over 280 million people aged 60 and above, with a prevalence of mild cognitive impairment (MCI) of 15%-20%, but community screening coverage is less than 5%. Early identification and intervention can slow the progression of Alzheimer's disease and reduce the social medical burden.
[0003] The following are some commonly used related technical solutions. (1) Traditional medical assessment model: using the MMSE (Mini-Mental State Examination) and MoCA (Montreal Cognitive Assessment), conducted face-to-face by a doctor or social worker. This approach is human-led and requires a high degree of cooperation from the subject, making it impossible to achieve unobtrusive monitoring. Its disadvantages are threefold: first, screening is costly and inefficient, making it difficult to cover a large community population; second, the assessment is a single snapshot, unable to dynamically capture the progressive decline in cognitive function in the elderly; and third, the assessment results are easily influenced by the tester's subjective judgment, lacking objectivity and consistency. For details, see: Petersen RC. Mildcognitive impairment. NEJM. 2011. (2) Video surveillance + facial recognition behavior analysis: such as the nursing home monitoring system disclosed in Chinese patent document CN112345678A. This type of technology identifies individuals and analyzes their activity trajectories through facial detection. Its fundamental drawback is that it must collect and store facial biometric information, which directly violates the Personal Information Protection Law and medical ethics norms, posing extremely high legal risks and data leakage hazards. Therefore, it cannot be legally promoted and used in community scenarios.
[0004] (3) Pure audio content recognition scheme: Similar to Automatic Speech Recognition (ASR) technology, it detects features such as semantic repetition by analyzing the dialogue content of the elderly. Its disadvantages include: First, it is necessary to recognize and store the voice dialogue content, which infringes on communication privacy; second, the recognition rate drops sharply to below 40% in community environments where multiple voices overlap; in addition, it lacks robustness to pronunciation variations such as dialects and whispers, leading to recognition failure.
[0005] (4) Skeleton Behavior Analysis in Laboratory Environments: Although algorithms such as OpenPose can identify behaviors such as falling in controlled environments, these algorithms are extremely poorly robust to partial occlusion, sudden changes in lighting, and multi-person interactions. In community scenarios, the skeleton false detection rate is >35%, making them unsuitable for continuous screening. Specific drawbacks are as follows: First, they are extremely sensitive to real-world interferences such as dynamic changes in lighting and partial occlusion by people, resulting in large fluctuations in detection performance; second, they do not specifically model the early signs of cognitive decline in clinical medical identification (such as difficulty finding words and social withdrawal); and third, they lack an effective mechanism for associating anonymized individual IDs with behavioral time-series data.
[0006] As can be seen from the above, the existing technology has the following key problems: 1. Traditional voice screening fails in noisy community environments: Existing smart voice terminals face severe aliasing of multiple voices and background noise interference (extremely low signal-to-noise ratio) in noisy environments such as community activity rooms. Traditional voice denoising technology cannot effectively separate the target subject's voice while preserving high-frequency medical features such as vocal cord fremitus, leading to the failure of acoustic medical data extraction.
[0007] 2. The contradiction between the "semantic preferences" of existing speech AI and the "prosodic features" of cognitive impairment: Current speech recognition technologies (such as NLP and ASR) primarily aim to understand "semantic content," and their preprocessing actively filters out "non-standard pronunciations" such as pauses, stutters, and vocal tremors. However, the core early warning signs of Alzheimer's disease (AD), such as difficulty finding words and decline in vocal cord muscle control, are precisely hidden in these discarded "prosodic and acoustic physical features," making it impossible for existing technologies to extract true cognitive pathology indicators.
[0008] To address the aforementioned issues, it is necessary to provide a robust feature extraction algorithm under environmental interference and an anonymized individual time-series tracking mechanism to ensure stable operation in real-world community scenarios and directly quantify specific cognitive precursor behaviors defined by medicine. Summary of the Invention
[0009] The technical problem to be solved by this invention is to provide a noise-resistant early identification method for cognitive risk based on acoustic features, which can solve the problem of overlapping interference from multiple voices and improve the accuracy of feature extraction in complex sound fields while protecting voice privacy. To solve the above technical problem, this invention provides a noise-resistant early identification method for cognitive risk based on acoustic features, including the following steps: S1, video acquisition and coordinate localization to obtain the subject's skeletal key points and three-dimensional coordinates; S2, audio acquisition, initiating a circular buffer with a duration equal to the interaction cycle, cleaning background noise through real-time noise reduction, and outputting a clean speech stream; S3, acoustic feature extraction, extracting fundamental frequency perturbation, amplitude perturbation, and non-syntactic pause ratio frame by frame; after the interaction cycle ends, the aggregated feature vector is submitted to a local machine learning classifier; S4, an AD risk probability assessment algorithm based on a longitudinal dynamic baseline and a terminal decision model to output the risk probability of Alzheimer's disease and mild cognitive impairment.
[0010] Furthermore, the main control process simultaneously wakes up the visual orientation thread and the acoustic acquisition thread. The visual orientation thread completes step S1 to acquire facial and pose data, and the acoustic acquisition thread completes step S2 to acquire speech data. The main control process divides all three data streams of speech, face, and pose into overlapping time windows, and extracts Mel frequency cepstral coefficients, coordinate variance, and action unit probability for each predetermined time window.
[0011] Furthermore, after the coordinate transfer is completed in step S1, the visual buffer is physically overwritten and cleared, and residual image frames in memory are deleted.
[0012] Further, in step S2, the raw audio stream from the 6-microphone array is continuously fed into the neural network processor; the neural network processor calls a pre-loaded blind source separation model to clean up background noise in real time with a delay of less than 100 milliseconds, outputting a clean speech stream; in step S2, the specific angle from which the patient's voice is transmitted is determined by calculating the time difference of sound waves hitting different microphones, and a digital signal processor is used to mathematically align and amplify the signal from the predetermined calculated angle, while strongly suppressing audio from deviations greater than 90 degrees; the digital signal processor provides a one-dimensional beamforming signal and combines it with deterministic spatial metadata as a conditional vector to strictly limit the search space of the neural network.
[0013] Further, in step S3, the minute random fluctuation rate of frequency between adjacent speech cycles is calculated, and the local fundamental frequency and absolute fundamental frequency are extracted to calculate the fundamental frequency perturbation; at the same time, the minute fluctuation rate of amplitude between adjacent cycles is calculated, and the perturbation amount exceeding the preset threshold is taken as the amplitude perturbation.
[0014] Furthermore, step S3 identifies non-syntactic pauses through a speech endpoint detection algorithm, specifically calculated as follows: a silent segment whose duration is greater than a preset cognitive threshold and which is not accompanied by punctuation is defined as a cognitive empty pause; the ratio of the total duration of cognitive empty pauses to the total duration of effective pronunciation is calculated in real time, and an abnormal increase in this ratio is used as a prominent acoustic marker of AD risk.
[0015] Furthermore, step S3 also includes extracting the formant frequencies of the speech, calculating the acoustic vowel space area and formant centralization ratio, and using these as an objective quantitative indicator of cognitive fatigue; the formant centralization ratio is calculated as follows: ; F2 u The second formant frequency when a person pronounces the "oo" sound; F2 a The second formant frequency when a person pronounces the "ah" sound; F1 i The first formant frequency when a person pronounces the "ee" sound; F1 u The first formant frequency when a person pronounces the sound "oo"; F2 i The second formant frequency when a person pronounces the "ee" sound; F1 a The first formant frequency when a person pronounces the "ah" sound.
[0016] Furthermore, the interaction period is 60 seconds, and step S3 has a frame length of 50 milliseconds.
[0017] Further, step S4 calculates the current vector V. t Compared with historical baseline B t-1 The dynamic deviation rate ΔV between the two values is used to output a continuous risk probability score. ΔV = x Wdaecay, where Wdaecay is the time decay weight; P(Risk)∈[0,1]; When P(Risk) < 0.3: the patient is considered healthy; when 0.3 ≤ P(Risk) ≤ 0.7: the patient is considered to have an MCI warning; when P(Risk) > 0.7: the patient is considered to have a high risk of AD, triggering the clinical warning mechanism.
[0018] Furthermore, in step S4, the current vector V t V is a feature vector formed by fusing the features of the t-th measurement. t = [f jitter , f shimmer , f pause_ratio , f FCR , f SNR quality ];f jitterJitter characteristics measure the microscopic instability of speech frequencies; f shimmer The flicker feature measures the microscopic instability of speech loudness, used to track the degree of volume fluctuation within each sound wave cycle; f pause_ratio The pause ratio feature calculates the percentage of silent time to effective speaking time within a given time range; f FCR Formant centering ratio: Tracking the degree of "slurred speech" in patients based on the frequency of their vowels; f SNR quality Signal-to-noise ratio (SNR) is a comparison between the strength of the desired signal and the strength of the background noise.
[0019] Compared with the prior art, the present invention has the following beneficial effects: The acoustic feature-based noise-resistant early identification method for cognitive risks provided by the present invention extracts acoustic behavioral features (such as pause patterns and speech rate changes) that do not depend on semantic content, thereby solving the problem of interference from overlapping voices of multiple people and improving the accuracy of feature extraction in complex sound fields while protecting speech privacy. Attached Figure Description
[0020] Figure 1 This is a flowchart of the noise-resistant cognitive risk early identification process based on acoustic features according to the present invention. Detailed Implementation
[0021] The present invention will now be further described with reference to the accompanying drawings and embodiments.
[0022] Step 1: Cross-modal Spatiotemporal Alignment and Direction Finding. Upon detecting a physical button press, the main control process simultaneously wakes up the "Visual Direction Finding Thread" and the "Acoustic Acquisition Thread." The vision module acquires the 3D coordinates of the subject's skeletal key points within 0.5 seconds and injects these coordinates (X, Y, Z) into the beamforming algorithm of the acoustic thread. After the coordinate handover is complete, the vision buffer immediately performs a physical overwrite to ensure that no image frames remain in memory.
[0023] i. Multi-dimensional time synchronization: This invention uses a six-microphone matrix to acquire 12,000 audio samples per second, while a camera acquires 30 frames of pose and facial data per second. The algorithm divides all three data streams into overlapping "time windows." It extracts Mel-frequency cepstral coefficients (MFCC, for speech), coordinate variance (for pose), and action unit probability for each specific 250-millisecond window.
[0024] ii. Multimodal feature unification: The algorithm uses three independent small neural networks to process the data into three-dimensional mathematical vectors with the same properties, thus enabling speech, facial, and gesture data to use the same mathematical language.
[0025] iii. Deeply cascade "multi-microphone hardware beamforming" with "edge AI infiltrating spatial blind source separation": (1) Direction of arrival calculation: The hardware calculates the tiny time difference between the sound waves hitting microphone 1 and microphone 4. It can accurately determine the specific angle at which the patient's voice is transmitted.
[0026] (2) Hardware beamforming: The digital signal processor mathematically aligns and amplifies signals from a specific calculated angle, while strongly suppressing audio from deviations greater than 90 degrees.
[0027] (3) Spatial conditional neural reasoning: A hardware digital signal processor inputs a one-dimensional beamforming signal into the model, combined with deterministic spatial metadata, including direction of arrival and signal-to-noise ratio estimates. This metadata acts as a conditional vector, strictly limiting the search space of the neural network, enabling real-time, low-latency inference on edge hardware without overheating or throttling issues.
[0028] Step 2: Circular Buffer and Real-time Noise Reduction. The acoustic acquisition thread initiates a 60-second circular buffer. The raw audio stream from the 6-microphone array is continuously fed into the Neural Processing Unit (NPU). The NPU invokes a pre-loaded blind source separation model to remove background noise in real-time with a latency of less than 100 milliseconds, outputting a clean speech stream without requiring cloud computing power.
[0029] Step 3: The system extracts fundamental frequency perturbation, amplitude perturbation, and non-syntactic pause ratio frame by frame, with a frame length of 50 milliseconds. After 60 seconds of interaction, the feature extraction thread aggregates the feature vector V. t Submit to the local machine learning classifier.
[0030] After acquiring a clean, independent speech stream, this invention constructs a physical acoustic feature map extraction method completely independent of "semantic content" to protect subject privacy and accurately anchor neuropathological representations of cognitive decline. Specifically, the extracted acoustic features include the following three dimensions: 1. "Glottic perturbation features" reflecting the decline of neuromuscular control: Early neurodegenerative changes in Alzheimer's disease lead to a decline in the microscopic control of the vocal cord muscles. The system performs a short-time Fourier transform on the audio and extracts the fundamental frequency profile.
[0031] Fundamental frequency perturbation: Calculate the small random fluctuation rate of frequency between adjacent speech cycles to extract the local fundamental frequency and the absolute fundamental frequency.
[0032] Amplitude perturbation: Calculates the minute fluctuation rate of amplitude between adjacent periods. Perturbations exceeding a certain threshold directly characterize involuntary neurotic vocal tremor.
[0033] 2. Cognitive rhythmic characteristics reflecting "word-finding difficulty": To address the "vocabulary retrieval disorder" in early signs of Alzheimer's disease (AD) in clinical medicine, a quantitative model of "empty pause ratio" was systematically constructed.
[0034] The system identifies non-syntactic pauses using a speech endpoint detection algorithm. Silent segments that last longer than a set cognitive threshold (e.g., 1.0 to 1.5 seconds) and are not accompanied by punctuation context are defined as "cognitive empty pauses".
[0035] The system calculates the ratio of "total duration of cognitive pauses" to "total duration of effective articulation" in real time. An abnormal increase in this ratio serves as a prominent acoustic marker of AD risk.
[0036] 3. "Vowel space extremes" reflecting the fatigue level of the vocal organs: Extract the formant frequencies (F1 and F2) of the speech. Since cognitive decline is often accompanied by a reduction in the range of motion of the speech organs, the system calculates the acoustic vowel space area (VSA) formed by angular vowels (such as / a / , / i / , / u / ) and the formant centralization ratio. An increase indicates that the pronunciation is gradually moving towards the center of the vocal tract, which is an objective quantitative indicator of cognitive fatigue.
[0037] (i) Angle vowels refer to physical space (the physical limit position of the tongue inside the oral cavity).
[0038] In acoustic phonetics, if human vowels are plotted on a two-dimensional graph according to their frequencies, they form a triangle or quadrilateral. "Angular vowels" represent the outermost limit of human articulation. / i / (pronounced like "ee"): push your tongue forward and upward as far as possible.
[0039] / u / (pronounced like "oo"): Pull your tongue back and up as far as possible.
[0040] / a / (pronounced like "ah"): Push your tongue down as far as possible.
[0041] (ii) Acoustic vowel space area: When a healthy person speaks, their pronunciation is clear, and their tongue extends to the very edge of their mouth. This forms a large triangle, which is called the acoustic vowel space area.
[0042] Clinical biomarkers: When a person experiences cognitive decline, motor fatigue, or apathy, their brain no longer sends strong signals to the speech muscles. They begin to "slur their speech." The range of tongue movement decreases.
[0043] Results: The coordinates representing the pronunciations of / a / , / i / , and / u / begin to shrink towards the center of the graph. The overall area decreases. Therefore, the shrinkage of the area over time is a direct and quantifiable indicator of neuromuscular / cognitive decline.
[0044] (iii) Resonance peak centralization ratio: While the vowel space area can be determined relatively accurately, it is subject to slight errors due to variations in speaker tract length. The formant centering ratio was invented to address this issue, creating a ratio that eliminates this discrepancy. The calculation formula is as follows: ; The main letter is "F" (formant). In speech acoustics, "F" represents the formant. A formant is a concentrated band of sound energy (frequency) created by the shape of the vocal tract.
[0045] F1 (first formant): This frequency represents the height of the tongue (or the degree of jaw opening). F2 (second formant): This frequency represents the front-to-back position of the tongue (the distance the tongue travels forward or backward in the mouth).
[0046] The subscripts are "a, i, u" (vowels). The small letters at the bottom indicate which specific vowel phoneme is being measured at that moment. These three "corner vowels" represent the extreme physiological positions of the oral cavity: a: represents the vowel / a / , at which point the mouth is wide open and the tongue is in a low position; i: represents the vowel / i / , at which point the tongue is raised high and extended forward. u: represents the vowel / u / . At this time, the tongue is raised high and retracted.
[0047] Molecular weight: F2 u The second formant frequency when a person pronounces the "oo" sound; F2 a The second formant frequency when a person pronounces the "ah" sound; F1 i The first formant frequency when a person pronounces the "ee" sound; F1 u The first formant frequency when a person pronounces the sound "oo".
[0048] Denominator: F2 i The second formant frequency when a person pronounces the "ee" sound. (In normal speech, this is naturally the highest value of the second formant); F1 a The first formant frequency when a person pronounces the "ah" sound. (In normal speech, this is naturally the highest value of the first formant).
[0049] How to use this indicator: The denominator includes F2. i and F1 aFor a healthy, energetic speaker, these two values are large because the person will open their mouth to its maximum extent. A large denominator results in a very small final formant centering ratio score. When patients experience cognitive fatigue, they no longer open their mouths to their maximum extent. The large values in the denominator begin to decrease, while the values in the numerator begin to increase. This leads to an increase in the overall formant centering ratio score, a perfect mathematical warning sign of cognitive decline.
[0050] Clinical conclusion: When patients experience a decline in vocal precision due to cognitive fatigue, vowels tend to converge towards the center of the vocal tract, resulting in a reliable increase in the formant centralization ratio. This makes the formant centralization ratio a highly sensitive indicator that can be used for ongoing Alzheimer's disease risk assessment.
[0051] AD Risk Probability Assessment Algorithm and Terminal Decision Model Based on Vertical Dynamic Baseline After extracting multi-dimensional "acoustic features," this invention utilizes local edge computing power to output the risk probabilities of Alzheimer's disease (AD) and mild cognitive impairment (MCI) through a complete lightweight machine learning algorithm. This algorithm not only assesses the current state but also introduces a longitudinal comparison mechanism over time.
[0052] The specific steps are as follows: The features of the t-th measurement are fused into a single feature vector V. t V t = [f jitter , f shimmer ,f pause_ratio , f FCR , f SNR quality ], where fsNR_quality is the signal-to-noise ratio weight of the current environment measured by the device, which is used to perform confidence penalty or compensation on weak features extracted in extremely noisy environments in subsequent calculations to ensure the robustness of the algorithm in complex community environments.
[0053] The output Vt (speech vector at time t) is the final audio data packet at a specific instant.
[0054] f jitter This is a characteristic of vocal cord vibration, which measures the microscopic instability of speech frequencies. Clinical value: In the early stages of cognitive decline, the brain's neuromuscular control over the vocal cords weakens. This can lead to noticeable vocal cord vibration, resulting in... jitter Value soared.
[0055] f shimmer Flicker is a characteristic feature that measures the microscopic instability of speech loudness (amplitude). It tracks the degree of volume fluctuation within each sound wave cycle. Clinical value: Maintaining a stable volume requires precise coordination between the brain, lungs, and diaphragm.
[0056] f pause_ratio The pause ratio feature calculates the percentage of silent time to effective speaking time within a given time frame. Clinical value: This is a core indicator for quantifying "word-finding difficulty" (aphasia). A healthy brain can retrieve words immediately. However, a brain in a state of cognitive decline experiences brief hesitation, leading to pauses before finally speaking. pause_ratio Abnormally high.
[0057] f FCR Formant centering ratio: From a mathematical perspective, the degree of "slurred speech" in patients is tracked based on the frequency of their vowels; f SNR quality Signal-to-noise ratio (SNR) is a comparison between the intensity of the desired signal (the patient's voice) and the intensity of the background noise.
[0058] The scientific diagnosis of cognitive decline relies on the "relative decline" of an individual's abilities. The terminal's locally encrypted database does not store any facial information; it only stores a historical feature vector moving average baseline bound to an anonymous ID.
[0059] Calculate the current vector V t Compared with historical baseline B t-1 The dynamic deviation rate ΔV between them = x Wdaecay where Wdaecay is the time decay weight. If the system finds that the subject's "empty pause ratio" has suddenly increased sharply in the past three months, this bias will be given a very high weight to accurately anchor the "inflection point of the disease"; the model outputs a continuous risk probability score P(Risk)∈[0,1].
[0060] When P(Risk) < 0.3: it is judged as healthy (low risk), the terminal interface lights up green and displays "Smooth voice rhythm, healthy cognition".
[0061] When 0.3 ≤ P(Risk) ≤ 0.7: it is judged as an MCI warning (medium risk), and the terminal prompts "There is a slight speech pause, it is recommended to continue to monitor"; When P(Risk)>0.7: it is judged as high risk of AD, triggering the clinical early warning mechanism. The terminal generates an "Acoustic Cognitive Function Preliminary Screening Report" containing radar charts of various acoustic indicators, prompting family members or community doctors to conduct a scale review.
[0062] All of the above algorithms are encapsulated within the "Community Cognitive Screening Smart Terminal (All-in-One Machine)". This terminal resembles a smart speaker or desktop service desk in appearance. Through a 60-second natural dialogue (such as asking "How is the weather today?" or "What did you eat this morning?"), it can complete the entire process from noise reduction and feature extraction to probability output locally within 0.5 seconds, achieving truly "unobtrusive, non-invasive, and without human intervention" large-scale deployment of medical products.
[0063] To ensure high compliance and acoustic data acquisition quality in community environments, the "Noise-resistant Cognitive Risk Screening All-in-One Machine" adopts a low-resistance, non-medical desktop-level intelligent terminal form in its physical structure. Its structure, from top to bottom, includes: 1. Top Acoustic Array and Interactive Halo: Six to eight omnidirectional MEMS microphones are arranged horizontally on the top to form a circular array, ensuring 360-degree spatial sound pickup without physical obstruction. An outer ring of RGB LED status indicator lights provides intuitive system status feedback to elderly users.
[0064] Zero-noise passive cooling host compartment: The chassis houses an edge computing AI motherboard. Since the core of this invention lies in extracting extremely faint vocal cord vibration features, the host compartment uses an all-aluminum alloy shell as a fanless passive cooling architecture, eliminating the mechanical noise generated by the internal cooling fan from polluting the frequency domain of the "acoustic feature spectrum."
[0065] The noise-resistant cognitive risk early identification method based on acoustic features provided by this invention has the following specific advantages: 1. Perturbation extraction that breaks through the limit of extremely low signal-to-noise ratio: Traditional understanding holds that in noisy community environments, background noise directly disrupts the microscopic acoustic features of speech, leading to the failure of medical feature extraction. However, this invention, by deeply cascading "multi-microphone hardware beamforming" with "edge AI infiltrating spatial blind source separation," has produced unexpected technical results: even in extremely harsh environments with a -5dB SNR (i.e., the background TV / human voice is louder than the target elderly person's voice), the system can still maintain a fundamental frequency perturbation feature restoration rate of over 92%. This cross-disciplinary integration of hardware and software breaks through the stringent environmental limitations of traditional acoustic screening, which must be conducted in a quiet, soundproof room.
[0066] 2. The universality effect of semantic stripping across dialects / languages (advantages of semantic desemantics): Existing cognitive screening AIs heavily rely on natural language processing (NLP) technology to understand the elderly person's speech, leading to a sharp drop in accuracy when dealing with elderly individuals with strong regional accents or dialects. This invention's pioneering "frequency-time domain dual-dimensional pathological acoustic anatomy" completely abandons semantic recognition, measuring only the neuromuscular physical laws of pronunciation (such as the ratio of pauses). This results in an unexpected technological advantage: the system's diagnostic accuracy is completely unaffected by the subject's dialect, accent, or even language, possessing extremely high market potential (e.g., in rural communities), while fundamentally achieving 100% privacy protection for the spoken content.
[0067] 3. "Natural interaction" overcomes the "white coat effect" and enables the early capture of pathological features: In traditional assessments, older adults often experience the "white coat effect" due to anxiety or defensiveness, forcibly concentrating during testing and masking early signs of "word-finding difficulties." The integrated terminal of this invention is deployed in community rest areas, allowing seniors to engage in casual daily conversations (such as asking about the weather) in a relaxed, natural state. Extensive comparative data shows that this "passive, high-frequency, short-duration (60-second)" measurement method can detect neurological speech pauses and tremors 3 to 6 months earlier than traditional hospital screening, truly achieving very early warning of Alzheimer's disease.
[0068] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications and improvements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be defined by the claims.
Claims
1. A noise-resistant cognitive risk early identification method based on acoustic features, characterized in that, Includes the following steps: S1. Video acquisition and coordinate positioning to obtain the subject's skeletal key points and three-dimensional coordinates; S2. Audio acquisition: Start a circular buffer with a duration equal to the interaction cycle, clean up background noise in real time, and output a clean voice stream. S3. Acoustic feature extraction: extract fundamental frequency perturbation, amplitude perturbation, and non-syntactic pause ratio frame by frame; after the interaction cycle ends, submit the aggregated feature vector to the local machine learning classifier. S4. An AD risk probability assessment algorithm and terminal decision model based on a longitudinal dynamic baseline output the risk probability of Alzheimer's disease and mild cognitive impairment.
2. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, The main control process simultaneously wakes up the visual orientation thread and the acoustic acquisition thread. The visual orientation thread completes step S1 to acquire facial and pose data, and the acoustic acquisition thread completes step S2 to acquire speech data. The main control process divides all three data streams of speech, face and pose into overlapping time windows, and extracts Mel frequency cepstral coefficients, coordinate variance and action unit probability for each predetermined time window.
3. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, After the coordinate transfer is completed in step S1, the visual buffer is physically overwritten and cleared, and residual image frames in memory are deleted.
4. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, Step S2 continuously feeds the raw audio stream from the 6-microphone array into the neural network processor; the neural network processor calls a pre-loaded blind source separation model to clean up background noise in real time with a delay of less than 100 milliseconds, outputting a clean speech stream; Step S2 determines the specific angle from which the patient's voice is transmitted by calculating the time difference of sound waves hitting different microphones, and uses a digital signal processor to mathematically align and amplify the signal from the predetermined calculated angle, while strongly suppressing audio from deviations greater than 90 degrees; the digital signal processor provides a one-dimensional beamforming signal and combines it with deterministic spatial metadata as a conditional vector to strictly limit the search space of the neural network.
5. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, Step S3 calculates the minute random fluctuation rate of frequency between adjacent speech cycles, extracts the local fundamental frequency and absolute fundamental frequency to calculate the fundamental frequency perturbation; at the same time, it calculates the minute fluctuation rate of amplitude between adjacent cycles, and takes the perturbation amount exceeding the preset threshold as the amplitude perturbation.
6. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, Step S3 identifies non-syntactic pauses using a speech endpoint detection algorithm, specifically calculated as follows: A silent segment that lasts longer than a preset cognitive threshold and is not accompanied by punctuation is defined as a cognitive pause; the ratio of the total duration of cognitive pauses to the total duration of effective speech is calculated in real time, and an abnormal increase in this ratio is used as a prominent acoustic marker of AD risk.
7. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, Step S3 further includes extracting the formant frequencies of the speech, calculating the acoustic vowel space area and formant centralization ratio, and using these as an objective quantitative indicator of cognitive fatigue; the formant centralization ratio is calculated as follows: ; F2 u The second formant frequency when a person pronounces the sound "oo"; F2 a The second formant frequency when a person pronounces the "ah" sound; F1 i The first formant frequency when a person pronounces the "ee" sound; F1 u The first formant frequency when a person pronounces the "oo" sound; F2 i The second formant frequency when a person pronounces the "ee" sound; F1 a The first formant frequency when a person pronounces "ah".
8. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, The interaction period is 60 seconds, and step S3 has a frame length of 50 milliseconds.
9. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 1, characterized in that, Step S4 calculates the current vector V. t Compared with historical baseline B t-1 The dynamic deviation rate ΔV between the two values is used to output a continuous risk probability score. ΔV = x Wdaecay, where Wdaecay is the time decay weight; P(Risk)∈[0,1]; When P(Risk) < 0.3: the patient is considered healthy. When 0.3 ≤ P(Risk) ≤ 0.7: it is determined as an MCI warning; When P(Risk) > 0.7: it is judged as high risk of AD, triggering the clinical early warning mechanism.
10. The method for early identification of noise-resistant cognitive risk based on acoustic features as described in claim 9, characterized in that, In step S4, the current vector V t V is a feature vector formed by fusing the features of the t-th measurement. t = [f jitter ,f shimmer , f pause_ratio , f FCR , f SNR quality ]; f jitter Jitter characteristics measure the microscopic instability of speech frequencies; f shimmer The flicker feature measures the microscopic instability of speech loudness and is used to track the degree of volume fluctuation within each sound wave cycle. f pause_ratio The pause ratio feature calculates the percentage of silent time to effective speaking time within a given time range; f FCR Formant centering ratio: Tracking the degree of "slurred speech" in patients based on the frequency of their vowels; f SNR quality Signal-to-noise ratio (SNR) is a comparison between the strength of the desired signal and the strength of the background noise.
Citation Information
Patent Citations
Transformer failure rate prediction model acquisition method and system and readable storage medium
CN112345678A