Intelligent companion robot interaction system based on multi-modal intent recognition

The intelligent companion robot interaction system, which utilizes multimodal intent recognition and synchronous processing of voice stimulation and facial video data, eliminates motion artifacts and enables diagnostic-level physiological sampling without wearing electrodes or remaining stationary, thereby improving the accuracy and reliability of cognitive state assessment.

CN122131922APending Publication Date: 2026-06-02XIAMEN MEIYA ZHONGMIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN MEIYA ZHONGMIN TECH CO LTD
Filing Date
2026-05-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

When existing intelligent companion robots process audio and video data uniformly in natural interaction scenarios, they are prone to introducing motion artifacts caused by vocalization, resulting in insufficient stability of physiological sampling and decreased accuracy of cognitive state assessment, leading to low reliability of non-contact cognitive-assisted diagnosis.

Method used

The intelligent companion robot interaction system adopts multimodal intent recognition. It outputs voice stimulation signals through the active interaction stimulation module, and simultaneously collects audio and facial video data through the multimodal acquisition module. The acoustic gating generation module identifies quasi-static physiological windows and generates visual mask gating signals. The cross-phase-locked masking module removes motion artifacts, the frequency domain reconstruction module performs spectrum reconstruction, and the cognitive state diagnosis module generates diagnostic results.

Benefits of technology

It effectively eliminates motion artifacts caused by vocalization, improves the signal-to-noise ratio of non-contact physiological signals, ensures the accuracy and stability of frequency domain analysis, and enhances the reliability of rapid identification and triage for patients with mild cognitive impairment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131922A_ABST
    Figure CN122131922A_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent robot interaction and multimodal signal processing technology, specifically to an intelligent companion robot interaction system based on multimodal intent recognition. The system includes: an active interaction stimulation module that outputs a voice stimulation signal to trigger the synchronous acquisition of audio data streams and facial video data streams; an acoustic gating generation module that identifies quasi-static physiological windows and generates visual mask gating signals; a cross-phase-locked mask module that locks the facial region of interest and acquires discrete physiological sampling time series; a frequency domain reconstruction module that extracts frequency domain physiological diagnostic indicators; and a cognitive state diagnosis module that generates cognitive state diagnostic results and feeds back updates to subsequent voice stimulation signals. This invention achieves non-contact physiological sampling during natural communication, providing a more stable and objective basis for cognitive state assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent robot interaction and multimodal signal processing technology, specifically to an intelligent companion robot interaction system based on multimodal intent recognition. Background Technology

[0002] In existing cognitive state assisted assessment systems, intelligent companion robots or terminal devices typically include a voice interaction unit, an audio and video acquisition unit, and a state analysis unit. The voice interaction unit outputs prompts to the test subject, the audio and video acquisition unit collects voice data and facial video data during the test subject's response process, and the state analysis unit then judges the test subject's response performance or physiological state based on the acquisition results. In existing solutions, the extraction of physiological information from facial videos often involves directly processing consecutive video frames or analyzing data based on a single modality. Furthermore, during the process of the subject speaking, pausing, hesitating, repeating, or slightly turning their head, the movement of the perioral muscles, facial deformation, environmental noise, and asynchronous timing can all easily interfere with the acquisition results. However, in natural interaction scenarios, directly processing all audio and video data in a unified manner can easily cause motion artifacts caused by vocalization to mix into facial physiological signals, resulting in insufficient physiological sampling stability, frequency domain analysis distortion, and decreased accuracy of cognitive state assessment, thus leading to the problem of low reliability of non-contact cognitive-assisted diagnosis. Summary of the Invention

[0003] The purpose of this invention is to provide an intelligent companion robot interaction system based on multimodal intent recognition, and to solve the following technical problems: This method enables diagnostic-level non-contact physiological sampling without requiring patients to wear electrodes or remain absolutely still. It effectively eliminates motion artifacts caused by vocalization to improve the signal-to-noise ratio of non-contact physiological signals and overcomes the limitation that frequency domain analysis must rely on completely continuous sampling. This provides a more stable and objective basis for cognitive status assessment, enabling rapid identification and triage of cognitive risk subjects such as suspected mild cognitive impairment patients.

[0004] The objective of this invention can be achieved through the following technical solutions: The intelligent companion robot interaction system based on multimodal intent recognition is characterized by including: an active interaction stimulation module, used to output voice stimulation signals to the test object through an audio output device to trigger multimodal data acquisition; The multimodal acquisition module is used to simultaneously acquire the audio data stream and facial video data stream of the tested object in response to the voice excitation signal; and to perform time stamp alignment processing on the audio data stream and the facial video data stream. The acoustic gating generation module is used to extract acoustic features from the audio data stream, identify quasi-static physiological windows, and generate visual mask gating signals based on the quasi-static physiological windows. The cross-phase-locked mask module is used to determine the region of interest (ROI) on the face from the facial video data stream and extract the original image mean time series of the ROI. The visual mask gating signal is multiplied with the original image mean time series by a mask to obtain a discrete physiological sampling time series. The frequency domain reconstruction module is used to input discrete physiological sampling time series into a preset periodogram estimation algorithm for spectrum reconstruction and to extract frequency domain physiological diagnostic indicators. The cognitive state diagnosis module is used to assess the state of the tested object based on frequency domain physiological diagnostic indicators, generate cognitive state diagnosis results, and feed the cognitive state diagnosis results back to the active interaction stimulation module to update the subsequent voice stimulation signals.

[0005] Optionally, the multimodal acquisition module simultaneously acquires the audio data stream and facial video data stream of the test object, specifically for: acquiring acoustic signals from the test object through a directional microphone array to obtain the audio data stream; Visual signals are acquired from the subject using a binocular camera module to obtain facial video data streams. The audio data stream and facial video data stream are time-stamp aligned to obtain the synchronized audio data stream and facial video data stream.

[0006] Optionally, the acoustic gating generation module extracts acoustic features from the audio data stream and identifies quasi-static physiological windows. Specifically, it is used to: perform speech activity detection and short-time Fourier transform processing on the audio data stream to obtain acoustic envelope data. Extracting short-time energy features and zero-crossing rate features from acoustic envelope data; When the short-term energy feature is lower than the preset pause energy threshold, it is determined to be an inter-word pause sequence; when the short-term energy feature is between the pause energy threshold and the preset burst energy threshold, and the zero-crossing rate feature is lower than the preset unvoiced zero-crossing rate threshold, it is determined to be a vowel duration sequence; the vowel duration sequence and the inter-word pause sequence are merged to obtain a quasi-static physiological window.

[0007] Optionally, the acoustic gating generation module generates a visual mask gating signal based on a quasi-static physiological window, specifically used for: performing binarization mapping processing on the quasi-static physiological window to generate an initial Boolean value sequence; Among them, the inter-word pause state and vowel persistence state in the initial Boolean value sequence are mapped to the first logical value, and the state in the initial Boolean value sequence where the short-time energy feature is greater than or equal to the preset burst energy threshold is defined as the violent vocalization state and mapped to the second logical value. Obtain the neuromuscular conduction delay constant pre-calibrated based on the multimodal stimulus response time of historical populations, convert the neuromuscular conduction delay constant into pre-compensation frame number and post-compensation frame number according to the facial video frame rate, perform delay compensation processing on the initial Boolean value sequence based on the pre-compensation frame number and post-compensation frame number, and generate a visual mask gating signal.

[0008] Optionally, the cross-phase-locked masking module determines the facial region of interest from the facial video data stream and extracts the original image mean time series of the facial region of interest, specifically for: Facial feature point tracking processing is performed on the facial video data stream to obtain the facial key point coordinate sequence; based on the facial key point coordinate sequence, the cheek and forehead regions of the tested object are locked, and the cheek and forehead regions are merged into the facial region of interest; Extract the original pixel values ​​from the region of interest on the face, and calculate the spatial mean of the original pixel values ​​to generate the original image mean time series.

[0009] Optionally, the cross-phase-locked masking module performs mask multiplication on the visual mask-gated signal and the original image mean time series to obtain a discrete physiological sampling time series, specifically used for: The visual mask gating signal is used as the prior weight signal; the prior weight signal and the original image mean time series are multiplied frame by frame in the time dimension, and the original image mean time series at the corresponding time is set to zero by the second logic value to remove motion artifact frames caused by violent sounding. The remaining data points after removing motion artifact frames are stitched together to obtain discrete physiological sampling time series.

[0010] Optionally, the frequency domain reconstruction module inputs the discrete physiological sampling time series into a preset periodogram estimation algorithm for spectrum reconstruction, extracting frequency domain physiological diagnostic indicators, specifically used for: The preset periodogram estimation algorithm is specifically a non-uniform sampling spectrum analysis algorithm; the discrete physiological sampling time series is input into the non-uniform sampling spectrum analysis algorithm, and the least squares method is used to fit the discrete data points in the discrete physiological sampling time series with a sine wave to obtain the continuous physiological signal waveform; the energy of the continuous physiological signal waveform is extracted from the preset low frequency band and the preset high frequency band respectively to obtain the sympathetic nerve activity index and the parasympathetic nerve activity index. The ratio of the sympathetic nerve activity index to the parasympathetic nerve activity index is calculated as a frequency domain physiological diagnostic index, wherein the frequency domain physiological diagnostic index includes heart rate variability data.

[0011] Optionally, the cognitive state diagnosis module assesses the state of the tested object based on frequency domain physiological diagnostic indicators and generates cognitive state diagnosis results, specifically used to: obtain a preset population health baseline threshold; The heart rate variability data is compared with the population health baseline threshold; if the heart rate variability data is greater than or equal to the population health baseline threshold, the cognitive state of the tested subject is determined to be normal, and a cognitive state diagnosis result containing a normal state label is generated. If the heart rate variability data is less than the population health baseline threshold, the cognitive state of the tested subject is determined to be in a state of decline, and a cognitive state diagnosis result containing a decline status label is generated.

[0012] Optionally, the active interaction incentive module is specifically used to: obtain the cognitive state diagnosis results and update the active interaction incentive strategy based on the cognitive state diagnosis results; Based on the updated active interaction stimulus strategy, the next round of voice stimulus signal is generated and output to the test object through the audio output device to trigger a new round of multimodal data acquisition.

[0013] Optionally, after outputting the next round of voice stimulation signal to the test object through the audio output device, the method further includes: recording the stimulation timestamp of the output of the next round of voice stimulation signal; Facial feature point tracking processing is performed on the facial video data stream, and the displacement vector of the facial key point coordinate sequence between adjacent facial video frames is calculated. When the magnitude of the displacement vector exceeds the preset micro-movement displacement threshold, the time of the corresponding video frame is recorded as a timestamp of facial micro-expression change. The difference between the excitation timestamp and the facial micro-expression change timestamp is calculated to obtain dynamic neuromuscular conduction delay data; the dynamic neuromuscular conduction delay data is fed back to the acoustic gating generation module to replace the pre-calibrated neuromuscular conduction delay constant for a new round of visual mask gating signal generation.

[0014] The beneficial effects of this invention are: 1) This invention extracts features from audio to identify quasi-static physiological windows and generates visual mask gating signals, thereby guiding the module to accurately remove facial image frames affected by intense vocalization; this mechanism transforms vocalization patterns into visual prior information, filtering out facial deformation interference when patients speak or hesitate from the source, effectively improving the signal-to-noise ratio of non-contact physiological signals; 2) This invention uses a directional microphone array and a binocular camera to synchronously acquire data and align the timestamps, effectively suppressing lateral noise and ensuring spatial and temporal consistency. At the same time, the system calculates the time difference between voice stimulation and facial micro-expressions to achieve self-correction of dynamic neuromuscular conduction delay, accurately adapting to the actual reaction speed of different patients and avoiding sampling deviation caused by fixed delay. 3) This invention addresses the discrete time series formed after artifact removal by introducing a non-uniform sampling spectrum analysis algorithm for processing. This method breaks the limitation of traditional frequency domain analysis which is extremely dependent on continuous video frames. It directly uses high-quality discrete sampling points to successfully fit continuous physiological waveforms, thereby accurately extracting frequency domain indicators such as heart rate variability and effectively preventing spectrum distortion in natural interactions. 4) This invention compares frequency domain diagnostic indicators such as heart rate variability with the population health baseline to generate status assessment results, and uses these results to update subsequent active interaction incentive strategies. This closed-loop mechanism can dynamically adjust the difficulty and speed of voice prompts according to the current cognitive risk status of the tested subject, so that the data collection task and the patient's status are adaptively matched, which greatly improves the reliability of multi-round diagnosis. Attached Figure Description

[0015] The invention will now be further described with reference to the accompanying drawings.

[0016] Figure 1 This is a block diagram of an intelligent companion robot interaction system based on multimodal intent recognition, as described in this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 An intelligent companion robot interaction system based on multimodal intent recognition includes: an active interaction stimulation module, which is used to output voice stimulation signals to the test object through an audio output device to trigger multimodal data acquisition; The multimodal acquisition module is used to simultaneously acquire the audio data stream and facial video data stream of the tested object in response to the voice excitation signal; and to perform time stamp alignment processing on the audio data stream and the facial video data stream. The acoustic gating generation module is used to extract acoustic features from the audio data stream, identify quasi-static physiological windows, and generate visual mask gating signals based on the quasi-static physiological windows. The cross-phase-locked mask module is used to determine the region of interest (ROI) on the face from the facial video data stream and extract the original image mean time series of the ROI. The visual mask gating signal is multiplied with the original image mean time series by a mask to obtain a discrete physiological sampling time series. The frequency domain reconstruction module is used to input discrete physiological sampling time series into a preset periodogram estimation algorithm for spectrum reconstruction and to extract frequency domain physiological diagnostic indicators. The cognitive state diagnosis module is used to assess the state of the tested object based on frequency domain physiological diagnostic indicators, generate cognitive state diagnosis results, and feed the cognitive state diagnosis results back to the active interaction stimulation module to update the subsequent voice stimulation signals.

[0019] This embodiment provides a diagnostic mechanism for an intelligent companion robot interaction system based on multimodal intent recognition; specifically, this embodiment deploys the system in the initial screening and follow-up scenarios of geriatric memory clinics to conduct non-contact cognitive state auxiliary assessments for patients suspected of having mild cognitive impairment. The system can be placed on the desktop of the examination room or next to the bed in the ward. Its core function is not to perform mechanical actions, but to actively provide structured voice prompts through audio output devices after medical staff initiate the examination, guide the subject to complete the response, and simultaneously collect voice and facial video during natural communication, and extract physiological representations related to autonomic nervous system regulation and cognitive load. Specifically, the active interactive stimulation module preferably uses a speaker or directional sound field device to output voice stimulation signals, such as date recall prompts, short word repetition prompts, and orientation question-and-answer prompts. The reason for using active prompts instead of passively waiting for natural speech is that cognitive screening requires repeatable and comparable stimulus conditions. Unified prompts enable data comparability between different time points and different test subjects, while also allowing the system to predict behaviors such as responses, pauses, hesitations, or restates that the test subject is about to exhibit, thus providing timing references for subsequent gating processing. After the voice stimulus is emitted, the multimodal acquisition module simultaneously records the audio data stream and facial video data stream of the tested object; the audio data reflects the intensity of the voice, the rhythm of the voice, and the pause pattern; the facial video data carries the subtle color fluctuations on the skin surface caused by changes in blood volume, as well as the facial muscle activity and micro-expression changes during the response. Because non-contact optical physiological extraction is easily interfered with by speaking and movement, the system does not directly treat all video frames as valid physiological frames. Instead, it first uses the prior information of audio to identify quasi-static physiological windows that are more suitable for extracting subtle changes in blood flow. After the acoustic gating generation module extracts acoustic features from the audio data, it converts the time periods that should be retained and the time periods that should be suppressed into visual mask gating signals. The cross-phase-locked masking module then locks the forehead and cheek regions from the facial video, extracts the original image mean sequence of these regions over time, and filters out high artifact frames when the vocal energy is greater than a preset burst energy threshold according to the gating signal, retaining only frames with an image mean change rate lower than a preset change rate threshold to form a discrete physiological sampling time series. Subsequently, the frequency domain reconstruction module performs spectral reconstruction on the discrete sequence under non-uniform sampling conditions to obtain frequency domain diagnostic indicators related to heart rate variability; the cognitive state diagnosis module classifies the state of the tested subject by combining the population baseline or the reference threshold set in the hospital, and feeds the results back to the active interaction stimulus module to adjust the difficulty, speed or interval of the next round of prompts. As a further example, suppose a screening round includes three speech segments S1, S2, and S3, where S1 is a system question, S2 is the subject's continuous response, and S3 is the subject's response after a pause; the video is continuously captured within these three segments, but the most suitable segments for extracting physiological information are likely concentrated in the continuous vowel end of S2 and the short pause segment before S3. The system does not calculate the entire video segment at once, but rather preserves these relatively stable time slices to form a discrete but reliable array of sampling points. The physical basis for doing so is that the movement of the jaw, perioral and cheek muscles during speech changes the light conditions on the skin surface, causing a disturbance stronger than changes in blood volume. When vowels are lengthened or there is a short pause, the muscle displacement is relatively weakened, making it easier to observe real physiological fluctuations. In terms of handling abnormal situations, if the subject refuses to answer throughout the process or the environmental noise exceeds the preset environmental noise threshold, making it impossible to reliably separate the subject's own voice from the audio data, the system will suspend the physiological extraction based on acoustic gating, only record that the stimulus was invalid for this round, and prompt medical staff to restart the examination; if the face is out of the field of view, severely occluded, or the illumination is insufficient on the video end, only the audio interaction log will be retained for this round, and no physiological diagnosis results will be output, so as to avoid using low-quality visual data for cognitive state judgment; In a memory clinic, a 72-year-old patient suspected of having mild cognitive impairment underwent a five-minute initial screening. The system prompted the subject to repeat the three words, and the subject began to answer. During the answering process, the speech rate was uneven, the pauses increased, and the facial movements that moved with the speech exceeded the preset movement displacement threshold around the mouth. The system uses acoustic gating to preserve video clips of pauses and vowel durations, and extracts relatively stable optical pulse information from the forehead and cheeks, ultimately generating an auxiliary report that includes heart rate variability trends and cognitive state labels; this report can serve as a reference for doctors to decide whether to further arrange scale assessments or neuropsychological tests. The purpose of this step is to achieve diagnostic-level non-contact physiological sampling using active interaction and cross-modal gating without requiring patients to wear electrodes or remain absolutely still, thus providing a more stable and objective basis for cognitive state assessment.

[0020] In a preferred embodiment of the present invention, the multimodal acquisition module simultaneously acquires the audio data stream and facial video data stream of the test object, specifically for: acquiring the audio data stream by acquiring acoustic signals from the test object through a directional microphone array; acquiring the facial video data stream by acquiring visual signals from the test object through a binocular camera module; and performing time-stamp alignment processing on the audio data stream and facial video data stream to obtain the synchronized audio data stream and facial video data stream.

[0021] This embodiment provides a mechanism for multimodal synchronous acquisition; specifically, in the aforementioned memory clinic scenario, relying solely on a single microphone or a single camera is prone to acquisition deviations due to factors such as reflected sound in the examination room, conversations with family members, and slight head turns by the patient. To address this, a directional microphone array and a binocular camera module were introduced, and unified timestamp alignment was performed on the two types of data streams to ensure that the time of subsequent sound occurrence and the time of facial movement could fall on the same time axis. Specifically, the directional microphone array has the functions of enhancing directivity, which is different from the overall acoustic gain, forming the main receiving direction towards the subject being tested, and suppressing lateral noise from the doctor's end, the caregiver's end, and the corridor environment. In cognitive screening scenarios, if the vocal energy and sound pressure level of the subject are lower than the preset vocal threshold, and if other people's prompts or background broadcasts are mixed in the environment, the system will mistakenly take the external sound source as a response sound, thus erroneously triggering visual gating. With an array, the voice directly in front can be retained according to the location of the sound source, improving the separability of the subject's own voice. The binocular camera module is used to acquire facial video data streams and assist in depth stabilization. Compared to ordinary monocular structures, binocular modules are more likely to determine whether the object being tested has shifted back and forth by less than a preset displacement threshold, whether it has deviated from the optimal observation distance, and whether the forehead and cheek areas are still within the effective imaging range. In non-contact physiological detection, changes in front and back distance will alter the imaging scale and reflection conditions, thereby reducing the stability of non-contact photoplethysmography (PPG) extraction. Therefore, binocular structures help improve the continuity of subsequent region of interest tracking. Timestamp alignment can be performed using a unified clock source; audio sampling frames and video frames can be written to buffers separately, with each data entry accompanied by a sampling time marker; the system rearranges data using the same reference clock so that the starting point of a sound in a certain audio segment can accurately correspond to a video frame within an adjacent time window; for example, if a plosive sound in the audio occurs in time slice T2, and the corresponding rapid lip closure and opening action in the video occurs near T2, the system can confirm through alignment that this is the same physiological event, rather than two unrelated events; Furthermore, if the microphone array detects multiple strong sound sources simultaneously, and the direction of the main sound source is unstable, the system can temporarily suspend the establishment of acoustic gating and only record the interference risk in this round; if the binocular camera causes depth distortion due to reflective glasses, masks, or drastic fluctuations in lighting, the system automatically degrades to two-dimensional face tracking mode and reduces the credibility of the physiological parameters in this round; if the difference between the audio and video timestamps exceeds the hospital's preset tolerance, this round will not enter the diagnostic process, but will trigger a data acquisition reset to avoid cross-modal mismatch. During the retelling task for the same outpatient, the family member, sitting to the side and slightly behind the patient, occasionally made a verbal prompt. The directional microphone array fixed the main receiving direction directly in front of the patient, thus significantly suppressing the family member's voice. The binocular camera detected that the patient leaned back slightly while thinking, but the face was still within the effective depth range. After unified timestamp alignment, the system was able to accurately match the patient's slow response with his / her facial micro-movements, without mistaking the family member's prompts for the patient's acoustic trigger signals. The purpose of this step is to establish a reliable temporal consistency and spatial directivity basis for cross-modal gating, thereby enabling accurate coupling of subsequent acoustic features with visual physiological signals.

[0022] In a preferred embodiment of the present invention, the acoustic gating generation module extracts acoustic features from the audio data stream and identifies quasi-static physiological windows. Specifically, it is used to: perform speech activity detection and short-time Fourier transform processing on the audio data stream to obtain acoustic envelope data. Short-term energy features and zero-crossing rate features are extracted from acoustic envelope data. When the short-term energy feature is lower than a preset pause energy threshold, it is determined to be an inter-word pause sequence. When the short-term energy feature is between the pause energy threshold and a preset burst energy threshold, and the zero-crossing rate feature is lower than a preset unvoiced zero-crossing rate threshold, it is determined to be a vowel duration sequence. The vowel duration sequence and the inter-word pause sequence are merged to obtain a quasi-static physiological window.

[0023] This embodiment provides a mechanism for identifying quasi-static physiological windows. Specifically, although audio and video synchronization was achieved in the previous embodiment, directly treating all speaking periods as equally effective would still present serious problems: certain syllables, especially consonants, plosives, and rapid linking, can cause sudden deformation of the perioral area and cheeks. Such deformations contaminate the optical physiological signal far more than vowel pauses. Therefore, this embodiment further extracts acoustic features from the audio data stream to identify quasi-static physiological windows that are more suitable for extracting non-contact physiological parameters. Specifically, speech activity detection is used to determine whether the subject actively makes a sound during a certain period of time, while short-time Fourier transform is used to observe the distribution of speech energy over time and frequency band; the focus here is not on semantic content, but on the facial movement state corresponding to the manner of speech. Segments with short-term energy greater than or equal to a preset high-energy threshold and an energy change rate greater than a preset mutation threshold often correspond to plosives, stressed consonants, or suddenly enhanced exhalation. These segments are usually accompanied by muscle contractions with a contraction amplitude greater than a preset action threshold. Segments with a high zero-crossing rate are more likely to contain noise components or voiceless consonants, and transient movements of the lips and jaw are more obvious. In contrast, although the vowel duration is still phonation, the position of the vocal organs is relatively stable, facial surface movements tend to be gentle, and subtle blood flow fluctuations are more easily observed in the forehead and cheeks. To ensure the extraction process has executable quantitative standards, the system presets a pause energy threshold, a burst energy threshold, and a zero-crossing rate threshold for unvoiced speech. The pause energy threshold is adaptively set based on the average background noise energy of the current detection environment, typically set to 1.2 to 1.5 times the average background noise energy. The burst energy threshold is dynamically calibrated based on the peak energy distribution during the subject's historical normal speech. The zero-crossing rate threshold for unvoiced speech is preset based on a statistical model of the short-term zero-crossing rate of typical elderly speech, ensuring the objectivity and feasibility of the threshold settings. The specific logical flow is as follows: when the short-term energy of a segment is lower than the pause energy threshold, the system directly determines that there is no obvious airflow and vocal cord vibration during that period, confirming it as a pause between words; when the short-term energy is between the pause energy threshold and the burst energy threshold, and the zero-crossing rate is lower than the zero-crossing rate threshold for unvoiced sounds, it indicates that the vocal cords are in a state of stable vibration and no strong fricative sound, and it is determined to be a vowel duration period; if the short-term energy is higher than the burst energy threshold, or the zero-crossing rate is higher than the zero-crossing rate threshold for unvoiced sounds, it is determined to be a violent vocalization or a voiceless consonant segment and is then separated. Through the above-mentioned logical judgment structure with multiple conditions in parallel, high-purity candidate quasi-static windows can be systematically obtained; the pause period between words is also an important window; the reason is that when the subject is thinking, recalling or waiting for the next utterance, the movement of the perioral and mandibular muscles will temporarily weaken, while the rhythmic changes generated by the cardiovascular system will continue, making this period a preferred window for extracting non-contact photoplethysmography pulse waves. However, a pause does not mean absolute stillness. If there is obvious nodding, coughing or turning of the head during the pause, the subsequent video gating still needs to be further screened. As a simplified example, suppose a response voice is divided into five consecutive segments A1 to A5, where A1 is "I", A2 is a short pause, A3 is "remember", A4 is "get", and A5 is the exhalation at the end of the sentence. If A1 and A3 are mainly characterized by vowel sustain, and A2 is a pause, then these three segments can be merged into a candidate quasi-static window; if A5 is accompanied by significant exhalation and mouth shape changes, then even if it is at the end of a sentence, it will not be given priority for inclusion. The system does not obtain a judgment on whether the content is correct or not, but rather time labels of which time periods are more suitable for physiological sampling; furthermore, if the subject has dysarthria, Parkinson's-like low voice or obvious wheezing, making the traditional vowel and pause boundary unclear, the system can relax the window recognition conditions and include low-change slow vocal segments into the candidate set, but at the same time reduce the confidence level of the physiological conclusion of this round. If low-frequency noise from an air conditioner or the sound of a wheelchair motor is present in the audio, causing false triggering of voice activity detection, the array directivity and silence baseline re-estimation will be prioritized to avoid misidentifying environmental noise as a period of speech. When the elderly patient performed a three-word delayed recall task, the patient first said the first word, paused for about one second, and then slowly said the second word; the system identified two vowel continuous segments and one word pause segment, and marked these three segments as quasi-static physiological windows; In contrast, when the patient tried the third word and couldn't remember the answer, he let out a noticeable sigh. The high-energy expiratory segment corresponding to this sigh was not retained as the preferred physiological window. The purpose of this step is to transform the physiological laws of vocalization in the audio into prior information for visual sampling timing, so as to achieve front-end avoidance of motion artifacts, rather than passive compensation at the back end.

[0024] In a preferred embodiment of the present invention, the acoustic gating generation module generates a visual mask gating signal based on a quasi-static physiological window, specifically used for: performing binarization mapping processing on the quasi-static physiological window to generate an initial Boolean value sequence; wherein, the word pause state and vowel persistence state in the initial Boolean value sequence are mapped to a first logic value, and the state in the initial Boolean value sequence where the short-time energy feature is greater than or equal to a preset burst energy threshold is defined as a violent vocalization state and mapped to a second logic value; Obtain the neuromuscular conduction delay constant pre-calibrated based on the multimodal stimulus response time of historical populations, convert the neuromuscular conduction delay constant into pre-compensation frame number and post-compensation frame number according to the facial video frame rate, perform delay compensation processing on the initial Boolean value sequence based on the pre-compensation frame number and post-compensation frame number, and generate a visual mask gating signal.

[0025] This embodiment provides a mechanism for generating visual mask gating signals; specifically, it is not enough to simply identify the quasi-static physiological window, because there is not complete synchronization between the sound already produced in the audio and the facial movement that begins to be obvious in the video. If you use audio tags directly to trim videos, there may be situations where the audio has been identified as high-risk vocalization, but the face has not yet made any obvious movements, or the audio has just ended and the muscle inertia has not yet recovered. Therefore, this embodiment introduces neuromuscular transmission delay compensation on the basis of binarization mapping to obtain a visual mask gating signal that better matches the real facial movement state; in detail, the system first maps the quasi-static window to an initial Boolean value sequence; The retainable state can be defined as the first logical value, and the state that should be blocked can be defined as the second logical value. Here, the first logical value usually corresponds to the inter-word pause state and the vowel continuation state, because the overall muscle displacement is relatively small in these two states. The second logical value corresponds to a violent vocalization state with a significant increase in energy in a short period of time, such as plosive sounds, sudden exhalation, coughing, throat clearing, and rapid denial. Binarization mapping acts like adding a temporal mask to a video frame: when the value is the first logical value, physiological information is allowed to enter the subsequent process, and when the value is the second logical value, it is suppressed first. However, there is a physiological delay in human reaction; after hearing the prompt, there is a time required for nerve conduction and muscle recruitment before facial expressions and verbal responses are made; similarly, after a strong sound is detected in the audio, local facial tissues may continue to shift for a short period of time. Therefore, the system will shift the initial Boolean value sequence forward or backward according to the pre-calibrated neuromuscular transmission delay constant. This is not a mathematical trick, but based on physiological facts: visible facial movements often lag behind nerve triggers and may still have a brief inertia after the speech ends. The pre-calibrated neuromuscular conduction delay constant is pre-calibrated and set by the system based on the statistical mean of the acoustic-visual multimodal stimulus response time of the historically healthy elderly population in the hospital. Its initial typical value range is set to 100 milliseconds to 300 milliseconds. At the specific algorithm execution level, this time delay compensation is manifested as the interval expansion masking processing of the second logical value, which is the state that needs to be masked to cause motion artifacts; since inhalation or mouth shape preparation displacement often accompanies the pronunciation, and the cheek muscle rebound lags behind the end of the audio after the pronunciation ends, the system converts the extracted neuromuscular transmission time delay constant into a specific number of pre-compensation frames and post-compensation frames based on the facial video frame rate. The system initiates a sliding window to traverse the initial Boolean value sequence. Once the continuous period of the second logical value is locked, the system will reverse the time of the starting point of the period by the corresponding number of preceding compensation frames and extend the time of the ending point by the corresponding number of following compensation frames. The logical values ​​within the adjacent extended intervals covered by this period will be uniformly and forcibly overwritten with the second logical value. This explicit data rewriting rule ensures that the preceding a priori action frames and the following inertial aftershock frames caused by sound emission are both isolated from the gating. In terms of specific algorithm rule definitions, for frame sequence numbers... initial Boolean sequence The system assigns the first logical value to frames that are in a word pause state or a vowel continuation state. The frame that experiences a short-term energy overrun is assigned the second logical value. Based on the number of pre-compensation frames With post-compensation frame count The mathematical operation of time delay compensation is as follows: once continuous delay is detected... The adjacent extended interval will soon be All state values ​​within are forcibly overwritten as This generates the final mask gating signal; As a simplified example, assume the initial sequence is [1, 1, 0, 0, 1] over five consecutive time slices, where 1 indicates that it can be retained and 0 indicates that it should be masked. If a short delay compensation is set based on the hospital's experience, several video frames adjacent to the third and fourth slices can also be marked as 0 to cover the actual muscle deformation propagation period. Although the resulting gated sequence has a wider mask coverage range, it is better able to avoid mistaking motion transition frames as valid physiological frames. Furthermore, if the subject has special neuromuscular conditions such as facial paralysis, Parkinson's mask face, or myasthenia gravis, the preset time delay may no longer be applicable. In this case, the system can call the individualized compensation parameters in the historical follow-up data. If individual historical data is missing, the default compensation value in the hospital will be maintained and the report will indicate that the neuromuscular time delay uses the general parameters. If the duration of audio signal interruption exceeds the preset signal interruption threshold and a stable initial Boolean value sequence cannot be formed, then the gating signal will not be generated in this round, and the process will fall back to conservative screening based solely on video quality thresholds. When the patient answered what day of the week it was, the system detected that the patient first made a slight twitch of the corner of the mouth before opening their mouth, uttered a short initial sound, and then entered a more stable vowel phase; without time delay compensation, the system may retain the transition frames before and after the initial sound of the word. After the compensation was introduced, several frames before and after the initial burst of the word were blocked, and only the relatively stable final vowel segments and pause segments were retained for subsequent physiological analysis. The purpose of this step is to establish a more physiologically consistent temporal relationship between acoustic triggering and real facial movements, thereby achieving more accurate visual gating.

[0026] In a preferred embodiment of the present invention, the cross-phase-locked masking module determines the region of interest in the face from the facial video data stream and extracts the original image mean time series of the region of interest in the face, specifically used for: performing facial feature point tracking processing on the facial video data stream to obtain the coordinate sequence of facial key points; Based on the facial key point coordinate sequence, the cheek and forehead regions of the tested object are located and merged into a facial region of interest; the original pixel values ​​in the facial region of interest are extracted and the spatial mean of the original pixel values ​​is calculated to generate the original image mean time series.

[0027] This embodiment provides a mechanism for determining facial regions of interest and extracting the time series mean of the original image. Specifically, after the visual gating signal is generated, if physiological extraction is still performed directly on the entire face image, two problems will be encountered: First, the area around the mouth and eyes is a high-frequency motion area, which leads to the introduction of vocalization and blinking artifacts. Secondly, areas such as hair, eyebrows, and eyeglass frames have unstable reflections and are not suitable as areas for blood flow observation; therefore, this embodiment introduces facial feature point tracking and merges the cheek area and forehead area into a facial region of interest. Specifically, facial feature point tracking can continuously locate stable anatomical landmarks such as the brow ridge, corner of the eye, nasal wing, corner of the mouth, and mandible. With the help of these key points, the system can still follow the changes in facial position and dynamically correct the boundary of the region of interest when the subject slightly turns, raises, or tilts their head back. The cheeks and forehead were chosen as the main observation areas because these two areas usually have a better skin exposure area and a more stable microcirculation. At the same time, they are far from the area of ​​significant deformation of the lips, making them more suitable for observing subtle color fluctuations caused by changes in subcutaneous blood volume. After extracting the original pixel values, spatial mean calculation is performed. The physical significance of this is that individual pixels are easily affected by noise, compression, specular reflection, or local shadows. However, by averaging the entire region of interest, the common slow physiological fluctuations within the region can be preserved, while scattered interference can be suppressed. What is obtained here is a time-varying sequence of the region's brightness or color mean, which can be used as the basic input for subsequent rPPG extraction. As a simplified example, suppose in the current video frame, the left cheek, right cheek, and forehead each form three small regions R1, R2, and R3. If R1 has slight reflection, R2 is relatively stable, and R3 is partially obscured by stray hairs on the forehead, the system can first confirm the boundaries of the three regions based on key points, then perform average statistics on the pixels of the three regions, and combine them according to quality scores. The resulting time series is more stable and less prone to failure due to local occlusion than using a single small region. Furthermore, if the patient is wearing a hat with a brim or their forehead is covered by thick bangs, the forehead area can be temporarily excluded from the combination, and only the cheek area is used; if the patient is wearing a mask and the visible area of ​​the cheeks is insufficient, the forehead area is retained first and the reliability of the result is reduced; if face tracking is lost midway, the system can maintain the area position of the previous frame for buffering within a short time window, but if the loss exceeds the preset number of frames, the extraction round is stopped to prevent reading pixel values ​​from the wrong position; When the patient was conducting location orientation questions, the patient slightly tilted their head due to thinking, causing local reflections in the glasses lenses. The system continuously updated the boundaries of the cheeks and forehead through facial key points, automatically avoiding the highly reflective areas of the glasses, preserving the stable skin areas in the center of the forehead and the middle of the cheeks, and stitching the original pixel averages of these areas into a time series to prepare for subsequent artifact removal processing. The purpose of this step is to obtain a visual baseline signal that is more correlated with changes in subcutaneous blood flow and less correlated with speech and occlusion interference in a real outpatient setting, thereby improving the reliability of subsequent frequency domain analysis.

[0028] In a preferred embodiment of the present invention, the cross-phase-locked masking module performs mask multiplication on the visual mask gating signal and the original image mean time series to obtain a discrete physiological sampling time series. Specifically, it is used to: use the visual mask gating signal as a priori weight signal; perform frame-by-frame dot product operation on the priori weight signal and the original image mean time series in the time dimension; set the original image mean time series at the corresponding time to zero through the second logic value to remove motion artifact frames caused by violent vocalization; and stitch together the remaining data points after removing motion artifact frames to obtain a discrete physiological sampling time series.

[0029] This embodiment provides a cross-phase-locked masking processing mechanism; specifically, simply extracting the region of interest is not enough to eliminate motion artifacts caused by speaking, because even if the cheeks and forehead are relatively stable, they will still produce obvious optical disturbances when the head and face move as a whole during loud talking. Traditional approaches often attempt to perform post-processing correction on all video frames. However, in natural communication scenarios, motion frequencies often overlap with heart rate frequencies, making reliable separation difficult during post-processing. Therefore, this embodiment adopts a more direct strategy: using the visual mask gating signal as a priori weight to filter out the mean time series of the original image frame by frame. In detail, the so-called mask multiplication is essentially to distinguish between usable and unusable frames according to time. If the gate signal is the first logic value, the image mean at the corresponding time is retained; if it is the second logic value, the image mean at the corresponding time is set to zero or directly discarded. The key point of doing this is not the calculation form, but to use audio prior to guide the visual module to remove frames that are more likely to be contaminated by vocalization from a physiological perspective, thereby avoiding their use in pulse wave analysis. The system splices the remaining data points to form a discrete physiological sampling time series, which is an important feature of this embodiment. The system does not strive to mechanically preserve every frame as continuous, but rather to preserve each frame as more reliable. From a diagnostic perspective, continuous sequences subject to strong motion interference are not superior to sparser but high-quality discrete sequences. Especially in cognitive screening scenarios, patients often speak intermittently due to hesitation, repetition, and talking to themselves, which provides several relatively stable time points for opportunistic sampling. As a simplified example, suppose the mean time series of a certain original image is [M1, M2, M3, M4, M5, M6], and the corresponding gating signal is [1, 1, 0, 0, 1, 0]. After processing, M1, M2, and M5 can be retained, while M3, M4, and M6 can be discarded. Although the final discrete sequence is no longer uniform and continuous, the three retained points are closer to the real physiological fluctuations, while the three removed points are more likely to contain disturbances caused by speaking actions. Furthermore, if the proportion of the second logical value is too high in a certain round of interaction, it indicates that the patient is continuously making loud noises, coughing, or turning their head significantly, resulting in too few frames that can be retained. In this case, the system does not force the output of frequency domain results, but instead prompts the user to re-initiate a shorter speech task, such as repeating words or numbers, to obtain more quasi-static windows. If there are insufficient effective frames in several consecutive rounds, the system marks the subject as unsuitable for non-contact physiological extraction and suggests using contact-based gold standard equipment for assisted examination. When the patient answered his home address, he spoke continuously for the first half of the speech, but paused for a duration longer than the preset pause time threshold in the second half. The system blocked a large number of the average video frames that were synchronized with the strong speech in the first half, and only retained a few data points in the pause segment and the elongated sound tail segment. Although the final sequence became discrete, these retained points have higher physiological reliability and can provide effective input for subsequent frequency domain reconstruction. The purpose of this step is to remove high-risk motion artifact frames from the entry point of the visual processing link, thereby achieving a substantial improvement in the signal-to-noise ratio of non-contact physiological signals.

[0030] In a preferred embodiment of the present invention, the frequency domain reconstruction module inputs the discrete physiological sampling time series into a preset periodogram estimation algorithm for spectrum reconstruction and extracts frequency domain physiological diagnostic indicators. Specifically, the preset periodogram estimation algorithm is a non-uniform sampling spectrum analysis algorithm; the discrete physiological sampling time series is input into the non-uniform sampling spectrum analysis algorithm, and the least squares method is used to perform sine wave fitting on the discrete data points in the discrete physiological sampling time series to obtain continuous physiological signal waveforms. Energy is extracted from the continuous physiological signal waveform in preset low-frequency bands and preset high-frequency bands to obtain sympathetic and parasympathetic activity indicators; the ratio of the sympathetic and parasympathetic activity indicators is calculated as a frequency domain physiological diagnostic indicator, wherein the frequency domain physiological diagnostic indicator includes heart rate variability data.

[0031] This embodiment provides a mechanism for frequency domain reconstruction and frequency domain physiological diagnostic index extraction; specifically, in the previous embodiment, the system actively discarded a large number of contaminated frames, which improved the quality of a single sampling point, but also brought new problems: the obtained time series is no longer a regular and equally spaced continuous sequence; If conventional frequency domain analysis methods that rely on uniform sampling are still used at this point, spectral distortion is likely to occur. Therefore, this embodiment uses a non-uniform sampling spectral analysis algorithm to reconstruct the spectrum of discrete physiological sampling time series. In detail, this type of algorithm is suitable for scenarios where the sampling point intervals are not completely consistent; its core idea is not to forcibly fill in the missing frames, but to directly use the already retained valid points to find the potential periodic physiological rhythms; in this embodiment, the system fits a continuous physiological waveform based on discrete data points, and then extracts energy distribution from the preset low frequency band and preset high frequency band to characterize the activity trends of the sympathetic and parasympathetic nervous systems. Since heart rate variability essentially reflects the autonomic nervous system's ability to regulate heart rhythm, even if the original video sequence has been gated and screened, as long as sufficient effective rhythm information is retained, a frequency domain index with diagnostic value can still be reconstructed. In the specific implementation of waveform reconstruction using the least squares method, for waveforms reconstructed from discrete sampling timestamps... and its corresponding image mean The set of discrete sampling points The system is based on a fundamental sinusoidal wave model: ; Based on this, continuous physiological signal waveforms are fitted to quantify and estimate each frequency. Corresponding amplitude and power spectral density, where For continuous time variables, For the instantaneous values ​​of the reconstructed continuous signal waveform, Pi This is the initial phase; the low and high frequencies mentioned here are not arbitrarily defined, but correspond to different physiological components in autonomic nervous system regulation. Low-frequency components are usually associated with sympathetic regulation and some blood pressure reflex activities, while high-frequency components are usually more closely associated with respiratory-related vagal nerve regulation. In cognitive impairment screening, persistent abnormal cognitive load, neurodegenerative changes, or emotional stress may manifest as decreased heart rate variability and imbalance in frequency band energy distribution. Therefore, these frequency domain indicators can serve as objective inputs for the auxiliary assessment of cognitive state; the specific frequency band energy extraction and calculation process has a rigorous physical indicator boundary division: the system defines the frequency components in the integral interval of 0.04Hz to 0.15Hz in the spectrum as the preset low frequency band, and quantifies and statistically analyzes the power spectral density integral value in this frequency band as an indicator of sympathetic nerve activity; The integration interval from 0.15Hz to 0.40Hz is defined as a preset high-frequency band, and parasympathetic nerve activity indicators are extracted. In the indicator calculation and encapsulation step, the system uses mathematical methods to calculate the ratio of low-frequency indicators to high-frequency indicators, i.e., the low-frequency / high-frequency ratio, to quantitatively evaluate the dynamic state of the tension balance of the autonomic nervous system between the sympathetic and parasympathetic nerves. This ratio, together with the specific energy values ​​of the integrals in each frequency band, is used as the core element to assemble a heart rate variability dataset. By utilizing this rigorous and clear frequency band boundary division and functional ratio extraction logic, the calculation correlation path when the underlying frequency domain is analyzed into cardiovascular index is clarified, and a highly standardized objective data correlation is achieved. As a further example, suppose the discrete sequence retains only six valid time points P1 to P6, which are not evenly spaced, but exhibit slow and repetitive fluctuations as a whole. The system does not need to use an interpolation algorithm to forcibly complete P1 to P6 into six consecutive equally spaced points, but directly estimates its periodic characteristics based on these real retained points. If the reconstructed waveform has an energy greater than the first preset reference threshold in the low-frequency band and an energy less than the second preset reference threshold in the high-frequency band, it can indicate an abnormal tendency in the autonomic nervous system balance. In terms of handling abnormal situations, if there are too few discrete data points, the time span is too short, or the retained points are concentrated in local segments, it is difficult to support stable spectrum estimation. In this case, the system will not output the heart rate variability result, but will only give a prompt of insufficient sampling. If a patient has significant arrhythmia, atrial fibrillation, or pacemaker intervention, causing the rhythm itself to lose its significance in the usual heart rate variability, the system should automatically reduce or mask the diagnostic weight of this indicator after accessing the medical record information. During the patient's delayed recall phase, the system obtained discrete sampling points from multiple rounds of brief pauses and slow vocalizations. Although these sampling points were not uniform and continuous, the frequency domain reconstruction module was still able to recover the continuous rhythm waveforms reflecting cardiovascular autoregulation and extract the corresponding low-frequency, high-frequency energy and heart rate variability data, providing a basis for the next stage of status assessment. The purpose of this step is to overcome the limitation that frequency domain analysis must rely on completely continuous sampling, thereby enabling the effective diagnostic utilization of discrete high-quality physiological data in natural interaction scenarios.

[0032] In a preferred embodiment of the present invention, the cognitive state diagnosis module assesses the state of the tested object based on frequency domain physiological diagnostic indicators and generates a cognitive state diagnosis result. Specifically, it is used to: obtain a preset population health baseline threshold; compare heart rate variability data with the population health baseline threshold; if the heart rate variability data is greater than or equal to the population health baseline threshold, the cognitive state of the tested object is determined to be normal, and a cognitive state diagnosis result containing a normal state label is generated; if the heart rate variability data is less than the population health baseline threshold, the cognitive state of the tested object is determined to be in a declining state, and a cognitive state diagnosis result containing a declining state label is generated.

[0033] This embodiment provides a mechanism for generating cognitive state diagnostic results. Specifically, after obtaining frequency domain physiological diagnostic indicators, if they only remain at the raw numerical level, medical staff will find it difficult to use them quickly in outpatient settings. Therefore, this embodiment compares heart rate variability data with a preset population health baseline threshold and outputs cognitive state labels that can be directly used for screening and triage. Specifically, the baseline threshold for population health can be established within the hospital based on age stratification, gender stratification, or samples of previously healthy individuals, or it can refer to a validated standard database. The reason for using a population baseline instead of a single fixed value is that autonomic nervous function is significantly affected by age, chronic diseases, medications, and circadian rhythms. Especially in geriatric memory clinics, simply applying the same numerical standard to all patients can easily lead to misjudgment. Therefore, this embodiment prefers to establish a relatively robust reference threshold range based on the elderly population, and then project the heart rate variability data of the tested subjects into the corresponding reference level. When heart rate variability is higher than or reaches the reference level, it can be considered that the autonomic nervous system's regulatory capacity has not shown significant abnormal decline. In combination with the application purpose of this system, a normal state label can be output. When heart rate variability is lower than the baseline threshold for population health, it suggests that there may be a risk related to cognitive decline, and a decline state label is output. The decline state here is an auxiliary diagnostic label and does not replace the clinical diagnosis. Instead, it provides doctors with a basis for further scheduling of scales, imaging or laboratory tests. As a simplified example; assuming the system establishes a reference baseline G for the same age group, and a patient's heart rate variability index obtained in this round is H; if H is not lower than G, the result label is biased towards normal; if H is lower than G, the result label is biased towards decline; the system can also add risk warning levels for mild or significant deviations, but its core is still to classify based on the baseline; Furthermore, if the patient has a fever, infection, severe anxiety, or has just taken medication that affects heart rate on the day of collection, heart rate variability may temporarily deviate. In this case, the system can receive medical history questionnaire information before collection and mark the results as needing to be interpreted in conjunction with the clinical background when there are interfering factors. If the population health baseline is missing, the age group is mismatched, or the current sample quality level is insufficient, the system will not force the output of normal / declining labels and will only return an uncertain result. During the outpatient screening of this 72-year-old patient, the system compared the reconstructed heart rate variability data with the baseline of healthy elderly people of the same age and found that his autonomic nervous regulation level was lower than the reference range. Therefore, it generated an auxiliary diagnostic result containing a label of declining status. Based on this, the doctor decided to add a brief mental status test and an extended memory scale to this outpatient visit instead of ending the consultation based solely on the patient's chief complaint. The purpose of this step is to transform complex frequency domain physiological information into screening tags that can be implemented in outpatient settings, thereby enabling rapid identification and triage of individuals at risk of cognitive impairment. In a preferred embodiment of the present invention, the active interaction stimulus module is specifically used for: acquiring cognitive state diagnosis results and updating the active interaction stimulus strategy based on the cognitive state diagnosis results; generating a next-round voice stimulus signal based on the updated active interaction stimulus strategy, and outputting the next-round voice stimulus signal to the tested object through the audio output device to trigger a new round of multimodal data acquisition.

[0034] This embodiment provides a mechanism for updating the active interactive incentive strategy. Specifically, if the system terminates immediately after outputting a cognitive state label, although a single round of screening can be completed, it is difficult to adapt to the needs of stratified assessment in outpatient settings. Especially for patients in borderline states, a single round of interaction may not be sufficient to stably reflect the true state due to tension, environmental interference, or short-term fatigue. Therefore, this embodiment dynamically adjusts the voice incentive signal for the next round based on the diagnostic results of the previous round, so that the interactive task matches the current state of the tested object. Specifically, when the previous round's results are relatively normal, the system can appropriately increase the task complexity, such as changing from word repetition to two-step instruction execution, or from date-oriented question-and-answer to short sentence recall, to increase the opportunities to observe executive function and working memory; when the results are relatively declining, the system will prioritize reducing the prompting speed, shortening sentence length, extending response waiting time, and reducing complex questions that may induce frustration, instead adopting gentler structured prompts. This update mechanism is not a simple optimization of human-computer dialogue, but serves the purpose of diagnosis: by adjusting the stimulus load, it induces a response pattern that is more suitable for the current patient state, thereby obtaining more discriminative subsequent multimodal data; Furthermore, the cognitive state feedback to the active incentive module can also change the collection rhythm; for example, for patients suspected of deterioration, a longer pause can be added after each round of questioning to artificially create more quasi-static physiological windows; for patients with a more stable state, the interval between rounds can be appropriately shortened to improve the efficiency of outpatient screening; in this way, the diagnostic results are not only used for terminal output, but also participate in the optimization of sampling conditions, forming a closed loop. As a simplified example; assuming the first round result is a decline, the strategy in the next round is switched from strategy package C1 to strategy package C2; ​​C1 may include normal speaking speed and standard length questions; C2 includes slowing down the speaking speed, simplifying vocabulary and extending the waiting time; the audio and video collected again after the switch are more conducive to observing the true response ability of this type of patient, rather than being masked by the excessive difficulty of the task; Furthermore, if the results of the previous round are uncertain or the quality level is low, the system will not make aggressive strategy updates, but will maintain the basic prompt template to avoid misclassification amplifying the bias; if the patient shows obvious agitation, refusal to answer or fatigue, the system can automatically enter a low-stimulation mode, retain only the shortest necessary questions and answers and prompt medical staff to take over manually; after the first round of screening for the patient, the system will give a deterioration status label. The active interaction incentive module changed the next round of prompts from "Please repeat these three unrelated words" to "Please first say whether it is daytime or nighttime today," and slowed down the speech rate and lengthened the intervals. Patients responded more completely under lower load conditions, and more pause windows were also generated, which further improved the quality of subsequent physiological sampling. The purpose of this step is to feed the diagnostic results back into the interaction design, thereby achieving adaptive matching between the data collection task and the patient's status, and improving the stability and repeatability of multi-round screening.

[0035] In a preferred embodiment of the present invention, after outputting the next round of voice excitation signal to the test object through the audio output device, the method further includes: recording the excitation timestamp of the output of the next round of voice excitation signal; performing facial feature point tracking processing on the facial video data stream, calculating the displacement vector of the facial key point coordinate sequence between adjacent facial video frames, and when the magnitude of the displacement vector exceeds the preset micro-movement displacement threshold, recording the time of the corresponding video frame as the facial micro-expression change timestamp. The difference between the excitation timestamp and the facial micro-expression change timestamp is calculated to obtain dynamic neuromuscular conduction delay data; the dynamic neuromuscular conduction delay data is fed back to the acoustic gating generation module to replace the pre-calibrated neuromuscular conduction delay constant for a new round of visual mask gating signal generation.

[0036] This embodiment provides a mechanism for dynamic neuromuscular conduction delay self-correction. Specifically, in the aforementioned implementation, visual gating relies on a pre-calibrated neuromuscular conduction delay constant. However, in real outpatient settings, the stimulus response delay may vary significantly for different patients, different disease stages, or even at different times for the same patient. If the same fixed time delay is used for a long time, the gating boundary may gradually deviate from the actual time of movement as the patient becomes fatigued, attention declines, or neurodegenerative disease progresses. Therefore, this embodiment further records the excitation timestamp and the facial micro-expression change timestamp to calculate dynamic neuromuscular conduction delay data, which is then used to replace the fixed constant in subsequent rounds. Specifically, when the audio output device issues the prompt for the next round, the system synchronously records the excitation timestamp. In the new round of multimodal acquisition, the system extracts the starting points of micro-expression changes related to stimuli from facial videos, such as slight eyebrow raising, corner of mouth twitching, eyelid changes, or jaw initiation before opening the mouth; these changes often precede complete vocalization and are visible markers that are closer to the starting point of neuromuscular response. Specifically, based on facial feature point tracking processing, the system calculates the displacement vector of the facial key point coordinate sequence between adjacent facial video frames. When the magnitude of the displacement vector exceeds the preset micro-motion displacement threshold, it is determined that the starting point of micro-expression change has been captured, and the time of the corresponding video frame is recorded as the timestamp of the facial micro-expression change. Using the time difference between the two as dynamic neuromuscular transmission delay data can more accurately reflect the patient's current reaction speed and muscle execution delay. This dynamic delay data, fed back to the acoustic gating generation module, can replace the original general fixed delay. In this way, subsequent rounds, when mapping audio signals to visual gating, no longer rely on a single empirical value, but on individualized delay parameters based on the patient's current state. This self-correction is especially important for patients with cognitive decline, significantly slowed reaction time, or neuromuscular dyskinesia, because the actual time of their facial movements may be continuously delayed. If the time delay is still treated as that of a normal healthy person, it is easy to mistakenly retain transition frames or mistakenly discard valid frames. As a further example, suppose the system uses the default delay D0 in the first round, records the excitation time T0 at the beginning of the second round, and captures the first identifiable micro-expression change on the face in the video at the time T1. The difference between the two can form a new delay Dn. To prevent abrupt changes in the gating boundary caused by abnormal actions from a single measurement, the system further applies a smooth transition rule when updating the delay parameter using the difference; specifically, based on the extracted instantaneous difference... Compared with the currently used delay constant The system calculates the updated dynamic neuromuscular conduction delay data. The structured computation rules are as follows: ; in, The preset smoothing weight coefficients; this quantization rule makes the parameter iteration process transparent and ensures the stable operation of the self-correction mechanism; if Dn is continuously greater than D0, the system will automatically use new time delay data in subsequent gating processing, so that the audio trigger and video mask boundary are closer to the real response; In terms of handling abnormal situations, if the timestamps of micro-expression changes cannot be reliably extracted, for example, if the patient has very few facial expressions, severe occlusion, or insufficient lighting, the system will maintain the effective time delay parameters of the previous round; if the dynamic time delay cannot be extracted for several consecutive rounds, it will revert to the preset default value and mark individualized time delay correction as not enabled in the report; if the system finds that the dynamic time delay fluctuation is abnormally large and exceeds the reasonable physiological range, it will be preferentially determined to be caused by abnormal acquisition or large patient movement, rather than being directly used to update the gating parameters. In the second round of simplified question and answer for this patient, after the system recorded the starting point of the prompt broadcast, it was found that the patient's face showed a slight eyebrow raising and mouth corner twitching only after a relatively long delay, which was significantly slower than the common response of healthy elderly people; the system updated the individualized dynamic delay accordingly and moved the shielding boundary back in the third round of gating, so as to more accurately eliminate the transition artifact frames before and after the patient's slow response. As multiple rounds of screening are conducted, the gating parameters are gradually adapted to the patient's actual response characteristics, ultimately improving the temporal consistency of the entire examination. The purpose of this step is to use the patient's own immediate response characteristics to perform closed-loop correction of the gating delay, thereby achieving cross-modal alignment that is more in line with the individual's physiological state and more stable subsequent diagnostic output.

[0037] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. An intelligent companion robot interaction system based on multimodal intent recognition, characterized in that, include: The active interactive stimulation module is used to output a voice stimulation signal to the test object through an audio output device to trigger multimodal data acquisition; The multimodal acquisition module is used to simultaneously acquire the audio data stream and facial video data stream of the tested object in response to the voice excitation signal; and to perform time stamp alignment processing on the audio data stream and the facial video data stream. The acoustic gating generation module is used to extract acoustic features from the audio data stream, identify quasi-static physiological windows, and generate visual mask gating signals based on the quasi-static physiological windows. The cross-phase-locked mask module is used to determine the region of interest (ROI) on the face from the facial video data stream and extract the original image mean time series of the ROI. The visual mask gating signal is multiplied with the original image mean time series by a mask to obtain a discrete physiological sampling time series. The frequency domain reconstruction module is used to input discrete physiological sampling time series into a preset periodogram estimation algorithm for spectrum reconstruction and to extract frequency domain physiological diagnostic indicators. The cognitive state diagnosis module is used to assess the state of the tested object based on frequency domain physiological diagnostic indicators, generate cognitive state diagnosis results, and feed the cognitive state diagnosis results back to the active interaction stimulation module to update the subsequent voice stimulation signals.

2. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 1, characterized in that, The multimodal acquisition module simultaneously acquires the audio data stream and facial video data stream of the test object, specifically for: acquiring acoustic signals from the test object through a directional microphone array to obtain the audio data stream; Visual signals are acquired from the subject using a binocular camera module to obtain facial video data streams. The audio data stream and facial video data stream are time-stamp aligned to obtain the synchronized audio data stream and facial video data stream.

3. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 2, characterized in that, The acoustic gating generation module extracts acoustic features from the audio data stream and identifies quasi-static physiological windows. Specifically, it is used to: perform speech activity detection and short-time Fourier transform processing on the audio data stream to obtain acoustic envelope data. Extracting short-time energy features and zero-crossing rate features from acoustic envelope data; When the short-term energy feature is lower than the preset pause energy threshold, it is determined to be an inter-word pause sequence; when the short-term energy feature is between the pause energy threshold and the preset burst energy threshold, and the zero-crossing rate feature is lower than the preset unvoiced zero-crossing rate threshold, it is determined to be a vowel duration sequence; the vowel duration sequence and the inter-word pause sequence are merged to obtain a quasi-static physiological window.

4. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 3, characterized in that, The acoustic gating generation module generates a visual mask gating signal based on a quasi-static physiological window, specifically used to: perform binarization mapping processing on the quasi-static physiological window to generate an initial Boolean value sequence; Among them, the inter-word pause state and vowel persistence state in the initial Boolean value sequence are mapped to the first logical value, and the state in the initial Boolean value sequence where the short-time energy feature is greater than or equal to the preset burst energy threshold is defined as the violent vocalization state and mapped to the second logical value. Obtain the neuromuscular conduction delay constant pre-calibrated based on the multimodal stimulus response time of historical populations, convert the neuromuscular conduction delay constant into pre-compensation frame number and post-compensation frame number according to the facial video frame rate, perform delay compensation processing on the initial Boolean value sequence based on the pre-compensation frame number and post-compensation frame number, and generate a visual mask gating signal.

5. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 4, characterized in that, The cross-phase-locked masking module determines the region of interest (ROI) of a face from the facial video data stream and extracts the original image mean time series of the ROI. Specifically, it is used for: Facial feature point tracking processing is performed on the facial video data stream to obtain the facial key point coordinate sequence; based on the facial key point coordinate sequence, the cheek and forehead regions of the tested object are locked, and the cheek and forehead regions are merged into the facial region of interest; Extract the original pixel values ​​from the region of interest on the face, and calculate the spatial mean of the original pixel values ​​to generate the original image mean time series.

6. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 5, characterized in that, The cross-phase-locked masking module performs mask multiplication on the visual mask gating signal and the original image mean time series to obtain a discrete physiological sampling time series, specifically used for: The visual mask gating signal is used as the prior weight signal; the prior weight signal and the original image mean time series are multiplied frame by frame in the time dimension, and the original image mean time series at the corresponding time is set to zero by the second logic value to remove motion artifact frames caused by violent sounding. The remaining data points after removing motion artifact frames are stitched together to obtain discrete physiological sampling time series.

7. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 6, characterized in that, The frequency domain reconstruction module inputs the discrete physiological sampling time series into a preset periodogram estimation algorithm for spectrum reconstruction, extracting frequency domain physiological diagnostic indicators, specifically used for: The preset periodogram estimation algorithm is specifically a non-uniform sampling spectrum analysis algorithm; the discrete physiological sampling time series is input into the non-uniform sampling spectrum analysis algorithm, and the least squares method is used to fit the discrete data points in the discrete physiological sampling time series to obtain the continuous physiological signal waveform; Energy extraction was performed on the preset low-frequency band and preset high-frequency band of the continuous physiological signal waveform to obtain sympathetic and parasympathetic activity indicators. The ratio of the sympathetic nerve activity index to the parasympathetic nerve activity index is calculated as a frequency domain physiological diagnostic index, wherein the frequency domain physiological diagnostic index includes heart rate variability data.

8. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 7, characterized in that, The cognitive state diagnosis module assesses the state of the tested object based on frequency domain physiological diagnostic indicators and generates cognitive state diagnosis results. Specifically, it is used to: obtain a preset group health baseline threshold. The heart rate variability data is compared with the population health baseline threshold; if the heart rate variability data is greater than or equal to the population health baseline threshold, the cognitive state of the tested subject is determined to be normal, and a cognitive state diagnosis result containing a normal state label is generated. If the heart rate variability data is less than the population health baseline threshold, the cognitive state of the tested subject is determined to be in a state of decline, and a cognitive state diagnosis result containing a decline status label is generated.

9. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 8, characterized in that, The active interaction incentive module is specifically used for: Acquire cognitive state diagnosis results and generate or update active interaction incentive strategies based on the cognitive state diagnosis results; generate the next round of voice incentive signals based on the updated active interaction incentive strategies, and output the next round of voice incentive signals to the test subject through the audio output device to trigger a new round of multimodal data acquisition.

10. The intelligent companion robot interaction system based on multimodal intent recognition according to claim 9, characterized in that, After outputting the next round of voice excitation signal to the test object through the audio output device, the method further includes: recording the excitation timestamp of the output of the next round of voice excitation signal; Facial feature point tracking processing is performed on the facial video data stream, and the displacement vector of the facial key point coordinate sequence between adjacent facial video frames is calculated. When the magnitude of the displacement vector exceeds the preset micro-movement displacement threshold, the time of the corresponding video frame is recorded as a timestamp of facial micro-expression change. The difference between the excitation timestamp and the facial micro-expression change timestamp is calculated to obtain dynamic neuromuscular conduction delay data; the dynamic neuromuscular conduction delay data is fed back to the acoustic gating generation module to replace the pre-calibrated neuromuscular conduction delay constant for a new round of visual mask gating signal generation.