Concentration state information acquisition method and device
By combining millimeter-wave radar and audio sensors and employing a multimodal fusion model, the privacy and convenience issues of existing technologies for children's attention assessment are resolved, achieving seamless and accurate attention assessment suitable for various learning scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies for assessing children's attention rely heavily on visual recognition, wearable devices, or task-based behavioral assessments, which raise concerns about privacy breaches and inconvenience, making it difficult to achieve efficient and adaptable attention assessments across various learning scenarios.
Using millimeter-wave radar and audio sensors for non-intrusive data acquisition, and through feature extraction and multimodal fusion models, combined with radar echo signals and environmental audio signals, attention state information is generated to achieve non-intrusive and accurate attention assessment.
It enables seamless acquisition of attention status information, improving the robustness, accuracy, and interpretability of assessments, and is suitable for scenarios such as home learning desktops, classrooms, and smart teaching aids.
Smart Images

Figure CN121647672A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of deep learning, signal processing and multimodal fusion technology. Background Technology
[0002] Children's attention span (the degree of concentration) is a core indicator in research on learning outcomes and psychological development. Its accurate assessment is crucial for optimizing children's learning strategies and enhancing the targeted nature of educational interventions. With the integration of intelligent education and sensing technologies, children's attention span assessment is gradually shifting from traditional manual observation to intelligent and data-driven approaches. It is widely applied in various scenarios such as home learning guidance, classroom teaching management, and the development of intelligent teaching aids, becoming a key entry point connecting educational needs with technological innovation. The market demand for efficient attention span assessment solutions adapted to different learning scenarios is increasingly urgent.
[0003] Currently, the industry has developed several mature technologies for assessing attention, providing diverse options for monitoring children's attention. Among these, visual recognition technology uses cameras to collect visual information such as children's gaze trajectory and body posture, indirectly assessing their attentional state. Wearable detection technology utilizes wearable terminals such as EEG (Electroencephalography) devices, heart rate monitors, and smartwatches to collect physiological signals or motion data, constructing a correlation model between attention and physiological characteristics. Task-based behavioral assessment technology designs specific tasks and quantifies children's attention levels based on their reaction time, operational stability, and other behavioral performance. These technologies have all demonstrated their assessment value in different application scenarios. Summary of the Invention
[0004] This disclosure provides a method, apparatus, device, storage medium, and program product for acquiring focus state information.
[0005] In a first aspect, embodiments of this disclosure propose a method for acquiring attention state information, comprising: extracting features from a user's radar echo signal to generate radar echo features; extracting features from the audio signal of the user's environment to generate audio features; and inputting the radar echo features and audio features into a multimodal fusion model to output the user's attention state information.
[0006] Secondly, embodiments of this disclosure propose a device for acquiring attention state information, comprising: a first extraction module configured to extract features from a user's radar echo signal to generate radar echo features; a second extraction module configured to extract features from the audio signal of the user's environment to generate audio features; and an evaluation module configured to input the radar echo features and audio features into a multimodal fusion model to output the user's attention state information.
[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described in the first aspect.
[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described in the first aspect.
[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0010] The key or essential features of the embodiments disclosed herein are not intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. Wherein: Figure 1 This is a flowchart of an embodiment of the method for obtaining attention state information according to the present disclosure; Figure 2 This is a flowchart of yet another embodiment of the attention state information acquisition method according to the present disclosure; Figure 3 This is a flowchart of a method for assessing a user's focused state without them noticing. Figure 4 This is a schematic diagram of a structure of an embodiment of the attention state information acquisition device according to the present disclosure; Figure 5 This is a block diagram of an electronic device used to implement the attention state information acquisition method of the embodiments of this disclosure. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0014] Figure 1 A flow 100 of an embodiment of a method for acquiring focus state information according to the present disclosure is shown. The method for acquiring focus state information includes the following steps: Step 101: Extract features from the user's radar echo signal to generate radar echo features.
[0015] In this embodiment, the execution entity of the focus state information acquisition method can extract features from the radar echo signal of the user's human body and generate radar echo features.
[0016] By deploying millimeter-wave radar sensing modules in user (e.g., children) learning scenarios, reflected signals from the user's head and torso area can be collected, i.e., radar echo signals. The deployment of these millimeter-wave radar sensing modules ensures coverage of the user's head and core torso area without physical contact, achieving seamless data collection. The core of the feature extraction process from these radar echo signals is to mine micro-motion-related features that reflect the user's focus state, providing crucial data support for subsequent focus assessment. It should be noted that the acquisition and subsequent application of the user's radar echo signals are all done with the user's knowledge and authorization.
[0017] Step 102: Extract features from the audio signal of the user's environment to generate audio features.
[0018] In this embodiment, the aforementioned execution entity can extract features from the audio signals of the user's environment and generate audio features.
[0019] Audio sensors can collect ambient audio signals from the user's learning environment, including but not limited to: the sound of pen rubbing, talking, turning pages, tapping, and walking—all sounds related to learning. By extracting features from these audio signals, the user's behavioral state can be captured from an environmental acoustic perspective, complementing radar echo features and improving the comprehensiveness and accuracy of attention assessment. It should be noted that the acquisition and subsequent application of the user's audio signals are done with the user's knowledge and authorization.
[0020] Audio features may include, but are not limited to: average energy, spectral centroid, burst rate, speech detection ratio, learning behavior sounds (such as writing sounds) and distraction sounds (such as speaking sounds, desktop tapping sounds), etc.
[0021] In some embodiments, the average energy can be obtained through the following steps: The first step is to segment the ambient audio signal using the second preset time window to generate a set of audio signals within the window.
[0022] The second preset time window can be adjusted according to the needs of the actual learning scenario in order to accurately capture short-term audio feature changes.
[0023] The second step is to calculate the average energy of the audio signal set within the window and generate the average energy.
[0024] The formula for calculating average energy is as follows: .in, This represents the audio temporal amplitude value of the nth sampling point within a short time window. It is a discrete-time sequence of ambient sound data obtained through microphone acquisition and A / D (Analog / Digital) conversion. For a sampling moment corresponding to a short-time analysis window, N is the total number of sampling points included in the analysis window.
[0025] Average energy reflects the overall intensity of ambient sound. When focusing on studying, the ambient sound energy is usually more stable, while when distracted, there may be sudden changes in energy.
[0026] In some embodiments, the spectral centroid can be obtained through the following steps: The first step is to calculate the amplitude spectrum of each audio signal within the window in the set of audio signals within the window.
[0027] By performing Fourier transform and other processing on the audio signal, the time-domain signal is converted into a frequency-domain signal, and the amplitude spectrum is obtained, thus reflecting the energy distribution of different frequency components.
[0028] The second step is to calculate the spectral centroid based on the amplitude spectrum of the audio signal within each window.
[0029] The formula for calculating the spectral centroid is as follows: Where k is the frequency index of the amplitude spectrum. This refers to the amplitude spectrum. The centroid of the spectrum can distinguish the frequency characteristics of a sound. For example, the centroid of the spectrum of learning-related sounds such as the sound of a pen rubbing or turning pages differs from that of distraction-related sounds such as speaking or knocking, providing a basis for judging behavioral states.
[0030] In some embodiments, the incident rate can be obtained through the following steps: The first step is to calculate the energy of each audio signal within the set of audio signals within the window.
[0031] By capturing changes in the intensity of audio signals through energy calculations, sudden sound events can be identified.
[0032] The second step is to compare the energy of the audio signal in each window with a preset energy threshold and count the number of sudden events.
[0033] Preset energy threshold ,in The mean energy represents the energy of all short-time windows within the analysis period. The statistical average value is used as a benchmark intensity to reflect the overall energy level of environmental sound; The standard deviation of energy represents the energy of the short-time window. Relative to mean energy The degree of fluctuation is used to characterize the amplitude of changes in sound energy; This is an adjustment coefficient, an empirically set proportional coefficient used to adjust the sensitivity of the sudden sound detection threshold. The higher the value, the less likely the system is to misinterpret ordinary sounds as sudden events.
[0034] When the audio signal energy in the window This was recorded as an emergency.
[0035] The third step is to calculate the ratio of the number of emergencies to the unit time to generate the emergency rate.
[0036] The incident rate reflects the frequency of sudden sounds in the environment. Distracting behaviors (such as suddenly banging on a table or playing with objects) often lead to an increased incident rate.
[0037] In some embodiments, the speech detection ratio can be obtained through the following steps: The first step is to calculate the energy and zero-crossing rate of each audio signal in the set of audio signals within the window.
[0038] Zero-crossing rate refers to the number of times an audio signal crosses the zero level per unit time. It can reflect the time-domain fluctuation characteristics of sound. The zero-crossing rate of speech signals usually has a specific range.
[0039] The second step is to compare the energy of the audio signal in each window with a preset energy threshold, and to compare the zero-crossing rate of the audio signal in each window with a preset zero-crossing rate threshold, and to count the number of voices.
[0040] When the energy of the audio signal within the window (E is the preset energy threshold) and zero-crossing rate When Z is the preset zero-crossing rate threshold, the window is determined to be a voice window, and the number of voice windows is accumulated. .
[0041] The third step is to calculate the ratio of the number of speech samples to the total number of windows to generate the speech detection ratio.
[0042] Speech detection ratio (N is the total number of windows) can reflect the proportion of speech signals in the environment. When users are focused on learning, the proportion of speech is usually low, while when they are distracted, the proportion of speech may increase due to talking.
[0043] In some embodiments, learning behavior sounds and distraction sounds can be obtained by inputting audio features into a classifier to generate learning labels or distraction labels.
[0044] The classifier can be a binary classifier, such as an SVM (Support Vector Machine) classifier or a lightweight CNN (Convolutional Neural Network) classifier. Audio features can include, but are not limited to, MFCC (Mel-Frequency Cepstral Coefficients), spectral centroid, average energy, burst rate, etc. Two labeled datasets are constructed: a learning sound dataset (containing writing sounds, page-turning sounds, reading sounds, etc.) and a distraction sound dataset (containing table-tapping sounds, object-playing sounds, walking sounds, etc.). The classifier is trained using the labeled datasets to obtain the discrimination boundary between learning sounds and distraction sounds. The real-time extracted audio features are input into the trained classifier, which outputs the corresponding learning label or distraction label. In addition, the output time-series labels are smoothed for 1-2 seconds to reduce misclassification caused by instantaneous noise and improve the reliability of the labels.
[0045] Step 103: Input the radar echo features and audio features into the multimodal fusion model and output the user's attention state information.
[0046] In this embodiment, the aforementioned execution entity can input radar echo features and audio features into a multimodal fusion model and output the user's attention state information.
[0047] The extracted radar echo features and audio features are input into a multimodal fusion model. Through deep fusion and analysis of the dual-modal features, the model can output the user's attention state information. This attention state information can be used to characterize the user's attention level, such as an attention score. The focus score provides a clear and intuitive quantification of a user's level of focus; a higher score indicates a higher level of focus.
[0048] If the score is below the attention rating threshold, it indicates distraction; if the score remains stable, it indicates a high level of focus. The system can provide feedback through smart teaching aids or learning terminals.
[0049] In some embodiments, the multimodal fusion model may include a cross-attention layer for cross-attention fusion of radar echo features and audio features. The input layer of the multimodal fusion module may incorporate temporal consistency regularization constraints for smoothing and denoising the attention curve.
[0050] Multimodal fusion models can employ improved cross-attention fusion models based on Transformers, with improvements including, but not limited to: By incorporating a Cross-Attention mechanism between radar echo features and audio features, the model can adaptively identify and focus on modal features that are more critical to attention assessment, thereby improving the targeting of the fusion. Introducing spectral entropy stability as an additional coding feature enhances the model's sensitivity to changes in signal characteristics when focus decreases, thereby improving the timeliness of evaluation. Adding time consistency regularization constraints to the output layer effectively suppresses the impact of instantaneous noise on the scoring results, making the focus curve smoother and the evaluation results more stable. The model structure is optimized for lightweight design, resulting in fewer parameters compared to the standard Transformer. This makes it more suitable for deployment at the edge (such as smart teaching aids and learning terminals) to meet real-time evaluation requirements.
[0051] The attention state information acquisition method provided in this disclosure achieves completely imperceptible acquisition of user attention state information through dual-modal data acquisition of millimeter-wave radar and audio sensors, avoiding the inconvenience of wearable devices and the privacy leakage problem of cameras. By extracting the spectral stability characteristics of radar echoes and the multi-dimensional characteristics of audio, it comprehensively captures the user's micro-motion state and environmental behavior. Combined with an improved multimodal fusion model, it significantly improves the robustness, accuracy and interpretability of attention assessment, and can be widely applied to various scenarios such as home learning desktops, classrooms, and smart teaching aids.
[0052] Figure 2 A flow 200 of another embodiment of the attention state information acquisition method according to the present disclosure is shown. This attention state information acquisition method includes the following steps: Step 201: Preprocess the radar echo signal to generate a micro-motion timing signal.
[0053] In this embodiment, the execution entity of the focus state information acquisition method can perform a series of preprocessing operations on the radar echo signal acquired by the millimeter-wave radar. The purpose is to filter noise, extract effective information, and transform the original signal into a micro-motion time-series signal that can be used for feature extraction, thus laying the foundation for subsequent focus-related feature analysis.
[0054] Preprocessing may include, but is not limited to: clutter removal, range-Doppler transformation, phase unwrapping, and time-frequency analysis.
[0055] In some embodiments, clutter removal may include the following steps: The first step is to calculate the average background amplitude of N consecutive frames of radar snapshot signals in the radar echo signal.
[0056] The reference background modeling method is used to calculate the mean background amplitude from N consecutive frames of radar snapshot signals. Where i is the snapshot index and k is the signal sampling point index. Let N be the signal value of the snapshot signal in the i-th frame at sampling point k. The range of N can be adjusted according to the actual scene to ensure the accuracy of background modeling.
[0057] The second step is to perform the following processing steps for each frame of radar snapshot signal: calculate the difference between the absolute value of the background amplitude and the mean value of the background amplitude of the radar snapshot signal in that frame; in response to determining that the difference is less than a preset amplitude threshold, set the radar snapshot signal in that frame to zero.
[0058] For the current frame signal Calculate the difference .right Set an amplitude threshold T, which can be adaptively adjusted based on sample statistics according to the actual scenario. If... If the signal is considered to be static noise generated by static objects such as desktops and walls, it is set to zero to filter out invalid interference.
[0059] The third step is to perform one-dimensional median filtering smoothing on the processed N consecutive frames of radar snapshot signals.
[0060] Median filtering further eliminates residual noise, making the signal more stable and obtaining a signal with complete clutter suppression, providing a cleaner input for subsequent processing.
[0061] In some embodiments, the distance-Doppler transformation may include the following steps: The first step is to organize the radar echo signal into the radar matrix according to the linear frequency modulated signal (chirp signal).
[0062] The input radar matrix is organized according to the linear frequency modulated signal (chirp). Where n is the number of snapshots, corresponding to the continuously transmitted chirp signal sequence, and m is the number of sampling points, corresponding to the sampled data of each chirp signal.
[0063] The second step is to perform a range-direction FFT (Fast Fourier Transform) and a velocity-direction FFT on each linear frequency modulated signal to generate a range-Doppler matrix.
[0064] Perform a distance-oriented FFT on each chirp signal: The distance dimension information of the signal is extracted by FFT through distance; Perform a speed-direction FFT in the slow time direction: The velocity dimension information of the signal is extracted by the velocity from the FFT, and finally a two-dimensional matrix containing distance and velocity information is obtained.
[0065] The third step is to perform logarithmic amplitude transformation on the range Doppler matrix.
[0066] The dynamic range of the original Doppler matrix is large, and direct analysis is easily affected by noise. By using logarithmic amplitude processing to compress the dynamic range of the data, the characteristics of weak signals are enhanced, while noise is suppressed, making it easier to extract the effective information in the matrix.
[0067] In some embodiments, phase unwrapping may include the following steps: The first step is to obtain the encapsulated phase sequence of the radar echo signal.
[0068] After the radar echo signal is processed as described above, a complex signal is obtained, the phase component of which is constrained by the periodicity of the trigonometric function. Within the interval, a discontinuous package phase sequence is formed. .
[0069] The second step is to perform the following processing steps for each package phase: calculate the increment of the current package phase relative to the previous package phase; if the increment is greater than π, subtract 2π from the current package phase; if the increment is less than -π, add 2π to the current package phase.
[0070] Calculate the increment by traversing the package phase sequence ,when When this occurs, it indicates the presence of a phase transition. ;when At that time, This operation eliminates phase jumps.
[0071] The third step is to perform cumulative correction on the processed wrapped phase sequence to generate the unwrapped phase.
[0072] By accumulating the phase sequences after jump correction, a continuous unwrapped phase is obtained. This restores the true phase change trend of the signal, providing an accurate phase basis for subsequent time-frequency analysis and micro-motion feature extraction.
[0073] In some embodiments, time-frequency analysis may include the following steps: The first step is to slide and truncate the radar echo signal according to the preset Hamming window and multiply it to generate multiple window segments.
[0074] The length L of the Hamming window can be set to 256-1024 points, and the overlap rate of the window is set to 50%. By sliding the Hamming window to truncate, the continuous signal is divided into multiple short-time window segments, which not only ensures the stability of the signal within each window segment, but also avoids information loss at the window boundary through overlap.
[0075] The second step, for each window segment, is to perform the following processing steps: Perform an FFT on the window segment to generate a time-frequency domain signal matrix. ,in, For frequency dimension, As a time dimension, the numerical values of the matrix elements represent moments. ,frequency The signal energy intensity at that point; take the absolute value of the time-frequency domain signal matrix. Alternatively, its logarithm can be used as a time-frequency plot; depending on actual needs, the time-frequency plot can be smoothed on the time axis or frequency axis to further reduce noise interference and make the time-frequency characteristics clearer.
[0076] Time-frequency analysis transforms the radar echo signal in the time domain into a two-dimensional time-frequency graph, clearly showing the frequency variation of the signal over time, thus providing a foundation for subsequent extraction of spectral stability features.
[0077] Step 202: Within the first preset time window, analyze the spectrum of the micro-motion timing signal to generate radar echo characteristics.
[0078] In this embodiment, the aforementioned execution entity performs in-depth analysis of the spectrum of the preprocessed micro-motion timing signal within a first preset time window, extracting core features that directly reflect the user's focus state. The first preset time window can range from 5 to 30 seconds and can be flexibly adjusted according to the actual needs of the learning scenario (such as short-term homework at home or long-term classroom lessons) to ensure that the features can accurately capture changes in focus state within that time period.
[0079] Radar echo characteristics may include, but are not limited to: spectral entropy, stability index, and temporal variability.
[0080] In some embodiments, spectral entropy can be obtained through the following steps: The first step is to perform an FFT on the micro-motion timing signal within the first preset time window to generate an amplitude spectrum.
[0081] The amplitude spectrum is obtained by converting the micro-motion time-series signal from the time domain to the frequency domain using FFT. The amplitude spectrum reflects the energy distribution of a signal at different frequency components.
[0082] The second step is to normalize the amplitude spectrum to generate a probability distribution.
[0083] Amplitude spectrum Normalization process yields the probability distribution. .in, The probability distribution is the sum of the energies of all frequency components of the amplitude spectrum. This reflects the energy proportion of each frequency component.
[0084] The third step is to calculate the spectral entropy based on the probability distribution.
[0085] The formula for calculating spectral entropy is: Spectral entropy is an indicator reflecting the degree of disorder in the distribution of signal energy. The more dispersed the energy distribution, the larger the spectral entropy H value; the more concentrated the energy distribution, the smaller the spectral entropy H value. When a user is focused, there are fewer and more regular micro-movements, the signal energy distribution is concentrated, and the spectral entropy is lower; when distracted, micro-movements are frequent and irregular, the signal energy distribution is dispersed, and the spectral entropy is higher.
[0086] In some embodiments, the stability index can be obtained by taking the reciprocal of the spectral entropy.
[0087] Stability Indicators This metric is positively correlated with focus; the smaller the spectral entropy H, the larger the stability index S, indicating higher user focus stability; conversely, the smaller the stability index S, the lower the focus stability. This transformation converts the abstract spectral entropy into a more intuitive quantitative indicator of focus stability.
[0088] In some embodiments, the temporal variability can be obtained through the following steps: The first step is to perform phase processing, time-frequency analysis, and spectral entropy calculation on the micro-motion time-series signal to generate a spectral stability sequence.
[0089] By employing a process of phase unwrapping, time-frequency analysis, spectral entropy calculation, and stability index conversion, a continuously time-varying spectral stability sequence is obtained. This sequence can dynamically reflect the time-varying trend of user focus stability.
[0090] The second step involves sliding the spectral stability sequence into segments with a preset window length and a preset step size to generate a set of stability values within the window.
[0091] The preset window length W can be set to 3-10 seconds, and the preset step size H can be set to 1-3 seconds. These values can be adjusted according to the actual evaluation accuracy requirements. The stability sequence is segmented by a sliding window to ensure that short-term fluctuations in focus can be captured.
[0092] The third step is to calculate the mean and standard deviation of the set of stability values within the window.
[0093] Calculate the mean U of all stability values within each sliding window. The mean U reflects the average level of focus stability within that window. Calculate the standard deviation S of the samples within the window. The standard deviation S reflects the degree of fluctuation in focus stability within that window.
[0094] The fourth step is to calculate the time series variability based on the mean and standard deviation.
[0095] The time series variability is calculated using the normalized standard deviation method, and the formula is as follows: This indicator takes into account both average stability and volatility, and can more comprehensively reflect the trend of fluctuations.
[0096] The fifth step is to generate a time-series variability sequence as the sliding window moves along the time axis.
[0097] The sliding window moves continuously along the time axis, outputting a time variability value with each move, ultimately forming a continuous time variability sequence. This sequence provides a dynamic basis for assessing the fluctuation characteristics of focus, including the temporal variability of user focus. The value is relatively small, but the temporal variability is relatively large when distracted.
[0098] Step 203: Extract features from the audio signal of the user's environment to generate audio features.
[0099] Step 204: Input the radar echo features and audio features into the multimodal fusion model and output the user's attention state information.
[0100] In this embodiment, the specific operations of steps 203-204 have been described. Figure 1 Steps 102-103 in the illustrated embodiments are described in detail and will not be repeated here.
[0101] The attention state information acquisition method provided in this disclosure effectively filters clutter and noise through multi-step refined preprocessing of radar echo signals, and accurately extracts the time-series signal reflecting the user's micro-motion state. By extracting core features such as spectral entropy, stability index, and temporal variability, it realizes the quantification of the user's attention stability and fluctuation trend. Combined with multi-dimensional analysis of audio features and deep fusion of multi-modal fusion models, it further improves the accuracy and reliability of attention assessment, and provides a complete technical implementation path for the imperceptible and accurate assessment of user attention.
[0102] Figure 3 A flowchart illustrating the method for non-perceptible assessment of user focus state is shown. This flowchart fully presents the logical relationship of the entire process of the focus state information acquisition method described in this disclosure, with each step sequentially connected and progressively advancing. Specific explanations are as follows: The process begins with the “Signal Acquisition 301” step, which uses millimeter-wave radar sensing modules and environmental audio sensors deployed in the user’s learning scenario to collect radar echo signals from the user’s head and torso area, as well as environmental sound signals, to achieve seamless acquisition of dual-modal data and provide raw data for subsequent processing. After signal acquisition, the system simultaneously enters two parallel stages: "millimeter-wave signal preprocessing 302" and "audio feature extraction 303". The "millimeter-wave signal preprocessing 302" stage performs clutter removal, range-Doppler transformation, phase unwrapping, and time-frequency analysis on the radar echo signal in sequence, transforming the original radar signal into a micro-motion time-series signal that can be used for feature extraction. The "audio feature extraction 303" stage performs event detection and feature calculation on the environmental audio signal, extracting features such as average energy, spectral centroid, burst event rate, and speech detection ratio, and generates learning / distraction labels through a binary classifier to complete the mining of audio dimension features. Subsequently, the micro-motion timing signal output from the "millimeter wave signal preprocessing 302" stage enters the "spectral stability feature extraction 304" stage. Within a preset time window, core features such as spectral entropy, stability index, and timing variability are calculated through spectral analysis to quantify the user's micro-motion state and focus stability. Next, the radar echo features output from the "Spectrum Stability Feature Extraction 304" step and the audio features output from the "Audio Feature Extraction 303" step are input into the "Multimodal Fusion and Attention Assessment 305" step. An improved Transformer cross-attention fusion model is used to deeply fuse and analyze the dual-modal features, outputting a quantitative user attention score. ; After the focus score is generated, the process proceeds to the "Focus Score Judgment 306" step: If the score is less than the preset threshold, the user is determined to be in a distracted state, and the "Output Distraction Tip 307" operation is executed; if the score remains stable at a high level, the user is determined to be in a highly focused state, and the "Output High Focus State 308" operation is executed. Finally, through the "Feedback and Intervention 309" step, the above status prompts are fed back to the user or guardian via smart teaching aids or learning terminals, realizing a closed loop of assessment and feedback, and providing support for timely adjustment of learning status and improvement of learning efficiency.
[0103] This flowchart clearly demonstrates the complete technical chain of this invention from data acquisition, preprocessing, feature extraction, fusion evaluation to result feedback. It highlights the core features of seamless acquisition, dual-modal fusion, and accurate evaluation. The logical connection between each link ensures the smoothness of the evaluation process and the reliability of the evaluation results.
[0104] Further reference Figure 4As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a focus state information acquisition device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0105] like Figure 4 As shown, the attention state information acquisition device 400 of this embodiment may include: a first extraction module 401, a second extraction module 402, and an evaluation module 403. The first extraction module 401 is configured to extract features from the user's radar echo signal to generate radar echo features; the second extraction module 402 is configured to extract features from the audio signal of the user's environment to generate audio features; and the evaluation module 403 is configured to input the radar echo features and audio features into a multimodal fusion model and output the user's attention state information.
[0106] In this embodiment, the specific processing of the first extraction module 401, the second extraction module 402, and the evaluation module 403 in the focus state information acquisition device 400, and the resulting technical effects, can be referred to respectively. Figure 1 The relevant descriptions of steps 101-103 in the corresponding embodiments will not be repeated here.
[0107] In some optional implementations of this embodiment, the first extraction module 401 includes: a preprocessing submodule configured to preprocess the radar echo signal to generate a micro-motion timing signal; and an analysis submodule configured to analyze the spectrum of the micro-motion timing signal within a first preset time window to generate radar echo features.
[0108] In some optional implementations of this embodiment, the preprocessing submodule is further configured to: calculate the average background amplitude of N consecutive frames of radar snapshot signals in the radar echo signal; for each frame of radar snapshot signal, perform the following processing steps: calculate the difference between the absolute value of the background amplitude of the radar snapshot signal and the average background amplitude; in response to determining that the difference is less than a preset amplitude threshold, set the radar snapshot signal of the frame to zero; and perform one-dimensional median filtering smoothing on the processed N consecutive frames of radar snapshot signals.
[0109] In some optional implementations of this embodiment, the preprocessing submodule is further configured to: organize the radar echo signal into the radar matrix according to the linear frequency modulated signal; for each linear frequency modulated signal, perform range-to-fast Fourier transform and velocity-to-fast Fourier transform to generate a range-Doppler matrix; and perform logarithmic amplitudeization on the range-Doppler matrix.
[0110] In some optional implementations of this embodiment, the preprocessing submodule is further configured to: acquire the wrapped phase sequence of the radar echo signal; for each wrapped phase, perform the following processing steps: calculate the increment of the current wrapped phase relative to the previous wrapped phase; if the increment is greater than π, subtract 2π from the current wrapped phase; if the increment is less than -π, add 2π to the current wrapped phase; perform cumulative correction on the processed wrapped phase sequence to generate an unwrapped phase.
[0111] In some optional implementations of this embodiment, the preprocessing submodule is further configured to: truncate and multiply the radar echo signal according to a preset Hamming window to generate multiple window segments; for each window segment, perform the following processing steps: perform a fast Fourier transform on the window segment to generate a time-frequency domain signal matrix; take the absolute value or logarithm of the time-frequency domain signal matrix as a time-frequency graph; and smooth the time-frequency graph on the time axis or frequency axis.
[0112] In some optional implementations of this embodiment, the analysis submodule is further configured to: perform a fast Fourier transform on the micro-motion timing signal within a first preset time window to generate an amplitude spectrum; normalize the amplitude spectrum to generate a probability distribution; and calculate the spectral entropy based on the probability distribution.
[0113] In some optional implementations of this embodiment, the analysis submodule is further configured to: take the reciprocal of the spectral entropy to generate a stability index.
[0114] In some optional implementations of this embodiment, the analysis submodule is further configured to: perform phase processing, time-frequency analysis, and spectral entropy calculation on the micro-motion time-series signal to generate a spectral stability sequence; slide segment the spectral stability sequence with a preset window length and a preset step size to generate a set of stability values within the window; calculate the mean and standard deviation of the set of stability values within the window; calculate the time-series variability based on the mean and standard deviation; and generate a time-series variability sequence as the sliding window moves along the time axis.
[0115] In some optional implementations of this embodiment, the second extraction module 402 is further configured to: segment the environmental audio signal using a second preset time window to generate a set of audio signals within the window; calculate the average energy of the set of audio signals within the window to generate average energy.
[0116] In some optional implementations of this embodiment, the second extraction module 402 is further configured to: calculate the amplitude spectrum of each audio signal in the set of audio signals in the window; and calculate the spectral centroid based on the amplitude spectrum of each audio signal in the window.
[0117] In some optional implementations of this embodiment, the second extraction module 402 is further configured to: calculate the energy of each audio signal in the set of audio signals in the window; compare the energy of each audio signal in the window with a preset energy threshold to count the number of sudden events; calculate the ratio of the number of sudden events to the number of unit events to generate a sudden event rate.
[0118] In some optional implementations of this embodiment, the second extraction module 402 is further configured to: calculate the energy and zero-crossing rate of each audio signal in the set of audio signals in the window; compare the energy of each audio signal in the window with a preset energy threshold, and compare the zero-crossing rate of each audio signal in the window with a preset zero-crossing rate threshold, and count the number of speech signals; calculate the ratio of the number of speech signals to the total number of windows, and generate a speech detection ratio.
[0119] In some optional implementations of this embodiment, the second extraction module 402 is further configured to: input audio features into a classifier to generate learning labels or distraction labels.
[0120] In some optional implementations of this embodiment, the multimodal fusion model includes a cross-attention layer for cross-attention fusion of radar echo features and audio features, and a time consistency regularization constraint is added to the input layer of the multimodal fusion module for smoothing and noise reduction of the attention curve.
[0121] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0122] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0123] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0124] like Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0125] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0126] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the focus state information acquisition method. For example, in some embodiments, the focus state information acquisition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the focus state information acquisition method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the focus state information acquisition method by any other suitable means (e.g., by means of firmware).
[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0128] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0129] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0131] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0132] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0133] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0134] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for acquiring focus state information, comprising: Extract features from the user's radar echo signal to generate radar echo features; The audio signal of the user's environment is used to extract features and generate audio features; The radar echo features and the audio features are input into a multimodal fusion model to output the user's attention state information.
2. The method according to claim 1, wherein, The step of extracting features from the user's radar echo signal to generate radar echo features includes: The radar echo signal is preprocessed to generate a micro-motion timing signal; Within a first preset time window, the spectrum of the micro-motion timing signal is analyzed to generate the radar echo characteristics.
3. The method according to claim 2, wherein, The preprocessing of the radar echo signal includes: Calculate the average background amplitude of N consecutive frames of radar snapshot signals in the radar echo signal; For each frame of radar snapshot signal, the following processing steps are performed: calculate the difference between the absolute value of the background amplitude of the radar snapshot signal and the mean value of the background amplitude; in response to determining that the difference is less than a preset amplitude threshold, set the radar snapshot signal of the frame to zero; One-dimensional median filtering is applied to smooth the processed N consecutive frames of radar snapshot signals.
4. The method according to claim 2, wherein, The preprocessing of the radar echo signal includes: The radar echo signal is organized into a linear frequency modulated signal and input into the radar matrix; For each linear frequency modulated signal, perform a range-to-fast Fourier transform and a velocity-to-fast Fourier transform to generate a range-Doppler matrix; The distance-Doppler matrix is logarithmically magnituded.
5. The method according to claim 2, wherein, The preprocessing of the radar echo signal includes: Obtain the encapsulated phase sequence of the radar echo signal; For each package phase, perform the following processing steps: calculate the increment of the current package phase relative to the previous package phase; if the increment is greater than π, subtract 2π from the current package phase; if the increment is less than -π, add 2π to the current package phase. The processed wrapped phase sequence is cumulatively corrected to generate the unwrapped phase.
6. The method according to claim 2, wherein, The preprocessing of the radar echo signal includes: The radar echo signal is truncated and multiplied by a preset Hamming window to generate multiple window segments; For each window segment, perform the following processing steps: perform a fast Fourier transform on the window segment to generate a time-frequency domain signal matrix; take the absolute value or logarithm of the time-frequency domain signal matrix as a time-frequency graph; smooth the time-frequency graph on the time axis or frequency axis.
7. The method according to claim 2, wherein, The step of analyzing the spectrum of the micro-motion timing signal within the first preset time window includes: Within the first preset time window, a fast Fourier transform is performed on the micro-motion timing signal to generate an amplitude spectrum; The amplitude spectrum is normalized to generate a probability distribution; Calculate the spectral entropy based on the probability distribution.
8. The method according to claim 7, wherein, The step of analyzing the spectrum of the micro-motion timing signal within the first preset time window includes: The stability index is generated by taking the reciprocal of the spectral entropy.
9. The method according to claim 8, wherein, The step of analyzing the spectrum of the micro-motion timing signal within the first preset time window includes: The micro-motion timing signal is subjected to phase processing, time-frequency analysis, and spectral entropy calculation to generate a spectral stability sequence; The spectral stability sequence is segmented by a preset window length and a preset step size to generate a set of stability values within the window. Calculate the mean and standard deviation of the stability values within the window; Calculate the time series variability based on the mean and the standard deviation; As the sliding window moves along the time axis, a time series of variability is generated.
10. The method according to claim 1, wherein, The step of extracting features from the audio signal of the user's environment to generate audio features includes: The environmental audio signal is segmented using a second preset time window to generate a set of audio signals within the window; The average energy is calculated for the set of audio signals within the window to generate the average energy.
11. The method according to claim 10, wherein, The step of extracting features from the audio signal of the user's environment to generate audio features includes: Calculate the amplitude spectrum of each audio signal within the window in the set of audio signals within the window; The spectral centroid is calculated based on the amplitude spectrum of the audio signal within each window.
12. The method according to claim 11, wherein, The step of extracting features from the audio signal of the user's environment to generate audio features includes: Calculate the energy of the audio signal within each window in the set of audio signals within the window; The energy of the audio signal in each window is compared with a preset energy threshold to count the number of sudden events. Calculate the ratio of the number of sudden events to the number of events per unit to generate the sudden event rate.
13. The method according to claim 12, wherein, The step of extracting features from the audio signal of the user's environment to generate audio features includes: Calculate the energy and zero-crossing rate of each audio signal in the set of audio signals within the window; The energy of the audio signal in each window is compared with a preset energy threshold, and the zero-crossing rate of the audio signal in each window is compared with a preset zero-crossing rate threshold to count the number of voices. Calculate the ratio of the number of voices to the total number of windows to generate the voice detection ratio.
14. The method according to claim 13, wherein, The step of extracting features from the audio signal of the user's environment to generate audio features includes: The audio features are input into a classifier to generate learning labels or distraction labels.
15. The method according to any one of claims 1-14, wherein, The multimodal fusion model includes a cross-attention layer for cross-attention fusion of the radar echo features and the audio features. The input layer of the multimodal fusion module incorporates time consistency regularization constraints to smooth and reduce noise in the attention curve.
16. A device for acquiring attention state information, comprising: The first extraction module is configured to extract features from the user's radar echo signal and generate radar echo features. The second extraction module is configured to extract features from the audio signal of the user's environment and generate audio features. The evaluation module is configured to input the radar echo features and the audio features into a multimodal fusion model and output the user's attention state information.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-15.
18. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method of any one of claims 1-15.
19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-15.