Audio feedback control method and system based on two-stage individualized inference and intervention response evaluation

CN122785984APending Publication Date: 2026-09-22UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611109448.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0009]本发明的目的就是为了克服上述现有技术存在的缺陷而提供一种基于二阶段个体化推理和干预响应评估的音频反馈控制方法及系统,以解决或部分解决RR间隔数据局部异常点难以识别、群体分类模型对个体差异适应能力不足、复杂个体化模型实时计算开销较大以及固定音频反馈不能根据生理响应动态调整的问题

Benefits of technology

(1)提高状态识别对个体差异的适应性:本发明根据用户历史融合特征嵌入向量及其状态标签建立并归一化个人类别原型,将随机森林分类器输出的全局预测概率与个体化匹配概率进行融合,使状态识别同时利用群体训练信息和个人校准信息,有助于降低用户个体差异对状态识别结果的影响,并避免针对每个用户重新训练完整分类模型。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122785984A_ABST
    Figure CN122785984A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of audio feedback control method and system based on two-stage individualized inference and intervention response evaluation, comprising: obtaining RR interval sequence and preprocessing;Original RR interval sliding window is constructed and multi-dimensional heart rate variability feature is extracted;Fusion feature embedding vector is obtained by convolutional neural network sequence feature extractor and heart rate variability feature encoder;Global prediction probability is output using random forest classifier, and individual class prototype is established according to user historical data in personal memory unit, and individual matching probability is obtained;First stage prediction probability is obtained by fusion, and class interval is calculated, when interval does not reach threshold value, context data is input into neural process model to obtain second stage individualized prediction probability, and effective state recognition result is determined according to this;According to the result, audio is played, and according to the change of state prediction probability or heart rate variability feature before and after intervention, audio output device continues to play, adjustment strategy or stop playing is controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wearable physiological signal processing and feedback control technology, and in particular to an audio feedback control method and system based on two-stage individualized reasoning and intervention response assessment. Background Technology

[0002] Heart rate variability (HRV) is a physiological parameter formed by continuous changes in heart rate intervals and can be used to characterize changes in human autonomic nervous activity. With the development of wearable heart rate monitoring devices, acquiring real-time RR interval sequences and identifying stress-related states through these devices has become a method for assisting in the monitoring of physiological states.

[0003] However, existing pressure-related state identification and feedback control technologies based on RR interval data still have the following problems.

[0004] First, the RR interval sequence output by heart rate monitoring devices may be affected by factors such as unstable device contact, user movement, and communication packet loss. While filtering using only a fixed physiological threshold can remove data points that significantly exceed the normal range, it is difficult to identify local anomalies that are within the reasonable physiological range but show significant abnormalities relative to the user's recent RR interval changes, thus affecting subsequent feature construction and status recognition results.

[0005] Second, existing methods typically only use artificially constructed heart rate variability features, or only input the original RR interval sequence into the machine learning model, without fully combining the local variation patterns of the original RR interval sequence with the statistical features of heart rate variability, which limits the feature expression ability.

[0006] Third, state classification models trained on group data are difficult to adequately adapt to the physiological differences among different users. If a uniform group model is used for all samples, uncertain prediction results are likely to occur for boundary samples; if computationally intensive individualized inference is performed on all samples, it will increase real-time processing overhead.

[0007] Fourth, existing audio feedback systems typically play fixed audio based on a single state recognition result, without utilizing newly collected RR interval data after audio playback to evaluate the user's actual physiological response. This makes it difficult to dynamically decide whether to continue playback, adjust the audio strategy, or stop playback based on the intervention response.

[0008] Therefore, there is an urgent need for a method and system that can perform robust preprocessing of RR interval sequences, integrate original sequence features with heart rate variability features, perform individualized two-stage inference for boundary samples, and dynamically control audio output based on physiological responses after audio playback. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the existing technology by providing an audio feedback control method and system based on two-stage individualized reasoning and intervention response assessment, so as to solve or partially solve the problems of difficulty in identifying local outliers in RR interval data, insufficient adaptability of group classification models to individual differences, large real-time computational overhead of complex individualized models, and inability of fixed audio feedback to be dynamically adjusted according to physiological response.

[0010] The objective of this invention can be achieved through the following technical solutions: One aspect of the present invention provides an audio feedback control method based on two-stage individualized reasoning and intervention response assessment, comprising the following steps: The user's RR interval sequence is obtained and processed through a sliding window to extract multidimensional heart rate variability features; Sequence feature embedding vectors are extracted based on sliding window, and heart rate variability feature embedding vectors are extracted based on the multidimensional heart rate variability features. Fusion feature embedding vectors are obtained through fusion processing. The fused feature embedding vector is classified to obtain the global predicted probability corresponding to multiple pressure-related state categories; Based on the pre-acquired user history fusion feature embedding vector and its state label, establish the personal category prototype corresponding to each stress-related state category, and obtain the individualized matching probability based on the similarity between the current fusion feature embedding vector and the normalized personal category prototype. The global prediction probability and the individualized matching probability are fused to obtain the first-stage prediction probability, and the class interval between the highest probability and the second highest probability in the first-stage prediction probability is calculated. When the category interval reaches a preset interval threshold, the effective state recognition result is determined based on the first stage prediction probability; when the category interval does not reach the preset interval threshold, the current fusion feature embedding vector, the user's historical fusion feature embedding vector, and the state label are input into the neural process model to obtain the second stage individualized prediction probability, and the effective state recognition result is determined based on the second stage individualized prediction probability. Based on the valid state recognition results, an audio feedback strategy is determined, and the audio output device is controlled to play the corresponding audio. During or after the audio playback, the newly generated RR interval sequence after the audio playback is acquired again. The state recognition result after intervention is obtained according to the aforementioned state recognition process. Based on the changes in the state recognition results before and after intervention, the state prediction probability, or the heart rate variability characteristics, the audio output device is controlled to continue playing, adjust the audio feedback strategy, or stop playing.

[0011] As a preferred technical solution, before the sliding window processing, the RR interval sequence is further preprocessed, and the preprocessing includes the following steps: RR interval data points that exceed the preset physiological reasonable range are identified as extreme outliers; Within a local statistical window consisting of multiple consecutive valid RR interval data points, calculate the median and median absolute deviation of the RR interval data in the local statistical window; Based on the deviation between the RR interval data point to be detected and the median, and the absolute deviation of the median, it is determined whether the RR interval data point to be detected is a local outlier. The extreme and local anomalies are filtered out.

[0012] As a preferred technical solution, sliding window processing is achieved through a pre-constructed RR interval sliding window, wherein the RR interval sliding window includes multiple consecutive valid RR interval data points, and the multidimensional heart rate variability features include at least two types of time-domain features, statistical features, frequency-domain features, and nonlinear features.

[0013] As a preferred technical solution, a sequence feature embedding vector is extracted by a convolutional neural network sequence feature extractor, and a heart rate variability feature embedding vector is extracted by a heart rate variability feature encoder. The convolutional neural network sequence feature extractor includes at least one one-dimensional convolutional layer, a non-linear activation layer, a pooling layer, and a feature mapping layer. The heart rate variability feature encoder includes at least two fully connected layers. The sequence feature embedding vector and the heart rate variability feature embedding vector are concatenated and feature mapped to obtain the fused feature embedding vector.

[0014] As a preferred technical solution, the classification of the fused feature embedding vector is achieved by a random forest classifier. The random forest classifier includes multiple decision trees, each of which outputs the classification result of the stress-related state category. The global prediction probability is obtained based on the voting ratio or the average probability of each category of the multiple decision trees.

[0015] As a preferred technical solution, the process of establishing the personal category prototype includes the following steps: RR interval sequences are acquired during the user's performance of a status-labeled personal calibration task; Multiple RR interval sliding windows are constructed based on the RR interval sequence, and corresponding fusion feature embedding vectors are obtained for each. The mean, weighted mean, or cluster center of the fusion feature embedding vectors belonging to the same stress-related state category are used as the individual category prototype for the corresponding category. The individual category prototypes are normalized. Store the fused feature embedding vector and its state label.

[0016] As a preferred technical solution, the global prediction probability and the individualized matching probability are weighted and fused to obtain the first-stage prediction probability. For stress-related state categories for which no personal category prototype has been established, the global prediction probability output by the random forest classifier is retained.

[0017] As a preferred technical solution, the category interval is the difference between the highest probability and the second highest probability in the first stage prediction probability. When the category interval is lower than the preset interval threshold, the current RR interval sliding window is determined as the boundary sample, and the neural process model is triggered to perform the second stage of individualized inference.

[0018] As a preferred technical solution, the stored user history fusion feature embedding vector and its state label are used as context data of the neural process model, and the current fusion feature embedding vector is used as target data. The neural process model outputs the individualized prediction probability of the current target data belonging to each stress-related state category.

[0019] Another aspect of the present invention provides an audio feedback control system based on two-stage individualized reasoning and intervention response assessment, for implementing the aforementioned audio feedback control method based on two-stage individualized reasoning and intervention response assessment, the system comprising: The RR interval data acquisition module is used to acquire the user's RR interval sequence from the heart rate monitoring device; A two-level anomaly filtering module is used to perform physiologically reasonable interval filtering and local statistical filtering on the RR interval sequence; The feature construction module is used to construct the original RR interval sliding window and extract multidimensional heart rate variability features; A convolutional neural network sequence feature extraction module is used to generate a sequence feature embedding vector based on the original RR interval sliding window; The heart rate variability feature encoding and fusion module is used to generate a heart rate variability feature embedding vector based on multidimensional heart rate variability features, and to fuse the sequence feature embedding vector and the heart rate variability feature embedding vector into a fused feature embedding vector. The random forest classification module is used to generate global prediction probabilities based on the fused feature embedding vector. The personal memory and personal category prototype module is used to store the user's historical fusion feature embedding vector and status label, establish and normalize the personal category prototype based on the user's historical fusion feature embedding vector, and calculate the individualized matching probability based on the similarity between the current fusion feature embedding vector and the normalized personal category prototype. The boundary sample judgment and neural process reasoning module is used to determine whether to call the neural process model based on the category interval of the first-stage prediction probability obtained by fusing the global prediction probability and the individualized matching probability, and outputs the second-stage individualized prediction probability when calling it to determine the effective state recognition result. The audio feedback control module is used to control the audio output device to play audio based on the valid status recognition result. The intervention response assessment module is used to evaluate the intervention response based on the newly acquired RR interval sequence after audio playback, and to control the audio output device to continue playback, adjust the audio feedback strategy, or stop playback.

[0020] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) Improve the adaptability of state recognition to individual differences: This invention establishes and normalizes the personal category prototype based on the user's historical fusion feature embedding vector and its state label, and fuses the global prediction probability output by the random forest classifier with the individualized matching probability, so that state recognition can utilize both group training information and personal calibration information at the same time, which helps to reduce the impact of individual differences of users on the state recognition results and avoids retraining the complete classification model for each user.

[0021] (2) Reduce the computational overhead of individualized reasoning: The present invention judges the uncertainty of the current sample based on the category interval between the highest probability and the second highest probability in the first stage prediction probability. Only when the category interval is lower than the preset interval threshold, the neural process model is called to perform the second stage of individualized reasoning, which helps to improve the individualized judgment ability of boundary samples while reducing the computational overhead of performing complex reasoning on non-boundary samples.

[0022] (3) Forming a closed loop of audio feedback based on physiological response: During or after the audio playback, the present invention reacquires the RR interval sequence, evaluates the intervention response based on the state recognition results before and after the intervention, the state prediction probability or the change in heart rate variability characteristics, and controls the audio output device to continue playing, adjust the audio feedback strategy or stop playing, so that the audio output can be dynamically adjusted according to the user's actual physiological response. Attached Figure Description

[0023] Figure 1 This is an overall flowchart of the audio feedback control method in the embodiment; Figure 2 This is a module architecture diagram of the pressure-related state recognition and audio feedback control system in an embodiment of the present invention; Figure 3 This is a flowchart of the two-stage anomaly filtering of RR interval data in an embodiment of the present invention; Figure 4This is a schematic diagram of the structure of bi-branch feature fusion and random forest classification in an embodiment of the present invention; Figure 5 This is a flowchart of the two-stage individualized reasoning in an embodiment of the present invention; Figure 6 This is a flowchart illustrating personal calibration, personal category prototyping, and personal memory unit updates in an embodiment of the present invention. Figure 7 This is a flowchart illustrating the closed-loop control process for audio feedback and intervention response evaluation in an embodiment of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0025] Example 1 To address the problems existing in the aforementioned prior art, this embodiment provides an audio feedback control method based on two-stage individualized reasoning and intervention response assessment. (See [link to relevant documentation]). Figure 1 The method includes the following steps: Step S1: Obtain the user's RR interval sequence.

[0026] In this embodiment, the heart rate monitoring device uses Polar H10, which receives heart rate data and RR interval sequences output by the heart rate monitoring device via Bluetooth Low Energy protocol. It should be noted that Polar H10 is only an optional implementation device and does not constitute a limitation on the type of heart rate monitoring device.

[0027] This embodiment directly uses the RR interval data output by the heart rate monitoring device, without performing raw ECG signal filtering and R wave detection.

[0028] Step S2: Filtering within the physiologically reasonable range and filtering for local statistical anomalies.

[0029] See Figure 3 The flowchart for two-stage anomaly filtering of RR interval data includes the following sub-steps: Step S201: Receive one RR interval data point.

[0030] The received RR interval data is parsed, timestamp recorded, and format converted. The RR interval data that meets the quality requirements is written into the data buffer in the order of collection.

[0031] Step S202: Determine if the data point is within the physiologically reasonable range. If yes, proceed to step S203; otherwise, mark the RR interval data point as an extreme outlier and filter it out.

[0032] Perform physiologically reasonable interval filtering. In one implementation, RR interval data points with intervals less than or equal to 300 milliseconds and greater than or equal to 2000 milliseconds are identified as extreme outliers and filtered out.

[0033] Step S203: Add a local statistics window.

[0034] Local statistical filtering is performed on the data filtered through physiologically reasonable intervals. A local statistical window is used to store multiple consecutive historical valid RR interval data points. In one implementation, the local statistical window stores a maximum of 15 historical valid RR interval data points.

[0035] Step S204: Calculate the window center and MAD.

[0036] When the window contains at least 5 historical valid RR interval data points, calculate the absolute deviation between the median and the middle of the window. The median of the RR interval data in the local statistical window is expressed as: in, This represents the median of the historical valid RR interval data within the local statistical window. This indicates the number of historical valid RR interval data points in the local statistics window.

[0037] The absolute deviation of the median is expressed as: in, Indicates the absolute deviation of the median. Represents the first in the local statistics window RR interval data points.

[0038] Step S205: Calculate the relative median deviation of the current point.

[0039] Calculate robustness metrics based on median absolute deviation: in, Indicates a robust measure. This indicates the preset minimum scale parameter, used to avoid incorrectly filtering out normal data points due to excessively small absolute deviation of the median.

[0040] Step S206: Determine if the deviation is too large. If not, mark it as valid RR data and proceed to step S207; otherwise, mark the RR interval data point as a local outlier and filter it out.

[0041] When the RR interval data points to be detected satisfy At that time, the RR interval data points to be detected are identified as local statistical outliers. Among them, Configurable local anomaly detection coefficients.

[0042] In one implementation method Set to 20 milliseconds. Set to 4. The above parameters are only one implementation method and do not constitute a limitation on the scope of protection of this invention.

[0043] Step S207: Write to the RR buffer.

[0044] Effective RR interval data filtered by two levels of anomalies are written into the RR interval buffer in chronological order, providing input for the construction of the original RR interval sliding window and the extraction of heart rate variability features.

[0045] Step S3: Construct the original RR sliding window and extract multidimensional heart rate variability (HRV) features.

[0046] Read multiple consecutive valid RR interval data points from the RR interval buffer to construct the original RR interval sliding window.

[0047] In one implementation, each original RR interval sliding window comprises 64 consecutive valid RR interval data points. Both the sliding window length and the sliding step size are configurable parameters.

[0048] Step S4: Convolutional Neural Network (CNN) sequence feature extraction and HRV feature encoding are fused.

[0049] See Figure 4 This is a schematic diagram of the structure of feature fusion and random forest classification using a two-branch system (i.e., sequence feature embedding branch and heart rate variability feature embedding branch).

[0050] (1) Sequence feature embedding branch.

[0051] The original RR-interval sliding window is normalized before being input into the convolutional neural network sequence feature extractor. The normalization process can be represented as: in, and These are the mean and standard deviation obtained from the training data, respectively. A preset constant is used to prevent the denominator from being zero.

[0052] A convolutional neural network sequence feature extractor is used to map the original RR interval sliding window into a sequence feature embedding vector; a heart rate variability feature encoder is used to map multidimensional heart rate variability features into a heart rate variability feature embedding vector.

[0053] In one implementation, the sequence feature embedding vector and the heart rate variability feature embedding vector are both 32-dimensional. The two feature embedding vectors are concatenated and then passed through a feature mapping layer to obtain a 64-dimensional fused feature embedding vector.

[0054] A convolutional neural network sequence feature extractor includes multiple one-dimensional convolutional layers, batch normalization layers, non-linear activation layers, pooling layers, and feature mapping layers.

[0055] In one implementation, the convolutional neural network sequence feature extractor sequentially includes: A one-dimensional convolutional layer with 16 output channels and a kernel size of 3, a batch normalization layer, a ReLU activation layer, a max pooling layer, and a Dropout layer; A one-dimensional convolutional layer with 32 output channels and a kernel size of 3, a batch normalization layer, a ReLU activation layer, a max pooling layer, and a Dropout layer; The system consists of a one-dimensional convolutional layer with 64 output channels and a kernel size of 3, a batch normalization layer, a ReLU activation layer, an adaptive average pooling layer, and a fully connected mapping layer.

[0056] The convolutional neural network sequence feature extractor outputs a 32-dimensional sequence feature embedding vector.

[0057] (2) Embedded branch of heart rate variability features.

[0058] Multidimensional heart rate variability features were extracted from the preprocessed RR interval sequences.

[0059] In one implementation, the multidimensional heart rate variability characteristics include the following 32 items: Mean RR interval, median RR interval, minimum RR interval, maximum RR interval, RR interval range, SDNN, RMSSD, SDSD, pNN20, pNN50, mean heart rate, standard deviation of heart rate, RR interval coefficient of variation, RR interval interquartile range, absolute deviation of RR interval median, RR interval skewness, RR interval kurtosis, 10th percentile of RR interval, 25th percentile of RR interval, 75th percentile of RR interval, 90th percentile of RR interval, very low frequency power, low frequency power, high frequency power, total power, low frequency normalized power, high frequency normalized power, low frequency power to high frequency power ratio, Poincaré plot SD1, Poincaré plot SD2, SD1 to SD2 ratio, and sample entropy.

[0060] Preferably, for frequency domain features requiring calculation of long data sequences, an independent long-term RR interval buffer is set up. When the amount of data in the long-term buffer reaches the preset frequency domain analysis condition, the frequency domain features are calculated; when the amount of data does not reach the preset condition, state identification can be temporarily suspended or a backup feature configuration that does not include frequency domain features can be used.

[0061] The multidimensional heart rate variability features are standardized and then input into the heart rate variability feature encoder.

[0062] The heart rate variability feature encoder consists of a first fully connected layer with an input dimension equal to the multidimensional heart rate variability feature dimension, and a second fully connected layer for outputting a heart rate variability feature embedding vector. A non-linear activation layer and an optional Dropout layer are placed between the first and second fully connected layers.

[0063] In one implementation, the first fully connected layer outputs 64-dimensional intermediate features, and the second fully connected layer outputs a 32-dimensional heart rate variability feature embedding vector.

[0064] The 32-dimensional sequence feature embedding vector and the 32-dimensional heart rate variability feature embedding vector are concatenated into a 64-dimensional vector, and then a 64-dimensional fused feature embedding vector is obtained through a fusion feature mapping layer.

[0065] Step S5: Random Forest (RF) global classification is matched with the individual prototype.

[0066] See Figure 4 After the convolutional neural network sequence feature extractor and heart rate variability feature encoder are trained, they are used to generate fused feature embedding vectors corresponding to the training samples, and then the fused feature embedding vectors are used to train a random forest classifier.

[0067] A random forest classifier consists of multiple decision trees, each trained on randomly selected training samples and feature subsets. During prediction, the global predicted probability is obtained based on the voting proportions or average class probabilities of the multiple decision trees for each stress-related state category. .

[0068] The number of decision trees, maximum depth, minimum number of leaf node samples, feature selection method, and class weights in a random forest are all configurable parameters.

[0069] During the training phase, the RR interval data with state labels are divided into multiple sample windows according to the preset window length and sliding step size. Two-level anomaly filtering, feature construction and feature fusion are performed on each sample window.

[0070] In one implementation, training data was collected using the WESAD dataset and a self-built Polar H10 dataset. The RR interval sequences were divided into samples using a 64-point window and an 8-point sliding step. State labels included relaxed or calm, excited or happy, and stressed or anxious. Training, validation, and test sets were divided according to the subjects. The random forest consisted of 300 decision trees, with a minimum leaf node sample size of 2. Features were selected using the square root method, and class-balanced weights were employed.

[0071] The training set, validation set, and test set are preferably divided according to the subjects, so that the data of the same subject does not appear in the training set and the test set at the same time, thereby reducing the impact of individual data leakage on the model evaluation results.

[0072] To enable the group training model to adapt to the physiological differences among different users, a personal memory unit and a personal category prototype are established for each user.

[0073] Personalized user calibration includes at least two calibration tasks with clearly defined status labels. In one implementation, calibration statuses include relaxed or calm states, excited or pleasant states, and stressed or anxious states.

[0074] During the user's personal calibration task, RR interval data is continuously collected, and two-level anomaly filtering is performed on the RR interval data.

[0075] When the number of valid RR interval data collected reaches a preset number or the calibration time reaches a preset duration, the calibration RR interval sequence is divided into multiple sliding windows, and the fusion feature embedding vector corresponding to each sliding window is calculated.

[0076] For multiple fused feature embedding vectors belonging to the same state category, calculate their mean or weighted mean: in, Indicates state category The corresponding personal category prototype, Indicates belonging to a state category The number of calibration samples, Indicates the first The fusion feature embedding vector of each calibration sample.

[0077] Normalize the individual category prototype: The personal memory unit is used to store the user's historical fusion feature embedding vector, corresponding status label, acquisition time, and data quality information.

[0078] In one implementation, the personal memory unit is set with a maximum sample capacity; when the number of stored samples exceeds the maximum capacity, the oldest historical sample is deleted, or samples that need to be retained are selected based on data quality, class balance, and representativeness.

[0079] Step S6: Fuse the predicted probabilities from the first stage and calculate the class interval.

[0080] See Figure 5 The flowchart for two-stage individualized reasoning includes the following sub-steps: Step S601, RF global prediction probability.

[0081] During state recognition, the current fused feature embedding vector is first input into a random forest classifier to obtain the global prediction probability: in, This represents the number of pressure-related state categories.

[0082] Step S602: Calculate the cosine similarity and individualized matching probability between the current fused feature embedding vector and the individual prototype; For a state category for which a personal category prototype has already been established, calculate the current fused feature embedding vector. Compared with the normalized individual category prototype Cosine similarity between : Convert the similarity of each state category into individualized matching probabilities: in, These are configurable temperature parameters.

[0083] Step S603: Weighted fusion of RF global prediction probability and individualized matching probability to calculate the first-stage prediction probability and the Top1-Top2 category interval.

[0084] The global prediction probability and the individualized matching probability are weighted and fused: in, For the first stage of probability prediction, Configurable fusion weights.

[0085] For state categories for which no individual category prototype has been established, the global prediction probability output by the random forest can be directly retained, and the probabilities of all categories after fusion can be normalized.

[0086] Step S7: Determine whether the category interval has reached the threshold. If yes, output the first-stage identification result. If no, call the neural process model (NP) to perform the second-stage individualized reasoning.

[0087] See Figure 5 Specifically, it includes the following steps: Step S701: Determine whether the category interval is less than the preset interval threshold; if not, output the first stage recognition result; if yes, execute step S702.

[0088] Let the highest probability in the first stage prediction probability be denoted as... The second most likely is recorded as The category interval is represented as: When the category interval is greater than or equal to the preset interval threshold, the final state recognition result is determined based on the prediction probability of the first stage.

[0089] When the category interval is less than the preset interval threshold, the current sample is identified as the boundary sample, and the neural process model is invoked to perform the second stage of individualized inference.

[0090] Step S702: Read the personal memory unit (Memory), construct the neural process model NP context data, use the current fused feature embedding vector (Embedding) as the target data, and output the individualized probability of the neural process model (NP).

[0091] The neural process model includes an encoder and a decoder. The encoder encodes contextual data, consisting of historical fusion feature embedding vectors and their state labels from individual memory units, into latent variable distribution parameters; the decoder generates second-stage individualized prediction probabilities for each stress-related state category based on the latent variables and the current fusion feature embedding vector.

[0092] Specifically, stress-related state categories include relaxed or calm states, excited or pleasant states, and stressed or anxious states.

[0093] Context data is represented as: in, Embed vectors for historical fusion features. Set its corresponding status label. For the number of contexts.

[0094] The current fused feature embedding vector is used as the target data input into the neural process model to obtain the second-stage individualized prediction probability. .

[0095] Step S703: Output the individualized results.

[0096] In one implementation, the output of the neural process model is further fused with the first-stage predicted probability: in, The second-stage fusion weights are configurable.

[0097] Step S8: Determine the final state recognition result.

[0098] The effective state identification result (i.e., the stress-related state category) is determined based on the state category with the highest probability in the final predicted probability, and the corresponding probability is used as the state identification confidence level.

[0099] Step S9: Match and play the audio feedback.

[0100] Audio feedback can be activated only when the final recognition result falls within a preset feedback trigger category. For example, the preset feedback trigger category can be set to stress or anxiety.

[0101] To reduce the risk of triggering audio feedback from a single unstable prediction, state stability judgment conditions can be set. An audio feedback session is initiated when the state recognition results for a preset number of consecutive preset number of iterations all fall within a preset feedback trigger category, and the corresponding confidence levels all reach a preset feedback trigger threshold.

[0102] Audio feedback strategies include at least one parameter among audio file type, playback volume, and playback duration.

[0103] Preferably, based on the status recognition result, a piano sound, flowing water sound, birdsong sound, forest ambient sound, binaural beats, or other audio files can be selected from a preset audio file directory and played through an audio playback device.

[0104] When the audio starts playing, record the pre-intervention state identification results, the predicted probability of each state category, heart rate variability characteristics, audio strategy identifier, playback start time, and cumulative position of RR interval data.

[0105] Step S10: Obtain the RR sequence after playback again.

[0106] After the audio playback reaches the preset minimum evaluation duration, the newly generated RR interval data since the start of audio playback is read from the RR interval buffer. The RR interval data used for post-intervention state recognition must not overlap with the RR interval data in the pre-intervention state recognition window.

[0107] When the newly generated valid RR interval data reaches the preset feedback evaluation window length, the post-intervention state identification results and state prediction probabilities are obtained according to the aforementioned two-level anomaly filtering, feature construction and two-stage reasoning methods.

[0108] Step S11: Assess the impact of the intervention.

[0109] The predicted probability of the pre-defined feedback trigger category before intervention is denoted as . The predicted probability of the corresponding category after intervention is denoted as The change in state probability is expressed as: The pre-intervention RMSSD is denoted as The RMSSD after intervention is recorded as The change in RMSSD is represented as: The system according to , The intervention response outcome is determined by the post-intervention state category and the post-intervention prediction confidence level.

[0110] Step S12: Determine if further feedback is needed. If yes, adjust the audio strategy and proceed to step S9; otherwise, end.

[0111] In one specific embodiment, when the predicted probability of the preset feedback trigger category decreases by a certain amount after intervention, it is determined to be an effective response, and the audio output device is controlled to reduce the playback volume or stop playback. When the predicted probability of the preset feedback trigger category decreases, but the decrease does not reach the first preset threshold, and the RMSSD increases, it is determined to be a partially effective response, and the audio output device is controlled to maintain or reduce the playback volume and continue playback.

[0112] If the change in state probability before and after intervention does not meet the preset requirements, it is determined to be an insufficient response. The audio output device is then controlled to switch audio files or adjust playback parameters before playback continues.

[0113] When the predicted probability of the preset feedback trigger category increases after intervention compared to before intervention, and the increase reaches the second preset threshold, it is determined to be a negative response, and the audio output device is controlled to stop the current audio playback and output a prompt.

[0114] When the prediction confidence level after intervention is lower than the preset response evaluation threshold, the audio feedback strategy will not be adjusted for the time being, and new RR interval data will be collected and re-evaluated.

[0115] The system is configured with a maximum number of feedback loops. Automatic audio feedback control will stop when the number of feedback loops reaches the preset limit, communication with the heart rate monitoring device is interrupted, the proportion of effective RR interval data is lower than the preset requirement, or the user manually stops the system.

[0116] For details, see Figure 7The flowchart for the closed-loop control of audio feedback and intervention response evaluation in this embodiment is as follows: First, the final state and confidence level of step S8 are obtained to determine whether the feedback triggering condition is met. If not, monitoring continues; if yes, subsequent steps are executed. The pre-intervention state probability, RMSSD, and timestamp are recorded. The audio type, volume, and duration are matched, the audio is played, and the system waits for the preset feedback evaluation window. The new RR sequence after audio playback is obtained, state recognition is re-executed, and the changes in stress probability and HRV are calculated to determine the response category. If the response is insufficient, the audio strategy is changed and playback continues; if the response is effective, the volume is reduced or playback is stopped; if the response is negative, the current audio is stopped and a prompt is given. The number of feedback attempts is checked; if so, automatic feedback is stopped.

[0117] See Figure 6 The flowchart for updating personal calibration, personal category prototype, and personal memory unit begins with the individual selecting or confirming a status label. At least a preset number of RR intervals are collected, followed by two-level anomaly filtering. Multiple 64-point windows are generated, and the RR sequence embedding vector and HRV embedding vector are calculated to obtain a fusion vector. The average of similar fusion vectors is calculated and normalized to obtain the personal category prototype, which is then added to the personal memory unit. The process then determines whether the neural process model update conditions are met. If not, the personal calibration data is directly saved. If yes, a personalized update of the neural process model is performed, saving the personal category prototype, personal memory unit update, and model parameters.

[0118] Example 2 Based on Example 1, this example provides an audio feedback control system based on two-stage individualized reasoning and intervention response assessment, which is used to implement the audio feedback control method based on two-stage individualized reasoning and intervention response assessment in Example 1.

[0119] See Figure 2The system can be divided into a data acquisition layer, a data processing layer, a state recognition layer, and a feedback control layer. The data acquisition layer includes a heart rate acquisition device, a BLE communication module, and an RR interval data acquisition module. The RR interval data acquisition module establishes a connection with the heart rate monitoring device via Bluetooth Low Energy communication and receives heart rate data and RR interval data. The data processing layer includes a two-level anomaly filtering module, a sliding window construction module, and an HRV feature extraction module. The two-level anomaly filtering module performs physiologically reasonable interval filtering and local statistical filtering on the RR interval sequence. The sliding window construction module constructs the original RR interval sliding window, and the HRV feature extraction module extracts multidimensional heart rate variability features based on the RR interval sliding window. The state recognition layer comprises a Convolutional Neural Network (CNN) sequence feature extraction module, a Heart Rate Variability (HRV) encoding module, a feature fusion module, a Random Forest classification module, a personal memory module, a personal category prototype matching module, a boundary sample judgment module, and a Neural Process Model (NP) individualized reasoning module. Specifically, the CNN sequence feature extraction module generates sequence feature embedding vectors based on the original RR interval sliding window; the HRV encoding module generates HRV feature embedding vectors based on multidimensional HRV features; the feature fusion module fuses the sequence feature embedding vectors and HRV feature embedding vectors into a fused feature embedding vector; the Random Forest classification module generates global prediction probabilities based on the fused feature embedding vectors; the personal memory module stores the user's historical fused feature embedding vectors and state labels; and the personal category prototype matching module establishes and normalizes personal category prototypes based on the user's historical fused feature embedding vectors, and calculates individualized matching probabilities based on the similarity between the current fused feature embedding vector and the normalized personal category prototype. The boundary sample judgment module determines whether to invoke the neural process model based on the category interval of the first-stage predicted probability obtained by fusing the global prediction probability and the individualized matching probability. The neural process model individualized inference module outputs the second-stage individualized predicted probability upon invocation to determine the effective state recognition result. The feedback control layer includes an intervention response evaluation module, an audio strategy matching module, a state result management module, and an audio playback control module. The audio playback control module controls the audio output device to play audio based on the effective state recognition result. The intervention response evaluation module evaluates the intervention response based on the newly acquired RR interval sequence after audio playback and controls the audio output device to continue playback, adjust the audio feedback strategy, or stop playback.

[0120] In one software implementation, the system is developed using the Python language. The audio output module is implemented using pygame, and the software runtime environment includes Python 3.10.0, pygame 2.6.1, and SDL 2.28.4.

[0121] Convolutional neural networks and neural process models can be implemented using deep learning frameworks, random forest classifiers can be implemented using machine learning frameworks, and Bluetooth Low Energy communication can be implemented using corresponding cross-platform Bluetooth communication libraries.

[0122] The software library, version number, and heart rate monitoring device model mentioned above are only one implementation method and do not constitute a limitation on the scope of protection of this invention.

[0123] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An audio feedback control method based on two-stage individualized reasoning and intervention response assessment, characterized in that, The steps include the following: The user's RR interval sequence is obtained and processed through a sliding window to extract multidimensional heart rate variability features; Sequence feature embedding vectors are extracted based on sliding window, and heart rate variability feature embedding vectors are extracted based on the multidimensional heart rate variability features. A fused feature embedding vector is obtained through fusion processing. The fused feature embedding vector is classified to obtain the global predicted probability corresponding to multiple pressure-related state categories; Based on the pre-acquired user history fusion feature embedding vector and its state label, establish the personal category prototype corresponding to each stress-related state category, and obtain the individualized matching probability based on the similarity between the current fusion feature embedding vector and the normalized personal category prototype. The global prediction probability and the individualized matching probability are fused to obtain the first-stage prediction probability, and the class interval between the highest probability and the second highest probability in the first-stage prediction probability is calculated. When the category interval reaches a preset interval threshold, the effective state recognition result is determined based on the first stage prediction probability; when the category interval does not reach the preset interval threshold, the current fusion feature embedding vector, the user's historical fusion feature embedding vector, and the state label are input into the neural process model to obtain the second stage individualized prediction probability, and the effective state recognition result is determined based on the second stage individualized prediction probability. Based on the valid state recognition results, an audio feedback strategy is determined, and the audio output device is controlled to play the corresponding audio. During or after the audio playback, the newly generated RR interval sequence after the audio playback is acquired again. The state recognition result after intervention is obtained according to the aforementioned state recognition process. Based on the changes in the state recognition results before and after intervention, the state prediction probability, or the heart rate variability characteristics, the audio output device is controlled to continue playing, adjust the audio feedback strategy, or stop playing.

2. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 1, characterized in that, Before the sliding window processing, the RR interval sequence is preprocessed, which includes the following steps: RR interval data points that exceed the preset physiological reasonable range are identified as extreme outliers; Within a local statistical window consisting of multiple consecutive valid RR interval data points, calculate the median and median absolute deviation of the RR interval data in the local statistical window; Based on the deviation between the RR interval data point to be detected and the median, and the absolute deviation of the median, it is determined whether the RR interval data point to be detected is a local outlier. The extreme and local anomalies are filtered out.

3. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 1, characterized in that, Sliding window processing is achieved through a pre-constructed RR interval sliding window, wherein the RR interval sliding window includes multiple consecutive valid RR interval data points, and the multidimensional heart rate variability features include at least two of the following: time domain features, statistical features, frequency domain features, and nonlinear features.

4. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 1, characterized in that, A sequence feature embedding vector is extracted using a convolutional neural network sequence feature extractor, and a heart rate variability feature embedding vector is extracted using a heart rate variability feature encoder. The convolutional neural network sequence feature extractor includes at least one one-dimensional convolutional layer, a non-linear activation layer, a pooling layer, and a feature mapping layer. The heart rate variability feature encoder includes at least two fully connected layers. The sequence feature embedding vector and the heart rate variability feature embedding vector are concatenated and feature mapped to obtain the fused feature embedding vector.

5. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 1, characterized in that, The classification of the fused feature embedding vector is achieved by a random forest classifier, which includes multiple decision trees. Each decision tree outputs the classification result of the stress-related state category, and the global prediction probability is obtained based on the voting ratio or the average probability of each category of the multiple decision trees.

6. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 5, characterized in that, The global prediction probability and the individualized matching probability are weighted and fused to obtain the first-stage prediction probability. For stress-related state categories for which no personal category prototype has been established, the global prediction probability output by the random forest classifier is retained.

7. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 1, characterized in that, The process of establishing the individual category prototype includes the following steps: RR interval sequences are acquired during the user's performance of a status-labeled personal calibration task; Multiple RR interval sliding windows are constructed based on the RR interval sequence, and corresponding fusion feature embedding vectors are obtained for each. The mean, weighted mean, or cluster center of the fusion feature embedding vectors belonging to the same stress-related state category are used as the individual category prototype for the corresponding category. The individual category prototypes are normalized. Store the fused feature embedding vector and its state label.

8. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 1, characterized in that, The category interval is the difference between the highest probability and the second highest probability in the first stage of prediction. When the category interval is lower than the preset interval threshold, the current RR interval sliding window is determined as the boundary sample, and the neural process model is triggered to perform the second stage of individualized inference.

9. The audio feedback control method based on two-stage individualized reasoning and intervention response assessment according to claim 8, characterized in that, The stored user history fusion feature embedding vector and its state label are used as context data for the neural process model, and the current fusion feature embedding vector is used as target data. The neural process model outputs the individualized prediction probability that the current target data belongs to each stress-related state category.

10. An audio feedback control system based on two-stage individualized reasoning and intervention response assessment, characterized in that, The system is used to implement the audio feedback control method based on two-stage individualized reasoning and intervention response assessment as described in any one of claims 1-9, the system comprising: The RR interval data acquisition module is used to acquire the user's RR interval sequence from the heart rate monitoring device; A two-level anomaly filtering module is used to perform physiologically reasonable interval filtering and local statistical filtering on the RR interval sequence; The feature construction module is used to construct the original RR interval sliding window and extract multidimensional heart rate variability features; A convolutional neural network sequence feature extraction module is used to generate a sequence feature embedding vector based on the original RR interval sliding window; The heart rate variability feature encoding and fusion module is used to generate a heart rate variability feature embedding vector based on multidimensional heart rate variability features, and to fuse the sequence feature embedding vector and the heart rate variability feature embedding vector into a fused feature embedding vector. The random forest classification module is used to generate global prediction probabilities based on the fused feature embedding vector. The personal memory and personal category prototype module is used to store the user's historical fusion feature embedding vector and status label, establish and normalize the personal category prototype based on the user's historical fusion feature embedding vector, and calculate the individualized matching probability based on the similarity between the current fusion feature embedding vector and the normalized personal category prototype. The boundary sample judgment and neural process reasoning module is used to determine whether to call the neural process model based on the category interval of the first-stage prediction probability obtained by fusing the global prediction probability and the individualized matching probability, and outputs the second-stage individualized prediction probability when calling it to determine the effective state recognition result. The audio feedback control module is used to control the audio output device to play audio based on the valid status recognition result. The intervention response assessment module is used to evaluate the intervention response based on the newly acquired RR interval sequence after audio playback, and to control the audio output device to continue playback, adjust the audio feedback strategy, or stop playback.