Earphone control method and system with voice interaction function

CN122160668APending Publication Date: 2026-06-05DONGGUAN OUMU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGGUAN OUMU TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing noise-canceling headphones lack the ability to perceive dynamic environmental risks during outdoor sports, resulting in untimely risk warnings or frequent false alarms. Furthermore, their simplistic safety strategies lead to conflicts in the auditory experience, making it difficult to achieve a refined and adaptive balance between ensuring safety and maintaining auditory comfort.

Method used

A multimodal feature vector acquisition method is adopted, combined with a contextual incremental update model and auditory particle swarm optimization decision-making, to assess environmental risks in real time and generate the optimal auditory enhancement strategy. The environmental sound source features and user status are obtained through blind source separation algorithm, inertial data and ultrasonic frequency band signals. The auditory strategy is optimized by particle swarm optimization algorithm to achieve a balance between safety and comfort.

Benefits of technology

It enables real-time risk assessment of dynamic environments and generates a continuous and smooth auditory strategy from background noise reduction to full transparency mode, ensuring motion safety and minimizing auditory interference, thereby improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122160668A_ABST
    Figure CN122160668A_ABST
Patent Text Reader

Abstract

The application provides an earphone control method and system with voice interaction function, comprising obtaining a multi-modal feature vector containing environmental sound source characteristics, user motion state parameters and physiological load index, based on the multi-modal feature vector, dynamic matching and risk modeling are performed through a context incremental update model containing a scene information library, a multi-dimensional risk vector is obtained, based on the multi-dimensional risk vector and the physiological load index, an auditory particle swarm is enhanced to make a decision, optimal auditory enhancement strategy parameters are obtained in real time, sound field reconstruction is performed based on the optimal auditory enhancement strategy parameters, and parameter optimization is performed on the context incremental update model, the application realizes real-time dynamic evaluation of environmental risks, and generates an optimal auditory strategy that minimizes auditory interference and maximizes auditory comfort under the premise of ensuring safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio interaction technology, and more specifically, to a headphone control method and system with voice interaction functionality. Background Technology

[0002] With the widespread adoption of wireless noise-canceling headphones, wearing them for music playback during outdoor sports has become a common practice. However, while these devices enhance the user's auditory experience during outdoor activities, they also increase safety risks due to auditory blockage. In balancing auditory experience and safety risk warning, existing technologies suffer from two main drawbacks: a lack of dynamic environmental safety risk perception capabilities and the conflict arising from a single safety strategy.

[0003] First, there is a lack of dynamic environmental safety risk perception capabilities. Existing solutions are mostly based on preset acoustic scenarios or simple volume thresholds for control, which cannot conduct real-time and accurate risk assessment of dynamically changing outdoor environments (such as entering a noisy intersection from a quiet park). For example, these methods have difficulty distinguishing between the sound of rapidly approaching vehicles and continuous environmental noise, resulting in untimely risk warnings or frequent false alarms.

[0004] Second, the conflict arising from a single safety strategy: traditional methods often adopt a binary decision of either turning on or off, such as fully transmitting ambient sound or playing warning sounds when a potential risk is detected. This often disrupts the continuity of audio and damages the auditory experience. Existing methods cannot make a fine-grained and adaptive trade-off between the conflicting goals of ensuring safety, minimizing auditory interference, and maintaining auditory comfort, making it difficult to meet the dual needs of outdoor sports users for safety and experience.

[0005] Therefore, there is an urgent need in this field for a noise-canceling headphone audio interaction technology that can dynamically assess environmental risks in real time and generate an optimal auditory strategy that minimizes auditory interference and maintains auditory comfort while ensuring safety. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a headphone control method and system with voice interaction functionality.

[0007] The first aspect of this invention provides a headphone control method with voice interaction function, comprising: Obtain a multimodal feature vector containing environmental sound source features, user motion state parameters, and physiological load index; Based on multimodal feature vectors, dynamic matching and risk modeling are performed through a contextual incremental update model that includes a scenario information database to obtain multidimensional risk vectors; Based on multidimensional risk vectors and physiological load index, the optimal auditory enhancement strategy parameters are obtained in real time through auditory particle swarm optimization decision-making. Sound field reconstruction is performed based on the optimal auditory enhancement strategy parameters, and the parameters of the context incremental update model are optimized.

[0008] According to a preferred embodiment, a multimodal feature vector comprising environmental sound source features, user motion state parameters, and physiological load index is obtained, including: Acquire raw multi-channel audio streams, inertial data, and ultrasonic frequency band echo signals; Based on the original multi-channel audio stream, a blind source separation algorithm is used to obtain multiple audio source signal streams. Event identification is performed on each audio source signal stream to obtain the corresponding event label and confidence level. Risk audio source signal streams are then filtered out, and time delay estimation is performed on the risk audio source signal streams to obtain the horizontal azimuth angle. User motion state parameters are obtained using a motion recognition strategy based on inertial data; The frequency domain index of heart rate variability is obtained based on echo signals in the ultrasonic band and used as a physiological load index. Multimodal feature vectors are obtained based on event labels, confidence levels, horizontal azimuth angles, user motion state parameters, and physiological load indices from the audio source signal stream.

[0009] According to a preferred embodiment, based on multimodal feature vectors, dynamic matching and risk modeling are performed through a context incremental update model containing a scenario information database to obtain a multidimensional risk vector, including: The context incremental update model includes a dimensionality reduction encoder, a context information base, reward statistics parameters, a risk assessment arm set, and an upper-level selection strategy; The multimodal feature vector is reduced in dimensionality using the dimensionality reduction encoder to obtain a low-dimensional context code. The low-dimensional context encoding is used as a query vector to perform a matching query in the scenario information database to obtain scenario cluster identifiers and basic risk prototypes. Based on the scenario cluster identifier, obtain the associated scenario-related arm subset, wherein the scenario-related arm subset contains multiple different risk assessment arms; Based on the upper-level selection strategy and reward statistics parameters, a risk assessment arm is selected from the scenario-related arm subset. Collision probability, threat level, uncertainty, expected contact time and spatial dispersion are obtained based on low-dimensional context encoding, multimodal feature vectors and basic risk prototypes, and a risk assessment arm is used to obtain multidimensional risk vectors.

[0010] According to a preferred embodiment, a risk assessment arm is selected from the scenario-related arm subset based on an upper-level selection strategy and reward statistics parameters, including: The reward statistics parameters include the historical average reward, number of times selected, and reward variance for each risk assessment arm under the scenario cluster identifier. Based on the reward statistics parameters, the upper confidence bound algorithm is used to obtain the selection confidence scores of all risk assessment arms in the scenario-related arm subset, and risk assessment arms are selected from the scenario-related arm subset based on the selection confidence scores.

[0011] According to a preferred embodiment, collision probability, threat level, uncertainty, expected contact time, and spatial dispersion are obtained based on low-dimensional context encoding, multimodal feature vectors, and a basic risk prototype. A risk assessment arm is then used to obtain a multidimensional risk vector, including: The collision probability is obtained based on a basic risk prototype using a Bayesian algorithm; The threat level is obtained based on event labels in the multimodal feature vector; The uncertainty is obtained by calculating the matching distance between the low-dimensional context code and the center of the scenario cluster; The estimated contact time is obtained based on the relative velocity of the risky audio source signal stream; The spatial dispersion is obtained by performing coherence analysis on all audio signal streams.

[0012] According to a preferred embodiment, based on a multidimensional risk vector and a physiological load index, the optimal auditory enhancement strategy parameters are obtained in real time through auditory particle swarm optimization decision-making, including: Using a multidimensional risk vector and a physiological load index as input vectors, an audio-safe multi-objective optimization problem is constructed. The audio-safe multi-objective optimization problem includes at least minimizing the safety risk function, minimizing the user interference function, and maximizing the auditory comfort function. Initialize the particle swarm, where the position of each particle represents a set of candidate auditory enhancement strategy parameters; For each particle representing the auditory enhancement strategy parameters, the corresponding safety risk function value, user interference function value, and auditory comfort function value are obtained by combining the multidimensional risk vector and physiological load index. Based on the safety risk function value, user interference function value, and auditory comfort function value, the Pareto dominance relation is used for iteration to update the individual historical best position of each particle and the global best guidance set of the population until the preset convergence condition is met. The optimal auditory enhancement strategy parameters are then selected from the global best guidance set using preset decision rules.

[0013] According to a preferred embodiment, the audio-safety multi-objective optimization problem includes at least minimizing a security risk function, minimizing a user interference function, and maximizing a listening comfort function, including: A safety risk objective function is constructed based on the collision probability, threat level, and expected contact time in the multidimensional risk vector. An interference objective function is constructed based on the physiological load index and the parameters of the candidate auditory enhancement strategies. A target function for auditory comfort is constructed based on the parameters of the candidate auditory enhancement strategies.

[0014] According to a preferred embodiment, sound field reconstruction is performed based on optimal auditory enhancement strategy parameters, and the parameters of the context incremental update model are optimized, including: Collect user inertial measurement data after sound field reconstruction, and obtain reward signals based on user inertial measurement data; Based on the reward signal, an incremental update algorithm is used to incrementally update the corresponding reward statistics parameters in the context model.

[0015] A second aspect of the present invention also provides a headphone control system with voice interaction function, comprising: The feature acquisition module is used to acquire a multimodal feature vector containing environmental sound source features, user motion state parameters, and physiological load index. The risk modeling module is used to obtain a multidimensional risk vector based on a multimodal feature vector and by updating the model through contextual incremental updates. The auditory enhancement module is used to obtain the optimal auditory enhancement strategy parameters in real time based on a multidimensional risk vector and a physiological load index through auditory particle swarm enhancement decision-making. The sound field reconstruction module is used to reconstruct the sound field based on the optimal auditory enhancement strategy parameters and to optimize the parameters of the context incremental update model.

[0016] Based on the above, this application embodiment introduces a context incremental update model based on a context gambling machine algorithm. The multimodal feature vectors acquired in real time are dynamically matched in the context information database contained in the context incremental update model. The optimal risk assessment arm is adaptively selected through the upper confidence bound algorithm, thereby outputting a multidimensional risk vector including collision probability, threat level, etc. in real time. The context incremental update model continuously learns and quickly identifies risk patterns in different environments, such as intersections and parks. This realizes the transformation from static assessment of environmental risk based on preset acoustic scenes or simple volume thresholds to real-time dynamic assessment of environmental risk based on context awareness.

[0017] On the other hand, by adopting auditory particle swarm optimization based on a multi-objective particle swarm optimization algorithm, the problem of single safety strategy and conflicting objectives is solved. Auditory particle swarm enhancement decision takes a multi-dimensional risk vector and the user's real-time physiological load as input, and optimizes multiple objective functions such as safety risk, user interference and auditory comfort in parallel. By searching the Pareto optimal solution set, the optimal strategy parameters are dynamically generated. It realizes a continuous and smooth strategy spectrum from background noise reduction and risk sound enhancement to full transparency mode based on the real-time risk level and user fatigue level, rather than a simple binary switch. Thus, while ensuring sports safety, auditory interference is minimized and auditory comfort is maintained. Attached Figure Description

[0018] Figure 1 The execution flowchart of a headphone control method with voice interaction function according to the present invention is shown.

[0019] Figure 2 A flowchart of the context incremental update model in a headphone control method with voice interaction function according to the present invention is presented.

[0020] Figure 3 A schematic diagram of a headphone control system with voice interaction function according to the present invention is shown. Detailed Implementation

[0021] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of this application can be combined with each other.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0024] like Figure 1 , Figure 2 , Figure 3 As shown: The first aspect of this invention provides a headphone control method with voice interaction function, comprising: X1: Obtain a multimodal feature vector containing environmental sound source features, user motion state parameters, and physiological load index.

[0025] First, the raw multi-channel audio stream, inertial data, and ultrasonic echo signal are acquired. The raw multi-channel audio stream can be acquired through a microphone array consisting of at least two microphones arranged on the headphones. The raw multi-channel audio stream is a mixed signal containing ambient sound and possible user voice. The inertial data can be acquired through the inertial measurement unit built into the headphones. The inertial measurement unit includes at least a three-axis accelerometer and a three-axis gyroscope, used to acquire the three-axis acceleration sequence and three-axis angular velocity sequence of the user's head in three-dimensional space. The ultrasonic echo signal is acquired through the ultrasonic transceiver module integrated into the headphones. The module includes at least an ultrasound transmitter and an ultrasound receiver. Specifically, the ultrasound transmitter emits ultrasound waves at a specific frequency, such as 40kHz, and the ultrasound receiver receives the echo signal reflected from the user's body surface as the ultrasound frequency band echo signal. Since physiological activities such as heartbeat and breathing cause tiny vibrations on the body surface, based on the Doppler effect, when ultrasound waves are projected onto the user's skin surface, the echo signal will undergo frequency shift due to these tiny vibrations. Therefore, by demodulating and analyzing the spectrum of the ultrasound frequency band echo signal, the heartbeat interval sequence reflecting the heartbeat cycle can be extracted, providing a data basis for subsequent calculation of the physiological load index.

[0026] Secondly, based on the original multi-channel audio stream, a blind source separation algorithm is used to obtain multiple audio source signal streams. Event identification is performed on each audio source signal stream to obtain the corresponding event label and confidence level. Risky audio source signal streams are then filtered out, and time delay estimation is performed on these risky streams to obtain the horizontal azimuth angle. Specifically, multiple independent audio source signal streams can be estimated and separated from the original multi-channel audio stream using blind source separation algorithms such as independent component analysis. Each audio source signal stream represents an independent sound source. Audio event identification is then performed on each separated audio source signal stream. The input stream contains a pre-trained audio event classification model. For example, this model is based on a one-dimensional convolutional neural network structure, consisting of sequentially stacked one-dimensional convolutional layers, pooling layers, and fully connected layers. The model is trained under supervised supervision on publicly available audio event datasets such as Audioset or ESC-50. The training data consists of a large number of labeled audio segments, including waveform data and their corresponding event category labels. The audio source signal stream is preprocessed into fixed-length segments, such as audio frames corresponding to 1 second in length. After being converted into single-channel waveform data, it is used as input to the audio event classification model. The output of the audio event classification model is a probability distribution vector. The category corresponding to the maximum value in the probability distribution vector can be used as the event label of the sound source signal stream, and the maximum value can be used as the confidence score of this classification. The event label can include car horns, human conversations, bicycle bells, wind sounds, etc. The confidence score is a probability measure of the certainty of the audio event classification model for this classification result. According to the preset risk event label library, the sound source signal stream whose event label belongs to the risk event label library is marked as a risk sound source signal stream. The risk event label library is a set of acoustic events predefined by industry technicians based on the application scenario involved in this method, namely the safety requirements of outdoor sports. For example, the risk event label library can include car horns, emergency alarms, and vehicle approach noise, etc., to identify the types of sound sources that may pose a potential threat to user safety. For each risk sound source signal stream, the horizontal azimuth angle of the risk sound source signal stream relative to the user's head can be obtained based on the time difference between its arrival at different microphones on the headphones, through time delay estimation algorithms such as generalized cross-correlation.

[0027] Simultaneously, a motion recognition strategy is employed based on inertial data to obtain user motion state parameters. Specifically, based on inertial data, i.e., the triaxial acceleration sequence and triaxial angular velocity sequence, preprocessing is performed, followed by fusion and noise reduction using complementary filtering or Kalman filtering algorithms, and the user's head attitude angles, such as pitch and roll angles, are calculated. The composite acceleration modulus is obtained based on the triaxial acceleration sequence and used as a reference indicator of motion intensity. Time-domain and frequency-domain features are extracted from the inertial data. The time-domain features may include mean, variance, zero-crossing rate, etc., and the frequency-domain features may include spectral centroid and energy entropy, etc. The time-domain and frequency-domain features are then combined. The input is a pre-trained motion classification model, which can be implemented by a one-dimensional convolutional neural network or a support vector machine. It can be trained based on a labeled motion pattern dataset, which can be obtained from public benchmark datasets such as PAMAP2 or RealWorldHAR datasets. Each inertial data sample is labeled with its corresponding real motion pattern. The motion classification model outputs a corresponding motion pattern classification result based on the input time-domain and frequency-domain features, and encapsulates the synthetic acceleration modulus, head posture angle and motion pattern classification result into user motion state parameters.

[0028] Simultaneously, the frequency domain index of heart rate variability is obtained based on ultrasound echo signals and used as a physiological workload index. Specifically, the ultrasound echo signal experiences frequency shifts due to periodic micro-vibrations on the skin surface of the ear caused by heartbeats. Using bandpass filtering and phase demodulation techniques, the user's heartbeat interval sequence can be extracted based on this frequency shift. Fast Fourier Transform or Lomb-Scargle spectral analysis is then performed on the heartbeat interval sequence to calculate the percentage of energy in the high-frequency band, such as 0.15Hz to 0.4Hz, relative to the total spectral energy. This yields the high-frequency power component of heart rate variability as the frequency domain index. This index is then converted into the user's current physiological workload index using a preset linear or nonlinear mapping function. This mapping function can be a linear function that normalizes the frequency domain index to a fixed interval, such as 0 to 1, or a nonlinear function such as the sigmoid function. The physiological workload index is a scalar value, positively correlated with the user's autonomic nervous system excitation level; a higher value indicates greater physiological workload, such as tension or fatigue. Finally, multimodal feature vectors are obtained based on the event labels, confidence levels, horizontal azimuth angles, user motion state parameters, and physiological load index of the audio source signal streams. The data obtained in step X1 is standardized and structured and encoded to form multimodal feature vectors. These multimodal feature vectors serve as a digital snapshot of the user's auditory environment and state at the current moment. The modal feature vectors contain information in multiple dimensions, including event labels and confidence levels of all audio source signal streams, horizontal azimuth angles of all risky audio source signal streams, user motion state parameters including motion velocity vectors, motion acceleration vectors, and motion pattern labels, and physiological load index.

[0029] In some possible embodiments, assuming the current environment is separated into 3 sound source signal streams, sound source signal stream 1 is identified as a car horn [risk label, confidence 0.92, horizontal azimuth 30 degrees], sound source signal stream 2 is wind sound [non-risk, confidence 0.85, horizontal azimuth not calculated], and sound source signal stream 3 is human voice conversation [non-risk, confidence 0.78, horizontal azimuth not calculated]. Based on inertial data, the calculated movement speed is 1.5 m / s, the direction is straight ahead, the movement mode label is walking, and the physiological load index is 0.65. Then the final generated multimodal feature vector can be [sound source signal stream 1, car horn, confidence 0.92, horizontal azimuth 30 degrees, sound source signal stream 2, wind sound, confidence 0.85, sound source signal stream 3, human voice, confidence 0.78, user speed 1.5, walking, physiological load 0.65].

[0030] In summary, step X1 constructs a joint feature vector that characterizes environmental risk attributes, user dynamic state, and user physiological state by fusing heterogeneous data from different types of sensors. The environmental risk attributes can be jointly characterized by event labels, confidence levels, and horizontal azimuth angles. The user dynamic state can be jointly characterized by motion velocity vectors, motion acceleration vectors, and motion pattern labels. The user physiological state can be characterized by physiological load indexes. This ensures that complete data support is provided for risk modeling in step X2 and optimization decision-making in step X3. For example, step X2 uses motion velocity vectors, motion acceleration vectors, and azimuth to calculate relative speed and expected contact time. Step X3 can dynamically adjust the interference threshold for hearing enhancement based on physiological load indexes and motion patterns.

[0031] X2: Based on multimodal feature vectors, dynamic matching and risk modeling are performed through a contextual incremental update model containing a contextual information base to obtain multidimensional risk vectors. The contextual incremental update model includes a dimensionality reduction encoder, a contextual information base, reward statistical parameters, a risk assessment arm set, and an upper-level selection strategy.

[0032] First, the multimodal feature vectors are dimensionality reduced using the dimensionality reduction encoder to obtain a low-dimensional context code. This encoder can be the encoder part of a pre-trained autoencoder or a principal component analysis model. The dimensionality reduction encoder compresses high-dimensional, sparse multimodal feature vectors into a low-dimensional, dense vector representation, i.e., a low-dimensional context code. This low-dimensional context code retains the most discriminative information from the multimodal feature vectors and, compared to the multimodal feature vectors themselves, can more accurately represent environmental risk attributes, user dynamic states, and user physiological states. It also serves as an index for subsequent matching in a context information database. For example, assuming the dimensionality reduction encoder is the encoder part of a neural network-based autoencoder, this autoencoder can contain a symmetric encoder-decoder structure. The encoder part consists of several fully connected layers stacked sequentially, each followed by a non-linear activation function, such as ReLU, with the number of neurons decreasing layer by layer. Ultimately, this maps the input multimodal feature vectors to a significantly reduced-dimensional vector. The bottleneck layer outputs the desired low-dimensional context code. The decoder then attempts to reconstruct the original high-dimensional feature vector from this low-dimensional code. The autoencoder can be trained unsupervised on a large number of synthetic multimodal feature vector samples to minimize reconstruction errors, such as by optimizing the mean squared error. This allows the encoder to learn the most effective low-dimensional representations from the data. The synthetic multimodal feature vector samples can be obtained through combined simulations using publicly available single-modal datasets. For example, they can be combined with publicly available audio event datasets such as Audioset, human activity recognition datasets such as RealWorldHAR, and physiological signal datasets such as PPG. Through time alignment and data fusion, synthetic multimodal training samples can be constructed. Alternatively, the dimensionality reduction encoder can be implemented using a principal component analysis model. By calculating the principal components of the synthetic multimodal feature vector samples and selecting the feature vectors corresponding to the first k principal components as projection bases, the new multimodal feature vectors are projected onto this low-dimensional subspace to obtain the low-dimensional context code.

[0033] Secondly, the low-dimensional context encoding is used as a query vector to perform a matching query in the scenario information database to obtain scenario cluster identifiers and basic risk prototypes. The scenario information database consists of historical low-dimensional context encoding, binary feedback signals associated with the historical low-dimensional context encoding, scenario clusters obtained by clustering the historical low-dimensional context encoding, and basic risk prototypes constructed for each scenario cluster. The historical low-dimensional context encoding originates from historical data accumulated during previous operations. In the initial deployment phase, when there is no historical data, the scenario information database can be initialized by a preset initial cluster set. This initial cluster set can be based on simulated multimodal feature vectors of typical motion scenarios, such as stationary, walking, and running, and multiple scenario clusters and their corresponding cluster centers can be obtained in advance through clustering. In the initial deployment phase, the scenario information database can be initialized by a simulated dataset generated based on typical motion scenarios. The simulated dataset includes simulated scenario clusters and their simulated cluster centers. The dataset can be obtained by clustering and combining publicly available datasets. For example, it can combine audio event datasets such as Audioset and human activity recognition datasets containing inertial data such as PAMAP2 to synthesize simulated multimodal feature vectors corresponding to typical motion scenarios such as stationary, walking, and running. Specifically, representative audio event labels, motion state parameters, and preset baseline physiological load indices are configured for each typical motion scenario. For example, audio event labels such as wind sound and distant traffic sound are configured for the typical motion scenario of walking. The motion state parameters are specifically the inertial feature statistics corresponding to typical motion scenarios obtained from the human activity recognition dataset used. By arranging and combining these elements, a large number of simulated multimodal feature vectors covering different typical motion scenarios are generated. Then, the corresponding simulated low-dimensional context codes are obtained through a dimensionality reduction encoder. These simulated low-dimensional context codes are clustered to obtain simulated scenario clusters and their simulated cluster centers.

[0034] In subsequent use of the headphones, step X1 will continuously generate historical multimodal feature vectors. These historical multimodal feature vectors are processed by the dimensionality reduction encoder in the context incremental update model to obtain the corresponding historical low-dimensional context code. When constructing the context information database, unsupervised clustering is performed on all collected historical low-dimensional context codes using clustering algorithms such as K-Means to obtain multiple context clusters. Each context cluster has a unique identity, namely the context cluster identifier. The cluster center of each context cluster is the typical context pattern of that context cluster. As the context cluster center, the distance between the current query vector and each context cluster center is calculated, such as Euclidean distance or cosine similarity. The context cluster with the smallest distance is determined as the context cluster to which the current query vector belongs, and the current query vector is assigned the context cluster identifier of that context cluster.

[0035] For each historical multimodal feature vector, the scenario information database records its risk feedback signal. This risk feedback signal indicates whether a real risk event occurred during the historical period corresponding to the historical multimodal feature vector. The risk feedback signal can include binary feedback signals and continuous feedback signals. Specifically, the binary feedback signal is a discrete "yes" or "no" signal; for example, the binary feedback signal can take values ​​of [risk occurred, risk did not occur]. In this case, the basic risk prototype represents the prior probability of a risk event occurring under that scenario cluster. For a given scenario cluster, the binary feedback signals corresponding to all historical low-dimensional context codes belonging to it are statistically analyzed. The probability of risk occurrence can be directly calculated using maximum likelihood estimation. For example, if a scenario cluster has 1000 historical low-dimensional context codes, and 50 of them are marked as risk occurrence, then the probability of risk occurrence is 0.05. To characterize the uncertainty of this estimation, Bayesian inference can be used. The basic risk prototype combines uniform prior data with observed data to obtain a beta distribution as a complete probability description. Its parameters are determined by the number of times the risk occurred and did not occur. The continuous feedback signal specifically refers to the risk feedback signal as a continuous value, such as the minimum distance or collision probability calculated retrospectively using more precise sensor data, such as GPS trajectory. The basic risk prototype represents the probability distribution of risk metrics under this scenario cluster. For a certain scenario cluster, the continuous risk metrics corresponding to all historical low-dimensional context codes within it are statistically analyzed. Gaussian distribution is commonly used to model its distribution. For example, if the mean μ and variance σ² of the historical risk metrics within the scenario cluster are statistically obtained, then the basic risk prototype is a Gaussian distribution with mean μ and variance σ². In summary, the basic risk prototype is specifically a probability model of risk learned from historical low-dimensional context codes, providing a starting point based on historical collective experience for subsequent real-time calculations.

[0036] Then, based on the scenario cluster identifier, the associated scenario-related arm subset is obtained. The scenario-related arm subset contains multiple different risk assessment arms. Based on the scenario cluster identifier, a preset scenario-assessment arm mapping is used to obtain the risk assessment arms associated with the scenario cluster identifier from the global risk assessment arm set and combine them into a scenario-related arm subset. The preset scenario-assessment arm mapping can be dynamically established through learning from historical data. For example, during the scenario information database construction phase, whenever a new scenario cluster is created, a default scenario-related arm subset can be assigned to it. This scenario-related arm subset can contain all risk assessment arms or select some risk assessment arms based on the characteristics of the scenario cluster. As the system runs, the scenario-assessment arm mapping can be dynamically adjusted according to the reward statistics parameters of each risk assessment arm under different scenario clusters, and risk assessment arms with consistently poor performance can be removed from the relevant scenario-related arm subset.

[0037] Each risk assessment arm is essentially an independent function or lightweight model. It receives input data—low-dimensional context encoding, multimodal feature vectors, and a basic risk prototype—and outputs the components of the multidimensional risk vector through a specific algorithm or parameterized calculation rule: collision probability, threat level, uncertainty, estimated contact time, and spatial dispersion. Different risk assessment arms utilize different algorithms or parameter configurations to adapt to different scenarios. For example, a risk assessment arm optimized for urban street walking scenarios might assign a higher weight to the rate of change of the horizontal azimuth angle of the vehicle sound source. A precise kinematic model based on relative speed and distance is used to calculate the collision probability and estimated contact time. Its threat level mapping table may assign a higher level to truck engine noise than car noise. Assuming a risk assessment arm optimized for a park running scenario, its internal algorithm may focus more on the spatial distribution analysis of multiple non-motorized vehicle sound sources, such as bicycle bells and skateboard roller skating sounds, and use an empirical model based on the number of sound sources and spatial dispersion to assess the overall risk level. Its method of calculating uncertainty may rely more on the coherence of the overall sound field of the environment than on the distance to the center of a single scenario cluster. In summary, each risk assessment arm represents a specific risk modeling perspective and calculation strategy.

[0038] Next, based on the upper-level selection strategy and reward statistics parameters, risk assessment arms are selected from the scenario-related arm subset. The reward statistics parameters include the historical average reward, the number of times each risk assessment arm is selected under the scenario cluster identifier, and the reward variance. The historical average reward is specifically the average reward signal obtained when the risk assessment arm is selected under the scenario cluster identifier. This reward signal originates from the feedback generated in step X4 based on the actual effect of the system decision, such as whether the user smoothly avoids the obstacle. The number of times the risk assessment arm is selected is specifically the number of times the risk assessment arm is selected under the scenario cluster identifier. The reward variance is specifically the dispersion of the reward signal obtained by the risk assessment arm under the scenario cluster identifier, used to characterize its performance stability. Based on the reward statistics parameters, the upper confidence bound algorithm is used to obtain the risk assessment arm subset... A selection confidence score is generated for each risk assessment arm. Based on this selection confidence score, risk assessment arms are selected from a subset of scenario-related arms. During each selection, an upper confidence bound algorithm is used to calculate a selection confidence score for each risk assessment arm in the subset of scenario-related arms. The selection confidence score includes the historical average reward and the number of exploration items. The specific calculation formula is: Selection Confidence Score = Historical Average Reward + c * sqrt[In(Total Selections) / (Number of Selections + 1)], where the total selections are the total number of times all risk assessment arms in the current scenario cluster have been selected, c is an adjustable exploration coefficient used to control the trade-off between exploration and optimization, In represents the natural logarithm, and sqrt represents the square root. All selection confidence scores are obtained, and the risk assessment arm with the highest confidence score is selected to obtain the multidimensional risk vector.

[0039] In summary, the upper-level selection strategy based on the multi-armed gambling machine concept tends to select the risk assessment arm with a good historical performance, i.e., a high average reward, while also giving the arm with fewer selection opportunities a certain chance to avoid getting trapped in local optima.

[0040] Finally, based on low-dimensional context encoding, multimodal feature vectors, and basic risk prototypes, collision probability, threat level, uncertainty, expected contact time, and spatial dispersion are obtained. A risk assessment arm is then used to obtain a multidimensional risk vector. The selected risk assessment arm will execute a specific set of algorithms encapsulated within it to calculate the five components of the multidimensional risk vector, namely collision probability, threat level, uncertainty, expected contact time, and spatial dispersion. Different risk assessment arms may use different mathematical models, parameter weights, or calculation processes to solve these components. The general calculation principles and examples of each component are described below. However, it should be noted that the specific implementation depends on the specific internal strategy of the selected risk assessment arm.

[0041] The collision probability is obtained based on the basic risk prototype using a Bayesian algorithm. The collision probability represents the estimated likelihood of a collision between the identified risky audio source signal stream and the user in the current scenario. The collision probability is a value between [0, 1], with a value closer to 1 indicating a higher probability of collision. It should be noted that the current scenario refers to the comprehensive environment and user state represented by the current multimodal feature vector. Specifically, the basic risk prototype is used as the prior distribution P (risk), and evidence parsed in real-time from the current multimodal feature vector, such as extremely high relative speed, extremely short expected contact time, and the horizontal orientation of the risky audio source, is used. The likelihood P (evidence|risk) is defined as the angle directly in front of the user's movement path, the user's movement state is running, and the direction intersects with the risk source. A posterior probability P (risk|evidence) is obtained using a Bayesian algorithm, and this posterior probability P (risk|evidence) is used as the current collision probability. The evidence can be quantified based on intermediate quantities such as the estimated contact time and relative velocity calculated in real time, using a preset likelihood function. This likelihood function defines the joint probability or probability density of various pieces of evidence, such as the estimated contact time and relative velocity, observed under the condition that the risk occurs. In some possible embodiments, this likelihood function can be based on the sum of multiple pieces of evidence. The weights are combined to construct the risk contribution function. For example, a risk contribution function can be defined for key evidence such as the expected contact time, relative velocity, and azimuth offset (i.e., the angle between the risk source direction and the user's movement direction). The risk contribution function maps the original evidence values ​​to risk scores in the interval [0, 1]. Then, all risk scores are combined into a total likelihood value through geometric averaging. The shape of the weights and risk contribution function can be learned from historical data or pre-set based on kinematic principles. The risk contribution function can be a linear function, a piecewise linear function, or a sigmoid function. For example, for evidence of expected contact time, an sigmoid function can be defined. When the time is greater than 10 seconds, the risk score is close to 0; when it is less than 2 seconds, the risk score is close to 1. The transition is smooth between 2 and 10 seconds. In some possible embodiments, assuming that the currently matched scenario cluster is walking on a city street, its basic risk prototype indicates that the prior collision probability is 0.05, that is, 5% of the cases in history have occurred. The evidence currently parsed in real time from the multimodal feature vector is that a car is approaching from the left at a 30-degree angle, the calculated relative speed is 8 m / s, the estimated contact time is 2.5 seconds, and the user's movement state is walking at a constant speed. The Bayesian algorithm integrates these strong risk evidences and may update the collision probability to 0.4.

[0042] The threat level is obtained based on event labels in the multimodal feature vector. Specifically, the threat level is a static threat severity classification based on the acoustic attributes of the event labels corresponding to the risky sound source streams. It reflects the inherent danger level of different sound source types and is unrelated to their dynamic movement. Specifically, the threat level can be obtained by querying a preset threat level mapping table based on the event labels of the risky sound source signal streams in the multimodal feature vector. The threat level mapping table is defined by the potential harm of acoustic events, such as: [car horn: 3], [emergency alarm sound: 3], [bicycle bell: 2], [large vehicle roar: 3], [human shouting: 1]. The confidence level corresponding to the event label can be used to fine-tune the threat level. For example, when the confidence level is lower than the confidence threshold, such as 0.7, the threat level is reduced by one level.

[0043] The uncertainty is obtained by calculating the matching distance between the low-dimensional context code and the scenario cluster center. The uncertainty represents the confidence level of the current scenario judgment or the degree of difference between the current scenario and existing historical experience patterns. The higher the uncertainty, the more unfamiliar, complex, or contradictory the current situation is, and the more conservative the decision should be. Specifically, the matching distance between the current low-dimensional context code and the corresponding scenario cluster center is obtained. The matching distance can be Euclidean distance. The larger the matching distance, the greater the difference between the current scenario and the historical typical scenario, that is, the higher the uncertainty. The matching distance can be mapped to the range of [0, 1] through a preset normalization function, such as the Sigmoid function, as an uncertainty index. For example, if the matching distance between the current low-dimensional context code and the matched park jogging scenario cluster center is very small, such as 0.05, then the uncertainty is low, such as 0.1. If the matching distance with the nearest urban walking scenario cluster center is large, such as 0.8, or the matching distance with the urban walking scenario cluster center is 0.4 while the matching distance with the transportation hub scenario cluster center is 0.42, the matching distance between the two scenario cluster centers is detailed. Both of these cases indicate high uncertainty, such as 0.75.

[0044] The estimated contact time is obtained based on the relative velocity of the risky sound source signal stream. The estimated contact time represents the time required for the risky sound source, represented by the risky sound source signal stream, to make contact with the user, i.e., to reach the same spatial location, assuming the current relative motion state remains unchanged. The shorter the estimated contact time, the higher the urgency. The estimated contact time can be obtained by estimating the relative velocity between the risky sound source and the user. Specifically, this relative velocity can be estimated by analyzing the Doppler frequency shift of the risky sound source signal stream, or by combining the motion velocity vector and the rate of change of the horizontal azimuth angle of the risky sound source stream over time for geometric estimation. The relative distance is then obtained through a sound source intensity attenuation model. This model is based on the physical principle that the energy attenuation of sound waves propagating in air is inversely proportional to the square of the distance. By calibrating the reference sound source intensity at a known distance, a linear relationship model between the sound pressure level and the logarithmic distance is established, thereby determining the relative distance based on the currently received wind... The relative distance to the dangerous sound source signal stream is estimated by measuring its sound pressure level. Then, the ratio of the relative distance to the relative velocity is calculated and used to calculate the estimated contact time. If a reliable relative distance cannot be obtained, the reciprocal of the relative velocity can be used as a substitute indicator positively correlated with the urgency of the time. For example, assuming the relative distance can be estimated, by analyzing the Doppler frequency shift of the dangerous sound source signal stream, its relative velocity is estimated to be 10 m / s. At the same time, using the sound source intensity attenuation model mentioned above, based on the sound pressure level difference between the sound stream and the reference quiet environment, the relative distance is estimated to be 25 meters. Then, the estimated contact time = 25 meters / 10 meters / second = 2.5 seconds. If, in a complex reverberation environment or when the sound source intensity is unknown, the relative distance cannot be reliably estimated, then the reciprocal of the relative velocity (1 / relative velocity) can be directly used as a substitute indicator positively correlated with the urgency of the time. If the relative velocity is 10 m / s, then the urgency index is 0.1 seconds / meter. The larger this value, the more urgent the situation.

[0045] The spatial dispersion is obtained by performing coherence analysis on all audio signal streams. Specifically, spatial dispersion refers to the degree of spatial dispersion of all separated sound source signal streams in the environment. A large spatial dispersion indicates that the sound sources come from multiple different directions, indicating a complex environment. A small spatial dispersion indicates that the sound sources are more concentrated in one direction. Spatial dispersion can be quantified by performing coherence analysis on all audio signal streams separated in step X1. Specifically, the cross-correlation coefficient or coherence function value of a specific frequency band between each two audio signal streams can be calculated as a coherence index, such as the main energy frequency band. The smaller the coherence index, the stronger the independence between the audio signal streams, that is, the more spatially dispersed the sound sources. Spatial dispersion can be used to calculate statistical measures of the coherence indices among all audio signal streams. For example, the variance or entropy of all coherence indices can be calculated as a scalar value of spatial dispersion. It should be noted that when the number of audio signal streams is less than 2, the spatial dispersion is directly defined as the minimum value, such as 0, to characterize the absolute concentration of sound source directions. For example, suppose there are 4 main sound source signal streams separated at a noisy intersection, coming from the front, left, right, and rear directions respectively. The cross-correlation coefficients between them are very small, and the spatial dispersion obtained by calculating the entropy is very large, such as 0.95. Suppose there is only one main sound source on a quiet straight road, with only one sound source signal stream, then the spatial dispersion is 0.

[0046] Finally, the obtained collision probability, threat level, uncertainty, expected contact time, and spatial dispersion are combined into a data structure, namely a multidimensional risk vector. For example, a specific multidimensional risk vector may be [0.4, 3, 0.75, 2.5, 0.95], which means that the collision probability is 40%, the threat level is 3, the uncertainty is 0.75, the expected contact time is 2.5 seconds, and the spatial dispersion is 0.95.

[0047] In summary, step X2 updates the model by incrementally updating the multimodal feature vectors obtained in step X1 through context. Based on the scenario information database that can represent historical experience and the reward statistics parameters that can provide real-time feedback, the most suitable risk assessment arm is dynamically selected to achieve targeted risk modeling in complex environments.

[0048] X3: Based on multidimensional risk vectors and physiological load indices, the optimal auditory enhancement strategy parameters are obtained in real time through auditory particle swarm optimization decision-making. Specifically, the multidimensional risk vector and physiological load index are used as input vectors to construct an audio-safety multi-objective optimization problem. The audio-safety multi-objective optimization problem includes at least minimizing the safety risk function, minimizing the user interference function, and maximizing the auditory comfort function. The minimizing safety risk function specifically reduces the safety threats in the user's environment. The minimizing user interference function specifically avoids discomfort or interference to the user caused by auditory enhancement, such as excessive amplification of warning sounds. The maximizing auditory comfort function specifically ensures that the overall sound field after processing sounds natural and comfortable.

[0049] It should be noted that the parameters of the auditory enhancement strategy... It is a structured parameter vector composed of multiple sub-parameters. Each sub-parameter controls a specific audio processing operation in the sound field reconstruction, including the risk source gain vector, noise suppression parameters, equalizer setting vector, and spatial filter parameters. The risk source gain vector is either a scalar or a vector, used to control the gain boost of the risk source signal streams identified in step X1 to directly enhance the audibility of the risk warning sound. The unit is typically decibels. If it is a scalar, the same gain is applied to all risk source signal streams; if it is a vector, an independent gain value can be assigned to each risk source signal stream. The noise suppression parameters are a set of parameters controlling the background noise suppression algorithm. They are used to reduce the masking effect of non-risk ambient sound on risky sound sources, while avoiding over-suppression that leads to a hollow or unnatural sound field. For example, the noise suppression parameters can be the over-subtraction factor in spectral subtraction, or the prior signal-to-noise ratio estimation parameters in Wiener filtering. The equalizer setting vector is a frequency domain gain vector that defines gain adjustment values ​​across multiple frequency bands to optimize overall sound quality and clarity. For example, it can boost mid-to-high frequencies to enhance speech intelligibility, or attenuate specific resonant frequencies to reduce harshness. Specifically, the frequency domain gain vector is an array of real numbers. Each element corresponds to a gain adjustment amount for a specific frequency band, usually in decibels (dB). The frequency band division can be based on a standard frequency scale, such as a 1 / 3 octave or a Buck scale. Each divided frequency band corresponds to a gain value in the frequency domain gain vector. The dimensions of the frequency domain gain vector, i.e., the number of frequency bands and the center frequency and bandwidth of each band, are preset. For example, a common implementation might use a 10-band equalizer based on a 1 / 3 octave, whose equalizer setting vector is a 10-dimensional vector [Gb1, Gb2, ..., Gb10], where Gb1 to Gb10 correspond to center frequencies of... A specific frequency band gain, such as [80Hz, 250Hz, ..., 12.5kHz], is defined. Each gain value can be adjusted within a preset range, for example, [-10dB, +10dB]. The spatial filter parameters are used to control the adaptive beamformer or spatial audio rendering algorithm, and may include the pointing angle of the main lobe, beamwidth, null depth, and direction. Through the spatial filter parameters, while enhancing risky sound sources in a specific direction, interference noise from other directions can be suppressed, and the spatial distribution of the sound field can be maintained as much as possible, allowing users to intuitively judge the location of the sound source. The auditory enhancement strategy parameters... The structure can be defined through preset templates, such as auditory enhancement strategy parameters. This can be represented as [risk source gain, noise suppression factor, equalizer low-frequency gain, equalizer mid-frequency gain, equalizer high-frequency gain, beam pointing angle, beamwidth]. For example, suppose the optimal solution selected through the decision rule corresponds to the auditory enhancement strategy parameters. The values ​​are [15, 0.8, -2, 0, 3, 30, 40]. Specifically, the risky sound source is boosted by 15dB, noise suppression with an intensity of 0.8 is applied, low frequencies are attenuated by 2dB on the equalizer, mid frequencies remain unchanged, high frequencies are boosted by 3dB, and a beamformer with a main lobe pointing at 30 degrees and a width of 40 degrees is used to spatially enhance the risky sound source.

[0050] A safety risk objective function is constructed based on the collision probability, threat level, and estimated contact time in the multidimensional risk vector. The security risk objective function Measurable parameters when using auditory enhancement strategies Afterwards, users face residual security risks. Constructed based on key components in the multidimensional risk vector R, for example Among them, the gain of risky sound sources ( ) is the parameter of the auditory enhancement strategy. The boost amount for identified risky sound sources, [w1, w2, w3] and α are preset weights and scaling factors. This indicates that auditory particle swarm enhancement decision-making tends to select auditory enhancement strategy parameters that can effectively enhance risky sound sources and impose higher penalties on high threat levels and short expected contact times, thereby reducing collision risk.

[0051] An interference objective function is constructed based on the physiological load index and the parameters of the candidate auditory enhancement strategies. The interference target function Measurable parameters of auditory enhancement strategies The perceived disturbance to users is based on the understanding that users under high physiological load, such as tension or fatigue, are more sensitive to disturbances. The disturbance objective function is... Highly correlated with the user's physiological workload index P, for example ,in, Based on Audio signal stream power estimation before and after application The difference in the spectral centroid or harmonic structure of the audio signal stream before and after processing can be quantified by comparing the differences. β(P) is a sensitivity coefficient that increases with P, such as β(P) = 1 + k*P. The overall loudness change and spectral unnaturalness are based on... Calculated acoustic features Ensure that a more conservative, less disruptive enhancement strategy is adopted when user load is high.

[0052] Construct an auditory comfort objective function based on candidate auditory enhancement strategy parameters. The auditory comfort objective function Measurable parameters of auditory enhancement strategies The overall auditory quality of the sound scene can include assessments of indicators such as the degree of background noise suppression, the preservation of sound field spatiality, and the naturalness of timbre. =-[Noise suppression distortion ( ) + spatial perception loss ( ]], of which, noise suppression distortion ( It can measure the energy of audio artifacts introduced by noise reduction processing, such as musical noise, and the degree of spatial loss. The reduction in interaural coherence (IACC) between the pre- and post-processing binaural signals can be evaluated. This is an objective that needs to be maximized, and therefore is usually transformed into minimization in multi-objective optimization. .

[0053] In summary, the audio-security multi-objective optimization problem can be used to find the optimal parameters for the auditory enhancement strategy. , so that the vector target [ , , It is optimal in the Pareto sense.

[0054] Secondly, the particle swarm is initialized, where the position of each particle represents a set of candidate auditory enhancement strategy parameters. Specifically, a population of N particles is initialized first, with each particle... Contains location ,speed and individual historical best position Three types of data, the location Let represent a set of candidate auditory enhancement strategy parameters. The vector, whose dimensions and range are determined by the parameter space settings, such as the gain range and frequency boundaries, and its position... The initial value is randomly generated within a preset parameter space range, and the speed... Represented as a vector indicating the direction and step size of parameter adjustment, with dimensions and position. Similarly, its initial value is usually randomly set within a certain range, speed The initial value is usually randomly set within a preset small range, for example , It can be associated with the range of each dimension of the parameter space, for example, set as a certain proportion of the width of that dimension, such as 10%, representing the individual's historical best position. This represents the position that the particle itself finds with the best overall evaluation of the objective function, and it is initialized as its initial position. At the same time, the auditory particle swarm enhancement decision maintains a globally optimal guidance set A, which is used to obtain and store the currently searched non-dominated Pareto optimal solution, i.e., the approximation of the Pareto front.

[0055] Then, for each particle representing the auditory enhancement strategy parameters, the corresponding safety risk function value, user interference function value, and auditory comfort function value are obtained by combining the multidimensional risk vector and physiological load index. Specifically, for each particle in the population... Extract its current position That is, auditory enhancement strategy parameters ,Will Substituting the multidimensional risk vector R into the safety risk objective function Calculate the security risk value, and Substitute the physiological load index P into the interference objective function. Calculate the user interference value, and Substitute into the auditory comfort objective function The auditory comfort value is calculated and transformed into a minimization objective. Each particle i obtains a three-dimensional objective vector during the iteration. .

[0056] Finally, based on the safety risk function value, user interference function value, and auditory comfort function value, the Pareto dominance relation is used for iteration to update the individual historical optimal position of each particle and the global optimal guidance set of the population until the preset convergence condition is met. From the global optimal guidance set, the optimal auditory enhancement strategy parameters are selected using preset decision rules. Specifically, for each particle i, its current three-dimensional target vector is compared. With the individual's historical best position Historical three-dimensional target vector The judgment is based on the Pareto dominance relationship. For example, if... None of the target values ​​are greater than If there is one inferior, and at least one better, then it is called... Dominate ,like Dominate Then update Set the current positions of all particles and all Add the solution to the candidate set. Based on the Pareto dominance relationship, select all non-dominated solutions from the candidate set and update the global optimal guiding set A. If the global optimal guiding set A exceeds a preset upper limit, prune it using methods such as crowding distance sorting to maintain the diversity of solutions and the distribution of the frontier. For each particle i, select a solution from the global optimal guiding set A using a selection strategy as its global guide for this iteration. The selection strategy can be roulette wheel, which assigns weights based on crowding distance, or random selection, to balance convergence and exploration, updating the velocity of each particle according to the standard PSO formula. and location , specific , ,in, These are inertial weights, c1 and c2 are learning factors, and r1 and r2 are random numbers. It should be noted that the updated position... The iterative process must continue within the parameter boundaries until a preset convergence condition is met. This convergence condition can be reaching the maximum number of iterations, such as 300, or the global optimal guiding set A becoming stable. For example, in multiple consecutive iterations, such as 10 iterations, the change in the generational distance index or hypervolume index of the Pareto front is less than a given threshold, such as 0.01. It should be noted that the generational distance index measures the degree of improvement of the new front relative to the old front, and the hypervolume index measures the volume of the space dominated by the front. Its stability indicates that the global optimal guiding set A is difficult to further optimize. After the iteration ends, a final solution is selected from the global optimal guiding set A, i.e., the approximate Pareto front, according to a preset decision rule, as the optimal auditory enhancement strategy parameter. For example, decision rules can be set by defining a decision scalar function. Obtain Calculate the decision scalar function value for all particles in the globally optimal guidance set A, and select the particle position of the particle with the smallest decision scalar function value as the optimal auditory enhancement strategy parameter. , As decision weights, , It reflects the emphasis on different objectives and can be preset according to the application scenario. For example, in scenarios with extremely high security requirements, it can be set to... Much larger and ,may be Safety should be the top priority.

[0057] X4: Reconstructs the sound field based on the optimal auditory enhancement strategy parameters and optimizes the parameters of the context incremental update model.

[0058] First, sound field reconstruction is performed based on the optimal auditory enhancement strategy parameters. Specifically, for the risky sound source signal stream identified and marked in step X1, the sound field is reconstructed according to the optimal auditory enhancement strategy parameters. The gain parameters of the risky sound source can be targeted to increase their amplitude. This can be achieved by applying an adaptive beamformer pointing at the horizontal azimuth angle of the risky sound source signal stream, thereby spatially focusing and amplifying the warning sound. Suppression of non-risky sound sources or background noise can be achieved through spectral subtraction and Wiener filtering to reduce the masking effect of environmental noise on the audibility of the risky sound source signal stream. This is based on the optimal auditory enhancement strategy parameters. The equalizer settings allow you to adjust the gain across the entire frequency range or specific frequency bands to optimize overall sound quality and clarity, ensuring the enhanced sound is natural, discernible, and not harsh, based on optimal auditory enhancement strategy parameters. The spatial filter parameters in the system, while suppressing noise and enhancing risky sound sources, aim to maintain or optimize the spatial sense of the sound scene as much as possible, avoiding the mixing of all sounds. This allows users to intuitively determine the location of the risky sound source signal stream through hearing. For example, assuming the optimal hearing enhancement strategy parameters... If the signal source signal stream from a 30-degree direction with the event label "car horn" is given a 15dB gain boost and low-frequency traffic noise is suppressed by 10dB, the corresponding signal processing algorithm will be invoked to generate an audio signal in real time that is 30 degrees in direction, with the horn sound prominent, the overall ambient noise reduced, and the spatial orientation preserved. This signal will then be played to the user through the headphones' speakers.

[0059] Secondly, after sound field reconstruction, user inertial measurement data is collected. A reward signal is obtained based on this data, where the reward signal is a scalar value used to quantify the decision-making effect. Specifically, an inertial data sequence containing triaxial acceleration and triaxial angular velocity sequences is recorded within a certain period after sound field reconstruction, such as the next 1 to 3 seconds. User feedback features reflecting user avoidance behavior or state stability are extracted from this inertial data sequence. These features include jerkiness, head turning angular velocity and angle, and gait regularity changes. Jerkiness is the derivative of triaxial acceleration, and its peak value can characterize sudden, violent movements, such as sudden stops and dodges. The head turning angular velocity and angle specifically indicate whether the user's head turns towards the risky sound source for observation or in the opposite direction for avoidance. The head turning angular velocity can be directly obtained from the data reflecting the horizontal turning axis in the triaxial angular velocity sequence, while the angle can be obtained by numerically integrating the angular velocity sequence or by using sensor fusion algorithms such as complementary filtering and Kalman filtering to fuse accelerometer and gyroscope data. The angle is used to determine whether the head turns towards the risky sound source. The system observes the direction of the source or turns to the opposite direction to avoid it. Specifically, the gait regularity changes are determined through periodic analysis of three-axis acceleration to assess whether the user's walking or running rhythm has been disrupted. If disrupted, it may be due to alertness or avoidance. Then, according to a preset reward rule, the user feedback characteristics are mapped to a scalar reward signal *r*. This reward rule aims to encourage strategies that prompt the user to adopt safe and smooth avoidance behaviors, and to punish strategies that cause user confusion or ineffective responses. The reward rule can specifically be set to include positive rewards, negative rewards, neutral rewards, and continuous rewards. Specifically, when a rapid and smooth avoidance behavior, such as a lateral step or pause, is detected, and the movement state quickly returns to stability (e.g., a reasonable peak in aggression, a head turn angle pointing towards the risk source or a safe direction, and a gait that returns to regularity within a short period), a positive reward is given, such as r=+1. When a high risk warning level is issued, such as a threat level of level 3 in the multidimensional risk vector R, or a collision probability greater than a high threshold, such as 0.7, but no avoidance action features are detected in the following period, the current optimal auditory enhancement strategy parameters are inferred. It may be ineffective. When violent, unbalanced limb movements are detected, such as abnormally high jerkiness and disordered angular changes, it indicates that the current optimal auditory enhancement strategy parameters are ineffective. If the risk warning is low, such as level 1 in the multidimensional risk vector R and low collision probability, such as 0.1, a neutral reward is given if no obvious avoidance action is detected in the following period of time. In addition, the reward can also be a continuous value, i.e., a continuous reward. For example, a continuous reward r∈[0,1] can be set based on the smoothness and timeliness of the avoidance action.

[0060] Finally, based on the reward signal, an incremental update algorithm is used to update the corresponding reward statistics parameters in the context incremental update model. Specifically, based on the scenario cluster identifier and the selected risk assessment arm recorded during the current decision-making process, the corresponding reward statistics parameters are found in the context incremental update model, and the number of selections, historical average reward, and reward variance are updated. Specifically, the number of selections N is updated to N+1, and the historical average reward Q is updated to... Update the second moment M of the reward variance to ,in The average reward before the update is used, and then the reward variance is calculated based on the updated second moment M. When N>1, .

[0061] In summary, these updated Q, N, and Var values ​​will directly affect the calculation of the upper confidence bound of the context incremental update model in the subsequent operation of the headphones. The risk assessment arm with a larger average reward Q will have a higher probability of being selected in the same or similar scenarios in the future, while the risk assessment arm with fewer selections N can get a chance to try through the exploration items.

[0062] A second aspect of the present invention provides a headphone control system with voice interaction function, comprising: The feature acquisition module is used to acquire a multimodal feature vector containing environmental sound source features, user motion state parameters, and physiological load index. The risk modeling module is used to obtain a multidimensional risk vector based on a multimodal feature vector and by updating the model through contextual incremental updates. The auditory enhancement module is used to obtain the optimal auditory enhancement strategy parameters in real time based on a multidimensional risk vector and a physiological load index through auditory particle swarm enhancement decision-making. The sound field reconstruction module is used to reconstruct the sound field based on the optimal auditory enhancement strategy parameters and to optimize the parameters of the context incremental update model.

[0063] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A headphone control method with voice interaction function, characterized in that, The method includes: Obtain a multimodal feature vector containing environmental sound source features, user motion state parameters, and physiological load index; Based on multimodal feature vectors, dynamic matching and risk modeling are performed through a contextual incremental update model that includes a scenario information database to obtain multidimensional risk vectors; Based on multidimensional risk vectors and physiological load index, the optimal auditory enhancement strategy parameters are obtained in real time through auditory particle swarm optimization decision-making. Sound field reconstruction is performed based on the optimal auditory enhancement strategy parameters, and the parameters of the context incremental update model are optimized.

2. The headphone control method with voice interaction function according to claim 1, characterized in that, Obtain a multimodal feature vector containing environmental sound source features, user motion state parameters, and physiological load index, including: Acquire raw multi-channel audio streams, inertial data, and ultrasonic frequency band echo signals; Based on the original multi-channel audio stream, a blind source separation algorithm is used to obtain multiple audio source signal streams. Event identification is performed on each audio source signal stream to obtain the corresponding event label and confidence level. Risk audio source signal streams are then filtered out, and time delay estimation is performed on the risk audio source signal streams to obtain the horizontal azimuth angle. User motion state parameters are obtained using a motion recognition strategy based on inertial data; The frequency domain index of heart rate variability is obtained based on echo signals in the ultrasonic band and used as a physiological load index. Multimodal feature vectors are obtained based on event labels, confidence levels, horizontal azimuth angles, user motion state parameters, and physiological load indices from the audio source signal stream.

3. The headphone control method with voice interaction function according to claim 1, characterized in that, Based on multimodal feature vectors, dynamic matching and risk modeling are performed through a contextual incremental update model that includes a scenario information database to obtain multidimensional risk vectors, including: The context incremental update model includes a dimensionality reduction encoder, a context information base, reward statistics parameters, a risk assessment arm set, and an upper-level selection strategy; The multimodal feature vector is reduced in dimensionality using the dimensionality reduction encoder to obtain a low-dimensional context code. The low-dimensional context encoding is used as a query vector to perform a matching query in the scenario information database to obtain scenario cluster identifiers and basic risk prototypes. Based on the scenario cluster identifier, obtain the associated scenario-related arm subset, wherein the scenario-related arm subset contains multiple different risk assessment arms; Based on the upper-level selection strategy and reward statistics parameters, a risk assessment arm is selected from the scenario-related arm subset. Collision probability, threat level, uncertainty, expected contact time and spatial dispersion are obtained based on low-dimensional context encoding, multimodal feature vectors and basic risk prototypes, and a risk assessment arm is used to obtain multidimensional risk vectors.

4. The headphone control method with voice interaction function according to claim 3, characterized in that, Based on the upper-level selection strategy and reward statistics parameters, risk assessment arms are selected from the scenario-related arm subset, including: The reward statistics parameters include the historical average reward, number of times selected, and reward variance for each risk assessment arm under the scenario cluster identifier. Based on the reward statistics parameters, the upper confidence bound algorithm is used to obtain the selection confidence scores of all risk assessment arms in the scenario-related arm subset, and risk assessment arms are selected from the scenario-related arm subset based on the selection confidence scores.

5. A headphone control method with voice interaction function according to claim 3, characterized in that, Collision probability, threat level, uncertainty, expected contact time and spatial dispersion are obtained based on low-dimensional context encoding, multimodal feature vectors and basic risk prototypes. A risk assessment arm is used to obtain a multidimensional risk vector, including: The collision probability is obtained based on a basic risk prototype using a Bayesian algorithm; The threat level is obtained based on event labels in the multimodal feature vector; The uncertainty is obtained by calculating the matching distance between the low-dimensional context code and the center of the scenario cluster; The estimated contact time is obtained based on the relative velocity of the risky audio source signal stream; The spatial dispersion is obtained by performing coherence analysis on all audio signal streams.

6. The headphone control method with voice interaction function according to claim 1, characterized in that, Based on a multidimensional risk vector and physiological load index, the optimal auditory enhancement strategy parameters are obtained in real time through auditory particle swarm optimization decision-making, including: Using a multidimensional risk vector and a physiological load index as input vectors, an audio-safe multi-objective optimization problem is constructed. The audio-safe multi-objective optimization problem includes at least minimizing the safety risk function, minimizing the user interference function, and maximizing the auditory comfort function. Initialize the particle swarm, where the position of each particle represents a set of candidate auditory enhancement strategy parameters; For each particle representing the auditory enhancement strategy parameters, the corresponding safety risk function value, user interference function value, and auditory comfort function value are obtained by combining the multidimensional risk vector and physiological load index. Based on the safety risk function value, user interference function value, and auditory comfort function value, the Pareto dominance relation is used for iteration to update the individual historical best position of each particle and the global best guidance set of the population until the preset convergence condition is met. The optimal auditory enhancement strategy parameters are then selected from the global best guidance set using preset decision rules.

7. A headphone control method with voice interaction function according to claim 6, characterized in that, The audio-safe multi-objective optimization problem includes at least minimizing the safety risk function, minimizing the user interference function, and maximizing the auditory comfort function, including: A safety risk objective function is constructed based on the collision probability, threat level, and expected contact time in the multidimensional risk vector. An interference objective function is constructed based on the physiological load index and the parameters of the candidate auditory enhancement strategies. A target function for auditory comfort is constructed based on the parameters of the candidate auditory enhancement strategies.

8. The headphone control method with voice interaction function according to claim 1, characterized in that, Sound field reconstruction is performed based on the optimal auditory enhancement strategy parameters, and the parameters of the context incremental update model are optimized, including: Collect user inertial measurement data after sound field reconstruction, and obtain reward signals based on user inertial measurement data; Based on the reward signal, an incremental update algorithm is used to incrementally update the corresponding reward statistics parameters in the context model.

9. An IoT-based anti-cheating monitoring system for electronic scales, applied to the method described in any one of claims 1 to 8, characterized in that, include: The feature acquisition module is used to acquire a multimodal feature vector containing environmental sound source features, user motion state parameters, and physiological load index. A risk modeling module is used to obtain a multidimensional risk vector based on a multimodal feature vector and by updating the model through contextual incremental updates. The auditory enhancement module is used to obtain the optimal auditory enhancement strategy parameters in real time based on a multidimensional risk vector and physiological load index through auditory particle swarm enhancement decision-making. The sound field reconstruction module is used to reconstruct the sound field based on the optimal auditory enhancement strategy parameters and to optimize the parameters of the context incremental update model.