A Reinforcement Learning-Based Adaptive Optimization Method and System for Active Noise Cancelling Headphone Parameters

Through a two-layer policy decision-making system combining multi-source perception and reinforcement learning, active noise-canceling headphones can predict future noise changes and optimize resource allocation, solving the problems of noise cancellation lag and resource waste in existing technologies, and achieving personalized dynamic noise cancellation effects and extended battery life.

CN122317486APending Publication Date: 2026-06-30DONGGUAN TAIMING ACOUSTICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-03
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing active noise-canceling headphones lack the ability to predict future noise changes when facing complex dynamic environments, resulting in delayed noise cancellation effects and high resource consumption, making it difficult to achieve personalized adaptive optimization.

Method used

By acquiring sound field signals and user motion signals through a multi-source sensing module, and combining feature extraction and fusion to generate a comprehensive situational coding, a two-layer policy decision system based on reinforcement learning is used for prediction and real-time parameter adjustment, thereby realizing the generation and optimization of dynamic policy time series diagrams.

Benefits of technology

It improves the predictive ability of active noise-canceling headphones in dynamic environments, reduces noise cancellation lag, optimizes resource allocation, extends battery life, and achieves personalized noise cancellation effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122317486A_ABST
    Figure CN122317486A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of active noise cancellation technology. Specifically, it discloses a reinforcement learning-based adaptive optimization method and system for active noise cancellation headphone parameters. The method includes: acquiring headphone microphone sound field signals and user motion signals from inertial sensors to obtain raw multi-source sensing data; extracting and fusing features to generate a current comprehensive situational awareness code; inferring predicted noise physical evolution segments based on the situational awareness code and predicted user motion intentions; performing a two-layer policy decision based on the situational awareness code and evolution segments, retrieving baseline policy elements through a policy memory network, exploring innovative policy time series diagrams through a reinforcement learning orchestrator, and generating a dynamic policy time series diagram through performance comparison and experience consolidation; and synthesizing and parsing a control command sequence to adjust the noise cancellation module's parameters in real time. This invention achieves active prediction and intelligent evolution of noise cancellation strategies, improving the auditory experience and reducing power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of active noise cancellation technology, and relates to a method and system for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning. Background Technology

[0002] Active noise-canceling headphones, as an important personal listening device, work by collecting ambient noise and generating inverse sound waves to cancel it out, providing a purer listening experience. As consumer demands for audio products continue to rise, active noise cancellation technology has matured. Its key components include ambient noise pickup microphones, error microphones, digital signal processors, and noise cancellation algorithm modules. Early products primarily relied on fixed parameters or switching between a few preset modes to achieve noise cancellation, such as airport mode and traffic mode. With technological advancements, more advanced systems have begun to attempt to adaptively adjust noise cancellation parameters through acoustic feature analysis, such as spectral energy distribution, to cope with different types of noise environments.

[0003] However, existing active noise-canceling headphones still have limitations. Most solutions rely on pre-set rules or single acoustic signal analysis to adjust noise-canceling parameters. This typically means that the headphones only passively respond to the current noise situation, with coarse-grained strategy adjustments and a lack of predictive ability for future environmental changes. For example, when a user suddenly moves from a quiet indoor environment to a noisy street, or experiences a sudden increase in wind noise while walking, existing systems often struggle to provide the most suitable noise cancellation effect immediately, easily resulting in delays in noise cancellation, poor handling of transient noise, or auditory discomfort during transitions. Furthermore, some systems, in optimizing noise cancellation, often neglect fine-grained management of computational resource consumption, leading to high power consumption and impacting battery life.

[0004] The main technical shortcomings of the aforementioned existing solutions lie in their insufficient adaptability and intelligence. Relying solely on single-dimensional acoustic information fails to fully consider the complex coupling between the user's own motion state and the external acoustic environment, leading to biased strategy decisions. Furthermore, the lack of a mechanism to predict the future evolution of noise makes the system sluggish and rigid when dealing with rapidly changing dynamic scenarios. More importantly, existing technologies generally lack an intelligent decision-making framework capable of continuous learning, self-optimization, and the generation of better strategies, inherently limiting their performance and making it difficult to continuously adapt to and meet users' increasingly personalized noise reduction needs during long-term use. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above objectives, the present invention proposes the following technical solution: an active noise cancellation headphone parameter adaptive optimization method based on reinforcement learning, comprising: S1, acquiring the sound field signal collected by the headphone microphone and the user motion signal collected by the inertial sensor to obtain the original multi-source perception data.

[0006] S2. Extract and fuse features from the original multi-source sensing data to generate the current comprehensive situation code.

[0007] S3. Based on the current comprehensive situation coding and the user's motion intention predicted from the user's motion signal, the predicted noise physical evolution segment is deduced.

[0008] S4. Based on the current comprehensive situation coding and prediction of the noise physical evolution segment, perform a two-layer policy decision and generate a dynamic policy time sequence diagram; the two-layer policy decision includes: retrieving and generating baseline policy elements through a policy memory network, exploring and generating innovative policy time sequence diagrams through a reinforcement learning policy orchestrator, and comparing the performance of the baseline policy elements and the innovative policy time sequence diagrams and solidifying the experience in real time.

[0009] S5. Based on the dynamic strategy timing diagram, synthesize the noise reduction algorithm control instruction sequence.

[0010] S6. Analyze the noise reduction algorithm control command sequence and adjust the parameters of the headphone's noise reduction module in real time.

[0011] The second aspect of the present invention provides an active noise cancellation headphone parameter adaptive optimization system based on reinforcement learning, comprising: a multi-source sensing module, which acquires sound field signals collected by the headphone microphone and user motion signals collected by an inertial sensor to obtain raw multi-source sensing data.

[0012] The feature fusion module extracts and fuses features from the original multi-source sensing data to generate the current comprehensive situation code.

[0013] The trend prediction module deduces the predicted noise physical evolution segment based on the current comprehensive situation coding and the user's motion intention predicted from the user's motion signal.

[0014] The two-layer policy decision module performs two-layer policy decisions based on the current comprehensive situation coding and the predicted noise physical evolution segment, and generates a dynamic policy time sequence diagram. The two-layer policy decision includes: retrieving and generating baseline policy elements through a policy memory network, exploring and generating innovative policy time sequence diagrams through a reinforcement learning policy orchestrator, and comparing the performance of the baseline policy elements and the innovative policy time sequence diagrams and solidifying the experience in real time.

[0015] The instruction synthesis module uses a dynamic strategy timing diagram to synthesize a noise reduction algorithm to control the instruction sequence.

[0016] The parameter adjustment module analyzes the noise reduction algorithm control command sequence and adjusts the noise reduction module parameters of the headphones in real time.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention can realize the improvement of noise reduction strategy from passive adaptation to active prediction and intelligent generation. By deeply integrating multimodal perception data and combining it with intelligent prediction of future noise evolution trends, the system gets rid of the limitation of traditional noise-canceling headphones that only adjust based on current environmental information. It can predict noise changes earlier and adjust the noise reduction algorithm parameters in advance before the noise actually arrives or changes. This forward-looking processing reduces the lag of noise reduction and improves the continuity and comfort of the noise reduction experience, so that users can enjoy a stable noise reduction effect in dynamic and complex environments.

[0018] (2) The reinforcement learning dual-layer policy hub constructed in this invention combines the mechanisms of "rapid experience matching" and "online policy exploration" to enhance the adaptive and self-evolving capabilities of the noise reduction system. The policy memory network provides a stable baseline policy that has been verified historically, ensuring the reliability of the system; while the reinforcement learning policy orchestrator continuously explores better policy combinations within a safe range. More importantly, through real-time performance comparison and experience solidification, successful exploration results can be absorbed and transformed into system memory in an instant, so that the noise reduction performance is no longer static and fixed, but continuously optimized with the accumulation of usage time and the diversity of scenarios, realizing the continuous evolution of personalized noise reduction capabilities for specific users and environments.

[0019] (3) This invention optimizes the configuration of noise reduction resources, effectively reducing power consumption while improving the noise reduction effect. Traditional noise reduction systems often tend to adopt a "large and comprehensive" or "static switching" strategy, which may lead to unnecessary consumption of computing resources even when high-intensity noise reduction is not required. The dynamic strategy timing diagram of this invention can perform fine-grained time-domain and frequency-domain scheduling of multiple noise reduction algorithm components according to the current and predicted situation, activating and adjusting their running priority and duration as needed. This lean resource management mechanism ensures that while maintaining the noise reduction effect, it effectively saves the processor's computing resources, thereby extending the battery life of the active noise-canceling headphones and improving the ease of use of the product. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1This is a schematic diagram of the implementation steps of the method of the present invention.

[0022] Figure 2 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Example 1

[0025] Please see Figure 1 As shown, the active noise cancellation headphone parameter adaptive optimization method based on reinforcement learning proposed in this invention includes: S1, acquiring the sound field signal collected by the headphone microphone and the user motion signal collected by the inertial sensor to obtain the original multi-source perception data.

[0026] In a preferred embodiment, the sound field signal collected by the headphone microphone and the user motion signal collected by the inertial sensor are obtained to obtain raw multi-source perception data, including: obtaining multi-channel audio signals collected by the headphone microphone to obtain sound field texture signals;

[0027] The user behavior primitive signals are obtained by acquiring acceleration and angular velocity signals collected by inertial sensors.

[0028] The sound field texture signal and the user behavior primitive signal are time-aligned and stitched together to form the original multi-source perception data.

[0029] Specifically, to more accurately perceive the environment and distinguish between valid information and noise, this invention refines the acquisition process of raw multi-source sensing data, enabling it to provide high-quality input for subsequent feature extraction and fusion. The engineering objective is to acquire multi-channel, time-domain aligned acoustic and motion information to construct raw multi-source sensing data that reflects the wearer's behavior and the texture of the external sound field.

[0030] To achieve this, firstly, multiple microphone units located inside and outside the active noise-canceling headphones collect ambient sound information from different locations in space. After analog-to-digital conversion and preprocessing, multi-channel acoustic signals are obtained. These acoustic signals include incoming ambient noise captured from the outside of the headphones and sounds near the user's ear canal captured from the inside of the headphones, together forming a sound field texture signal that reflects the location and type of external sound sources and the user's auditory perception. This can be represented as an instantaneous sound pressure level vector:

[0031]

[0032] in, Represents the time of the i-th microphone The acquired instantaneous sound pressure scalar; This refers to the number of microphones, typically between 2 and 6, depending on the headset model; the superscript outside the square brackets... This operation represents the transpose of a vector or matrix. It converts a row vector into a column vector, facilitating dimension matching and computation in subsequent neural network models or matrix equations.

[0033] Secondly, the inertial measurement unit (IMU) integrated inside the headphones collects motion data of the user's head in three-dimensional space. The IMU typically includes a three-axis accelerometer and a three-axis gyroscope. The accelerometer outputs the translational acceleration vector of the wearer's head. The gyroscope outputs the angular velocity vector of the head rotation. These acceleration and angular velocity data, after processing such as raw data filtering (e.g., Kalman filtering to remove sensor noise) and coordinate system transformation, yield user behavior primitive signals that characterize changes in user head posture, movement speed, and direction. User behavior primitive signals It can be represented as an acceleration vector. and angular velocity vector A sequence that changes over time:

[0034]

[0035] in, The sampling frequency for acceleration and angular velocity is typically set in the range of 200Hz to 1000Hz.

[0036] Finally, to ensure contextual consistency between acoustic and motion information, time alignment and data stitching are performed. This includes simultaneously initiating acoustic and inertial sensor acquisition, or subsequent time synchronization via a shared timestamp system. The synchronized sound field texture signal and user behavior primitive signal are logically stitched together in the time domain to form a continuous, multi-dimensional raw multi-source sensory data stream. This data stream integrates the correlation between acoustic events and user physical actions; for example, a sudden head turn by a user may be accompanied by changes in wind noise or footsteps while walking. Raw multi-source sensory data It is a direct combination of sound field texture signal and user behavior primitive signal in the time dimension, denoted as .

[0037] Through the above engineering steps, raw multi-source perception data was obtained, including sound field texture signals collected by the headphone microphone and user behavior primitive signals collected by the inertial sensor. Its temporal integrity and data dimensionality provide a data foundation for subsequent feature extraction and fusion.

[0038] S2. Extract and fuse features from the original multi-source sensing data to generate the current comprehensive situation code.

[0039] In a preferred embodiment, feature extraction and fusion are performed on the original multi-source sensing data to generate a current comprehensive situation code, including: inputting the original multi-source sensing data into a feature encoder network to extract a primary feature vector;

[0040] The primary feature vector is input into an adversarial discriminator network used to distinguish between comfortable and uncomfortable acoustic scene features for feature modulation, resulting in a high-level feature vector related to auditory comfort.

[0041] State abstraction processing is performed on the high-level feature vectors to generate a current comprehensive situation code that represents the coupling relationship between the current sound field interference and the user's activity state.

[0042] Specifically, to transform high-dimensional raw multi-source sensory data into a low-dimensional yet information-dense representation, this invention designs a multi-stage feature extraction and fusion process. The engineering objective is to modulate a current comprehensive situational encoding from complex and mixed raw signals that accurately describes the user's current integrated acoustic environment and activity state, and is directly related to auditory comfort, providing a standardized input for subsequent prediction and decision-making.

[0043] First, the acquired raw multi-source sensor data is input into a pre-defined feature encoder. This feature encoder typically employs a parallel neural network structure: a one-dimensional convolutional neural network is used to process user behavior primitive signals to extract motion pattern features; simultaneously, the sound field texture signal is converted into a Mel spectrogram, and a two-dimensional convolutional neural network is used to extract acoustic distribution features. The outputs of these parallel branches are concatenated in deeper layers and further processed through fully connected layers, ultimately outputting a fixed-dimensional primary feature vector. This vector is a preliminary, unoptimized fusion feature, typically set between 128 and 512 dimensions, capturing the fundamental relationship between acoustics and motion.

[0044] Next, to improve the directionality of the features, the obtained primary feature vectors are input into a pre-defined adversarial discriminator for feature modulation. This adversarial discriminator is not merely a simple classifier; it includes a feature recalibration module trained to distinguish features originating from comfortable and uncomfortable acoustic scenarios. Specifically, the discriminator employs a multilayer perceptron (MLP) structure, comprising an input layer, two hidden layers, and an output layer. The feature modulation process includes the following steps:

[0045] Extracting attention weights: The output vector of the second-to-last fully connected layer of the adversarial discriminator is truncated and passed through a sigmoid activation function, mapping each element of the vector to the (0,1) interval, generating a channel attention weight vector with the same dimension as the initial feature vector. The values ​​in this weight vector represent the sensitivity of the corresponding feature channel to "auditory discomfort".

[0046] Perform a weighted activation operation: Use the Hadamard product (i.e., element-wise multiplication) to transform the primary feature vectors. With channel attention weight vector The calculation is performed, and the formula is expressed as: .

[0047] Through the aforementioned channel-by-channel feature suppression and enhancement, the system automatically suppresses feature components unrelated to the background environment and amplifies feature components most relevant to negative experiences such as noise interference and ear pressure imbalance, thereby generating a high-level feature vector with a quantitative mapping relationship to auditory comfort. .

[0048] Finally, state abstraction is performed on the high-level feature vector of the continuous value. The purpose of this step is to discretize the continuous feature space to facilitate processing by reinforcement learning models. In practice, this is achieved using a pre-defined codebook or vector quantization techniques to discretize the high-level feature vector. The code is mapped to a finite number of pre-trained codewords, each codeword representing a typical integrated situation. The codeword index output by this mapping process is the final generated encoding of the current integrated situation. It is an integer, for example, ranging from 0 to 1023. It represents the coupling relationship between the current sound field interference and the user's activity state in a highly condensed form, and can be directly used as the input state for subsequent prediction and decision-making models.

[0049] S3. Based on the current comprehensive situation coding and the user's motion intention predicted from the user's motion signal, the predicted noise physical evolution segment is deduced.

[0050] In a preferred embodiment, the predicted noise physical evolution segment is deduced based on the current integrated situation coding and the user's motion intention predicted from the user's motion signal, including: inputting the user's motion signal into a time series prediction model to predict the user's motion intention;

[0051] After the current comprehensive situation code is converted into a continuous situation feature matrix through a preset feature mapping codebook, it is used as the user's motion intention as an initial condition and input into the acoustic physical model used to characterize the relationship between motion and changes in acoustic features.

[0052] The predicted noise physical evolution segment is obtained by forward simulation calculation using an acoustic physical model.

[0053] Specifically, to enable the noise reduction system to be forward-looking, thus shifting from a passive response to active pre-adjustment, this invention designs a noise evolution prediction process based on physical models and motion intentions. The engineering objective is to simulate, at the computational level, the physical changes of the noise field over a short period of time in the future, based on the current known overall situation and predictions of the user's upcoming actions, thereby providing predictive input with a time dimension for strategic decision-making.

[0054] First, based on the collected user motion signals, a time series prediction model is used to predict the user's actions within a preset future time window. This model is typically a lightweight long short-term memory network that takes the user's behavior primitive signals from the past approximately 500 milliseconds as input, learns historical motion patterns, and outputs a predicted sequence of acceleration and angular velocity for the next 1 to 2 seconds. This sequence represents the user's motion intention.

[0055] Next, the current integrated situation code generated in the previous process and the predicted user motion intention are used as initial conditions and driving forces, and input into the preset differentiable acoustic physics model. Specifically, since the current integrated situation code is a discrete codeword index, the system internally pre-sets a situation feature mapping codebook corresponding to the state abstraction processing stage. Before inputting the acoustic physics model, the system performs a reverse lookup in the codebook based on this index to extract the continuous situation feature matrix (i.e., the initial continuous sound pressure field distribution matrix) uniquely bound to this index. Subsequently, this continuous situation feature matrix is ​​combined with the user motion intention as the initial boundary condition of the continuous physical field of the partial differential wave equation in the acoustic physics model, driving the subsequent forward simulation calculation. This model is specifically constructed as a physical information neural network architecture, which embeds the acoustic wave equation as a constraint term into the network structure. The differentiable acoustic physics model specifically contains two parallel computational branches:

[0056] Data-driven branch: A 3-layer Long Short-Term Memory (LSTM) network is used to capture the statistical regularity of environmental noise changes over time.

[0057] Physical constraint branch: This branch incorporates a discretization operator for the differentiable convection wave equation. This operator substitutes the user's motion velocity as a medium flow velocity parameter into the equation to simulate the physical mechanisms of the Doppler effect and wind noise generation.

[0058] The model's "differentiable" property and training process are as follows: During the offline training phase, the model's loss function... It is designed to consist of two parts: . Data loss: the mean square error between the network prediction and the actual acquired acoustic data; Physical residual loss: The residual value of the partial differential equation obtained by substituting the predicted sound field output by the network into the acoustic wave equation; These are preset physical loss weights used to balance the strength of data-driven and physical constraints. The backpropagation algorithm simultaneously minimizes data error and physical residuals, ensuring the model learns the correct parameters. It not only fits the observation data but also strictly follows the physical laws of sound wave propagation in a moving medium.

[0059] Finally, the evolution of noise is deduced through forward simulation calculations within this differentiable acoustic physical model. This process can be formally represented as an iterative update of the state:

[0060]

[0061] in, The acoustic situation feature vector represents the k-th prediction time step in the future simulation process; k represents the iteration step index of the forward simulation calculation, which is determined by the preset future prediction window length and step size. The ratio is obtained; The time step for forward simulation is set according to the system's real-time requirements and computing resources, for example, a value of 20ms. The historical acoustic situation feature vector representing the k-1th prediction time step in the future simulation process is obtained based on the results of the previous iteration. This represents the motion feature vector at the k-1th corresponding time step extracted from the predicted user motion intent; This represents the initial motion vector at the current moment extracted from the predicted user's motion intent; Represents a differentiable acoustic physical model The internal learnable parameters are obtained by training on specific acoustic environment and motion-related data in the offline phase; This represents a differentiable acoustic physical model (mapping function) used to characterize the relationship between motion and changes in acoustic features. It is obtained by modeling physical laws described mathematically, and its function is to iteratively update the execution state, translating motion intent into the evolution trend of the sound field. This is achieved by updating the state in steps of tens of milliseconds within a preset future time window. By iteratively executing the update formula, the system can generate a set of temporal features describing the evolution trajectory of noise characteristics (such as energy spectrum and spatial directivity) over the next 1 to 2 seconds. This set is the final output, a predicted segment of the noise's physical evolution. This segment provides crucial information about "what the noise will become" for subsequent policy decisions.

[0062] S4. Based on the current comprehensive situation coding and prediction of the noise physical evolution segment, perform a two-layer policy decision and generate a dynamic policy time sequence diagram; the two-layer policy decision includes: retrieving and generating baseline policy elements through a policy memory network, exploring and generating innovative policy time sequence diagrams through a reinforcement learning policy orchestrator, and comparing the performance of the baseline policy elements and the innovative policy time sequence diagrams and solidifying the experience in real time.

[0063] In a preferred embodiment, performing a two-layer policy decision and generating a dynamic policy timeline includes: inputting the current integrated situation code and the predicted noise physical evolution segment into the policy memory network to retrieve the baseline policy element;

[0064] The current integrated situation encoding and the predicted noise physical evolution segment are input into the reinforcement learning policy orchestrator to explore within the policy neighborhood defined by the baseline policy element and generate an innovative policy time sequence graph.

[0065] The performance evaluation tool is used to evaluate the noise reduction performance generated by the execution of the baseline strategy element and the innovative strategy time sequence diagram in real time, and to compare and select the performance.

[0066] Based on the results of the performance comparison, either the baseline strategy element or the innovative strategy time series diagram is selected as the final output dynamic strategy time series diagram.

[0067] If the innovation strategy sequence diagram wins in the effectiveness comparison, the current overall situation code, the innovation strategy sequence diagram and its effectiveness gain will be used as new experience to write back and update the strategy memory network.

[0068] Specifically, to achieve an organic combination of rapid and stable response and continuous innovation and optimization of noise reduction strategies, this invention designs a two-layer strategy decision-making process. The engineering objective is to coordinate reflexive rapid strategy matching with exploratory online strategy generation. Through real-time comparison and experience consolidation mechanisms, while ensuring the stability of noise reduction effects, it continuously absorbs better noise reduction strategies, ultimately generating an optimal dynamic strategy time sequence diagram.

[0069] First, two branches are executed in parallel: The first branch inputs the current integrated situation code and the predicted noise physical evolution segment into a pre-defined policy memory network. This network is a fast inference model trained on large-scale offline data, internally storing a large amount of historical experience mappings of "situation-policy-effectiveness". It immediately retrieves the historical best policy that best matches the current input and outputs it as a safe and reliable baseline policy element. The second branch provides the same input to a reinforcement learning-based policy orchestrator. This orchestrator aims to maximize long-term noise reduction benefits, but its exploration behavior is constrained within the policy neighborhood defined by the baseline policy element, avoiding the generation of aggressive policies that degrade user experience. It generates an innovative policy time series diagram describing how multiple noise reduction algorithm units novelly cooperate in the time-frequency domain.

[0070] Next, the system enters the real-time performance comparison phase. Using a pre-set performance evaluator, within a very short time window (e.g., 50 to 100 milliseconds), the control commands corresponding to the baseline policy elements and innovative policy timing diagrams are simulated or briefly executed. This evaluator monitors in real-time the residual noise signal energy fed back from the error microphone inside the earphone and the original computational load percentage of the noise reduction processor.

[0071] The evaluator first performs dimensionless processing on both: dividing the original residual noise energy value by the system's preset maximum tolerable noise energy benchmark value, and dividing the processor load by the maximum available computing power percentage, mapping them to a uniform (0,1) interval to obtain the normalized residual noise energy. Power consumption with normalization .

[0072] Subsequently, a pre-defined unified comprehensive scoring function is invoked to calculate the performance score for each individual. This comprehensive score... The specific calculation formula is as follows: ,in, This is the overall score; the lower the value, the higher the overall effectiveness of the strategy. and These are preset weighting coefficients summing to 1, representing the system's emphasis on noise reduction depth and energy efficiency, respectively. They can be dynamically adjusted based on macroscopic conditions such as the remaining battery power of the headphones (e.g., setting...). , The system compares the combined scores of the two strategies. The strategy with the lower score is selected as the final winner in this decision, and its corresponding strategy diagram is the final output dynamic strategy sequence diagram.

[0073] Finally, the key experience consolidation step is executed. The comparison results indicate the score of the innovation strategy timeline diagram. If the innovation strategy is lower, meaning it wins, the system will not only adopt that strategy but also immediately initiate a write-back process. This process records the successful exploration, forming a new experience tuple containing the current overall situation code, the timeline of the winning innovation strategy, and the effectiveness gain (i.e., the cost difference between the two strategies). This tuple is then asynchronously updated in the knowledge base of the policy memory network. This process allows successful "innovations" to be instantly transformed into quickly accessible "memories," thereby enabling the baseline level of the entire decision-making system to be continuously and adaptively improved.

[0074] In a further preferred embodiment, the policy memory network is a graph neural network; it stores the triple relationship of situation code-policy time sequence graph-performance as node information; its retrieval of baseline policy element includes: performing nearest neighbor search based on the current integrated situation code to obtain candidate nodes, sorting them according to the performance information of the candidate nodes, and selecting the policy time sequence graph associated with the optimal node as the baseline policy element.

[0075] The action space of the reinforcement learning policy orchestrator is a control timing diagram that generates and describes the activation, sleep, and priority of the denoising algorithm units; the state space of the reinforcement learning policy orchestrator includes the current integrated situation encoding and the predicted noise physical evolution segment; the reward function of the reinforcement learning policy orchestrator is determined by a combination of denoising performance and system computational power consumption.

[0076] Specifically, to achieve fast and accurate retrieval of the policy memory network, this invention specifies its internal data structure and working mechanism. The engineering objective is to organize discrete, massive amounts of historical denoising experience into a network with topological relationships, enabling the system to efficiently find the most relevant historical optimal solution as a benchmark for decision-making when facing new situations through structured search and sorting.

[0077] To address this, the policy memory network is implemented using a graph neural network structure. In this structure, each node represents a fixed historical denoising experience, and the edges between nodes represent the similarity of the environmental situations corresponding to different experiences. The core information stored in each node is a situation code-policy time series graph-performance triple relationship. Specifically, the feature vector of each node is generated by its corresponding situation code, and the node itself also stores the time series graph of the policy that was proven to be optimal at that time and the performance score achieved by that policy.

[0078] When the system needs to retrieve baseline policy elements, the retrieval process consists of two consecutive steps: nearest neighbor search and performance ranking. First, the current integrated situation input is encoded into a feature vector. Then, a nearest neighbor search is performed in the node feature space of the graph neural network. This search aims to find the M historical experience nodes most similar to the current input situation, with the similarity measured... This can be achieved by calculating the Euclidean distance between feature vectors:

[0079]

[0080] in, This represents the similarity distance between the current integrated situation coding feature vector and the feature vector of the candidate node in the graph; b represents the component number of the vector. This represents the total number of dimensions of the feature vector; This represents the b-th component of the current integrated situation coding feature vector; This represents the b-th component of the feature vector of the n-th candidate node in the graph. The system will select... The M nodes with the smallest values ​​are selected as candidates, and the value of M is usually between 3 and 10.

[0081] Next, the system proceeds to the performance ranking step. For the selected M candidate nodes, the system reads their associated performance scores from storage. This score is a quantitative indicator determined by historical comparisons; a higher score indicates that the associated policy time series graph performs better under similar situations. The system sorts these M candidate nodes in descending or ascending order based on their performance scores. Finally, the policy time series graph associated with the highest-ranked node is extracted and used as the baseline policy meta-output for this decision. This process ensures that the system not only finds the most similar historical scenario but also selects the most effective response from it.

[0082] In a further implementation, to enable the system to possess the ability to innovate strategies beyond preset rules, this invention clearly defines the core mechanism of the reinforcement learning policy orchestrator. The engineering objective is to construct an intelligent decision-making core capable of autonomously exploring and optimizing multi-algorithm collaborative strategies by defining its perceptual inputs, action outputs, and learning objectives, thereby generating novel and efficient noise reduction schemes.

[0083] The working principle of a policy orchestrator is its accurate perception of the environment, i.e., its state space. This state space consists of two parts: the first part is the current comprehensive situational encoding, representing the current human-machine-environment coupling state; the second part is a predicted segment of the physical evolution of noise, characterizing the noise change trend in the short term. These two parts together constitute a high-dimensional state vector, which serves as the input to the policy orchestrator's neural network, providing it with comprehensive information about "what is happening now" and "what will happen in the future" for its decision-making.

[0084] The decision output of the strategy orchestrator, i.e., the action space, is innovatively defined as generating a control timing diagram describing the time-frequency cooperation relationships between denoising algorithm units. This action does not adjust a single parameter, but rather selects several components from a pre-defined library of denoising algorithm components. These components include, for example, an adaptive filtering component for steady-state noise, a spectral subtraction component for wind noise, and a beamforming component for human voice. The orchestrator then assigns activation and sleep sequences and operational priorities to these components within a very short time window (e.g., 200 to 500 milliseconds). The priority is directly mapped to the proportion of computational resources allocated to that component by the system. Therefore, each action output by the orchestrator is a refined and dynamic blueprint for the denoising process.

[0085] The orchestrator's learning and optimization are driven by a comprehensive reward function. This function aims to guide the policy exploration towards the Pareto optimal frontier, which balances high noise reduction efficiency with low system power consumption. To ensure global consistency in the system evaluation framework, the reinforcement learning reward function directly takes the negative of the comprehensive score output by the performance evaluator. At the end of each decision cycle, the reward function value is calculated based on the performance evaluator's output.

[0086]

[0087] in, This represents the immediate reward obtained in the d-th decision cycle, and is a negative value. The smaller the absolute value, the better the strategy. This represents the discrete decision cycle index in the reinforcement learning decision-making process. The comprehensive score calculated by the representative effectiveness evaluator in the d-th decision cycle; This represents the normalized residual noise energy during the d-th decision cycle; γ represents the normalized computational power consumed during the d-th decision cycle; γ and δ use preset weights consistent with the global performance evaluator. By maximizing long-term cumulative rewards, the policy orchestrator can autonomously learn how to balance noise reduction effects and resource consumption under different situations, generating the optimal innovation policy sequence diagram.

[0088] In a further preferred embodiment, the performance evaluator performs real-time evaluation by monitoring the signal of the noise-canceling error microphone inside the headphones and calculating the residual noise energy value.

[0089] Obtain the processor load during the noise reduction algorithm execution to get the power consumption value;

[0090] The noise reduction performance is obtained by weighting the residual noise energy value and the calculated power consumption value.

[0091] Specifically, to provide an objective and real-time quantitative evaluation basis for the two-layer strategy decision-making module, this invention specifies in detail the implementation method of the performance evaluator. The engineering objective is to convert the physical effect and computational cost of the noise reduction strategy—two different dimensions—into a single, comparable comprehensive score in real time, thereby driving the strategy comparison and learning process.

[0092] The performance evaluator first continuously monitors the noise-canceling error microphone located in the user's ear canal to assess the actual noise reduction effect. It samples the output signal of the error microphone at fixed time windows, such as every 50 milliseconds. A Fast Fourier Transform is applied to the signal within this window to obtain its energy spectrum distribution within the key auditory frequency band, typically 20 Hz to 8 kHz. By integrating this energy spectrum, a scalar value representing the total energy of the residual noise is calculated. Simultaneously, the evaluator queries the system's built-in performance monitoring unit to obtain the number of processor cycles currently used to execute the noise reduction algorithm and converts it into a normalized percentage of processor load.

[0093] As described in the real-time performance comparison phase above, after acquiring and completing the dimensionless processing of the data, the performance evaluator calculates the final comprehensive score through a weighted summation function. The final generated comprehensive score, as a low-dimensional, dimensionless scalar value, simplifies the evaluation dimensions of the system and directly quantifies the instantaneous "cost-effectiveness" of the current strategy. This unified scalarization processing method not only provides an intuitive comparison standard for decision-making by comparison modules at the same level, but also provides a unified and unambiguous driving signal for reward calculation by upper-level reinforcement learning modules, ensuring a logical closed loop from bottom-level monitoring to top-level learning optimization.

[0094] S5. Based on the dynamic strategy timing diagram, synthesize the noise reduction algorithm control instruction sequence.

[0095] In a preferred embodiment, a noise reduction algorithm control instruction sequence is synthesized based on a dynamic strategy timing diagram, including: parsing the dynamic strategy timing diagram to determine the start time, runtime, and collaborative logic of each noise reduction algorithm unit.

[0096] Based on the startup time, runtime, and collaborative logic, these are converted into parameter adjustment instructions that the noise reduction module can execute.

[0097] The parameter adjustment instructions are arranged in chronological order to form a sequence of noise reduction algorithm control instructions with time-series labels.

[0098] Specifically, in order to transform the abstract policy blueprint output by the two-layer policy decision module into low-level instructions that can be directly executed by the noise reduction hardware, this invention specifies a set of instruction synthesis processes from logic to physical logic. The engineering objective is to parse the dynamic policy timing diagram, map the time-frequency cooperation relationships therein into specific hardware parameter adjustment instructions, and finally arrange them into a sequence of noise reduction algorithm control instructions with precise timing.

[0099] First, a parser decomposes the input dynamic policy timing diagram item by item. This parser identifies the identifiers of each defined denoising algorithm unit, as well as the start time relative to the current moment, estimated runtime, and cooperative logic indicating its resource allocation priority for each unit. The start time and runtime are typically quantized in milliseconds. After parsing these logical instructions, the system enters the parameter conversion stage. The system queries a pre-defined parameter mapping library to convert the parsed algorithm unit identifiers and their cooperative logic into specific parameter values ​​that can be directly received by the physical registers or memory addresses in the denoising module. These parameter adjustment instructions include, but are not limited to: updating the tap coefficients of the adaptive filter, setting the gain or attenuation value for a specific frequency band, adjusting the learning step size of the minimum mean square error algorithm, or switching specific hardware acceleration units on or off.

[0100] Finally, the instruction sequence arrangement step begins. The system sorts all the discrete parameter adjustment instructions generated in the previous step according to their corresponding start times in an output buffer. Each instruction is assigned a strict timestamp, ultimately forming a continuous instruction stream with time-stamped tags, which is the final noise reduction algorithm control instruction sequence. This sequence is then sent to the hardware instruction queue to await execution.

[0101] S6. Analyze the noise reduction algorithm control command sequence and adjust the parameters of the headphone's noise reduction module in real time.

[0102] In a preferred embodiment, parsing the noise reduction algorithm control instruction sequence and adjusting the noise reduction module of the headphones in real time includes: unpacking the frame structure of the noise reduction algorithm control instruction sequence and separating the algorithm unit logic identifier, timing trigger tag and control parameter payload.

[0103] The preset hardware register mapping table is invoked to convert the logical identifier of the algorithm unit into the physical address offset of the processor inside the noise reduction module, and the control parameter payload is formatted into target machine code data.

[0104] Based on the timing trigger tag, the target machine code data is pushed into the hardware instruction execution queue, and configuration data including filter tap coefficients, frequency band gain values ​​and algorithm convergence step size are written to the memory unit corresponding to the physical address offset.

[0105] Specifically, to achieve the transformation and low-latency delivery of noise reduction control logic from high-level strategies to the underlying hardware execution environment, this invention designs a real-time parameter injection process based on register mapping. The engineering objective is to ensure that the physical characteristics of the noise reduction module can strictly follow the scheduling intent of the dynamic strategy timing diagram for real-time evolution through standardized instruction parsing and timing alignment mechanisms.

[0106] First, the hardware abstraction layer (HAL) parser unpacks the noise reduction algorithm control command sequence output by S5 into a frame structure. Since the command sequence exists in the form of serial data frames, the parser needs to identify the frame header and extract key fields: among them, the algorithm unit logical identifier is used to identify specific algorithm components, such as feedback noise reduction units or residual noise suppression units; the timing trigger tag defines the absolute or relative timestamp when the parameters take effect; and the control parameter payload contains the specific quantized values ​​generated by the strategy decision.

[0107] Next, the system enters the address mapping and formatting stage. Since the logical identifiers of the noise reduction algorithm units cannot be directly recognized by the hardware, the system calls a pre-built hardware register mapping table. This mapping table defines the correspondence between logical identifiers and physical register addresses in the digital signal processor (DSP) or application-specific integrated circuit (ASIC) within the noise reduction module. The parser converts the logical identifiers into physical address offsets according to the mapping table. Simultaneously, the control parameter payloads are reorganized according to hardware fixed-point or floating-point format requirements, transforming them into target machine code data that can be directly read by the hardware.

[0108] Finally, the execution instructions are issued and parameters are injected. Based on the order in which the timing tags are triggered, the system sequentially pushes the formatted target machine code data into the hardware instruction execution queue. When the system clock reaches the time specified by the tag, the instruction scheduling unit triggers an atomic write operation, injecting the configuration data into the corresponding physical address offset storage unit. This configuration data directly affects the execution parameters at the hardware level, such as: updating filter tap coefficients to change the noise reduction frequency response curve; adjusting the frequency band gain value to control the attenuation depth of a specific noise band; and adjusting the algorithm convergence step size to optimize the steady-state error and convergence speed during the adaptive filtering process.

[0109] Through the above analysis and adjustment process, the system completes a closed loop from abstract strategy to changes in physical acoustic characteristics, ensuring that the noise reduction module can execute complex dynamic optimization schemes with millisecond-level response accuracy.

[0110] It is important to note that, to accommodate the limited computing resources on the active noise-canceling headphone side and achieve the technical effect of "reduced power consumption," the system of this invention preferably adopts an asymmetric computing architecture with end-device collaboration. Specifically: multi-source data acquisition in S1 and real-time parameter adjustment in S6 (i.e., low-level control and execution) are executed on the end side by the low-power digital signal processor (DSP) built into the headphone; while the high-computational-power processes involved in S2 to S5 (including feature encoding, differentiable acoustic physical model derivation, graph neural network retrieval, and reinforcement learning online policy orchestration, etc., AI large-scale model calculations) are all offloaded to the main processor of the paired smart terminal (such as a smartphone) via a low-latency wireless data link (such as Bluetooth). Furthermore, all neural network models running on the smart terminal have undergone INT8 model quantization and network pruning. After the smart terminal completes its calculations, it only sends a very small amount of "noise reduction algorithm control instruction sequence with time-series labels" to the headphone buffer and triggers execution. Through the above-mentioned offloading of computing power and lightweight design, this invention transfers heavy computing, ensuring that the headphone can achieve advanced intelligent noise reduction while maintaining extremely low hardware load and operating power consumption.

[0111] Example 2

[0112] Please see Figure 2 As shown, based on Embodiment 1, the second aspect of the present invention provides an active noise cancellation headphone parameter adaptive optimization system based on reinforcement learning, comprising: a multi-source perception module, a feature fusion module, a trend prediction module, a two-layer policy decision module, an instruction synthesis module, and a parameter adjustment module.

[0113] The multi-source sensing module acquires the sound field signal collected by the headphone microphone and the user motion signal collected by the inertial sensor to obtain the raw multi-source sensing data.

[0114] The feature fusion module extracts and fuses features from the original multi-source sensing data to generate the current comprehensive situation code.

[0115] The trend prediction module deduces the predicted noise physical evolution segment based on the current comprehensive situation coding and the user's movement intention predicted from the user's movement signal.

[0116] The two-layer policy decision module performs two-layer policy decision-making based on the current comprehensive situation coding and the predicted noise physical evolution segment, and generates a dynamic policy time sequence diagram. The two-layer policy decision-making includes: retrieving and generating baseline policy elements through a policy memory network, exploring and generating innovative policy time sequence diagrams through a reinforcement learning policy orchestrator, and comparing the performance of the baseline policy elements and the innovative policy time sequence diagrams and solidifying the experience in real time.

[0117] The instruction synthesis module synthesizes a noise reduction algorithm to control the instruction sequence based on a dynamic strategy timing diagram.

[0118] The parameter adjustment module parses the noise reduction algorithm control command sequence and adjusts the noise reduction module of the headphones in real time.

[0119] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A reinforcement learning-based adaptive optimization method for active noise cancellation headphone parameters, characterized in that, include: S1. Obtain the sound field signal collected by the headphone microphone and the user motion signal collected by the inertial sensor to obtain the raw multi-source sensing data; S2. Extract and fuse features from the original multi-source sensing data to generate the current comprehensive situation code; S3. Based on the current comprehensive situation coding and the user's motion intention predicted from the user's motion signal, the predicted noise physical evolution segment is deduced; S4. Based on the current comprehensive situation coding and prediction of the noise physical evolution segment, perform two-level policy decision-making and generate a dynamic policy timing diagram; The two-layer policy decision-making process includes: retrieving and generating baseline policy elements through a policy memory network, exploring and generating innovative policy time series diagrams through a reinforcement learning policy orchestrator, and comparing the performance of the baseline policy elements and the innovative policy time series diagrams and solidifying the experience in real time. S5. Based on dynamic strategy timing diagram, synthesize noise reduction algorithm control instruction sequence; S6. Analyze the noise reduction algorithm control command sequence and adjust the parameters of the headphone's noise reduction module in real time.

2. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 1, characterized in that, The sound field signal collected by the headphone microphone and the user motion signal collected by the inertial sensor are acquired to obtain raw multi-source sensing data, including: Acquire multi-channel audio signals from the headphone microphone to obtain sound field texture signals; The user behavior primitive signals are obtained by acquiring acceleration and angular velocity signals collected by inertial sensors. The sound field texture signal and the user behavior primitive signal are time-aligned and stitched together to form the original multi-source perception data.

3. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 1, characterized in that, Feature extraction and fusion are performed on the raw multi-source sensing data to generate the current comprehensive situation code, including: The raw multi-source sensing data is input into the feature encoder network to extract the primary feature vector; The primary feature vector is input into an adversarial discriminator network used to distinguish between comfortable and uncomfortable acoustic scene features for feature modulation, resulting in a high-level feature vector related to auditory comfort. State abstraction processing is performed on the high-level feature vectors to generate a current comprehensive situation code that represents the coupling relationship between the current sound field interference and the user's activity state.

4. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 1, characterized in that, Based on the current integrated situation coding and the user's motion intention predicted from the user's motion signals, the predicted noise physical evolution segment is deduced, including: The user's motion signal is input into the time series prediction model to predict the user's motion intention; After the current comprehensive situation code is converted into a continuous situation feature matrix through a preset feature mapping codebook, it is used as the user's motion intention as an initial condition and input into the acoustic physical model used to characterize the relationship between motion and changes in acoustic features. The predicted noise physical evolution segment is obtained by forward simulation calculation using an acoustic physical model.

5. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 1, characterized in that, Perform two-level policy decisions and generate a dynamic policy sequence diagram, including: The current integrated situation code and the predicted noise physical evolution segment are input into the policy memory network to retrieve the baseline policy element; The current integrated situation encoding and the predicted noise physical evolution segment are input into the reinforcement learning policy orchestrator to explore within the policy neighborhood defined by the baseline policy element and generate an innovative policy time sequence graph. The performance evaluation tool is used to evaluate the noise reduction performance generated by the execution of the baseline strategy element and the innovative strategy time sequence diagram in real time, and to compare and select the performance. Based on the results of the performance comparison, either the baseline strategy element or the innovative strategy time series diagram is selected as the final output dynamic strategy time series diagram. If the innovation strategy sequence diagram wins in the effectiveness comparison, the current overall situation code, the innovation strategy sequence diagram and its effectiveness gain will be used as new experience to write back and update the strategy memory network.

6. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 5, characterized in that, The strategy memory network is a graph neural network; Its storage situation coding-policy timing diagram-performance triplet relationship serves as node information; Its retrieval baseline strategy element includes: performing a nearest neighbor search based on the current integrated situation coding to obtain candidate nodes, sorting the candidate nodes according to their performance information, and selecting the strategy time sequence diagram associated with the optimal node as the baseline strategy element; The action space of the reinforcement learning policy orchestrator is a control timing diagram that generates and describes the activation, sleep, and priority of the denoising algorithm units; the state space of the reinforcement learning policy orchestrator includes the current integrated situation encoding and the predicted noise physical evolution segment; the reward function of the reinforcement learning policy orchestrator is determined by a combination of denoising performance and system computational power consumption.

7. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 5, characterized in that, The performance evaluator performs real-time evaluations in the following ways: The signal from the noise-canceling error microphone inside the headphones is monitored, and the residual noise energy value is calculated. Obtain the processor load during the noise reduction algorithm execution to get the power consumption value; The noise reduction performance is obtained by weighting the residual noise energy value and the calculated power consumption value.

8. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 1, characterized in that, Based on the dynamic policy timing diagram, the control instruction sequence of the noise reduction algorithm is synthesized, including: parsing the dynamic policy timing diagram and determining the start time, runtime and cooperative logic of each noise reduction algorithm unit. Based on the startup time, runtime, and collaborative logic, these are converted into parameter adjustment instructions that the noise reduction module can execute. The parameter adjustment instructions are arranged in chronological order to form a sequence of noise reduction algorithm control instructions with time-series labels.

9. The method for adaptive optimization of active noise cancellation headphone parameters based on reinforcement learning according to claim 1, characterized in that, The noise reduction algorithm control command sequence is analyzed to adjust the parameters of the headphone's noise reduction module in real time, including: The frame structure of the noise reduction algorithm control instruction sequence is unpacked to separate the algorithm unit logic identifier, timing trigger tag and control parameter payload; The preset hardware register mapping table is invoked to convert the logical identifier of the algorithm unit into the physical address offset of the processor inside the noise reduction module, and the control parameter payload is formatted into target machine code data. Based on the timing trigger tag, the target machine code data is pushed into the hardware instruction execution queue, and configuration data including filter tap coefficients, frequency band gain values ​​and algorithm convergence step size are written to the memory unit corresponding to the physical address offset.

10. A reinforcement learning-based adaptive optimization system for active noise cancellation headphone parameters, characterized in that, include: The multi-source sensing module acquires the sound field signal collected by the headphone microphone and the user motion signal collected by the inertial sensor to obtain the raw multi-source sensing data; The feature fusion module extracts and fuses features from the raw multi-source sensing data to generate the current comprehensive situation code. The trend prediction module, based on the current comprehensive situation coding and the user's motion intention predicted from the user's motion signal, deduces the predicted noise physical evolution segment; The two-layer policy decision module executes two-layer policy decisions based on the current integrated situation coding and the predicted noise physical evolution segment, and generates a dynamic policy timing diagram. The two-layer policy decision-making process includes: retrieving and generating baseline policy elements through a policy memory network, exploring and generating innovative policy time series diagrams through a reinforcement learning policy orchestrator, and comparing the performance of the baseline policy elements and the innovative policy time series diagrams and solidifying the experience in real time. The instruction synthesis module, based on a dynamic strategy timing diagram, synthesizes a noise reduction algorithm to control the instruction sequence. The parameter adjustment module analyzes the noise reduction algorithm control command sequence and adjusts the noise reduction module parameters of the headphones in real time.