Earphone rapid environment adaptation noise reduction method based on meta-learning

By using a meta-learning approach, combining feature extraction networks and adaptive strategy optimization networks, the problem of poor adaptability of existing technologies in complex environments is solved, enabling the headphones to quickly adapt to the environment and achieve efficient noise reduction, thus improving the user's call experience.

CN120812460APending Publication Date: 2025-10-17COSONIC INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510815599.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing noise reduction technologies are poorly adaptable to complex and ever-changing environments, making it difficult to quickly adapt to environmental changes, resulting in a decline in call quality. Furthermore, it is difficult to achieve a balance between noise reduction and maintaining the naturalness of speech, and traditional methods perform poorly in various complex scenarios.

Method used

A meta-learning-based approach is adopted to achieve rapid environmental adaptation and noise reduction of headphones by defining a meta-training dataset, executing meta-training, and performing online optimization, combining feature extraction networks and adaptation strategy optimization networks.

Benefits of technology

It achieves rapid adaptation and efficient noise reduction in complex environments, maintains stable call quality and natural voice, maintains excellent performance in various scenarios, and can continuously learn and improve to adapt to users' personal usage habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812460A_ABST
    Figure CN120812460A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of earphone rapid environment adaptation noise reduction methods, in particular to an earphone rapid environment adaptation noise reduction method based on meta learning, which comprises the following steps of: firstly, defining a meta training data set, comprising the steps of earphone sampling, meta-training sample setting, meta-training data set disruption, meta-training data set grouping and meta-training parameter setting; secondly, meta training is executed to achieve environment adaptation noise reduction model training based on a meta learning method, and the training process comprises the steps of calculating the weight of an initial feature extraction network, training the feature extraction network, optimizing an adaptation strategy, and calculating the weight of the initial feature extraction model and the initial weight of an adaptation strategy optimization model; jointly training the feature extraction network and the adaptive strategy optimization network; and finally, online optimization is carried out to realize extraction and optimization of the environment adaptation noise reduction model based on the meta-learning method, parameters can be rapidly adjusted in the face of a new environment, and the adaptation time is greatly shortened.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of earphone rapid environment adaptation noise reduction methods, and particularly relates to an earphone rapid environment adaptation noise reduction method based on meta-learning. BACKGROUND

[0002] With the popularity of smart devices, people have increasingly high requirements for the noise reduction performance of earphones. Maintaining clear call quality in noisy environments has become a key feature of modern communication devices. However, existing noise reduction techniques still face many challenges when confronted with complex and changing environments.

[0003] Traditional noise reduction methods mainly rely on fixed-parameter algorithms such as the minimum mean square error (MMSE) short-time spectral amplitude estimator. These methods perform well in static noise environments, but often fall short when faced with complex and dynamically changing environments. They usually need a long time to adapt to new noise environments, leading to users often encountering sudden call quality degradation when the environment changes. In addition, these methods often struggle to strike a good balance between reducing noise and maintaining natural speech, sometimes introducing noticeable speech distortion.

[0004] In recent years, machine learning methods, particularly deep learning techniques, have made significant progress in the field of speech processing. However, these methods usually require a large amount of training data and computing resources, making it difficult to run in real-time on resource-constrained mobile devices. More importantly, they often lack the ability to quickly adapt to new environments, failing to meet users' real-time needs in different scenarios.

[0005] Another major problem with existing technology is the difficulty of accommodating multiple complex scenarios. For example, a noise reduction algorithm optimized for office environments may not perform well outdoors or in vehicles. This limitation seriously affects user experience, especially when frequently switching between different usage environments.

[0006] In the face of these challenges, the industry urgently needs a new method that can quickly adapt to different environments while maintaining efficient noise reduction performance. The ideal solution should be able to guarantee noise reduction effectiveness while quickly adapting to environmental changes and maintaining stable performance in various complex scenarios. SUMMARY

[0007] The earphone rapid environment adaptation noise reduction method based on meta-learning proposed by the present application is designed to address the above problems. This method skillfully combines the flexibility of meta-learning and the stability of traditional noise reduction algorithms, achieving the organic unification of rapid environment adaptation and efficient noise reduction.

[0008] The application provides a headset rapid environment adaptation noise reduction method based on meta-learning, which comprises the following steps: firstly, defining a meta-training data set, including headset sampling, setting a meta-training sample, shuffling the meta-training data set, grouping the meta-training data set, and setting a meta-training parameter; secondly, performing meta-training to realize environment adaptation noise reduction model training based on the meta-learning method, the training process comprising calculating an initial feature extraction network weight, training the feature extraction network, optimizing an adaptation strategy, calculating an initial feature extraction model weight and an adaptation strategy optimization model initial weight, and jointly training the feature extraction network and the adaptation strategy optimization network; and finally, performing online optimization to realize environment adaptation noise reduction model extraction and optimization based on the meta-learning method, the optimization process comprising collecting a call environment sample and data, data preprocessing, meta-optimizing the feature extraction network, and calculating an optimal weight.

[0009] Preferably, in the step of defining the meta-training data set, when the headset is sampled, the headset is 30-40 cm away from the microphone, the call environment is a closed and undisturbed environment, and the user's speaking speed is moderate; in the step of setting the meta-training sample, a prefabricated meta-training data set is used, the sample format is {environment normal sample, environment interference sample}, wherein the environment normal sample and the environment interference sample each account for 50%, the normal sample is sample data collected when the user is in a normal call environment, and the interference sample is sample data collected when the user is 5-10 m away from the use environment; in the step of grouping the meta-training data set, the meta-training data set is randomly divided into three groups: 1 group meta-training data set, 2 group meta-training data set, and 3 group meta-training data set, and the data amount ratio is 2:1:1; and in the step of setting the meta-training parameter, the meta-training is 100 rounds, each sub-task is trained for 10 rounds, the training sub-tasks are normal group 1 training and interference group 1 training, the normal group is trained for 10 rounds normally first, and then trained for 10 rounds interferingly, the interference group is trained for 10 rounds interferingly first, and then trained for 10 rounds normally, the meta-training learning rate is 0.00001, and the momentum parameter is 0.9.

[0010] Preferably, in the step of performing meta-training, when the initial feature extraction network weight is calculated, the initial weight is obtained by minimizing a loss function L, and the loss function L is defined as:

[0011] L=||F(x;θ)-F(x)|| 2

[0012] wherein F(x;θ) is feature extraction of a neural network, x is an input sample, and θ is an initial weight;

[0013] When the feature extraction network is trained, the 1 group meta-training data set is used for training, and the target function is:

[0014]

[0015] wherein L(θt ) is the training objective function of F(x; θ), θ t is the weight at time t, θ t -1 is the weight at time t-1, Δθ t -1 is the weight update change value at time t-1, μ is the momentum parameter, is the gradient at time t, D t is the training data at time t.

[0016] Preferably, in the execution of the meta-training step: when optimizing the adaptation strategy, use 2 sets of meta-training data sets to train the adaptation strategy optimization network, and the objective function is:

[0017] L adapt (w) = ∑(y-f(x; w0)) 2 + λ||w-w0|| 2

[0018] Where y is the sample label, f(x; w0) is the corresponding network output, w0 is the initial weight, w is the optimized weight, and λ is the regularization parameter;

[0019] When calculating the initial feature extraction model weight and the initial weight of the adaptation strategy optimization model, first train the strategy optimization network to 3 rounds to fix the strategy optimization model weight w1, and then substitute the trained strategy optimization model weight into the calculation of the initial feature extraction model weight F(x; θ); When training the feature extraction network and the adaptation strategy optimization network jointly, use 1 set of meta-training data set to train, and the objective function is:

[0020] L joint (θ, w) = L(θ) + αL adapt (w)

[0021] Where L(θ) is the loss function of the feature extraction network, L adapt (w) is the loss function of the adaptation strategy optimization network, and α is the balance parameter.

[0022] Preferably, in the online optimization step: the data preprocessing includes frame division and windowing of the original data, double-end sampling of the speech data, sampling length of 25ms, window of 20ms, and sampling step of 5ms; Fourier transform is performed on each frame of speech data to obtain time-frequency data, the maximum and minimum frequencies are set, and the energy of all frequency points in the corresponding frequency range of the time-frequency data is extracted; the energy of all frequency points in the time-frequency data is first summed and then averaged; when the meta-optimization feature extraction network is optimized, first collect 50 frames of speech data, then substitute the feature data obtained by preprocessing into the feature extraction network to calculate the weight of the feature extraction network; when calculating the weight of the strategy optimization model, the following formula is used:

[0023] w i=min(w i-1 +Δw i , w max )

[0024] Among them, w i-1 is the historical weight, Δw i is the corresponding weight change value, w max is the maximum weight.

[0025] Preferably, the feature extraction network meets the following requirements: the input data is a feature vector of 25*1 dimension, which is a feature vector of 25 frames, including the maximum frequency energy, the minimum frequency energy, the intermediate frequency energy, the average energy ratio of the first 20 frames, the average energy ratio of the middle 10 frames, the average energy ratio of the last 20 frames, and three historical average energies;

[0026] The output is a 10-dimensional feature vector, including the target features, which include maximum noise energy, minimum noise energy, average noise energy, noise ratio, and three historical average noise energies; the training target is the mean square error between the 10-dimensional output target feature and the label.

[0027] Preferably, the adaptive strategy optimization network meets the following requirements: the input data is a 10-dimensional feature vector; the output is the probability corresponding to the 4-dimensional output actions a, b, c, d; the training target is the maximum mean square error between the predicted probability and the actual probability corresponding to a, b, c, d.

[0028] Preferably, the data preprocessing also includes the following steps: summing the data energy of 0-6000 Hz to obtain the maximum frequency energy; summing the data energy of 500-50 Hz to obtain the minimum frequency energy; averaging the data energy of 0-50 Hz, 500-150 Hz and 260-6000 Hz to obtain the historical average noise energy; averaging the data energy of 500-50 Hz, 1500-1800 Hz and 800-2500 Hz to obtain the intermediate frequency noise energy.

[0029] Preferably, when the calculation strategy optimizes the model weight, the weight update value Δw i Meet the following requirements:

[0030]

[0031] in, is the gradient of the loss function, α is the learning rate; and the sum of the weight change value and the maximum value of the historical weight change cannot exceed the maximum weight; and the sum of the weight change value and the maximum value of the historical weight change cannot exceed the maximum weight.

[0032] Preferably, different conversation scenes are also set in the meta-training data to further improve the earphone noise reduction effect of the environment-adaptive noise reduction method in various conversation scenes, and the different conversation scenes include:

[0033] An office scene, the microphone distance from the person is 50-150 cm;

[0034] A classroom scene, the microphone distance from the person is 15-30 m;

[0035] An elevator scene, the microphone distance from the person is 5-15 m;

[0036] A corridor scene, the microphone distance from the person is 6-15 m, and the walking speed is 10 km / h;

[0037] A conference room scene, the microphone distance from the person is 1.5-3 m;

[0038] A park scene, the microphone distance from the person is 8-12 m, and the walking speed is 5 km / h;

[0039] A shopping mall scene, the microphone distance from the person is 3-5 m;

[0040] A running scene, the microphone distance from the person is 6-10 m, and the walking speed is 5-10 km / h;

[0041] A sports field scene, the microphone distance from the person is 10-15 m, and the walking speed is 3-5 km / h.

[0042] The beneficial effects of the present application mainly include the following aspects:

[0043] By adopting the meta-learning framework, the present application solves the poor adaptability problem of the traditional method. Meta-learning enables the model to "learn how to learn", so that it can quickly adjust the parameters when facing new environments, greatly shortening the adaptation time. This feature ensures that users can enjoy continuous high-quality call experience when the environment changes.

[0044] Another innovative point of the present application is its multi-scene training strategy. By setting various typical scenes such as office, public transportation, and outdoor in the meta-training data, the present method can maintain excellent performance in various complex environments. This comprehensive training strategy solves the limitation of traditional methods that are difficult to consider multiple scenes.

[0045] In addition, the present application ingeniously designs a cooperative mechanism of the feature extraction network and the adaptive strategy optimization network. The cooperation of the two networks not only improves the noise reduction effect, but also ensures the naturalness and intelligibility of the speech. The feature extraction network can capture the subtle features of environmental noise, and the adaptive strategy optimization network dynamically adjusts the noise reduction strategy according to these features. This design effectively solves the contradiction between noise reduction and speech fidelity.

[0046] It is worth mentioning that the online optimization mechanism of the present application enables the model to continuously learn and improve. This not only improves the long-term performance of the system, but also enables the earphone to adapt to the personal usage habits and preferences of the user. This personalized noise reduction experience is difficult to achieve with traditional methods.

[0047] In summary, the present application successfully solves the contradiction between the adaptability, multi-scene compatibility and noise reduction effect of existing noise reduction methods in complex environments by innovatively combining meta-learning, multi-scene training, network collaboration and online optimization. It not only significantly improves the noise reduction performance, but also greatly improves the user's call experience in various complex environments. The application of this method is expected to promote the technological innovation of intelligent earphones and other audio devices, and bring users a more intelligent, efficient and personalized auditory experience. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 The overall method logic diagram of the present application.

[0049] Figure 2 The definition meta-training data set logic diagram of the present application.

[0050] Figure 3 The execution meta-training logic diagram of the present application.

[0051] Figure 4 The online optimization logic diagram of the present application.

[0052] Figure 5 The data preprocessing logic diagram of the present application.

[0053] Figure 6 The meta-optimization feature extraction network logic diagram of the present application.

[0054] Figure 7 The calculation strategy optimization model weight logic diagram of the present application. DETAILED DESCRIPTION

[0055] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined invention purpose, the specific implementation, structure, features and effects of the preferred embodiments of the present application are described in detail below. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0057] Example 1

[0058] ReferenceFigures 1-7 The present invention relates to a meta-learning-based method for rapid environmental adaptation and noise reduction for headphones. This method innovatively applies the principles of meta-learning to achieve rapid adaptive noise reduction in various complex environments. The following describes specific implementations of the present invention in detail.

[0059] First, the core of this invention is to use the principle of meta-learning to achieve rapid adaptive noise reduction for headphones in call environments. This method includes three main steps: defining a meta-training dataset, performing meta-training, and performing online optimization.

[0060] When defining the meta-training dataset, we first performed headphone sampling. Ideally, the headphones were kept within 30-40 cm of the microphone to ensure clear speech signals and simulate real-world usage scenarios. Furthermore, we chose to perform sampling in a closed, interference-free environment to ensure high-quality baseline data. Furthermore, the user's speaking speed was kept moderate, which helped to obtain more stable and representative speech samples.

[0061] Next, we set the meta-training samples. In one embodiment of the present invention, a prefabricated meta-training dataset is used, and its sample format is {normal environment samples, interference environment samples}. It is worth noting that normal environment samples and interference environment samples each account for 50%. This balanced design enables the model to learn the characteristics of normal environment and interference environment at the same time. Specifically, normal samples are sample data collected when the user is in a normal call environment, while interference samples are sample data collected when the user is 5-10m away from the use environment. This design effectively simulates various situations that may be encountered in actual use and enhances the generalization ability of the model.

[0062] We then randomly shuffled the meta-training samples. This step aims to eliminate any possible sequential correlation between samples and improve the robustness of the model. We then randomly divided the meta-training dataset into three groups: 1 meta-training dataset, 2 meta-training datasets, and 3 meta-training datasets, with a data size ratio of 2:1:1. This grouping allows us to use them separately for model training, validation, and testing, effectively preventing overfitting.

[0063] Finally, we set the meta-training parameters. In this invention, the meta-training is set to 100 rounds, with 10 rounds of training for each sub-task. The training sub-tasks include normal group 1 training and interference 1 training, where the normal group is trained normally for 10 rounds first, and then trained with interference for 10 rounds, while the interference group is trained with interference for 10 rounds first, and then trained normally for 10 rounds. This alternating training method can make the model better adapt to the dynamic changes of the environment. At the same time, we set the meta-training learning rate to 0.00001 and the momentum parameter to 0.9. These parameters are the optimal values obtained through a large number of experiments, which can achieve a good balance between training speed and model stability.

[0064] Secondly, we enter the execution of the meta-training phase. The purpose of this phase is to realize the training of the environment-adaptive noise reduction model based on the meta-learning method. First, we calculate the initial feature extraction network weight. This step is achieved by minimizing the loss function L, which is defined as:

[0065] L = ||F(x; θ) - F(x) || 2

[0066] Where F(x; θ) is the feature extraction of the neural network, x is the input sample, and θ is the initial weight. The design of this loss function enables the model to learn the key features of the input sample.

[0067] Then we train the feature extraction network. In this step, we use 1 group of meta-training data set for training, and the objective function is:

[0068]

[0069] Where L(θ t ) is the training objective function of F(x; θ), θ t is the weight at time t, θ t -1 is the weight at time t-1, Δθ t -1 is the weight update change value at time t-1, μ is the momentum parameter, is the gradient at time t, and D t is the training data at time t. The design of this objective function takes into account the historical gradient information, which helps to speed up the training process and avoid local optimal solution.

[0070] Next, we optimize the adaptation strategy. In this step, we use 2 groups of meta-training data sets to train the adaptation strategy optimization network, and the objective function is:

[0071] L adapt (w) = ∑(y - f(x; w0)) 2 + λ||w - w0|| 2

[0072] where y is the sample label, f(x; w0) is the corresponding network output, w0 is the initial weight, w is the optimized weight, and λ is the regularization parameter. The design of this objective function aims to enable the model to quickly adapt to new environments while preventing overfitting through the regularization term.

[0073] Then, we calculate the initial feature extraction model weight and the adaptive strategy optimization model initial weight. Specifically, we first train the strategy optimization network to 3 rounds to fix the strategy optimization model weight w1, and then substitute the trained strategy optimization model weight into the calculation of the initial feature extraction model weight F(x; θ). The purpose of this step is to prepare for the subsequent joint training.

[0074] Finally, we jointly train the feature extraction network and the adaptive strategy optimization network. In this step, we use a set of meta-training data set to train, and the objective function is:

[0075] L joint (θ, w) = L(θ) + αL adapt (w)

[0076] where L(θ) is the loss function of the feature extraction network, L adapt (w) is the loss function of the adaptive strategy optimization network, and α is the balance parameter. This joint training method can make the two networks cooperate with each other and improve the effect of environmental adaptation and noise reduction.

[0077] Third, we enter the online optimization phase. The purpose of this phase is to realize the extraction and optimization of the environmental adaptation and noise reduction model based on the meta-learning method. First, we collect the conversation environment samples and data. Then, we perform data preprocessing. Specifically, we frame and window the original data, and sample the speech data at both ends with a sampling length of 25ms, a window of 20ms, and a sampling step of 5ms. The selection of these parameters is based on the experience value of speech signal processing, which can effectively capture the short-time features of speech.

[0078] Next, we perform Fourier transform on each frame of speech data to obtain time-frequency data. We set the maximum and minimum frequencies and extract the energy of all frequency points in the corresponding frequency range in the time-frequency data. Then, we sum and average the energy of all frequency points in the time-frequency data. This processing method can effectively extract the frequency domain features of speech signals.

[0079] In the preprocessing stage, we also performed a series of energy calculations. First, we summed the data energy from 0 to 6000 Hz to get the maximum frequency energy. This frequency range covers most of the frequency range of human voice. Then we summed the data energy from 500 to 50 Hz to get the minimum frequency energy. This range mainly contains some low-frequency noise. Then we averaged the data energy from frequency 0 to 50 Hz, and frequency 500 to 150 Hz and frequency 260 to 6000 Hz to get the historical average noise energy. Finally, we averaged the data energy from frequency 500 to 50 Hz, frequency 1500 to 1800 Hz and frequency 800 to 2500 Hz to get the intermediate frequency noise energy. These carefully designed frequency segments can effectively distinguish between speech signals and various types of noise.

[0080] Next, we perform the meta-optimization feature extraction network. In this step, we first collect 50 frames of speech data, and then substitute the feature data obtained by preprocessing to calculate the feature extraction network weight. The choice of 50 frames is based on the best value obtained by experiment, which can ensure sufficient information and not cause excessive calculation.

[0081] Then we calculate the strategy optimization model weight. We use the following formula:

[0082] w i =min(w i-1 +Δw i ,w max )

[0083] where w i-1 is the historical weight, Δw i is the corresponding weight change value, and w max is the maximum weight. This formula considers both historical information and limits the upper limit of the weight to prevent instability caused by unlimited weight increase.

[0084] It is worth noting that the weight update value Δw i satisfies the following requirements:

[0085]

[0086] where is the gradient of the loss function, and α is the learning rate. We set the learning rate α to 0.001, which is the optimal value obtained after many experiments, which can achieve a good balance between convergence speed and stability. At the same time, we require that the sum of the weight change value and the maximum historical weight change value cannot exceed the maximum weight, which further guarantees the stability of the model.

[0087] In one embodiment of the invention, our feature extraction network has the following characteristics: the input data is a 25*1 dimensional feature vector, which is a 25-frame feature vector including maximum frequency energy, minimum frequency energy, intermediate frequency energy, average energy ratio of the first 20 frames, average energy ratio of the middle 10 frames, average energy ratio of the last 20 frames, and three historical average energies. This design fully considers the time-frequency characteristics of speech signals and can effectively capture the changes in environmental noise.

[0088] The output of the feature extraction network is a 10-dimensional feature vector, including target features such as maximum noise energy, minimum noise energy, average noise energy, noise ratio, and three historical average noise energies. These features comprehensively describe the noise characteristics of the current environment, providing strong support for subsequent noise reduction processing.

[0089] The training target of the feature extraction network is the mean square error between the 10-dimensional output target features and the labels. This design of training target enables the model to accurately predict various features of environmental noise.

[0090] On the other hand, our adaptive strategy optimization network also has its unique features. First, its input data is a 10-dimensional feature vector, which matches the output dimension of the feature extraction network. Second, its output is the probability of the 4-dimensional output actions a, b, c, d. These 4 actions can be understood as 4 different noise reduction strategies, and the network determines which strategy to use by outputting their probabilities. Finally, the training target of the adaptive strategy optimization network is the maximum mean square error between the predicted probabilities of a, b, c, d and the actual probabilities. This design enables the network to accurately predict the most suitable noise reduction strategy for the current environment.

[0091] An important feature of the invention is the setting of different conversation scenarios in the meta-training data, which further improves the earphone noise reduction effect of the environmental adaptive noise reduction method in various conversation scenarios. These scenarios include:

[0092] 1. Office scenario, microphone distance from person 50-150 cm. This simulates the daily office environment, which may have background noise such as keyboard typing and printer operation.

[0093] 2. Classroom scenario, microphone distance from person 15-30 m. This simulates a large classroom or lecture hall environment, which may have distant echoes and reverberations.

[0094] 3. Elevator scenario, microphone distance from person 5-15 m. This simulates the conversation environment in a closed space, which may have mechanical operation noise.

[0095] 4. Corridor scenario, microphone distance from person 6-15m, walking speed 10km / h. This simulates a moving call scenario, possibly with footstep noise and wind noise.

[0096] 5. Conference room scenario, microphone distance from person 1.5-3m. This simulates a multi-person meeting environment, possibly with multiple people talking at the same time.

[0097] 6. Park scenario, microphone distance from person 8-12m, walking speed 5km / h. This simulates an outdoor environment, possibly with wind noise and distant environmental sounds.

[0098] 7. Mall scenario, microphone distance from person 3-5m. This simulates a noisy indoor environment, possibly with various background music and human voices.

[0099] 8. Running scenario, microphone distance from person 6-10m, walking speed 5-10km / h. This simulates a call scenario in motion, possibly with rapid breathing sounds and mechanical running noise.

[0100] 9. Sports field scenario, microphone distance from person 10-15m, walking speed 3-5km / h. This simulates an open outdoor sports environment, possibly with wind noise and distant shouting.

[0101] These diverse scenario settings enable our model to adapt to various complex real-world environments, greatly improving the practicality and generalization ability of the invention method.

[0102] In summary, the invention applies meta-learning to the field of earphone noise reduction in an innovative way, combining complex data processing, multi-network collaboration, online optimization, and other technologies to achieve rapid environmental adaptation noise reduction for earphones. It not only considers various practical use scenarios, but also integrates advanced technologies from multiple disciplines, demonstrating significant technological innovation and practical value. This method is expected to greatly improve the user experience of earphones.

[0103] To verify the superiority of the invention, we designed a set of examples and comparative examples and tested them on the commonly used speech noise reduction dataset DEMAND (Diverse Environments Multichannel Acoustic Noise Database). This dataset contains 16 different noise environments, making it ideal for evaluating the performance of our environmental adaptation noise reduction method.

[0104] Example 1 uses the meta-learning-based earphone rapid environmental adaptation noise reduction method of the invention. We first performed meta-training according to the method described in the invention, and then tested it in different noise environments of the DEMAND dataset.

[0105] The comparative example 1 uses a traditional fixed parameter noise reduction algorithm, specifically the widely used Minimum Mean Square Error (MMSE) short-time spectral amplitude estimator. This method performs well in static noise environments but may struggle in complex and changing environments.

[0106] We selected the following key indicators to evaluate the noise reduction performance:

[0107] 1. SNR Improvement: Measures the improvement in signal quality after noise reduction.

[0108] 2. PESQ (Perceptual Evaluation of Speech Quality): Evaluates the overall quality of the speech after noise reduction.

[0109] 3. STOI (Short-Time Objective Intelligibility): Evaluates the intelligibility of the speech after noise reduction.

[0110] 4. Environment Adaptation Time: Measures the time required for the algorithm to adapt to a new environment.

[0111] The detection methods for these indicators are as follows:

[0112] SNR Improvement is calculated by comparing the signal-to-noise ratio of the original signal and the signal after noise reduction. PESQ and STOI are obtained by comparing the noise-reduced speech with the clean reference speech using standard evaluation tools. Environment Adaptation Time is measured by recording the time required for the algorithm to achieve stable performance in a new environment.

[0113] We tested our method in 16 noise environments from the DEMAND dataset, with each environment repeated 10 times and averaged. The test results are shown in the following table: Indicator Example 1 Comparative Example 1 SNR Improvement (dB) 12.3 8.7 PESQ 3.8 3.2 STOI 0.92 0.85 Ambient Adaptation Time (seconds) 0.5 N / A

[0114] From the test results, it can be seen that the method of the present invention is significantly better than the traditional method in all indicators. In particular, our method improves the SNR by 41.4%, PESQ by 18.8%, and STOI by 8.2%. These data fully demonstrate the superiority of the present invention in improving speech quality and intelligibility.

[0115] More importantly, our method exhibits excellent environmental adaptation capability. When the environment changes suddenly, the method of the present invention can adapt within an average of 0.5 seconds and quickly achieve the best noise reduction effect. In contrast, the traditional method has no environmental adaptation capability, and its performance fluctuates greatly in different environments.

[0116] Such test results fully demonstrate the innovativeness and practical value of the present application. First, the significantly improved SNR Improvement indicates that our method can effectively suppress noise in various complex environments, thanks to the application of meta-learning framework and multi-scene training data. Second, the improvement of PESQ and STOI proves that our method not only reduces noise, but also maintains the naturalness and intelligibility of speech, which is due to our carefully designed feature extraction network and adaptive strategy optimization network.

[0117] Finally, the fast environmental adaptation time highlights the core advantage of the present application. In real life, the user's communication environment often changes, and the ability to quickly adapt to new environments is crucial to user experience. Our method can complete adaptation within half a second, which means that users can hardly feel the impact of environmental changes and can always enjoy the best call quality.

[0118] In summary, the meta-learning-based earphone fast environmental adaptation noise reduction method of the present application not only performs excellently in various objective indicators, but more importantly, it solves the pain points that traditional methods are difficult to cope with complex and variable environments, and provides users with a more intelligent and efficient noise reduction experience. This method is expected to be widely applied in future smart earphones and other audio devices, bringing users a better auditory experience.

[0119] It should be noted that: the above only describes the preferred embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for rapid environmental adaptation and noise reduction of headphones based on meta-learning, characterized in that: The method includes the following steps: first, defining a meta-training data set, including headphone sampling, setting meta-training samples, shuffling the meta-training data set, grouping the meta-training data set, and setting meta-training parameters; second, performing meta-training to realize environmental adaptation noise reduction model training based on a meta-learning method, and the training process includes calculating the initial feature extraction network weights, training the feature extraction network, optimizing the adaptation strategy, calculating the initial feature extraction model weights and the initial weights of the adaptation strategy optimization model, and jointly training the feature extraction network and the adaptation strategy optimization network; finally, performing online optimization to realize environmental adaptation noise reduction model extraction and optimization based on a meta-learning method, and the optimization process includes collecting call environment samples and data, data preprocessing, meta-optimizing the feature extraction network, and calculating the optimal weights.

2. The method according to claim 1, characterized in that In the step of defining the meta-training data set: when sampling with the earphone, the earphone is 30-40 cm away from the microphone, the call environment is a closed and interference-free environment, and the user speaks at a moderate speed; when setting the meta-training samples, a prefabricated meta-training data set is used, and the sample format is {normal environment sample, interference environment sample}, where the normal environment sample and the interference environment sample each account for 50%, the normal sample is the sample data collected when the user is in a normal call environment, and the interference sample is the sample data collected when the user is 5-10 m away from the use environment; When grouping the meta-training dataset, the meta-training dataset is randomly divided into three groups: 1 meta-training dataset, 2 meta-training datasets, and 3 meta-training datasets, with the data volume ratio of 2:1:1; When setting the meta-training parameters, the meta-training is 100 rounds, each subtask is trained for 10 rounds, the training subtasks are 1 normal group training and 1 interference group training, the normal group is first trained normally for 10 rounds, and then interference training for 10 rounds, the interference group is first trained with interference for 10 rounds, and then normal training for 10 rounds, the meta-training learning rate is 0.00001, and the momentum parameter is 0.

9.

3. The method according to claim 1, characterized in that In the meta-training step, when calculating the initial feature extraction network weights, the initial weights are obtained by minimizing the loss function L, which is defined as: L=||F(x;θ)-F(x)|| 2 Among them, F(x;θ) is the feature extraction of the neural network, x is the input sample, and θ is the initial weight; When training the feature extraction network, one set of training data is used for training, and the objective function is: Among them, L(θ t ) is the training objective function of F(x;θ), θ t is the weight at time t, θ t -1 is the weight at time t-1, Δθ t -1 is the weight update change value at time t-1, μ is the momentum parameter, is the gradient at time t, D t is the training data at time t.

4. The method according to claim 3, characterized in that In the step of performing meta-training: when optimizing the adaptation strategy, two sets of meta-training data sets are used to train the adaptation strategy optimization network, and the objective function is: L adapt (w)=∑(yf(x;w0)) 2 +λ||w-w0|| 2 Where y is the sample label, f(x;w0) is the corresponding network output, w0 is the initial weight, w is the optimized weight, and λ is the regularization parameter; When calculating the initial feature extraction model weights and the initial weights of the adaptive strategy optimization model, first train the strategy optimization network for 3 rounds and then fix the strategy optimization model weight w1. Then substitute the trained strategy optimization model weights into the calculation of the initial feature extraction model weights F(x;θ). When jointly training the feature extraction network and the adaptive strategy optimization network, use one set of training data sets for training, and the objective function is: L joint (θ,w)=L(θ)+αL adapt (w) Among them, L(θ) is the loss function of the feature extraction network, L adapt (w) is the loss function of the adaptive strategy optimization network, and α is the balance parameter.

5. The method according to claim 1, wherein In the online optimization step, data preprocessing includes framing and windowing the original data, performing double-ended sampling on the voice data, with a sampling length of 25ms, a window of 20ms, and a sampling step of 5ms; performing Fourier transform on each frame of voice data to obtain time-frequency data, setting the maximum frequency and minimum frequency, and extracting the energy of all frequency points within the corresponding frequency range in the time-frequency data; summing the energy of all frequency points in the time-frequency data and then taking the average; when meta-optimizing the feature extraction network, first collect 50 frames of voice data, then substitute the feature data obtained by preprocessing, and calculate the feature extraction network weight; When calculating the strategy optimization model weights, the following formula is used: In i =min(in i-1 +Δw i ,In max ) Among them, w i-1 is the historical weight, Δw i is the corresponding weight change value, w max is the maximum weight.

6. The method according to any one of claims 1 to 5, characterized in that The feature extraction network meets the following requirements: the input data is a 25*1 dimension feature vector, which is a feature vector of 25 frames, including the maximum frequency energy, the minimum frequency energy, the middle frequency energy, the average energy ratio of the first 20 frames, the average energy ratio of the middle 10 frames, the average energy ratio of the last 20 frames, and three historical average energies; The output is a 10-dimensional feature vector, including the target features, which include maximum noise energy, minimum noise energy, average noise energy, noise ratio, and three historical average noise energies; The training target is the mean square error between the 10-dimensional output target features and the label.

7. The method according to any one of claims 1 to 6, characterized in that The adaptive strategy optimization network meets the following requirements: the input data is a 10-dimensional feature vector; the output is the probability corresponding to the 4-dimensional output actions a, b, c, and d; the training target is the maximum mean square error between the predicted probability and the actual probability corresponding to a, b, c, and d.

8. The method according to claim 5, characterized in that The data preprocessing also includes the following steps: summing the data energy of 0-6000 Hz to obtain the maximum frequency energy; summing the data energy of 500-50 Hz to obtain the minimum frequency energy; averaging the data energy of 0-50 Hz, 500-150 Hz and 260-6000 Hz to obtain the historical average noise energy; averaging the data energy of 500-50 Hz, 1500-1800 Hz and 800-2500 Hz to obtain the intermediate frequency noise energy.

9. The method according to claim 5, characterized in that When the calculation strategy optimizes the model weight, the weight update value Δw i Meet the following requirements: in, is the gradient of the loss function, α is the learning rate; and the sum of the weight change value and the maximum value of the historical weight change cannot exceed the maximum weight; and the sum of the weight change value and the maximum value of the historical weight change cannot exceed the maximum weight.

10. The method according to any one of claims 1 to 9, characterized in that The method also includes setting different call scenarios in the meta-training data to further improve the headphone noise reduction effect of the environmental adaptation noise reduction method in various call scenarios. The different call scenarios include: In an office environment, the microphone should be 50-150cm away from the person. In classroom scenarios, the microphone is 15-30 meters away from the user; In elevator scenarios, the microphone is 5-15 meters away from the person; In the corridor scenario, the microphone is 6-15 meters away from the person, and the walking speed is 10 km / h; In a conference room, the microphone should be 1.5-3 meters away from the person; In the park scenario, the microphone is 8-12 meters away from the person, and the walking speed is 5 km / h; In shopping malls, the microphone is 3-5 meters away from people; In the treadmill scenario, the microphone is 6-10 meters away from the person, and the walking speed is 5-10 km / h. In a sports field scenario, the microphone is 10-15 meters away from the person, and the walking speed is 3-5 km / h.