Noise reduction system for TWS wireless earphones and method thereof
By capturing changes in the user's center of gravity data using deep learning technology, and combining convolutional neural networks and diffusion models to optimize the audio signal of TWS wireless earbuds, the problem of frequency shift deviation during movement is solved, thus improving the sound effect.
Patent Information
- Application Number
- CN202310082831.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-02-08
AI Technical Summary
When users use TWS wireless earbuds during exercise, ambient noise and user movement cause frequency shift deviation, affecting the sound quality.
By employing deep learning-based artificial intelligence technology, the audio signal of TWS wireless earbuds is optimized by capturing changes in the user's center of gravity data, and noise reduction is performed using convolutional neural networks and diffusion models.
It effectively compensates for frequency shift deviation and improves the sound quality of TWS wireless earbuds in sports mode to meet user needs.
Smart Images

Figure CN116193314B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of noise reduction technology for wireless headphones, and more specifically, to a noise reduction system and method for TWS wireless headphones. Background Technology
[0002] TWS stands for True Wireless Stereo. TWS technology is based on Bluetooth chip technology. Its working principle involves the phone connecting to the main earbuds, which then wirelessly connect to the secondary earbuds, achieving true wireless separation of the left and right audio channels. In other words, audio data is first transmitted from the phone to the main earbuds, and then from the main earbuds to the secondary earbuds.
[0003] In scenarios where users use TWS wireless earbuds during exercise, the surrounding environment is noisy, and the user's exercise mode causes frequency shifts and other deviations in the sound data during transmission. This makes it impossible for users wearing TWS wireless earbuds to hear satisfactory sound effects.
[0004] Therefore, an optimized noise reduction solution for TWS wireless earbuds is needed. Summary of the Invention
[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a noise reduction system and method for TWS wireless earbuds. This system utilizes deep learning-based artificial intelligence technology to capture changes in a user's motion patterns from their center of gravity data, thereby optimizing and compensating for frequency shift deviations that occur during the use of the TWS wireless earbuds. This allows for sufficient noise reduction optimization of the audio signal, ensuring that the sound effects of the TWS wireless earbuds meet the needs of users in motion mode.
[0006] Accordingly, according to one aspect of this application, a noise cancellation system for TWS wireless earbuds is provided, comprising:
[0007] The signal receiving module is used to acquire the audio signal propagating to the left earphone within a predetermined time period;
[0008] The user motion data monitoring module is used to acquire the center of gravity data of users at multiple predetermined time points within the predetermined time period.
[0009] A primary noise reduction module is used to pass the audio signal through an automatic codec-based noise reducer to obtain a noise-reduced audio signal;
[0010] An audio waveform feature extraction module is used to pass the noise-reduced audio signal through a convolutional neural network model as a filter to obtain an audio waveform feature vector.
[0011] The motion pattern feature extraction module is used to arrange the center of gravity data of the users at multiple predetermined time points into a center of gravity input vector according to the time dimension, and then pass it through the multi-scale neighborhood feature extraction module to obtain the center of gravity temporal feature vector.
[0012] A multimodal association module is used to perform association encoding on the centroid time-series feature vector and the audio waveform feature vector to obtain a multimodal association feature matrix;
[0013] The feature optimization module is used to perform high-dimensional data manifold optimization on the multimodal association feature matrix to obtain an optimized multimodal association feature matrix; and
[0014] The secondary noise reduction module is used to pass the optimized multimodal correlation feature matrix through a noise reduction generator based on a diffusion model to obtain an optimized audio signal.
[0015] In the noise reduction system for TWS wireless earphones described above, the automatic codec includes a sound feature encoder and a sound feature decoder.
[0016] In the aforementioned noise reduction system for TWS wireless earphones, the primary noise reduction module includes: a sound signal encoding unit for inputting the audio signal into the sound feature encoder of the noise reducer, wherein the sound feature encoder uses a convolutional layer to explicitly spatially encode the audio signal to obtain sound features; and a sound feature decoding unit for inputting the sound features into the sound feature decoder of the noise reducer, wherein the sound feature decoder uses a deconvolutional layer to deconvolve the sound features to obtain the noise-reduced audio signal.
[0017] In the aforementioned noise reduction system for TWS wireless earphones, the audio waveform feature extraction module is further configured to: perform the following operations during the forward propagation of each layer of the convolutional neural network model: convolution processing on the input data to obtain a convolutional feature map; performing mean pooling on the convolutional feature map based on the local feature matrix to obtain a pooled feature map; and performing nonlinear activation on the pooled feature map to obtain an activation feature map; wherein the output of the last layer of the convolutional neural network model is the audio waveform feature vector, and the input of the first layer of the convolutional neural network model is the noise-reduced audio signal.
[0018] In the noise reduction system for TWS wireless earphones described above, the multi-scale neighborhood feature extraction module includes: a first convolutional layer and a second convolutional layer that run in parallel with each other, and a multi-scale fusion layer connected to the first convolutional layer and the second convolutional layer, wherein the first convolutional layer and the second convolutional layer each use one-dimensional convolutional kernels with different scales.
[0019] In the aforementioned noise reduction system for TWS wireless earphones, the motion pattern feature extraction module includes: a first-scale encoding unit, used to perform one-dimensional convolutional encoding on the centroid input vector using the first convolutional layer of the multi-scale neighborhood feature extraction module to obtain a first-scale centroid feature vector; wherein, the formula is:
[0020]
[0021] Where, a is the width of the first convolutional kernel in the x-direction, F(a) is the parameter vector of the first convolutional kernel, G(xa) is the local vector matrix operated with the convolutional kernel function, w is the size of the first convolutional kernel, X represents the centroid input vector, and Cov1(X) represents one-dimensional convolutional encoding of the centroid input vector; the second scale encoding unit is used to perform one-dimensional convolutional encoding on the centroid input vector using the second convolutional layer of the multi-scale neighborhood feature extraction module to obtain the second scale centroid feature vector; wherein, the formula is:
[0022]
[0023] Where b is the width of the second convolution kernel in the x direction, F(b) is the parameter vector of the second convolution kernel, G(xb) is the local vector matrix operated with the convolution kernel function, m is the size of the second convolution kernel, X represents the centroid input vector, and Cov2(X) represents one-dimensional convolution encoding of the centroid input vector; and a multi-scale feature fusion unit is used to concatenate the first-scale centroid feature vector and the second-scale centroid feature vector using the multi-scale fusion layer of the multi-scale neighborhood feature extraction module to obtain the centroid temporal feature vector.
[0024] In the aforementioned noise reduction system for TWS wireless earphones, the multimodal association module is further configured to: perform association encoding on the centroid timing feature vector and the audio waveform feature vector using the following formula to obtain a multimodal association feature matrix; wherein, the formula is:
[0025]
[0026] in V represents the transpose of the audio waveform feature vector. b Let M represent the centroid temporal feature vector, and let M represent the multimodal correlation feature matrix. This represents matrix multiplication.
[0027] In the aforementioned noise reduction system for TWS wireless earphones, the feature optimization module includes: a feature matrix expansion unit, used to expand the multimodal correlation feature matrix into multimodal correlation feature vectors; and a vector optimization unit, used to perform vector norming Hilbert probability spacerization on the multimodal correlation feature vectors using the following formula to obtain optimized multimodal correlation feature vectors; wherein, the formula is:
[0028]
[0029] Where V is the multimodal association feature vector, and ||V||2 represents the L2 norm of the multimodal association feature vector. v represents the square of the L2 norm of the multimodal correlation feature vector. i It is the i-th eigenvalue of the multimodal correlation feature vector, exp(·) represents the vector exponentiation operation, which means calculating the natural exponent function value raised to the power of each eigenvalue in the vector, and v i ' is the i-th eigenvalue of the optimized multimodal association feature vector; and, dimension reconstruction unit, used to perform dimension reconstruction on the optimized multimodal association feature vector to obtain the optimized multimodal association feature matrix.
[0030] According to another aspect of this application, a noise reduction method for TWS wireless earphones is also provided, comprising:
[0031] Acquire the audio signal transmitted to the left earphone within a predetermined time period;
[0032] Obtain the center of gravity data of users at multiple predetermined time points within the predetermined time period;
[0033] The audio signal is passed through an automatic codec-based noise reduction device to obtain a noise-reduced audio signal;
[0034] The denoised audio signal is passed through a convolutional neural network model as a filter to obtain an audio waveform feature vector;
[0035] After arranging the centroid data of the users at the multiple predetermined time points into a centroid input vector according to the time dimension, the centroid temporal feature vector is obtained by the multi-scale neighborhood feature extraction module.
[0036] The centroid timing feature vector and the audio waveform feature vector are correlated and encoded to obtain a multimodal correlation feature matrix;
[0037] The multimodal association feature matrix is optimized using high-dimensional data manifold optimization to obtain an optimized multimodal association feature matrix; and
[0038] The optimized multimodal correlation feature matrix is passed through a diffusion-based noise reduction generator to obtain the optimized audio signal.
[0039] In the above-described noise reduction method for TWS wireless earphones, the automatic codec includes a sound feature encoder and a sound feature decoder.
[0040] In the above-described noise reduction method for TWS wireless earphones, the step of passing the audio signal through an automatic codec-based noise reducer to obtain a noise-reduced audio signal includes: inputting the audio signal into the sound feature encoder of the noise reducer, wherein the sound feature encoder uses a convolutional layer to explicitly spatially encode the audio signal to obtain sound features; and inputting the sound features into the sound feature decoder of the noise reducer, wherein the sound feature decoder uses a deconvolutional layer to deconvolve the sound features to obtain the noise-reduced audio signal.
[0041] In the above-described noise reduction method for TWS wireless earphones, the step of passing the noise-reduced audio signal through a convolutional neural network model as a filter to obtain an audio waveform feature vector includes: using each layer of the convolutional neural network model to perform the following in the forward pass of the layer: convolution processing on the input data to obtain a convolutional feature map; performing mean pooling on the convolutional feature map based on the local feature matrix to obtain a pooled feature map; and performing nonlinear activation on the pooled feature map to obtain an activation feature map; wherein, the output of the last layer of the convolutional neural network model is the audio waveform feature vector, and the input of the first layer of the convolutional neural network model is the noise-reduced audio signal.
[0042] In the above-mentioned noise reduction method for TWS wireless earphones, the multi-scale neighborhood feature extraction module includes: a first convolutional layer and a second convolutional layer that run in parallel with each other, and a multi-scale fusion layer connected to the first convolutional layer and the second convolutional layer, wherein the first convolutional layer and the second convolutional layer each use one-dimensional convolutional kernels with different scales.
[0043] In the above-mentioned noise reduction method for TWS wireless earphones, the step of arranging the centroid data of the users at multiple predetermined time points into a centroid input vector according to the time dimension and then obtaining a centroid temporal feature vector through a multi-scale neighborhood feature extraction module includes: using the first convolutional layer of the multi-scale neighborhood feature extraction module to perform one-dimensional convolutional encoding on the centroid input vector to obtain a first-scale centroid feature vector using the following formula; wherein, the formula is:
[0044]
[0045] Where, a is the width of the first convolutional kernel in the x-direction, F(a) is the parameter vector of the first convolutional kernel, G(xa) is the local vector matrix operated with the convolutional kernel function, w is the size of the first convolutional kernel, X represents the centroid input vector, and Cov1(X) represents the one-dimensional convolutional encoding of the centroid input vector; the second convolutional layer of the multi-scale neighborhood feature extraction module performs one-dimensional convolutional encoding on the centroid input vector using the following formula to obtain the second-scale centroid feature vector; wherein, the formula is:
[0046]
[0047] Where b is the width of the second convolution kernel in the x direction, F(b) is the parameter vector of the second convolution kernel, G(xb) is the local vector matrix operated with the convolution kernel function, m is the size of the second convolution kernel, X represents the centroid input vector, and Cov2(X) represents one-dimensional convolution encoding of the centroid input vector; and the multi-scale fusion layer of the multi-scale neighborhood feature extraction module concatenates the first-scale centroid feature vector and the second-scale centroid feature vector to obtain the centroid temporal feature vector.
[0048] In the above-described noise reduction method for TWS wireless earphones, the step of associating and encoding the centroid timing feature vector and the audio waveform feature vector to obtain a multimodal association feature matrix includes: associating and encoding the centroid timing feature vector and the audio waveform feature vector using the following formula to obtain a multimodal association feature matrix; wherein, the formula is:
[0049]
[0050] in V represents the transpose of the audio waveform feature vector. b Let M represent the centroid temporal feature vector, and let M represent the multimodal correlation feature matrix. This represents matrix multiplication.
[0051] In the above-mentioned noise reduction method for TWS wireless earphones, the step of performing high-dimensional data manifold optimization on the multimodal correlation feature matrix to obtain an optimized multimodal correlation feature matrix includes: expanding the multimodal correlation feature matrix into multimodal correlation feature vectors; and performing vector norming Hilbert probability spaceization on the multimodal correlation feature vectors using the following formula to obtain the optimized multimodal correlation feature vectors; wherein, the formula is:
[0052]
[0053] Where V is the multimodal association feature vector, and ||V||2 represents the L2 norm of the multimodal association feature vector. v represents the square of the L2 norm of the multimodal correlation feature vector. i It is the i-th eigenvalue of the multimodal correlation feature vector, exp(·) represents the vector exponentiation operation, which means calculating the natural exponent function value raised to the power of each eigenvalue in the vector, and v i ' is the i-th eigenvalue of the optimized multimodal association feature vector; and the optimized multimodal association feature vector is reconstructed in dimensions to obtain the optimized multimodal association feature matrix.
[0054] According to another aspect of this application, an electronic device is provided, comprising: a processor; and a memory storing computer program instructions, which, when executed by the processor, cause the processor to perform the noise reduction method for TWS wireless earphones as described above.
[0055] According to another aspect of this application, a computer-readable medium is provided having computer program instructions stored thereon, which, when executed by a processor, cause the processor to perform the noise reduction method for TWS wireless earphones as described above.
[0056] Compared to existing technologies, the noise reduction system and method for TWS wireless earbuds provided in this application utilize deep learning-based artificial intelligence technology to capture changes in the user's motion patterns from the user's center of gravity data, thereby optimizing and compensating for frequency shift deviations generated during the use of TWS wireless earbuds. This allows for sufficient noise reduction optimization of the audio signal, ensuring that the sound effects of the TWS wireless earbuds meet the needs of users in motion mode. Attached Figure Description
[0057] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0058] Figure 1 This is a block diagram of a noise reduction system for TWS wireless earphones according to an embodiment of this application.
[0059] Figure 2 This is a schematic diagram of the architecture of a noise reduction system for TWS wireless earphones according to an embodiment of this application.
[0060] Figure 3This is a block diagram of a motion pattern feature extraction module in a noise cancellation system for TWS wireless earphones according to an embodiment of this application.
[0061] Figure 4 This is a block diagram of a feature optimization module in a noise reduction system for TWS wireless earphones according to an embodiment of this application.
[0062] Figure 5 This is a flowchart of a noise reduction method for TWS wireless earphones according to an embodiment of this application.
[0063] Figure 6 This is a block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0064] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0065] Application Overview
[0066] As mentioned above, in scenarios where users use TWS wireless earbuds during exercise, the surrounding environment is noisy, and the user's activity mode causes frequency shifts and other deviations in sound data propagation. This results in unsatisfactory sound quality for TWS wireless earbud users. Therefore, an optimized noise reduction solution for TWS wireless earbuds is needed.
[0067] Accordingly, considering that users in sports mode experience significant motion displacement during TWS wireless earbuds, leading to frequency shift deviations, optimization and compensation are needed to ensure satisfactory sound quality for active users. Furthermore, recognizing that users' movement states vary during exercise, capturing their movement patterns is crucial for accurate sound data compensation. The primary correlation between sound quality changes and user movement state lies in the shift in the user's center of gravity. In other words, the shift in the user's center of gravity differs depending on their sports mode, resulting in varying frequency shift deviations and requiring different levels of compensation. Therefore, this application aims to comprehensively optimize the audio signal of TWS wireless earbuds based on the implicit feature distribution information of the audio signal and the temporal variation characteristics of the user's center of gravity data. The challenge in this process lies in how to extract the correlation feature distribution information between the implicit features of the audio signal of the TWS wireless earphone and the temporal variation features of the user's center of gravity data, so as to optimize the audio signal of the TWS wireless earphone and make the sound effect of the TWS wireless earphone meet the needs of users in sports mode.
[0068] In recent years, deep learning and neural networks have been widely applied in fields such as computer vision, natural language processing, and text signal processing. Furthermore, deep learning and neural networks have demonstrated near-human or even superior performance in areas such as image classification, object detection, semantic segmentation, and text translation.
[0069] The development of deep learning and neural networks has provided new ideas and solutions for mining the correlation feature distribution information between the implicit features of the audio signals of TWS wireless earbuds and the temporal variation features of the user's center of gravity data. Those skilled in the art will know that deep neural network models based on deep learning can be trained using appropriate strategies, such as backpropagation algorithms with gradient descent, to adjust the parameters of the deep neural network model so that it can simulate complex nonlinear relationships between things. This is clearly suitable for simulating and mining the correlation feature distribution information between the implicit features of the audio signals of TWS wireless earbuds and the temporal variation features of the user's center of gravity data.
[0070] Specifically, in the technical solution of this application, firstly, the audio signal transmitted to the left earphone within a predetermined time period is acquired. Specifically, the audio signal of the left earphone is the audio signal transmitted from the mobile phone to the main earphone. Next, considering that noise such as ambient noise may be generated when the mobile phone transmits the audio signal to the left earphone, thereby reducing the sound quality and effect of the audio signal transmitted to the left earphone, the technical solution of this application, in order to filter out this noise and improve the accuracy of audio signal feature extraction, passes the audio signal through a noise reduction device based on an automatic codec to obtain a noise-reduced audio signal. Specifically, the automatic codec includes a voice feature encoder and a voice feature decoder, wherein the voice feature encoder uses a convolutional layer to explicitly spatially encode the audio signal to obtain audio features, and the voice feature decoder uses a deconvolutional layer to deconvolve the audio features to obtain the noise-reduced audio signal.
[0071] Furthermore, considering that the audio signal is represented as a waveform in the time domain, in the technical solution of this application, a convolutional neural network model that performs well in extracting hidden features of images is used as a filter to perform feature mining on the denoised audio signal, so as to extract the high-dimensional hidden feature distribution information in the waveform of the denoised audio signal, thereby obtaining the audio waveform feature vector.
[0072] Then, considering that the main correlation between changes in headphone sound effects and changes in user movement patterns lies in changes in the user's center of gravity during movement, the optimization of audio signals should primarily focus on changes in the user's center of gravity. In order to accurately capture changes in the user's movement patterns and thus accurately optimize the audio signals, in the technical solution of this application, firstly, the user's center of gravity data at multiple predetermined time points within the predetermined time period is obtained.
[0073] Next, since the user's center of gravity data is uncertain in time—that is, due to different changes in the user's movement patterns, the changes in their center of gravity in the time dimension also differ—the user's center of gravity data exhibits different pattern state change characteristics across different time period spans within the predetermined time period. Therefore, in the technical solution of this application, in order to fully and accurately capture the user's center of gravity changes, the user's center of gravity data at multiple predetermined time points is further arranged into a center of gravity input vector according to the time dimension, and then subjected to feature mining through a multi-scale neighborhood feature extraction module to extract the dynamic multi-scale neighborhood association features of the user's center of gravity data across different time spans within the predetermined time period, thereby obtaining a center of gravity temporal feature vector.
[0074] Furthermore, the centroid timing feature vector and the audio waveform feature vector are correlated and encoded to obtain a multimodal correlation feature matrix. This establishes the correlation feature distribution information between the user's multi-scale dynamic change features of the centroid and the temporal hidden features of the audio. Utilizing the correlation between multimodal features helps improve the accuracy of subsequent audio signal optimization. Accordingly, in a specific example of this application, the multimodal correlation feature matrix can be obtained by multiplying the transpose of the audio waveform feature vector by the centroid timing feature vector.
[0075] Then, in order to generate an optimized audio signal based on the multimodal fusion correlation features between the user's multi-scale dynamic change features of the center of gravity and the temporal hidden features of the audio, thereby optimizing the sound effect of the TWS wireless earbuds, the optimized multimodal correlation feature matrix is further processed by a diffusion model-based noise reduction generator to obtain the optimized audio signal. Specifically, in a specific example of this application, the diffusion model-based generator includes a forward diffusion process and a backward generation process. The forward diffusion process gradually adds Gaussian noise to the multimodal correlation feature matrix until it becomes random noise, while the backward generation process is a noise reduction process that gradually removes the random noise until the optimized audio signal is generated. It should be understood that since the overall structure and principle of the diffusion model are not complex, it can obtain a powerful generative model through large-scale training of a rich feature space. Furthermore, the diffusion model is a mapping of each point on a normal distribution to real data, providing better interpretability. In this way, the audio signal can be sufficiently optimized for noise reduction to compensate for frequency shift deviations, thereby ensuring that the sound effect of the TWS wireless earbuds meets the needs of users in sports mode.
[0076] Specifically, in the technical solution of this application, when the centroid temporal feature vector and the audio waveform feature vector are correlated and encoded to obtain the multimodal correlation feature matrix, the centroid temporal feature vector and the audio waveform feature vector are multiplied positionally to obtain the feature values at the corresponding positions of the multimodal correlation feature matrix. However, since the audio waveform feature vector and the centroid temporal feature vector respectively express the image semantics of the audio signal waveform and the temporal multi-scale neighborhood correlation of the centroid data, their feature distributions are not consistent. Therefore, after positional multiplication, the overall feature distribution of the multimodal correlation feature matrix will have a distribution deviation from the individual feature distributions of the centroid temporal feature vector and the audio waveform feature vector. This results in poor dependence of the multimodal correlation feature matrix on the specific distribution corresponding to natural speech, affecting the accuracy of the generated optimized audio signal.
[0077] Therefore, the multimodal correlation feature matrix is first expanded into a multimodal correlation feature vector, for example, denoted as V. Then, the multimodal correlation feature vector V is normalized using Hilbert probability space, specifically expressed as follows:
[0078]
[0079] Here, ||V||2 represents the L2 norm of the multimodal correlation feature vector V. It represents its square, that is, the inner product of the multimodal correlation feature vector V itself, v i It is the i-th feature value of the multimodal correlation feature vector V, and v i ' is the i-th eigenvalue of the optimized multimodal correlation feature vector V'.
[0080] Here, the normed Hilbert probability spaceization of the vector is used to perform a probabilistic interpretation of the multimodal associated feature vector V within a Hilbert space that defines the vector inner product, through the norming of the multimodal associated feature vector V itself. This reduces the hidden perturbation of the specific distribution representation of the multimodal associated feature vector V to the distribution representation of the overall Hilbert space topology, thereby improving the robustness of the feature distribution of the multimodal associated feature vector V to converge to the natural distribution. Simultaneously, the establishment of a metric-induced probability space structure enhances the long-range dependence of the feature distribution of the multimodal associated feature vector V on the natural distribution across the generator. Thus, restoring the multimodal associated feature vector V to the multimodal associated feature matrix improves the accuracy of the optimized audio signal generated by the multimodal associated feature matrix. This allows for accurate optimization of the audio signal of TWS wireless earbuds based on the actual user's motion mode state and environmental noise conditions, ensuring that the sound effects of the TWS wireless earbuds meet the needs of users in motion mode.
[0081] Based on this, this application provides a noise reduction system for TWS wireless earbuds, comprising: a signal receiving module for acquiring audio signals propagating to the left earbud within a predetermined time period; a user motion data monitoring module for acquiring center-of-gravity data of the user at multiple predetermined time points within the predetermined time period; a primary noise reduction module for passing the audio signal through a noise reducer based on an automatic codec to obtain a noise-reduced audio signal; an audio waveform feature extraction module for passing the noise-reduced audio signal through a convolutional neural network model as a filter to obtain an audio waveform feature vector; and a motion pattern feature extraction module for... The centroid data of users at multiple predetermined time points are arranged into a centroid input vector according to the time dimension and then processed by a multi-scale neighborhood feature extraction module to obtain a centroid temporal feature vector; a multimodal association module is used to associate and encode the centroid temporal feature vector and the audio waveform feature vector to obtain a multimodal association feature matrix; a feature optimization module is used to perform high-dimensional data manifold optimization on the multimodal association feature matrix to obtain an optimized multimodal association feature matrix; and a secondary noise reduction module is used to process the optimized multimodal association feature matrix through a noise reduction generator based on a diffusion model to obtain an optimized audio signal.
[0082] After introducing the basic principles of this application, various non-limiting embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0083] Exemplary System
[0084] Figure 1 This is a block diagram of a noise cancellation system for TWS wireless earphones according to an embodiment of this application. Figure 1As shown, a noise reduction system 100 for TWS wireless earphones according to an embodiment of this application includes: a signal receiving module 110 for acquiring audio signals propagating to the left earphone within a predetermined time period; a user motion data monitoring module 120 for acquiring center-of-gravity data of the user at multiple predetermined time points within the predetermined time period; a primary noise reduction module 130 for passing the audio signal through a noise reducer based on an automatic codec to obtain a noise-reduced audio signal; an audio waveform feature extraction module 140 for passing the noise-reduced audio signal through a convolutional neural network model as a filter to obtain an audio waveform feature vector; and a motion pattern feature extraction module 15. 0, used to arrange the centroid data of users at multiple predetermined time points into a centroid input vector according to the time dimension, and then use a multi-scale neighborhood feature extraction module to obtain a centroid time-series feature vector; multimodal association module 160, used to associate and encode the centroid time-series feature vector and the audio waveform feature vector to obtain a multimodal association feature matrix; feature optimization module 170, used to perform high-dimensional data manifold optimization on the multimodal association feature matrix to obtain an optimized multimodal association feature matrix; and a secondary noise reduction module 180, used to pass the optimized multimodal association feature matrix through a noise reduction generator based on a diffusion model to obtain an optimized audio signal.
[0085] Figure 2 This is a schematic diagram of the architecture of a noise cancellation system for TWS wireless earphones according to an embodiment of this application. Figure 2 As shown, firstly, the audio signal propagating to the left earphone within a predetermined time period is acquired, and simultaneously, the centroid data of users at multiple predetermined time points within the predetermined time period is acquired. Next, the audio signal is passed through a noise reduction device based on an automatic codec to obtain a noise-reduced audio signal. Then, the noise-reduced audio signal is passed through a convolutional neural network model acting as a filter to obtain an audio waveform feature vector. Simultaneously, the centroid data of users at multiple predetermined time points are arranged according to the time dimension into a centroid input vector, which is then passed through a multi-scale neighborhood feature extraction module to obtain a centroid temporal feature vector. Subsequently, the centroid temporal feature vector and the audio waveform feature vector are correlated and encoded to obtain a multimodal correlation feature matrix. Then, the multimodal correlation feature matrix is optimized using a high-dimensional data manifold to obtain an optimized multimodal correlation feature matrix. Finally, the optimized multimodal correlation feature matrix is passed through a noise reduction generator based on a diffusion model to obtain an optimized audio signal.
[0086] As mentioned in the background section, in scenarios where users use TWS wireless earbuds during exercise, the surrounding environment is noisy, and the user's activity mode causes frequency shifts and other deviations in sound data propagation. This results in unsatisfactory sound quality for TWS wireless earbud users. Therefore, an optimized noise reduction solution for TWS wireless earbuds is needed.
[0087] Accordingly, considering that users in sports mode experience significant motion displacement during TWS wireless earbuds, leading to frequency shift deviations, optimization and compensation are needed to ensure satisfactory sound quality for active users. Furthermore, recognizing that users' movement states vary during exercise, capturing their movement patterns is crucial for accurate sound data compensation. The primary correlation between sound quality changes and user movement state lies in the shift in the user's center of gravity. In other words, the shift in the user's center of gravity differs depending on their sports mode, resulting in varying frequency shift deviations and requiring different levels of compensation. Therefore, this application aims to comprehensively optimize the audio signal of TWS wireless earbuds based on the implicit feature distribution information of the audio signal and the temporal variation characteristics of the user's center of gravity data. The challenge in this process lies in how to extract the correlation feature distribution information between the implicit features of the audio signal of the TWS wireless earphone and the temporal variation features of the user's center of gravity data, so as to optimize the audio signal of the TWS wireless earphone and make the sound effect of the TWS wireless earphone meet the needs of users in sports mode.
[0088] In recent years, deep learning and neural networks have been widely applied in fields such as computer vision, natural language processing, and text signal processing. Furthermore, deep learning and neural networks have demonstrated near-human or even superior performance in areas such as image classification, object detection, semantic segmentation, and text translation.
[0089] The development of deep learning and neural networks has provided new ideas and solutions for mining the correlation feature distribution information between the implicit features of the audio signals of TWS wireless earbuds and the temporal variation features of the user's center of gravity data. Those skilled in the art will know that deep neural network models based on deep learning can be trained using appropriate strategies, such as backpropagation algorithms with gradient descent, to adjust the parameters of the deep neural network model so that it can simulate complex nonlinear relationships between things. This is clearly suitable for simulating and mining the correlation feature distribution information between the implicit features of the audio signals of TWS wireless earbuds and the temporal variation features of the user's center of gravity data.
[0090] In the noise reduction system 100 for TWS wireless earbuds described above, the signal receiving module 110 is used to acquire the audio signal transmitted to the left earbud within a predetermined time period. Specifically, here, the audio signal from the left earbud is the audio signal transmitted from the mobile phone to the main earbud.
[0091] In the aforementioned noise reduction system 100 for TWS wireless earphones, the user motion data monitoring module 120 is used to acquire the user's center of gravity data at multiple predetermined time points within the predetermined time period. Considering that the primary correlation between changes in earphone sound effects and changes in the user's motion mode lies in the change of the user's center of gravity during movement, audio signal optimization primarily needs to focus on the user's center of gravity changes. To accurately capture changes in the user's motion mode and thus accurately optimize the audio signal, the technical solution of this application first acquires the user's center of gravity data at multiple predetermined time points within the predetermined time period.
[0092] In the aforementioned noise reduction system 100 for TWS wireless earphones, the primary noise reduction module 130 is used to pass the audio signal through an automatic codec-based noise reducer to obtain a noise-reduced audio signal. Considering that when a mobile phone transmits an audio signal to the left earphone, noise such as ambient noise may be generated, thereby reducing the sound quality and effect of the audio signal transmitted to the left earphone, the technical solution of this application, in order to filter out this noise and improve the accuracy of audio signal feature extraction, passes the audio signal through an automatic codec-based noise reducer to obtain a noise-reduced audio signal. Specifically, the automatic codec here includes a sound feature encoder and a sound feature decoder, wherein the sound feature encoder uses a convolutional layer to explicitly spatially encode the audio signal to obtain audio features, and the sound feature decoder uses a deconvolutional layer to deconvolve the audio features to obtain the noise-reduced audio signal.
[0093] More specifically, in this embodiment, the first-level noise reduction module 130 first inputs the audio signal into the sound feature encoder of the noise reducer through the sound signal encoding unit, wherein the sound feature encoder uses a convolutional layer to explicitly spatially encode the audio signal to obtain sound features; then, through the sound feature decoding unit, the sound features are input into the sound feature decoder of the noise reducer, wherein the sound feature decoder uses a deconvolutional layer to deconvolve the sound features to obtain the noise-reduced audio signal.
[0094] In the noise reduction system 100 for TWS wireless earphones described above, the audio waveform feature extraction module 140 is used to pass the denoised audio signal through a convolutional neural network model acting as a filter to obtain an audio waveform feature vector. Considering that the audio signal is represented as a waveform in the time domain, the technical solution of this application uses a convolutional neural network model acting as a filter, which has excellent performance in extracting hidden features from images, to perform feature mining on the denoised audio signal, thereby extracting high-dimensional hidden feature distribution information from the waveform of the denoised audio signal, and thus obtaining the audio waveform feature vector.
[0095] Specifically, in this embodiment, the audio waveform feature extraction module 140 is further configured to: use each layer of the convolutional neural network model to perform the following during the forward propagation of the layers: convolution processing on the input data to obtain a convolutional feature map; performing mean pooling on the convolutional feature map based on the local feature matrix to obtain a pooled feature map; and performing nonlinear activation on the pooled feature map to obtain an activation feature map; wherein the output of the last layer of the convolutional neural network model is the audio waveform feature vector, and the input of the first layer of the convolutional neural network model is the denoised audio signal.
[0096] In the noise reduction system 100 for TWS wireless earphones described above, the motion mode feature extraction module 150 is used to arrange the center of gravity data of the users at multiple predetermined time points into a center of gravity input vector according to the time dimension, and then obtain the center of gravity temporal feature vector through the multi-scale neighborhood feature extraction module.
[0097] Because user center of gravity data is uncertain in time—that is, due to different changes in user movement patterns, the changes in their center of gravity in the time dimension also differ—the user's center of gravity data exhibits different pattern state change characteristics across different time period spans within the predetermined time period. Therefore, in the technical solution of this application, to fully and accurately capture the changes in the user's center of gravity, after acquiring the user's center of gravity data at multiple predetermined time points within the predetermined time period, the user's center of gravity data at these multiple predetermined time points is further arranged into a center of gravity input vector according to the time dimension. This vector is then processed through a multi-scale neighborhood feature extraction module to extract the dynamic multi-scale neighborhood association features of the user's center of gravity data across different time spans within the predetermined time period, thereby obtaining a center of gravity temporal feature vector. In a specific embodiment of this application, the multi-scale neighborhood feature extraction module includes: a first convolutional layer and a second convolutional layer that run in parallel, and a multi-scale fusion layer connected to the first convolutional layer and the second convolutional layer, wherein the first convolutional layer and the second convolutional layer each use one-dimensional convolutional kernels with different scales.
[0098] Figure 3 This is a block diagram of a motion pattern feature extraction module in a noise cancellation system for TWS wireless earphones according to an embodiment of this application. Figure 3 As shown, the motion pattern feature extraction module 150 includes: a first scale encoding unit 151, used to perform one-dimensional convolutional encoding on the centroid input vector using the first convolutional layer of the multi-scale neighborhood feature extraction module to obtain a first-scale centroid feature vector; wherein, the formula is:
[0099]
[0100] Where, a is the width of the first convolutional kernel in the x-direction, F(a) is the parameter vector of the first convolutional kernel, G(xa) is the local vector matrix operated with the convolutional kernel function, w is the size of the first convolutional kernel, X represents the centroid input vector, and Cov1(X) represents one-dimensional convolutional encoding of the centroid input vector; the second scale encoding unit 152 is used to perform one-dimensional convolutional encoding on the centroid input vector using the second convolutional layer of the multi-scale neighborhood feature extraction module to obtain the second scale centroid feature vector; wherein, the formula is:
[0101]
[0102] Where b is the width of the second convolution kernel in the x direction, F(b) is the parameter vector of the second convolution kernel, G(xb) is the local vector matrix operated with the convolution kernel function, m is the size of the second convolution kernel, X represents the centroid input vector, and Cov2(X) represents one-dimensional convolution encoding of the centroid input vector; and the multi-scale feature fusion unit 153 is used to concatenate the first-scale centroid feature vector and the second-scale centroid feature vector using the multi-scale fusion layer of the multi-scale neighborhood feature extraction module to obtain the centroid temporal feature vector.
[0103] In the noise reduction system 100 for TWS wireless earphones described above, the multimodal association module 160 is used to perform association encoding on the centroid timing feature vector and the audio waveform feature vector to obtain a multimodal association feature matrix. That is, by performing association encoding on the centroid timing feature vector and the audio waveform feature vector to obtain the multimodal association feature matrix, the correlation feature distribution information between the user's multi-scale dynamic change features of the centroid and the temporal hidden features of the audio is established. This allows for the utilization of the correlation relationship between multimodal features to improve the accuracy of subsequent audio signal optimization. Accordingly, in a specific example of this application, the multimodal association feature matrix can be obtained by multiplying the transpose of the audio waveform feature vector by the centroid timing feature vector.
[0104] More specifically, in this embodiment, the centroid timing feature vector and the audio waveform feature vector are correlated and encoded using the following formula to obtain a multimodal correlation feature matrix; wherein, the formula is:
[0105]
[0106] in V represents the transpose of the audio waveform feature vector. b Let M represent the centroid temporal feature vector, and let M represent the multimodal correlation feature matrix. This represents matrix multiplication.
[0107] In the noise reduction system 100 for TWS wireless earphones described above, the feature optimization module 170 is used to perform high-dimensional data manifold optimization on the multimodal correlation feature matrix to obtain an optimized multimodal correlation feature matrix. Specifically, in the technical solution of this application, when the centroid temporal feature vector and the audio waveform feature vector are correlated and encoded to obtain the multimodal correlation feature matrix, the centroid temporal feature vector and the audio waveform feature vector are multiplied positionally to obtain the feature values at the corresponding positions of the multimodal correlation feature matrix. However, since the audio waveform feature vector and the centroid temporal feature vector respectively express the image semantics of the audio signal waveform and the temporal multi-scale neighborhood correlation of the centroid data, their feature distributions are not consistent. Therefore, after positional multiplication, the overall feature distribution of the multimodal correlation feature matrix will have a distribution deviation from the individual feature distributions of the centroid temporal feature vector and the audio waveform feature vector, resulting in poor dependence of the multimodal correlation feature matrix on the specific distribution corresponding to natural speech, affecting the accuracy of the generated optimized audio signal.
[0108] Therefore, the multimodal correlation feature matrix is first expanded into a multimodal correlation feature vector, for example, denoted as V. Then, the multimodal correlation feature vector V is normalized using Hilbert probability space, specifically expressed as follows:
[0109]
[0110] Where V is the multimodal association feature vector, and ||V||2 represents the L2 norm of the multimodal association feature vector. v represents the square of the L2 norm of the multimodal correlation feature vector. i It is the i-th eigenvalue of the multimodal correlation feature vector, exp(·) represents the vector exponentiation operation, which means calculating the natural exponent function value raised to the power of each eigenvalue in the vector, and v i ' is the i-th feature value of the optimized multimodal correlation feature vector.
[0111] Here, the normed Hilbert probability spaceization of the vector is used to perform a probabilistic interpretation of the multimodal associated feature vector V within a Hilbert space that defines the vector inner product, through the norming of the multimodal associated feature vector V itself. This reduces the hidden perturbation of the specific distribution representation of the multimodal associated feature vector V to the distribution representation of the overall Hilbert space topology, thereby improving the robustness of the feature distribution of the multimodal associated feature vector V to converge to the natural distribution. Simultaneously, the establishment of a metric-induced probability space structure enhances the long-range dependence of the feature distribution of the multimodal associated feature vector V on the natural distribution across the generator. Thus, restoring the multimodal associated feature vector V to the multimodal associated feature matrix improves the accuracy of the optimized audio signal generated by the multimodal associated feature matrix. This allows for accurate optimization of the audio signal of TWS wireless earbuds based on the actual user's motion mode state and environmental noise conditions, ensuring that the sound effects of the TWS wireless earbuds meet the needs of users in motion mode.
[0112] Figure 4 This is a block diagram of a feature optimization module in a noise reduction system for TWS wireless earphones according to an embodiment of this application. Figure 4 As shown, the feature optimization module 170 includes: a feature matrix expansion unit 171, used to expand the multimodal association feature matrix into multimodal association feature vectors; and a vector optimization unit 172, used to perform vector norming Hilbert probability spacerization on the multimodal association feature vectors according to the following formula to obtain optimized multimodal association feature vectors; wherein, the formula is:
[0113]
[0114] Where V is the multimodal association feature vector, and ||V||2 represents the L2 norm of the multimodal association feature vector. v represents the square of the L2 norm of the multimodal correlation feature vector. i It is the i-th eigenvalue of the multimodal correlation feature vector, exp(·) represents the vector exponentiation operation, which means calculating the natural exponent function value raised to the power of each eigenvalue in the vector, and v i ' is the i-th eigenvalue of the optimized multimodal association feature vector; and, dimension reconstruction unit 173 is used to perform dimension reconstruction on the optimized multimodal association feature vector to obtain the optimized multimodal association feature matrix.
[0115] In the aforementioned noise reduction system 100 for TWS wireless earbuds, the secondary noise reduction module 180 is used to obtain an optimized audio signal by passing the optimized multimodal correlation feature matrix through a diffusion model-based noise reduction generator. To generate an optimized audio signal based on the multimodal fusion correlation features between the user's centroid multi-scale dynamic change features and the temporal hidden features of the audio, thereby optimizing the sound effects of the TWS wireless earbuds, the optimized multimodal correlation feature matrix is further passed through a diffusion model-based noise reduction generator to obtain the optimized audio signal. Specifically, in a specific example of this application, the diffusion model-based generator includes a forward diffusion process and a backward generation process. The forward diffusion process gradually adds Gaussian noise to the multimodal correlation feature matrix until it becomes random noise, while the backward generation process is a noise reduction process that gradually removes the random noise until the optimized audio signal is generated. It should be understood that, since the overall structure and principle of the diffusion model are not complex, it can obtain a model with powerful generative capabilities through large-scale training on a rich feature space. Furthermore, since each point in the diffusion model is a mapping of real data on a normal distribution, it has better interpretability. This allows for sufficient noise reduction optimization of the audio signal to compensate for frequency shift deviations, thereby ensuring that the sound effects of TWS wireless earbuds meet the needs of users in sports mode.
[0116] In summary, a noise reduction system 100 for TWS wireless earbuds based on embodiments of this application is explained. This system utilizes deep learning-based artificial intelligence technology to capture changes in the user's motion patterns from the user's center of gravity data, thereby optimizing and compensating for frequency shift deviations generated during the use of the TWS wireless earbuds. This allows for sufficient noise reduction optimization of the audio signal, ensuring that the sound effects of the TWS wireless earbuds meet the needs of users in motion mode.
[0117] As described above, the noise reduction system 100 for TWS wireless earbuds according to embodiments of this application can be implemented in various terminal devices, such as servers for noise reduction of TWS wireless earbuds. In one example, the noise reduction system 100 for TWS wireless earbuds according to embodiments of this application can be integrated into a terminal device as a software module and / or a hardware module. For example, the noise reduction system 100 for TWS wireless earbuds can be a software module in the operating system of the terminal device, or it can be an application developed for the terminal device; of course, the noise reduction system 100 for TWS wireless earbuds can also be one of many hardware modules of the terminal device.
[0118] Alternatively, in another example, the noise cancellation system 100 for TWS wireless earbuds and the terminal device can also be separate devices, and the noise cancellation system 100 for TWS wireless earbuds can be connected to the terminal device via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.
[0119] Exemplary methods
[0120] Figure 5 This is a flowchart of a noise reduction method for TWS wireless earphones according to an embodiment of this application. Figure 5 As shown, a noise reduction method for TWS wireless earphones according to an embodiment of this application includes: S110, acquiring an audio signal propagating to the left earphone within a predetermined time period; S120, acquiring centroid data of users at multiple predetermined time points within the predetermined time period; S130, passing the audio signal through a noise reduction device based on an automatic codec to obtain a noise-reduced audio signal; S140, passing the noise-reduced audio signal through a convolutional neural network model as a filter to obtain an audio waveform feature vector; S150, arranging the centroid data of users at the multiple predetermined time points according to the time dimension into a centroid input vector, and then passing it through a multi-scale neighborhood feature extraction module to obtain a centroid temporal feature vector; S160, performing correlation encoding on the centroid temporal feature vector and the audio waveform feature vector to obtain a multimodal correlation feature matrix; S170, performing high-dimensional data manifold optimization on the multimodal correlation feature matrix to obtain an optimized multimodal correlation feature matrix; and S180, passing the optimized multimodal correlation feature matrix through a noise reduction generator based on a diffusion model to obtain an optimized audio signal.
[0121] In one example, in the noise reduction method for TWS wireless earphones described above, the automatic codec includes a sound feature encoder and a sound feature decoder.
[0122] In one example, in the above-described noise reduction method for TWS wireless earphones, the step of passing the audio signal through an automatic codec-based noise reducer to obtain a noise-reduced audio signal includes: inputting the audio signal into the sound feature encoder of the noise reducer, wherein the sound feature encoder uses a convolutional layer to explicitly spatially encode the audio signal to obtain sound features; and inputting the sound features into the sound feature decoder of the noise reducer, wherein the sound feature decoder uses a deconvolutional layer to deconvolve the sound features to obtain the noise-reduced audio signal.
[0123] In one example, in the above-described noise reduction method for TWS wireless earphones, the step of passing the noise-reduced audio signal through a convolutional neural network model as a filter to obtain an audio waveform feature vector includes: using each layer of the convolutional neural network model to perform the following in the forward pass of the layer: convolution processing on the input data to obtain a convolutional feature map; performing mean pooling on the convolutional feature map based on the local feature matrix to obtain a pooled feature map; and performing nonlinear activation on the pooled feature map to obtain an activation feature map; wherein the output of the last layer of the convolutional neural network model is the audio waveform feature vector, and the input of the first layer of the convolutional neural network model is the noise-reduced audio signal.
[0124] In one example, in the noise reduction method for TWS wireless earphones described above, the multi-scale neighborhood feature extraction module includes: a first convolutional layer and a second convolutional layer that run in parallel with each other, and a multi-scale fusion layer connected to the first convolutional layer and the second convolutional layer, wherein the first convolutional layer and the second convolutional layer each use one-dimensional convolutional kernels with different scales.
[0125] In one example, in the noise reduction method for TWS wireless earphones described above, the step of arranging the centroid data of the users at multiple predetermined time points into a centroid input vector according to the time dimension and then obtaining a centroid temporal feature vector through a multi-scale neighborhood feature extraction module includes: using the first convolutional layer of the multi-scale neighborhood feature extraction module to perform one-dimensional convolutional encoding on the centroid input vector using the following formula to obtain a first-scale centroid feature vector; wherein, the formula is:
[0126]
[0127] Where, a is the width of the first convolutional kernel in the x-direction, F(a) is the parameter vector of the first convolutional kernel, G(xa) is the local vector matrix operated with the convolutional kernel function, w is the size of the first convolutional kernel, X represents the centroid input vector, and Cov1(X) represents the one-dimensional convolutional encoding of the centroid input vector; the second convolutional layer of the multi-scale neighborhood feature extraction module performs one-dimensional convolutional encoding on the centroid input vector using the following formula to obtain the second-scale centroid feature vector; wherein, the formula is:
[0128]
[0129] Where b is the width of the second convolution kernel in the x direction, F(b) is the parameter vector of the second convolution kernel, G(xb) is the local vector matrix operated with the convolution kernel function, m is the size of the second convolution kernel, X represents the centroid input vector, and Cov2(X) represents one-dimensional convolution encoding of the centroid input vector; and the multi-scale fusion layer of the multi-scale neighborhood feature extraction module concatenates the first-scale centroid feature vector and the second-scale centroid feature vector to obtain the centroid temporal feature vector.
[0130] In one example, in the noise reduction method for TWS wireless earphones described above, the step of associating and encoding the centroid timing feature vector and the audio waveform feature vector to obtain a multimodal association feature matrix includes: associating and encoding the centroid timing feature vector and the audio waveform feature vector using the following formula to obtain a multimodal association feature matrix; wherein, the formula is:
[0131]
[0132] in V represents the transpose of the audio waveform feature vector. b Let M represent the centroid temporal feature vector, and let M represent the multimodal correlation feature matrix. This represents matrix multiplication.
[0133] In one example, in the noise reduction method for TWS wireless earphones described above, the step of performing high-dimensional data manifold optimization on the multimodal correlation feature matrix to obtain an optimized multimodal correlation feature matrix includes: expanding the multimodal correlation feature matrix into multimodal correlation feature vectors; and performing vector normed Hilbert probability spacerization on the multimodal correlation feature vectors using the following formula to obtain the optimized multimodal correlation feature vectors; wherein, the formula is:
[0134]
[0135] Where V is the multimodal association feature vector, and ||V||2 represents the L2 norm of the multimodal association feature vector. v represents the square of the L2 norm of the multimodal correlation feature vector. i It is the i-th eigenvalue of the multimodal correlation feature vector, exp(·) represents the vector exponentiation operation, which means calculating the natural exponent function value raised to the power of each eigenvalue in the vector, and v i ' is the i-th eigenvalue of the optimized multimodal association feature vector; and the optimized multimodal association feature vector is reconstructed in dimensions to obtain the optimized multimodal association feature matrix.
[0136] In summary, the noise reduction method for TWS wireless earbuds according to the embodiments of this application is explained. It utilizes deep learning-based artificial intelligence technology to capture changes in the user's motion patterns from the user's center of gravity data, thereby optimizing and compensating for frequency shift deviations generated during the use of the TWS wireless earbuds. This allows for sufficient noise reduction optimization of the audio signal, ensuring that the sound effects of the TWS wireless earbuds meet the needs of users in motion mode.
[0137] Exemplary electronic devices
[0138] Below, for reference Figure 6 This describes an electronic device according to embodiments of the present application. Figure 6 This is a block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device 10 includes one or more processors 11 and memory 12.
[0139] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0140] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute the program instructions to implement the functions in the noise reduction methods for TWS wireless earphones described in the various embodiments of this application above, and / or other desired functions. Various content, such as the audio signal of the left earphone, the user's center of gravity data, etc., may also be stored in the computer-readable storage medium.
[0141] In one example, the electronic device 10 may also include an input device 13 and an output device 14, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0142] The input device 13 may include, for example, a keyboard, a mouse, etc.
[0143] The output device 14 can output various information to the outside, including optimized audio signals. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0144] Of course, for the sake of simplicity, Figure 6 Only some of the components of the electronic device 10 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 10 may include any other suitable components depending on the specific application.
[0145] Exemplary computer program products and computer-readable storage media
[0146] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the functions in the noise reduction methods for TWS wireless earphones according to various embodiments of this application described in the "Exemplary Methods" section of this specification.
[0147] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0148] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the functions in the noise reduction methods for TWS wireless earphones according to various embodiments of this application described in the "Exemplary Methods" section above.
[0149] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0150] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.
[0151] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0152] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.
[0153] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0154] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A noise reduction system for TWS wireless earphones, characterized in that, include: The signal receiving module is used to acquire the audio signal propagating to the left earphone within a predetermined time period; The user motion data monitoring module is used to acquire the center of gravity data of users at multiple predetermined time points within the predetermined time period. A primary noise reduction module is used to pass the audio signal through an automatic codec-based noise reducer to obtain a noise-reduced audio signal; An audio waveform feature extraction module is used to pass the noise-reduced audio signal through a convolutional neural network model as a filter to obtain an audio waveform feature vector. The motion pattern feature extraction module is used to arrange the center of gravity data of the users at multiple predetermined time points into a center of gravity input vector according to the time dimension, and then pass it through the multi-scale neighborhood feature extraction module to obtain the center of gravity temporal feature vector. A multimodal association module is used to perform association encoding on the centroid time-series feature vector and the audio waveform feature vector to obtain a multimodal association feature matrix; The feature optimization module is used to perform high-dimensional data manifold optimization on the multimodal association feature matrix to obtain an optimized multimodal association feature matrix; The secondary noise reduction module is used to pass the optimized multimodal correlation feature matrix through a noise reduction generator based on a diffusion model to obtain an optimized audio signal; The feature optimization module includes: The feature matrix expansion unit is used to expand the multimodal correlation feature matrix into multimodal correlation feature vectors; The vector optimization unit is used to perform vector norming and Hilbert probability spacerization on the multimodal association feature vectors according to the following formula to obtain the optimized multimodal association feature vectors. The formula is as follows: in It is the multimodal correlation feature vector, The L2 norm of the multimodal correlation feature vector is represented by the following: The square of the L2 norm of the multimodal correlation feature vector is represented by the following: It is the first of the multimodal correlation feature vectors 1 eigenvalue, The vector exponentiation operation represents calculating the value of a natural exponential function raised to the power of the eigenvalues at each position in the vector. It is the first of the optimized multimodal correlation feature vectors One eigenvalue; The dimension reconstruction unit is used to reconstruct the dimensions of the optimized multimodal association feature vector to obtain the optimized multimodal association feature matrix.
2. The noise reduction system for TWS wireless earphones according to claim 1, characterized in that, The automatic codec includes a sound feature encoder and a sound feature decoder.
3. The noise reduction system for TWS wireless earphones according to claim 2, characterized in that, The primary noise reduction module includes: A sound signal encoding unit is configured to input the audio signal into the sound feature encoder of the noise reduction unit, wherein the sound feature encoder uses a convolutional layer to explicitly spatially encode the audio signal to obtain sound features; and A sound feature decoding unit is used to input the sound features into the sound feature decoder of the noise reducer, wherein the sound feature decoder uses a deconvolution layer to perform deconvolution processing on the sound features to obtain the noise-reduced audio signal.
4. The noise reduction system for TWS wireless earphones according to claim 3, characterized in that, The audio waveform feature extraction module is further used for: Each layer of the convolutional neural network model is used in the forward propagation of the layer as follows: The input data is convolved to obtain a convolutional feature map; The convolutional feature map is subjected to mean pooling based on the local feature matrix to obtain a pooled feature map; as well as The pooled feature map is nonlinearly activated to obtain an activated feature map; The output of the last layer of the convolutional neural network model is the audio waveform feature vector, and the input of the first layer of the convolutional neural network model is the denoised audio signal.
5. The noise reduction system for TWS wireless earphones according to claim 4, characterized in that, The multi-scale neighborhood feature extraction module includes: a first convolutional layer and a second convolutional layer that run in parallel with each other, and a multi-scale fusion layer connected to the first convolutional layer and the second convolutional layer, wherein the first convolutional layer and the second convolutional layer use one-dimensional convolutional kernels with different scales.
6. The noise reduction system for TWS wireless earphones according to claim 5, characterized in that, The motion pattern feature extraction module includes: The first scale encoding unit is used to perform one-dimensional convolutional encoding on the centroid input vector using the first convolutional layer of the multi-scale neighborhood feature extraction module to obtain the first scale centroid feature vector; The formula is as follows: Where a is the width of the first convolution kernel in the x-direction, For the first convolution kernel parameter vector, Let w be the local vector matrix that operates with the convolution kernel function, w be the size of the first convolution kernel, and X be the centroid input vector. This indicates that a one-dimensional convolutional encoding is performed on the centroid input vector; The second scale encoding unit is used to perform one-dimensional convolutional encoding on the centroid input vector using the second convolutional layer of the multi-scale neighborhood feature extraction module to obtain the second scale centroid feature vector; The formula is as follows: Where b is the width of the second convolution kernel in the x-direction, For the second convolution kernel parameter vector, Let m be the local vector matrix that operates with the convolution kernel function, m be the size of the second convolution kernel, and X be the centroid input vector. This indicates that a one-dimensional convolutional encoding is performed on the centroid input vector; and The multi-scale feature fusion unit is used to concatenate the first-scale centroid feature vector and the second-scale centroid feature vector using the multi-scale fusion layer of the multi-scale neighborhood feature extraction module to obtain the centroid temporal feature vector.
7. The noise reduction system for TWS wireless earphones according to claim 6, characterized in that, The multimodal association module is further used for: The centroid timing feature vector and the audio waveform feature vector are correlated and encoded using the following formula to obtain a multimodal correlation feature matrix; The formula is as follows: = in This represents the transpose of the feature vector of the audio waveform. This represents the centroid time-series feature vector. This represents the multimodal correlation feature matrix. This represents matrix multiplication.
8. A noise reduction method for TWS wireless earphones, characterized in that, include: Acquire the audio signal transmitted to the left earphone within a predetermined time period; Obtain the center of gravity data of users at multiple predetermined time points within the predetermined time period; The audio signal is passed through an automatic codec-based noise reduction device to obtain a noise-reduced audio signal; The denoised audio signal is passed through a convolutional neural network model as a filter to obtain an audio waveform feature vector; After arranging the centroid data of the users at the multiple predetermined time points into a centroid input vector according to the time dimension, the centroid temporal feature vector is obtained by the multi-scale neighborhood feature extraction module. The centroid timing feature vector and the audio waveform feature vector are correlated and encoded to obtain a multimodal correlation feature matrix; The multimodal association feature matrix is optimized by performing high-dimensional data manifold optimization to obtain an optimized multimodal association feature matrix; The optimized multimodal correlation feature matrix is passed through a diffusion-based noise reduction generator to obtain the optimized audio signal; The step of performing high-dimensional data manifold optimization on the multimodal association feature matrix to obtain an optimized multimodal association feature matrix includes: expanding the multimodal association feature matrix into multimodal association feature vectors; and performing vector norming Hilbert probability spaceization on the multimodal association feature vectors using the following formula to obtain the optimized multimodal association feature vectors; wherein the formula is: in It is the multimodal correlation feature vector, The L2 norm of the multimodal correlation feature vector is represented by the following: The square of the L2 norm of the multimodal correlation feature vector is represented by the following: It is the first of the multimodal correlation feature vectors 1 eigenvalue, The vector exponentiation operation represents calculating the value of a natural exponential function raised to the power of the eigenvalues at each position in the vector. It is the first of the optimized multimodal correlation feature vectors The optimized multimodal association feature vector is reconstructed in terms of dimensions to obtain the optimized multimodal association feature matrix.
9. The noise reduction method for TWS wireless earphones according to claim 8, characterized in that, The automatic codec includes a sound feature encoder and a sound feature decoder.
Citation Information
Patent Citations
Noise reduction method and system for high-performance TWS Bluetooth audio chip and electronic equipment
CN113851142A
Intelligent control system and method for offshore wind turbine blade hoisting
CN115481677A