System and method for online adaptive training of large end-to-end audio noise reduction model

By using an end-to-end online adaptive training system for large audio noise reduction models, the neural network is automatically trained using hardware resource status and user feedback. This solves the problem of complex and time-consuming neural network model construction and achieves high-efficiency audio noise reduction model adaptability and design efficiency.

CN121963771APending Publication Date: 2026-05-01HANGZHOU BROADLINK ELECTRONICS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU BROADLINK ELECTRONICS TECH
Filing Date
2025-12-12
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The existing technology for building neural network models based on audio noise reduction algorithms is complex and time-consuming, resulting in low efficiency in the design of audio noise reduction products.

Method used

This paper provides an end-to-end audio noise reduction large-scale online adaptive training system. The system constructs a neural network based on the hardware resource status of the application product through a model training device, generates training audio data using recordings and test corpora in the business environment, and automatically trains the neural network based on user feedback during use until convergence.

Benefits of technology

It improves the training efficiency of audio noise reduction models, lowers the training threshold, quickly adapts to different speech environments, and builds targeted noise reduction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963771A_ABST
    Figure CN121963771A_ABST
Patent Text Reader

Abstract

The invention discloses a system and a method for online adaptive training of an end-to-end audio noise reduction large model. According to the system, a model training device is used for constructing a corresponding neural network according to a hardware resource state of an application product; generating training audio data by obtaining the business environment recording and the test corpus; training a neural network according to the training audio data until the neural network converges; an application product for applying the converged neural network to the application product; and in the use process, recording collected during use and a user feedback result are used as sample data to be fed back to the model training device, so that the model training device automatically trains the neural network according to the sample data. The method has the technical effects that the training efficiency of the audio noise reduction model can be improved, the training threshold of the noise reduction model can be reduced, different voice environments can be quickly adapted, and the targeted noise reduction model can be constructed.
Need to check novelty before this filing date? Find Prior Art

Description

System and method for online adaptive training of large end-to-end audio denoising models Technical Field

[0001] This invention relates to the field of computer technology applications, and in particular to a system and method for online adaptive training of a large end-to-end audio denoising model. Background Technology

[0002] Current audio noise reduction algorithms require collecting extensive ambient data, constructing neural network models, designing related algorithms, and conducting supervised training. This entire process necessitates continuous experimentation and parameter adjustments, making it extremely time-consuming.

[0003] There is currently no effective solution to the problem that the process of building neural network models based on audio noise reduction algorithms in existing technologies requires long and complex experimental and debugging steps, resulting in low design efficiency of audio noise reduction products. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention aims to provide a system and method for end-to-end online adaptive training of a large audio denoising model, thereby at least resolving the issue of low design efficiency in audio denoising products due to the lengthy and complex experimental and debugging processes required in the existing neural network model construction based on audio denoising algorithms.

[0005] The technical solution of this invention is implemented as follows: This invention provides an end-to-end audio noise reduction large-scale online adaptive training system, comprising: a model training device and an application product. The model training device is used to construct a corresponding neural network based on the hardware resource status of the application product; generate training audio data by acquiring recordings from the business environment and test corpora; train the neural network based on the training audio data until the neural network converges; the application product is used to apply the converged neural network to the application product; during use, the recordings collected during use and user feedback results are fed back to the model training device as sample data, so that the model training device automatically trains the neural network based on the sample data.

[0006] Optionally, the model training device is also used to construct a neural network with a corresponding number of layers based on the hardware resource status of the application product; to perform speech synthesis by acquiring recordings of the business environment and test corpora to generate training audio data; to train the neural network based on the training audio data, to monitor the error of the neural network output, and to determine that the neural network training has converged when the error reaches a preset standard, thereby obtaining a converged neural network; wherein, the training audio data includes: noisy audio and noise-free audio; the noisy audio is obtained from recordings of the business environment.

[0007] This invention provides a method for end-to-end audio denoising large model online adaptive training, applied to a system for end-to-end audio denoising large model online adaptive training, comprising: constructing a corresponding neural network based on the hardware resource status of the application product; generating training audio data by acquiring recordings from the business environment and test corpora; training the neural network based on the training audio data to obtain a converged neural network; receiving sample data during the use of the converged neural network; and automatically training the neural network based on the sample data.

[0008] Optionally, constructing a corresponding neural network based on the hardware resource status of the application product includes: constructing a neural network with a corresponding number of layers based on the hardware resource status of the application product.

[0009] Optionally, generating training audio data by acquiring business environment recordings and test corpora includes: performing speech synthesis by acquiring business environment recordings and test corpora to generate training audio data; wherein, the training audio data includes: noisy audio and noise-free audio; the noisy audio is obtained from business environment recordings.

[0010] Further, optionally, training the neural network based on the training audio data to obtain a converged neural network includes: training the neural network based on the training audio data, monitoring the error of the neural network output, and determining that the neural network training has converged when the error reaches a preset standard, thereby obtaining a converged neural network.

[0011] Optionally, if the sample data includes: collected recordings and user feedback results, the neural network can be automatically trained based on the recordings and user feedback results.

[0012] This invention provides a system and method for end-to-end online adaptive training of a large audio denoising model. A model training device is used to construct a corresponding neural network based on the hardware resource status of the application product; training audio data is generated by acquiring recordings from the business environment and test corpora; the neural network is trained based on the training audio data until it converges; the application product applies the converged neural network to the application product; during use, the recordings collected during use and user feedback results are fed back to the model training device as sample data, enabling the model training device to automatically train the neural network based on the sample data. This improves the training efficiency of the audio denoising model, lowers the training threshold for the denoising model, and quickly adapts to different speech environments, achieving the technical effect of constructing a targeted denoising model. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their descriptions, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 is a schematic diagram of an end-to-end audio denoising large model online adaptive training system provided by an embodiment of the present invention; Figure 2 is a schematic diagram of another end-to-end audio denoising large model online adaptive training system provided by an embodiment of the present invention; Figure 3 is a flowchart illustrating an end-to-end audio denoising large model online adaptive training method provided by an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this invention are used to distinguish different objects, rather than to limit a specific order.

[0016] It should also be noted that the various embodiments of the present invention described below can be executed individually or in combination with each other, and the embodiments of the present invention do not impose specific limitations in this regard.

[0017] This invention provides a system for end-to-end audio denoising large-scale online adaptive training. Figure 1 is a schematic diagram of such a system. As shown in Figure 1, the system includes a model training device 12 and an application product 14. The model training device 12 is used to construct a corresponding neural network based on the hardware resource status of the application product 14; generate training audio data by acquiring recordings from the business environment and test corpora; train the neural network based on the training audio data until the neural network converges; and apply the converged neural network to the application product 14. During use, the system feeds back the recordings and user feedback results collected during use as sample data to the model training device 12, so that the model training device 12 can automatically train the neural network based on the sample data.

[0018] Optionally, the model training device 12 is also used to construct a neural network with a corresponding number of layers based on the hardware resource status of the application product 14; to perform speech synthesis by acquiring recordings of the business environment and test corpora to generate training audio data; to train the neural network based on the training audio data, to monitor the error of the neural network output, and to determine that the neural network training has converged when the error reaches a preset standard, thereby obtaining a converged neural network; wherein, the training audio data includes: noisy audio and noiseless audio; the noisy audio is obtained from recordings of the business environment.

[0019] In this embodiment, the hardware resources can be the computing power of the application product; based on the hardware resource status of the application product 14, the neural network with the corresponding number of layers can be constructed by configuring the corresponding number of neural network layers according to the different computing power of the embedded system and the PC system.

[0020] In this application embodiment, the preset standard is adjusted according to different application products, mainly referring to the difference from the original sound.

[0021] Specifically, the end-to-end audio denoising large model online adaptive training system provided in this application aims to improve the training efficiency of audio denoising models and reduce the training threshold for denoising models, quickly adapt to different speech environments, and build targeted denoising models. Figure 2 is a schematic diagram of another end-to-end audio denoising large model online adaptive training system provided in this embodiment of the invention. As shown in Figure 2, the specific steps are as follows: Step 1, construct neural networks of different layers according to hardware resources; based on the recordings of the business environment and the test corpus, generate training audio data through speech synthesis technology, which is divided into clean audio and noisy audio. The noise model is extracted through recording.

[0022] Step 2: Use an open-source training framework to train the neural network, monitor the error in the model output, stop training once the standard is reached, and release the model version.

[0023] Step 3: After the model runs on-site for a period of time, the training data is reconstructed based on the collected recordings and user feedback, and published to the training system to automatically retrain the model and output a new version.

[0024] Through the above Steps 1 to 3, adaptive training and version release of the model version are achieved.

[0025] The end-to-end audio noise reduction large model online adaptive training system provided in this application avoids the need for manual algorithm design, data preparation, and training execution.

[0026] This invention provides an end-to-end online adaptive training system for a large audio denoising model. The system utilizes a model training device to construct a corresponding neural network based on the hardware resource status of the application product. It generates training audio data by acquiring recordings from the business environment and test corpora. The neural network is trained using this training audio data until it converges. The application product then applies the converged neural network to the application product. During use, the system feeds back the collected recordings and user feedback as sample data to the model training device, enabling the device to automatically train the neural network based on the sample data. This improves the training efficiency of the audio denoising model, lowers the training threshold, and allows for rapid adaptation to different speech environments, achieving the technical effect of building targeted denoising models.

[0027] This invention provides a method for end-to-end audio denoising large model online adaptive training. Figure 3 is a flowchart illustrating the method for end-to-end audio denoising large model online adaptive training provided by this invention. As shown in Figure 3, the system applied to end-to-end audio denoising large model online adaptive training, the method for end-to-end audio denoising large model online adaptive training provided by this application includes: step S302, constructing a corresponding neural network according to the hardware resource status of the application product; optionally, step S302, constructing a corresponding neural network according to the hardware resource status of the application product includes: constructing a neural network with a corresponding number of layers according to the hardware resource status of the application product.

[0028] Step S304: Generate training audio data by acquiring business environment recordings and test corpora; Optionally, generating training audio data by acquiring business environment recordings and test corpora in step S304 includes: performing speech synthesis by acquiring business environment recordings and test corpora to generate training audio data; wherein, the training audio data includes: noisy audio and noise-free audio; the noisy audio is obtained from business environment recordings.

[0029] Step S306: Train the neural network based on the training audio data to obtain a converged neural network; Optionally, step S306, training the neural network based on the training audio data to obtain a converged neural network, includes: training the neural network based on the training audio data, monitoring the error of the neural network output, and when the error reaches a preset standard, determining that the neural network training has converged to obtain a converged neural network.

[0030] Step S308: Receive sample data during the process of using the convergent neural network; Step S310: Automatically train the neural network based on the sample data.

[0031] Optionally, in step S310, if the sample data includes: the collected recordings and user feedback results, the neural network is automatically trained based on the recordings and user feedback results.

[0032] This invention provides an end-to-end online adaptive training method for a large audio denoising model. The method involves constructing a corresponding neural network based on the hardware resource status of the application product; generating training audio data by acquiring recordings from the business environment and test corpora; training the neural network using the training audio data to obtain a converged neural network; receiving sample data during the use of the converged neural network; and automatically training the neural network based on the sample data. This improves the training efficiency of the audio denoising model, lowers the training threshold, and enables rapid adaptation to different speech environments, thus achieving the technical effect of constructing a targeted denoising model.

[0033] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention.

Claims

1. A system for end-to-end audio noise reduction large model online adaptive training, characterized in that, include: A model training device and an application product are provided. The model training device is used to construct a corresponding neural network based on the hardware resource status of the application product; generate training audio data by acquiring recordings of the business environment and test corpora; train the neural network based on the training audio data until the neural network converges; the application product is used to apply the converged neural network to the application product; during use, the recordings collected during use and user feedback results are fed back to the model training device as sample data, so that the model training device can automatically train the neural network based on the sample data.

2. The system for end-to-end audio denoising large model online adaptive training according to claim 1, characterized in that, The model training device is further configured to construct a neural network of corresponding layers based on the hardware resource status of the application product; perform speech synthesis by acquiring recordings of the business environment and test corpora to generate the training audio data; train the neural network based on the training audio data, monitor the error of the neural network output, and determine that the neural network training has converged when the error reaches a preset standard, thereby obtaining the converged neural network; wherein, the training audio data includes: noisy audio and noise-free audio; the noisy audio is obtained from the recordings of the business environment.

3. A method for online adaptive training of a large end-to-end audio noise reduction model, characterized in that, A system for end-to-end audio noise reduction large-scale online adaptive training includes: constructing a corresponding neural network based on the hardware resource status of the application product; generating training audio data by acquiring recordings from the business environment and test corpora; training the neural network based on the training audio data to obtain a converged neural network; receiving sample data during the use of the converged neural network; and automatically training the neural network based on the sample data.

4. The method for end-to-end audio denoising large model online adaptive training according to claim 3, characterized in that, The step of constructing a corresponding neural network based on the hardware resource status of the application product includes: constructing a neural network with a corresponding number of layers based on the hardware resource status of the application product.

5. The method for end-to-end audio denoising large model online adaptive training according to claim 3 or 4, characterized in that, The step of generating training audio data by acquiring business environment recordings and test corpora includes: performing speech synthesis by acquiring business environment recordings and test corpora to generate the training audio data; wherein, the training audio data includes: noisy audio and noise-free audio; the noisy audio is obtained from the business environment recordings.

6. The method for end-to-end audio denoising large model online adaptive training according to claim 5, characterized in that, The step of training the neural network based on the training audio data to obtain a converged neural network includes: training the neural network based on the training audio data, monitoring the error of the neural network output, and determining that the neural network training has converged when the error reaches a preset standard, thereby obtaining a converged neural network.

7. The method for end-to-end audio denoising large model online adaptive training according to claim 6, characterized in that, When the sample data includes: collected recordings and user feedback results, the neural network is automatically trained based on the recordings and user feedback results.