Audio noise reduction method and device, electronic equipment and storage medium
By using multimodal audio large models to identify and classify noise types and automatically select noise reduction strategies, the problem of poor noise reduction effect in complex noise environments in the existing technology is solved, and a more efficient and robust audio noise reduction effect is achieved.
Patent Information
- Application Number
- CN202510130052.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-05
AI Technical Summary
The existing audio noise reduction algorithm has limited effect when processing complex and non-stationary noise, and the subjectivity of manual noise reduction algorithm is high, which may lead to poor noise reduction effect.
By obtaining multimodal data, including audio data and associated data, the trained multimodal audio model recognizes and classifies noise types, and automatically selects the optimal noise reduction strategy based on preset matching rules.
It improves the pertinence and efficiency of the noise reduction effect, reduces manual intervention, and enhances the robustness of the system in complex environments.
Smart Images

Figure CN119964540A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of neural network technology, and in particular to an audio noise reduction method, device, electronic device and storage medium. Background Art
[0002] Audio noise reduction is a commonly used technology in teleconferences, video calls or listening to music. Audio noise reduction can be used to remove or reduce background noise, making the target sound (such as human voice, music) clearer and more identifiable, improving the user's listening experience. Current audio noise reduction algorithms mainly include traditional signal processing algorithms and algorithms based on machine learning and deep learning. Traditional signal processing algorithms work well when dealing with simple noise environments, but have limited processing capabilities for complex and non-stationary noises; while algorithms based on machine learning and deep learning perform well when dealing with complex noise environments, but have high computational complexity and high demand for computing resources. Currently, it is necessary to manually determine which noise reduction algorithm to use. Manually determining the noise reduction algorithm is highly subjective and may result in poor noise reduction effects due to inappropriate algorithm selection. Summary of the invention
[0003] The present application provides an audio noise reduction method, device, electronic device and storage medium to solve the problem of poor noise reduction effect.
[0004] In a first aspect, the present application provides an audio noise reduction method, the method comprising:
[0005] Acquire multimodal data, wherein the multimodal data includes audio data and associated data of the audio data, wherein the associated data is used to assist in understanding the content of the audio data;
[0006] Processing the multimodal data through the trained multimodal audio large model, and outputting the noise type in the audio data;
[0007] Determine a noise reduction strategy corresponding to the noise type according to a preset matching rule;
[0008] The audio data is denoised using the denoising strategy.
[0009] Optionally, the multimodal data is processed by a trained multimodal audio large model, and the noise type in the output audio data includes:
[0010] Extracting multiple modal features of the multimodal data through the trained multimodal audio large model, wherein the multiple modal features include audio features and associated features of the associated data;
[0011] An alignment network is used to integrate the multiple modal features into the same vector space;
[0012] In the vector space, an attention mechanism is used to capture the interaction relationship between the audio feature and each of the associated features to obtain a fused multimodal embedding vector, wherein the multimodal embedding vector fuses the audio feature and all the associated features;
[0013] A fully connected layer is used to extract global semantic features of the multimodal embedding vector, and the noise type is determined according to the global semantic features, wherein the global semantic features can reflect the overall situation of the audio scene.
[0014] Optionally, an attention mechanism is used to capture the interactive relationship between the audio feature and each of the associated features, and the fused multimodal embedding vector is obtained, including:
[0015] Generate an interaction score based on the interaction relationship between the audio feature and the associated feature at each moment;
[0016] Performing a weighted summation of the interaction scores of different associated features at the same moment to obtain a context vector, wherein the context vector is used to indicate the audio association information at each moment;
[0017] Combining the audio feature with the context vector to generate a fused feature;
[0018] The multimodal embedding vector is obtained by integrating the fusion features at all moments.
[0019] Optionally, extracting a global semantic feature of the multimodal embedding vector using a fully connected layer, and determining the noise type according to the global semantic feature includes:
[0020] Using a fully connected layer to perform dimensionality reduction processing on the multimodal embedding vector to obtain concentrated features;
[0021] Extracting multiple semantic features from the concentrated features layer by layer through the fully connected layer to generate global semantic features;
[0022] The global semantic feature is matched with a pre-built knowledge base to determine the noise type, wherein the knowledge base contains various noise types and their corresponding semantic feature templates.
[0023] Optionally, the associated data includes text and video, and the multiple modal features of the multimodal data extracted by the trained multimodal audio large model include:
[0024] Extracting time domain features and frequency domain features of the audio data through the trained multimodal audio large model;
[0025] The text features of the text and the visual features of the video are extracted through the trained multimodal audio model.
[0026] Optionally, the noise type includes at least one type of noise, and determining the noise reduction strategy corresponding to the noise type according to a preset matching rule includes:
[0027] If the noise type is simple noise, determining a noise reduction algorithm in a signal processing algorithm library according to the matching rule;
[0028] If the noise type is complex noise, determining a noise reduction algorithm in a deep learning algorithm library according to the matching rule;
[0029] If the noise type includes simple noise and complex noise, a combined noise reduction algorithm in the signal processing algorithm library and the deep learning algorithm library is adopted according to the matching rule.
[0030] Optionally, the training process of the multimodal audio large model includes:
[0031] Acquire a training sample, wherein the training sample includes a sample audio, a noise label corresponding to the sample audio, and sample-related data, wherein the sample-related data is used to assist in understanding the content of the sample audio;
[0032] Using the training samples to train the initial multimodal audio large model to obtain an output noise classification result;
[0033] If the noise classification result is inconsistent with the noise label, the parameters of the initial multimodal audio model are adjusted until the noise classification result is consistent with the noise label, thereby obtaining a trained multimodal audio model.
[0034] In a second aspect, the present application provides an audio noise reduction device, the device comprising:
[0035] An acquisition module, configured to acquire multimodal data, wherein the multimodal data includes audio data and associated data of the audio data, wherein the associated data is used to assist in understanding the content of the audio data;
[0036] A processing module, used for processing the multimodal data through a trained multimodal audio model, and outputting the noise type in the audio data;
[0037] A determination module, used to determine a noise reduction strategy corresponding to the noise type according to a preset matching rule;
[0038] The noise reduction module is used to perform noise reduction on the audio data by adopting the noise reduction strategy.
[0039] In a third aspect, the present application provides an electronic device comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus.
[0040] In a fourth aspect, the present application further provides a computer storage medium storing computer executable instructions, wherein the computer executable instructions are used to execute the audio noise reduction method described in any one of the above items of the present application.
[0041] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art: the associated data can help to better understand the audio content, so that the multimodal audio large model can combine the associated data and the audio data, accurately identify and classify the noise type in the audio data, and automatically select the optimal noise reduction strategy according to the matching rules, thereby improving the noise reduction effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0044] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0045] Figure 1 A schematic diagram of the hardware structure of an audio noise reduction system provided in an embodiment of the present application;
[0046] Figure 2 A flow chart of a method for audio noise reduction provided in an embodiment of the present application;
[0047] Figure 3 A schematic diagram of the structure of an audio noise reduction device provided in an embodiment of the present application;
[0048] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0050] The disclosure below provides many different embodiments or examples to realize the different structures of the present application. In order to simplify the disclosure of the present application, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present application. In addition, the present application can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0051] In order to solve the problem that the audio noise reduction algorithm cannot be automatically determined in the prior art, this application combines a large audio model with multimodal data to accurately identify and classify noise types and dynamically select the optimal noise reduction strategy.
[0052] The application scenarios of noise reduction in this application include but are not limited to: online meetings, public broadcasting, driving environments, film and television shooting, or listening to music.
[0053] Optionally, in the embodiment of the present application, the above audio noise reduction method can be applied to Figure 1 In the hardware environment composed of the terminal 101 and the server 103 shown in FIG. Figure 1 As shown, the server 103 is connected to the terminal 101 via a network, and can be used to process multimodal data input by the terminal. A database 105 can be set on the server or independently of the server to provide data storage services for the server 103. The above-mentioned network includes but is not limited to: a wide area network, a metropolitan area network or a local area network, and the terminal 101 includes but is not limited to a PC, a mobile phone, a tablet computer, etc.
[0054] An audio noise reduction method in an embodiment of the present application can be executed by the server 103 or by the terminal 101.
[0055] The following will be combined with specific implementation methods to provide an audio noise reduction method provided by the embodiment of the present application in detail, taking application to a server as an example. Figure 2 As shown, the specific steps are as follows:
[0056] Step 201: Acquire multimodal data, wherein the multimodal data includes audio data and associated data of the audio data, and the associated data is used to assist in understanding the content of the audio data;
[0057] Step 202: Process the multimodal data using the trained multimodal audio large model, and output the noise type in the audio data;
[0058] Step 203: Determine a noise reduction strategy corresponding to the noise type according to a preset matching rule;
[0059] Step 204: adopt a noise reduction strategy to reduce noise on the audio data.
[0060] In the intelligent audio processing system, the terminal collects multimodal data and performs pre-processing operations such as removing obvious interference, framing, and windowing, and then uploads it to the server. Multimodal data refers to a data set that contains multiple different types of information at the same time, such as audio, text, and video, which together describe an event or scene and provide complementary contextual information. Specifically, audio data is the main input source, which can be voice, music, or other environmental sounds recorded by a microphone. Associated data is used to assist in understanding the content of the audio, including but not limited to text descriptions and visual information. Text descriptions are transcribed into text by automatic speech recognition technology, or contextual information provided by the user; visual information, such as video frames or images, can provide additional clues about the audio recording environment, such as objects in the scene, character movements, etc.
[0061] After the server obtains the above multimodal data, it processes it using the trained multimodal audio model. This model has been trained with a large amount of audio and its related data and has strong semantic understanding and feature extraction capabilities. The functions of the multimodal audio model include: integrating audio and related data, deeply understanding the specific context of the audio, such as distinguishing between human voices, background music or environmental noise, and accurately identifying and classifying various types of noise in the audio based on a comprehensive understanding of the audio content, such as traffic noise, mechanical noise and crowd noise.
[0062] Based on the noise type output by the model, the server will refer to the preset matching rules to determine the most appropriate noise reduction strategy. These rules are usually formulated by domain experts and take into account the characteristics of different noise types and the available noise reduction algorithms. For example, for low-frequency noise (such as wind noise and air conditioning operation), adaptive filters are preferred; for high-frequency sharp noise (such as keyboard tapping and alarm sounds), spectral subtraction is more suitable; for complex mixed noise (such as traffic + crowd noise), it is recommended to use a combined algorithm to suppress the low-frequency part first and then process the remaining high-frequency components.
[0063] Finally, the original audio data is processed according to the selected noise reduction strategy. During this process, the server may dynamically adjust parameters to optimize the effect, such as fine-tuning filter settings or updating noise models based on real-time feedback. The server can also perform post-processing operations on the audio signal after noise reduction, such as smoothing and format conversion, and output a clear, high-quality audio signal.
[0064] For example, a user recorded a speech video on a busy city street, and the audio was mixed with traffic noise, crowd noise, and other ambient sounds. The system first obtains this audio and its associated text "recording conversation audio in a public place with human voice and traffic noise in the background" as multimodal data. Then the multimodal audio large model analyzes the audio content and text content, and identifies the main noise types as traffic noise and crowd noise. According to the preset rules, the system chooses to suppress low-frequency traffic noise with an adaptive filter first, and then process the high-frequency crowd noise through a deep learning model. In the final output audio, the speaker's voice is clearer and the background noise is greatly reduced, providing a better listening experience.
[0065] In this application, the associated data can help to better understand the audio content, so that the multimodal audio model can combine the associated data and audio data, accurately identify and classify the noise type in the audio data, and automatically select the optimal noise reduction strategy according to the matching rules, thereby improving the noise reduction effect.
[0066] As an optional implementation, in step 202, the multimodal data is processed by the trained multimodal audio large model, and the noise types in the output audio data include the following:
[0067] Step S11: extracting multiple modal features of multimodal data through the trained multimodal audio large model, wherein the multiple modal features include audio features and associated features of associated data;
[0068] Step S12: using an alignment network to integrate multiple modal features into the same vector space;
[0069] Step S13: In the vector space, an attention mechanism is used to capture the interaction between the audio feature and each associated feature to obtain a fused multimodal embedding vector, wherein the multimodal embedding vector fuses the audio feature and all associated features;
[0070] Step S14: extracting global semantic features of the multimodal embedding vector using a fully connected layer, and determining the noise type based on the global semantic features, wherein the global semantic features can reflect the overall situation of the audio scene.
[0071] In an intelligent audio processing system, it is first necessary to extract rich features from the input multimodal data. These features include not only audio features, but also related features related to the audio (such as text features, video features, etc.) to assist in understanding the content of the audio. The following describes the features of various modalities.
[0072] Audio features: Extract time domain features (such as short-time energy and zero-crossing rate) and frequency domain features (such as spectrogram and Mel-frequency cepstrum coefficients) from audio signals. These features can reflect the time changes and frequency distribution of audio.
[0073] Text features: semantic information extracted from audio transcriptions or contextual information provided by users. These features help understand the specific context of the audio, and are particularly important in tasks such as speech recognition and semantic analysis. For example, users may provide information about the recording location, such as "city square" or "conference room," to help the system predict the type of noise that may occur.
[0074] Visual features: Visual information about the recording environment extracted from video frames or images. These features provide additional contextual clues to help more accurately judge the audio background. For example, they can detect major objects in video frames, such as vehicles, pedestrians, buildings, etc., and help the system infer possible noise sources.
[0075] As important components of multimodal data, text features and visual features not only help the system better understand the audio content, but also provide important clues for accurately determining the type of noise. By combining this rich correlation information, the system can make more informed noise reduction decisions.
[0076] In order to effectively fuse modal features from different sources in the same framework, the system uses an alignment network to map these features to the same vector space. This step ensures that information from different modalities can be compared and interacted in a unified space, avoiding information loss or confusion that may be caused by direct splicing, and facilitating meaningful fusion of information from different modalities in the same dimension.
[0077] In a unified vector space, the system uses an attention mechanism to capture the interactive relationship between audio features and each associated feature. The attention mechanism can dynamically focus on key information points, highlight the associated features that have the greatest impact on the current audio clip, and generate a fused multimodal embedding vector. This multimodal embedding vector not only contains the key information of the original audio, but also incorporates important background information from other modalities.
[0078] Finally, the system further processes the multimodal embedding vector through a fully connected layer to extract its global semantic features. These global semantic features can fully reflect the overall situation of the audio scene and are used to determine the type of noise present in it, such as traffic noise, crowd noise, etc.
[0079] In this application, a multimodal audio model is first used to extract multiple features of audio and related data, and these features are integrated into the same vector space through an alignment network. Then, in a unified vector space, an attention mechanism is used to capture the interactive relationship between audio and each associated feature, and a fused multimodal embedding vector is generated. Finally, the global semantic features of the embedded vector are extracted through a fully connected layer to accurately determine the type of noise. This method not only improves the accuracy of noise reduction, but also enhances the robustness of the system in complex environments.
[0080] As an optional implementation, in step S13, an attention mechanism is used to capture the interactive relationship between the audio feature and each associated feature, and the fused multimodal embedding vector includes the following contents:
[0081] Step S131: generating an interaction score based on the interaction relationship between the audio feature and the associated feature at each moment;
[0082] Step S132: performing weighted summation on the interaction scores of different associated features at the same time to obtain a context vector, wherein the context vector is used to indicate information associated with the audio at each time;
[0083] Step S133: Combining the audio feature with the context vector to generate a fusion feature;
[0084] Step S134: Obtain a multimodal embedding vector by integrating the fusion features at all times.
[0085] In the intelligent audio processing system, in order to better understand the audio content and improve the noise reduction effect, the system needs to dynamically evaluate the interactive relationship between audio features and related features (such as text descriptions and visual information) at each moment. The specific steps are as follows:
[0086] For each time point, the system calculates the correlation between the audio feature and all associated features (such as text features and visual features). These correlations are quantified as interaction scores, which are used to reflect the degree of influence of different associated features on the audio content at that moment. For example, if a video frame shows a large number of pedestrians flowing, and there is crowd noise in the audio at the same time, the interaction score between the two will be high.
[0087] In order to integrate the information of multiple associated features, the system performs a weighted summation of all interaction scores at the same time to generate a context vector, which integrates the important information of all associated features at the current time point and indicates the audio background corresponding to that moment. The system calculates attention weights based on the interaction scores to ensure that those associated features that are considered more important (i.e., with higher interaction scores) are given greater weights. The system uses these attention weights to perform a weighted summation of the associated features to form a context vector, which not only contains the information of the original associated features, but also integrates their interactions with the audio features, providing rich background information for subsequent processing.
[0088] In order to more comprehensively represent the audio content at each moment, the system combines the original audio features with the context vector generated above to form a new fusion feature. The most direct method is to simply concatenate the audio features and the context vector to form a higher-dimensional feature representation.
[0089] Finally, the system integrates the fusion features at all time points to generate a multimodal embedding vector. This embedding vector integrates the multimodal information of the entire audio clip and provides a basis for further extracting global semantic features. The system can complete the integration process in a variety of ways, such as taking the average, selecting the maximum value, or using a recurrent neural network to capture temporal dependencies. The final generated multimodal embedding vector not only contains the key information of the original audio, but also incorporates important background information from other modalities, allowing it to fully reflect the overall situation of the audio scene.
[0090] In this application, by dynamically evaluating the interaction between audio features and related features, the system can more accurately capture the audio content and its background environment, and then generate a multimodal embedding vector based on the interaction relationship. The system not only improves the audio quality and processing efficiency, but also demonstrates how to use data from multiple sources to enhance the intelligence level of audio processing.
[0091] As an optional implementation, in step S14, extracting the global semantic features of the multimodal embedding vector using a fully connected layer, and determining the noise type according to the global semantic features includes:
[0092] Step S141: using a fully connected layer to perform dimensionality reduction processing on the multimodal embedding vector to obtain concentrated features;
[0093] Step S142: extracting multiple semantic features from the concentrated features layer by layer through the fully connected layer to generate global semantic features;
[0094] Step S143: Match the global semantic feature with a pre-built knowledge base to determine the noise type, wherein the knowledge base contains various noise types and their corresponding semantic feature templates.
[0095] After completing feature extraction, feature alignment, and interactive relationship capture of multimodal data, the system generates a multimodal embedding vector that combines audio features and all related features. In order to further optimize processing efficiency and highlight key information, the system then uses a fully connected layer to reduce the dimensionality of this high-dimensional embedding vector.
[0096] The fully connected layer maps the high-dimensional multimodal embedding vector into a lower-dimensional space through learning to generate concentrated features. This process not only reduces the computational complexity, but also enhances the criticality of feature representation by removing redundant information. Despite the reduction in dimensionality, the fully connected layer ensures that the concentrated features still retain the core information in the original embedding vector through carefully designed network structure and parameter adjustment, providing a solid foundation for subsequent analysis.
[0097] In order to capture deeper semantic information from the condensed features, the system abstracts layer by layer through the fully connected layer, and gradually extracts semantic features at different levels. Each layer focuses on a specific type of semantic pattern, and finally forms a global semantic feature that comprehensively reflects the overall situation of the audio scene. The first layer may focus on basic sound characteristics (such as frequency distribution, pitch changes), while higher layers focus on more complex semantic understanding (such as voice content, environmental description, etc.). This hierarchical processing method enables the system to gradually extract more representative and explanatory features. After multiple layers of processing, the system generates a global semantic feature vector, which not only contains the key information of the original audio, but also integrates important background information from other modalities, so that it can comprehensively reflect the overall situation of the audio scene.
[0098] Finally, the system matches the generated global semantic features with the pre-built knowledge base to determine the noise type that best matches the current audio scene. This knowledge base contains various common noise types and their corresponding semantic feature templates, providing a reliable reference standard for the system. The system calculates the similarity between the global semantic features and each noise type template in the knowledge base, and selects the closest template as the matching result. For example, if a large number of traffic-related words or visual elements are detected, the system may determine that traffic noise exists. In some cases, the system may dynamically adjust the judgment of the noise type in combination with real-time feedback (such as user evaluation or device perception) to ensure accuracy.
[0099] In this application, the multimodal embedding vector is reduced in dimension through the fully connected layer to generate concentrated features, remove redundant information, ensure the retention of key features, and thus improve the accuracy of subsequent analysis. Multiple semantic features in the concentrated features are extracted layer by layer to generate global semantic features that fully reflect the overall situation of the audio scene. This enables the system to more accurately capture the audio content and its background environment, reduce misjudgment, match the global semantic features with the pre-built knowledge base, ensure that the determination of the noise type is based on rich templates and empirical data, and further improve the accuracy of classification.
[0100] As an optional implementation, in step 203, the noise type includes at least one type of noise, and determining the noise reduction strategy corresponding to the noise type according to the preset matching rule includes:
[0101] If the noise type is simple noise, the noise reduction algorithm in the signal processing algorithm library is determined according to the matching rule;
[0102] If the noise type is complex noise, the noise reduction algorithm in the deep learning algorithm library is determined according to the matching rule;
[0103] If the noise type includes simple noise and complex noise, a combined noise reduction algorithm in the signal processing algorithm library and the deep learning algorithm library is used according to the matching rules.
[0104] In intelligent audio processing systems, determining the appropriate noise reduction algorithm is a key step to ensure audio quality. By analyzing global semantic features and matching a pre-built knowledge base, the system can accurately identify the noise type and select the most appropriate noise reduction strategy based on different noise characteristics.
[0105] If the system determines that the noise type is simple noise (such as low-frequency wind noise, air conditioner running sound, etc.), it selects an appropriate noise reduction algorithm from the signal processing algorithm library according to the preset matching rules. The signal processing algorithm library includes but is not limited to: classical filters, such as Wiener filtering, Kalman filtering, spectral subtraction and adaptive filtering. For simple noise, the signal processing algorithm usually has a small amount of calculation and fast speed, and can achieve efficient real-time processing on devices with limited resources.
[0106] If the system determines that the noise type is complex noise (such as high-frequency sharp noise, non-steady-state noise or mixed noise), the corresponding noise reduction algorithm is selected from the deep learning algorithm library according to the preset matching rules. The deep learning algorithm library includes but is not limited to: convolutional neural network, recurrent neural network and long short-term memory network, transformer model. The deep learning algorithm can learn and identify complex noise patterns and provide higher noise reduction accuracy, especially in complex environments that are difficult to handle with traditional methods.
[0107] If the system determines that the noise type includes simple noise and complex noise, the combined noise reduction algorithm in the signal processing algorithm library and the deep learning algorithm library is used according to the preset matching rules. First, the signal processing algorithm is used to suppress simple noise (such as low-frequency traffic noise), and then the remaining complex noise (such as crowd noise) is processed by the deep learning algorithm. This phased approach can solve different types of problems in a targeted manner. The present application can improve the audio noise reduction effect by selecting the most appropriate noise reduction algorithm according to the noise type.
[0108] As an optional implementation, the training process of the multimodal audio large model includes: obtaining training samples, wherein the training samples include sample audio, noise labels corresponding to the sample audio, and sample-related data, and the sample-related data is used to assist in understanding the content of the sample audio; using the training samples to train the initial multimodal audio large model to obtain an output noise classification result; if the noise classification result is inconsistent with the noise label, using optimization algorithms such as the cross entropy loss function and the mean square error loss function to continuously adjust the model parameters to improve the accuracy and generalization ability of the model until the noise classification result is consistent with the noise label, thereby obtaining a trained multimodal audio large model.
[0109] The beneficial effects achieved by this application are as follows:
[0110] 1. It realizes intelligent noise reduction processing. By combining with the multimodal audio large model, it can accurately determine the type of noise in the audio, improving the pertinence and efficiency of noise reduction.
[0111] 2. Optimized the utilization of computing resources, avoided unnecessary complex calculations, and improved the real-time performance and resource utilization efficiency of the system.
[0112] 3. The robustness of the system is enhanced, and good noise reduction effect can be maintained even when the training data is insufficient or the noise type is complex.
[0113] Based on the same technical concept, the present application provides an audio noise reduction device, such as Figure 3 As shown, the device comprises:
[0114] An acquisition module 301 is used to acquire multimodal data, wherein the multimodal data includes audio data and associated data of the audio data, and the associated data is used to assist in understanding the content of the audio data;
[0115] A processing module 302 is used to process the multimodal data using the trained multimodal audio model and output the noise type in the audio data;
[0116] A determination module 303 is used to determine a noise reduction strategy corresponding to a noise type according to a preset matching rule;
[0117] The noise reduction module 304 is used to reduce the noise of the audio data by adopting a noise reduction strategy.
[0118] Optionally, the processing module 302 is used to:
[0119] Extracting multiple modal features of multimodal data through the trained multimodal audio large model, wherein the multiple modal features include audio features and associated features of associated data;
[0120] An alignment network is used to integrate multiple modal features into the same vector space;
[0121] In the vector space, the attention mechanism is used to capture the interactive relationship between the audio features and each associated feature, and a fused multimodal embedding vector is obtained, where the multimodal embedding vector combines the audio features and all associated features;
[0122] A fully connected layer is used to extract the global semantic features of the multimodal embedding vector, and the noise type is determined based on the global semantic features, where the global semantic features can reflect the overall situation of the audio scene.
[0123] Optionally, the processing module 302 is used to:
[0124] Generate an interaction score based on the interaction relationship between the audio features and the associated features at each moment;
[0125] The interaction scores of different associated features at the same time are weighted and summed to obtain a context vector, where the context vector is used to indicate the audio association information at each time;
[0126] Combine the audio features with the context vector to generate fused features;
[0127] The multimodal embedding vector is obtained by integrating the fusion features at all times.
[0128] Optionally, the processing module 302 is used to:
[0129] A fully connected layer is used to reduce the dimension of the multimodal embedding vector to obtain concentrated features;
[0130] Through the fully connected layer, multiple semantic features in the concentrated features are extracted layer by layer to generate global semantic features;
[0131] The global semantic features are matched with a pre-built knowledge base to determine the noise type, wherein the knowledge base contains various noise types and their corresponding semantic feature templates.
[0132] Optionally, the processing module 302 is used to:
[0133] Extract the time domain features and frequency domain features of audio data through the trained multimodal audio model;
[0134] The trained multimodal audio model is used to extract text features from text and visual features from video.
[0135] Optionally, the noise type includes at least one type of noise, and the determination module 303 is used to:
[0136] If the noise type is simple noise, the noise reduction algorithm in the signal processing algorithm library is determined according to the matching rule;
[0137] If the noise type is complex noise, the noise reduction algorithm in the deep learning algorithm library is determined according to the matching rule;
[0138] If the noise type includes simple noise and complex noise, a combined noise reduction algorithm in the signal processing algorithm library and the deep learning algorithm library is used according to the matching rules.
[0139] Optionally, the training process of the multimodal audio large model includes:
[0140] Obtaining a training sample, wherein the training sample includes a sample audio, a noise label corresponding to the sample audio, and sample-related data, where the sample-related data is used to assist in understanding the content of the sample audio;
[0141] The initial multimodal audio model is trained using training samples to obtain output noise classification results;
[0142] If the noise classification result is inconsistent with the noise label, the parameters of the initial multimodal audio model are adjusted until the noise classification result is consistent with the noise label, thereby obtaining a trained multimodal audio model.
[0143] like Figure 4 As shown, an embodiment of the present application provides an electronic device, including a processor 401, a communication interface 402, a memory 403 and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404.
[0144] The memory 403 is used to store computer programs.
[0145] In one embodiment of the present application, the processor 401 is used to implement the audio noise reduction method provided by any one of the aforementioned method embodiments when executing the program stored in the memory 403.
[0146] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the audio noise reduction method provided in any of the aforementioned method embodiments are implemented.
[0147] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0149] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "include", "comprise", "contain", and "have" are inclusive, and therefore specify the existence of stated features, steps, operations, elements and / or parts, but do not exclude the existence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not interpreted as necessarily requiring them to be performed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0150] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features applied for herein.
Claims
1. An audio noise reduction method, characterized in that: The method comprises: Acquire multimodal data, wherein the multimodal data includes audio data and associated data of the audio data, wherein the associated data is used to assist in understanding the content of the audio data; Processing the multimodal data through the trained multimodal audio large model, and outputting the noise type in the audio data; Determine a noise reduction strategy corresponding to the noise type according to a preset matching rule; The audio data is denoised using the denoising strategy.
2. The method according to claim 1, characterized in that: The multimodal data is processed by the trained multimodal audio large model, and the noise types in the outputted audio data include: Extracting multiple modal features of the multimodal data through the trained multimodal audio large model, wherein the multiple modal features include audio features and associated features of the associated data; An alignment network is used to integrate the multiple modal features into the same vector space; In the vector space, an attention mechanism is used to capture the interaction relationship between the audio feature and each of the associated features to obtain a fused multimodal embedding vector, wherein the multimodal embedding vector fuses the audio feature and all the associated features; A fully connected layer is used to extract global semantic features of the multimodal embedding vector, and the noise type is determined according to the global semantic features, wherein the global semantic features can reflect the overall situation of the audio scene.
3. The method according to claim 2, characterized in that The attention mechanism is used to capture the interactive relationship between the audio feature and each of the associated features, and the fused multimodal embedding vector includes: Generate an interaction score based on the interaction relationship between the audio feature and the associated feature at each moment; Performing a weighted summation of the interaction scores of different associated features at the same moment to obtain a context vector, wherein the context vector is used to indicate the audio association information at each moment; Combining the audio feature with the context vector to generate a fused feature; The multimodal embedding vector is obtained by integrating the fusion features at all moments.
4. The method according to claim 2, characterized in that: The global semantic features of the multimodal embedding vector are extracted using a fully connected layer, and the noise type is determined according to the global semantic features, including: Using a fully connected layer to perform dimensionality reduction processing on the multimodal embedding vector to obtain concentrated features; Extracting multiple semantic features from the concentrated features layer by layer through the fully connected layer to generate global semantic features; The global semantic feature is matched with a pre-built knowledge base to determine the noise type, wherein the knowledge base contains various noise types and their corresponding semantic feature templates.
5. The method according to claim 2, characterized in that: The associated data includes text and video, and the multiple modal features of the multimodal data extracted by the trained multimodal audio model include: Extracting time domain features and frequency domain features of the audio data through the trained multimodal audio large model; The text features of the text and the visual features of the video are extracted through the trained multimodal audio model.
6. The method according to claim 1, characterized in that The noise type includes at least one type of noise, and determining the noise reduction strategy corresponding to the noise type according to the preset matching rule includes: If the noise type is simple noise, determining a noise reduction algorithm in a signal processing algorithm library according to the matching rule; If the noise type is complex noise, determining a noise reduction algorithm in a deep learning algorithm library according to the matching rule; If the noise type includes simple noise and complex noise, a combined noise reduction algorithm in the signal processing algorithm library and the deep learning algorithm library is adopted according to the matching rule.
7. The method according to claim 1, characterized in that The training process of the multimodal audio large model includes: Acquire a training sample, wherein the training sample includes a sample audio, a noise label corresponding to the sample audio, and sample-related data, wherein the sample-related data is used to assist in understanding the content of the sample audio; Using the training samples to train the initial multimodal audio large model to obtain an output noise classification result; If the noise classification result is inconsistent with the noise label, the parameters of the initial multimodal audio model are adjusted until the noise classification result is consistent with the noise label, thereby obtaining a trained multimodal audio model.
8. An audio noise reduction device, characterized in that: The device comprises: An acquisition module, configured to acquire multimodal data, wherein the multimodal data includes audio data and associated data of the audio data, wherein the associated data is used to assist in understanding the content of the audio data; A processing module, used for processing the multimodal data through a trained multimodal audio model, and outputting the noise type in the audio data; A determination module, used to determine a noise reduction strategy corresponding to the noise type according to a preset matching rule; The noise reduction module is used to perform noise reduction on the audio data by adopting the noise reduction strategy.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-7 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and apparatus for real-time sound enhancement
CN117121103A
Audio noise processing method and system based on big data
CN118411998A
Earphone noise reduction method based on improved Transform model, earphone and storage medium
CN119172682A
Method and apparatus for real-time sound enhancement
US20230298593A1
Artificial intelligence-based audio processing method and apparatus, electronic device, computer readable storage medium, and computer program product
WO2022116825A1
Cited By
Voice signal processing method and device, equipment, storage medium and computer program product
CN120656470A
Environmental atmosphere control method and device based on data analysis, equipment and medium
CN120730586A
Environment atmosphere control method and device based on data analysis, equipment and medium
CN120730586B
Multi-modal voice communication anti-noise method based on deep learning
CN121011159A