Audio processing method and system based on reinforcement learning, and digital loudspeaker

Through an audio processing method based on reinforcement learning, the digital speaker system can identify the audio content type in real time and dynamically adjust the sound effects, solving the problems of delayed and inaccurate sound effect adjustment in existing technologies and improving the intelligence and adaptability of sound effect processing.

CN120783784AInactive Publication Date: 2025-10-14GUANGDONG TIANHONGSHENG OPTOELECTRONICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510769707.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing digital speaker systems lack the ability to perceive audio content types and playback environments, resulting in delayed or inaccurate sound adjustment responses in various usage scenarios, and are unable to fully utilize content recognition results to dynamically drive the switching and adjustment of enhancement strategies.

Method used

An audio processing method based on reinforcement learning is adopted. By collecting audio signals and converting them into audio frequency domain correction signals to suppress distortion, the energy dynamic trend vector is extracted, and the content is recognized using a pre-trained audio content classification model. The sound effect processing strategy template is dynamically adjusted to adapt to different audio scenarios.

Benefits of technology

It enables high-precision, real-time sound adjustment of input audio by digital speakers, improves the pertinence and responsiveness of sound processing, and optimizes the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783784A_ABST
    Figure CN120783784A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and system based on reinforcement learning and a digital loudspeaker, and relates to the technical field of audio processing, and the method comprises the steps: collecting an audio signal received by the digital loudspeaker; the method comprises the following steps: converting an audio signal into an audio frequency domain correction signal, performing distortion suppression on the audio frequency domain correction signal to obtain an audio distortion weighing signal, and further extracting an energy dynamic trend vector in the audio distortion weighing signal; acquiring a pre-trained audio content classification model based on reinforcement learning, and performing content classification on the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning so as to obtain an audio content type of the audio signal received by the digital loudspeaker; and performing dynamic sound effect enhancement on the audio signal received by the digital loudspeaker according to the audio content type. According to the invention, the audio content identification result can be fully utilized to dynamically drive the switching and adjustment of the enhancement strategy, so that the sound effect adjustment response capability of the digital loudspeaker to the input audio is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and more specifically, to an audio processing method, system and digital speaker based on reinforcement learning. Background Art

[0002] In recent years, artificial intelligence (AI) technology has achieved breakthroughs in fields such as speech recognition and image processing. Some research has begun to incorporate reinforcement learning (RL) mechanisms into audio processing to achieve adaptive system adjustments and optimization driven by user behavior feedback. RL-based audio processing is a cutting-edge approach that incorporates intelligent learning mechanisms into audio signal analysis and enhancement. With the widespread adoption of consumer audio devices, especially digital speakers, users' demands for sound quality are increasing. Modern digital speaker systems typically feature high sampling rates and low-distortion output capabilities, and integrate a variety of digital signal processing modules (such as equalizers, dynamic compressors, and noise reduction processors) to enhance the clarity, layering, and spatial perception of audio playback.

[0003] However, existing technologies usually rely on static or manually configured parameters and lack the ability to perceive the audio content type and playback environment, resulting in poor performance in a variety of usage scenarios. For example, in a voice call scenario, enhancing low frequencies or excessive reverberation may weaken semantic clarity; while in a music playback scenario, if the sound effects are not dynamically adjusted according to the rhythm and spectral characteristics, the immersive experience may be affected. In addition, current systems often implement content classification and sound effect processing separately, and fail to fully utilize the content recognition results to dynamically drive the switching and adjustment of enhancement strategies, resulting in delayed or inaccurate sound effect adjustment responses. Therefore, how to fully utilize the audio content recognition results to dynamically drive the switching and adjustment of enhancement strategies to improve the digital speakers' ability to adjust the sound effects of input audio is a difficult problem facing the industry. Summary of the Invention

[0004] The present application provides an audio processing method, system and digital speaker based on reinforcement learning, which can make full use of the audio content recognition results to dynamically drive the switching and adjustment of the enhancement strategy, so as to improve the digital speaker's ability to adjust the sound effect of the input audio.

[0005] In a first aspect, the present application provides an audio processing method based on reinforcement learning, the audio processing method comprising the following steps:

[0006] Collecting audio signals received by digital speakers;

[0007] converting the audio signal into an audio frequency domain correction signal, performing distortion suppression on the audio frequency domain correction signal to obtain an audio distortion trade-off signal, and then extracting an energy dynamic trend vector from the audio distortion trade-off signal;

[0008] Obtaining a pre-trained audio content classification model based on reinforcement learning, and performing content classification on the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning, thereby obtaining an audio content type of the audio signal received by the digital speaker;

[0009] Dynamic sound effect enhancement is performed on the audio signal received by the digital speaker according to the audio content type.

[0010] In this embodiment, a digital signal monitor is used to collect the audio signal received by the digital speaker.

[0011] In this embodiment, converting the audio signal into an audio frequency domain corrected signal specifically includes:

[0012] Performing baseline drift correction on the audio signal to obtain an audio correction signal;

[0013] Perform frequency domain conversion on the audio correction signal to obtain an audio frequency domain correction signal.

[0014] In this embodiment, performing distortion suppression on the audio frequency domain correction signal to obtain the audio distortion trade-off signal specifically includes:

[0015] Performing convolution on the audio frequency domain correction signal to obtain a convolution-combined audio frequency domain correction signal;

[0016] Determining the noise estimation spectrum in the audio frequency domain corrected signal after convolution;

[0017] Determining a noise suppression spectrum based on the spectrum of the audio frequency domain correction signal after convolution, the noise estimation spectrum, and a preset distortion equalization factor;

[0018] The noise suppression spectrum is converted into an audio distortion trade-off signal.

[0019] In this embodiment, extracting the energy dynamic trend vector from the audio distortion tradeoff signal specifically includes:

[0020] generating an energy dynamic trend curve corresponding to the audio distortion trade-off signal;

[0021] Performing multi-dimensional feature extraction on the energy dynamic trend curve to obtain all energy dynamic trend features;

[0022] An energy dynamic trend vector in the audio distortion trade-off signal is determined according to all energy dynamic trend features.

[0023] In this embodiment, the audio content classification model based on reinforcement learning is a self-learning model.

[0024] In this embodiment, content classification of the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning is to input the energy dynamic trend vector as input data into the audio content classification model based on reinforcement learning to perform content classification on the audio signal.

[0025] In this embodiment, dynamically enhancing the audio signal received by the digital speaker according to the audio content type specifically includes:

[0026] Acquire a corresponding sound effect processing strategy template based on the audio content type;

[0027] Acquiring content characteristics of an audio signal received by the digital speaker and environmental data of an environment in which the digital speaker is located;

[0028] Dynamically regulating the sound effect processing strategy template using the content features and the environmental data, thereby obtaining a sound effect processing dynamic strategy;

[0029] The sound effect is enhanced on the audio signal received by the digital speaker according to the dynamic sound effect processing strategy.

[0030] In a second aspect, the present application provides an audio processing system based on reinforcement learning, for performing an audio processing method based on reinforcement learning, the audio processing system comprising:

[0031] An audio acquisition module, used to collect audio signals received by the digital speaker;

[0032] a feature extraction module, configured to convert the audio signal into an audio frequency domain correction signal, perform distortion suppression on the audio frequency domain correction signal to obtain an audio distortion trade-off signal, and further extract an energy dynamic trend vector from the audio distortion trade-off signal;

[0033] a content classification module, configured to obtain a pre-trained audio content classification model based on reinforcement learning, perform content classification on the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning, and thereby obtain the audio content type of the audio signal received by the digital speaker;

[0034] The sound effect enhancement module is used to dynamically enhance the sound effect of the audio signal received by the digital speaker according to the audio content type.

[0035] In a third aspect, the present application provides a digital speaker, which includes the above-mentioned audio processing system based on reinforcement learning.

[0036] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:

[0037] The method collects audio signals received by a digital speaker; converts the audio signals into audio frequency domain correction signals, performs distortion suppression on the audio frequency domain correction signals to obtain audio distortion trade-off signals, and then extracts energy dynamic trend vectors from the audio distortion trade-off signals; obtains a pre-trained audio content classification model based on reinforcement learning, performs content classification on the audio signals according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning, and then obtains the audio content type of the audio signals received by the digital speaker; and performs dynamic sound enhancement on the audio signals received by the digital speaker according to the audio content type.

[0038] It can be seen that in this application, first, the audio signal is converted into an audio frequency domain correction signal, and its distortion is suppressed to obtain an audio distortion trade-off signal, and the energy dynamic trend vector is further extracted, which can effectively remove the non-ideal factors introduced by equipment characteristics, signal interference or environmental noise, and extract more expressive energy change laws. Through fine and multi-dimensional content modeling, a clearer and more recognizable state input is provided for the reinforcement learning model, which helps the model to accurately judge the content type or scene state of the current audio; then, the obtained energy dynamic trend vector is input into the pre-trained audio content classification model based on reinforcement learning for content recognition, which can give full play to the autonomous strategy optimization of the reinforcement learning model in complex environments. It takes advantage of the advantages of digitalization and dynamic decision-making, thereby achieving high-precision and real-time discrimination of audio content types; finally, it dynamically enhances the audio signal received by the digital speaker according to the audio content type, which can realize intelligent adaptation and precise control of different audio scenes. By first identifying and classifying the audio signal content, and then calling the matching sound enhancement strategy template based on the classification result, and dynamically adjusting the strategy parameters in combination with the playback environment and real-time audio characteristics, the sound effect processing process is no longer static and fixed, but has adaptive capabilities, which significantly improves the pertinence and responsiveness of the sound effect enhancement, and achieves the purpose of continuously optimizing the auditory experience in complex environments, thereby significantly improving the digital speaker's sound effect adjustment responsiveness and intelligence level for the input audio.

[0039] In summary, the technical solution adopted in this application can make full use of the audio content recognition results to dynamically drive the switching and adjustment of the enhancement strategy, so as to improve the digital speaker's ability to adjust the sound effects of the input audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the accompanying drawings in the following description only relate to the part of the present application, and other drawings can be obtained by those of ordinary skill in the art without any creative effort.

[0041] Figure 1 is an exemplary flowchart of an audio processing method based on reinforcement learning provided by the present application;

[0042] Figure 2 is a flowchart of determining an audio distortion trade-off signal provided by the present application;

[0043] Figure 3 is a flowchart of dynamic sound effect enhancement provided by the present application;

[0044] Figure 4 is a module structure diagram of an audio processing system based on reinforcement learning provided by the present application. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments only relate to part of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative effort are within the scope of protection of the present application.

[0046] The embodiments of the present application provide an audio processing method, system and digital speaker based on reinforcement learning. The core is to collect an audio signal received by a digital speaker; convert the audio signal into an audio frequency domain correction signal, suppress distortion of the audio frequency domain correction signal to obtain an audio distortion trade-off signal, and then extract an energy dynamic trend vector in the audio distortion trade-off signal; obtain a pre-trained audio content classification model based on reinforcement learning, classify the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning, and then obtain the audio content type of the audio signal received by the digital speaker; and dynamically enhance the audio effect of the audio signal received by the digital speaker according to the audio content type. The above scheme can fully utilize the audio content recognition result to dynamically drive the switching and adjustment of the enhancement strategy, so as to improve the audio effect adjustment response capability of the digital speaker to the input audio.

[0047] Embodiment one, in order to better understand the above technical solutions, the above technical solutions will be described in detail in the following with reference to the drawings in the specification and specific implementation manners, referring to Figure 1As shown in the figure, the figure is an example flowchart of an audio processing method based on reinforcement learning according to the embodiment of the present application, the audio processing method comprising the following steps:

[0048] In step S1, an audio signal received by a digital speaker is collected.

[0049] In implementation, the digital signal listener can be used to collect the audio signal received by the digital speaker; it should be noted that the digital signal listener (Digital Signal Listener) refers to a device or logic module capable of listening to and collecting audio data stream in a digital audio transmission channel, which is usually embedded in the audio signal path or connected in parallel on the digital audio bus, and is used for non-intrusive acquisition of audio sample data, i.e. the audio signal received by the digital speaker.

[0050] In step S2, the audio signal is converted into an audio frequency domain correction signal, distortion suppression is performed on the audio frequency domain correction signal to obtain an audio distortion trade-off signal, and then an energy dynamic trend vector in the audio distortion trade-off signal is extracted.

[0051] In the embodiment, the audio signal is converted into an audio frequency domain correction signal in the following manner:

[0052] Baseline drift correction is performed on the audio signal to obtain an audio correction signal;

[0053] Frequency domain conversion is performed on the audio correction signal to obtain an audio frequency domain correction signal.

[0054] In implementation, first, in digital audio signal collection, due to unstable power supply, digital interface synchronous jitter, audio pre-stage circuit bias drift, etc., the original audio signal may have low-frequency jitter, DC offset or slow baseline offset phenomenon, which will affect the accuracy of subsequent spectral analysis and feature extraction, therefore, baseline drift correction can be performed on the audio signal, i.e. an IIR or FIR high-pass filter with low cutoff frequency is designed to remove the low-frequency drift component in the audio signal and retain the useful frequency band, so that the obtained signal can be used as an audio correction signal; then, frequency domain conversion can be performed on the audio correction signal, i.e. the audio correction signal can be divided into equal-length frames (such as 1024 points per frame, 512-point frame shift), in order to avoid spectral leakage, window processing (such as Hamming window, Hanning window) is required to smooth the boundary, fast Fourier transform can be applied to each frame, so that the final frequency domain data can be used as an audio frequency domain correction signal, which reflects the structural characteristics of the input audio in the frequency domain and has good signal-to-noise ratio and analysis stability.

[0055] Preferably, in this embodiment, the audio frequency domain correction signal is subjected to distortion suppression to obtain an audio distortion trade-off signal, referring to Figure 2 As shown in FIG. 1 , this figure is a schematic diagram of a process for determining an audio distortion tradeoff signal in some embodiments of the present application. In this embodiment, determining the audio distortion tradeoff signal can be implemented using the following steps:

[0056] In step S21, the audio frequency domain correction signal is convoluted to obtain a convoluted audio frequency domain correction signal;

[0057] In step S22, the noise estimation spectrum in the audio frequency domain corrected signal after convolution is determined;

[0058] In step S23, a noise suppression spectrum is determined based on the spectrum of the audio frequency domain correction signal after convolution, the noise estimation spectrum, and a preset distortion equalization factor;

[0059] In step S24, the noise suppression spectrum is converted into an audio distortion trade-off signal.

[0060] In a specific implementation, first, the audio frequency domain correction signal can be convolved, that is, the audio frequency domain correction signal is convolved with a preset smoothing filter kernel, so that a convolution-combined audio frequency domain correction signal can be obtained. Through the frequency domain convolution operation, the current audio frame is compounded with the smoothing filter kernel to enhance the effective frequency domain structure and weaken abnormal fluctuations, thereby improving the ability to distinguish between noise and distortion; then, the noise estimation spectrum in the convolution-combined audio frequency domain correction signal can be determined. In actual speech or music signals, noise and useful signals often overlap in the frequency domain. In order to achieve accurate suppression, the noise spectrum energy of the current frame needs to be estimated. The minimum mean method or spectral subtraction estimation model can be used to obtain a reference noise spectrum from the silent segment or low energy segment, that is, the noise estimation spectrum in the convolution-combined audio frequency domain correction signal; secondly, the noise suppression spectrum can be determined based on the spectrum of the convolution-combined audio frequency domain correction signal, the noise estimation spectrum and the preset distortion equalization factor, wherein the noise suppression spectrum represents the audio spectrum after noise suppression. In actual implementation, the noise suppression spectrum can be determined by the following formula:

[0061]

[0062] Among them, F(k) represents the noise suppression spectrum, S(k) represents the spectrum of the audio frequency domain correction signal after convolution, X(k) represents the noise estimation spectrum, and α represents the preset distortion equalization factor; finally, the noise suppression spectrum can be converted into an audio distortion trade-off signal, wherein the audio distortion trade-off signal represents the audio signal after achieving the best compromise between noise suppression and sound quality preservation, and the noise suppression spectrum can be inverse fast Fourier transformed, so that the transformed signal is used as the audio distortion trade-off signal. It should be noted that through distortion suppression, the noise and distortion components in the audio frequency domain correction signal are accurately modeled and suppressed, and a signal with a high signal-to-noise ratio and retaining the main semantic components is output, namely, the audio distortion trade-off signal.

[0063] In this embodiment, the energy dynamic trend vector in the audio distortion tradeoff signal may be extracted in the following manner:

[0064] generating an energy dynamic trend curve corresponding to the audio distortion trade-off signal;

[0065] Performing multi-dimensional feature extraction on the energy dynamic trend curve to obtain all energy dynamic trend features;

[0066] An energy dynamic trend vector in the audio distortion trade-off signal is determined according to all energy dynamic trend features.

[0067] In specific implementation, first, an energy dynamic trend curve corresponding to the audio distortion trade-off signal can be generated, wherein the energy dynamic trend curve is a curve used to represent the envelope trend of the overall energy of the audio distortion trade-off signal. The upper and lower envelope envelope method can be used to generate the envelope curve of the audio distortion trade-off signal, so that the envelope curve can be used as the energy dynamic trend curve corresponding to the audio distortion trade-off signal; then, multi-dimensional feature extraction can be performed on the energy dynamic trend curve to obtain all energy dynamic trend features. It should be noted that the energy dynamic trend features in this application include mean energy and variance energy, extreme value position distribution, energy change rate, energy change variance, kurtosis and skewness of the energy curve, energy rhythm density and envelope consistency index; finally, the energy dynamic trend vector in the audio distortion trade-off signal can be determined based on all energy dynamic trend features, that is, the feature vector composed of all energy dynamic trend features can be used as the energy dynamic trend vector in the audio distortion trade-off signal.

[0068] It should be noted that converting the audio signal into an audio frequency domain correction signal and suppressing its distortion to obtain an audio distortion trade-off signal, and further extracting the energy dynamic trend vector therefrom, can effectively remove non-ideal factors introduced by device characteristics, signal interference or environmental noise, and extract a more expressive energy change pattern. As an important feature that describes the energy fluctuations, rhythm strength and dynamic change pattern of audio on the time axis, the energy dynamic trend vector complements the frequency domain features and can comprehensively characterize the essential attributes of the audio content. Through sophisticated and multi-dimensional content modeling, a clearer and more identifiable state input is provided for the reinforcement learning model, which helps the model to accurately judge the content type or scene state of the current audio.

[0069] In step S3, a pre-trained audio content classification model based on reinforcement learning is obtained, and the audio signal is content-classified according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning, thereby obtaining the audio content type of the audio signal received by the digital speaker.

[0070] It should be noted that in this application, the audio content classification model based on reinforcement learning is a self-learning model; in specific implementation, the audio content classification model based on reinforcement learning can be constructed through processes such as building an environment, defining a model structure, designing a reward mechanism, training and updating strategies, and saving and deploying models. The audio content classification model is a self-learning model, the core of which is to gain experience through continuous interaction with the environment to improve the accuracy and generalization ability of audio content classification.

[0071] In this embodiment, content classification of the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning is to input the energy dynamic trend vector as input data into the audio content classification model based on reinforcement learning to perform content classification on the audio signal.

[0072] It should be noted that the audio content classification model based on reinforcement learning has learned to judge the corresponding audio content type (such as speech, pop music, classical music, environmental noise, etc.) from the feature vector during the training stage. The extracted energy dynamic trend vector is input as the state input into the pre-trained audio content classification model based on reinforcement learning. The audio content classification model will select the classification action according to the input state, thereby identifying the content type of the audio signal.

[0073] In addition, it should be noted that inputting the obtained energy dynamic trend vector into a pre-trained audio content classification model based on reinforcement learning for content recognition can give full play to the advantages of the reinforcement learning model in autonomous strategy optimization and dynamic decision-making in complex environments, thereby achieving high-precision and real-time discrimination of audio content types, so that the digital speaker can not only passively receive audio signals, but also actively understand the semantic attributes of the audio (such as speech, music, noise, etc.), and dynamically switch the most suitable sound enhancement strategy according to the recognition results.

[0074] In step S4, dynamic sound enhancement is performed on the audio signal received by the digital speaker according to the audio content type.

[0075] Preferably, in this embodiment, the audio signal received by the digital speaker is dynamically enhanced according to the audio content type, referring to Figure 3 As shown in FIG, this figure is a schematic diagram of the process of performing dynamic sound enhancement in some embodiments of the present application. In this embodiment, dynamic sound enhancement can be implemented using the following steps:

[0076] In step S41, a corresponding sound effect processing strategy template is obtained based on the audio content type;

[0077] In step S42, the content characteristics of the audio signal received by the digital speaker and the environmental data of the environment in which the digital speaker is located are obtained;

[0078] In step S43, the sound effect processing strategy template is dynamically regulated using the content features and the environment data, thereby obtaining a sound effect processing dynamic strategy;

[0079] In step S44, the sound effect is enhanced on the audio signal received by the digital speaker according to the dynamic sound effect processing strategy.

[0080] In specific implementation, first, the corresponding sound effect processing strategy template can be obtained based on the audio content type, that is, after identifying the audio content type (such as voice, music, ambient sound, mixed type), the system automatically calls the sound effect processing strategy template pre-matched with the audio content type. The sound effect processing strategy template predefines a set of core processing parameters, such as frequency band gain curve, compression threshold, reverberation degree, stereo width, etc.; then, the content characteristics of the audio signal received by the digital speaker and the environmental data of the environment in which the digital speaker is located can be obtained, wherein the content characteristics include the dynamic range, spectrum distribution, energy fluctuation trend, rhythmic index, etc. of the audio signal, and the environmental data include the current playback environment parameters collected from the built-in or supporting sensors of the speaker (such as background noise level and environmental reverberation degree, etc.); secondly, the content characteristics and environmental data can be used The sound effect processing strategy template is dynamically regulated, that is, at this stage, the system introduces content features and environmental data, and adaptively adjusts the parameters of the sound effect processing strategy template. For example, in a high-noise environment, the voice band gain is enhanced and the noise threshold is increased. Under close-range listening conditions, the stereo expansion is compressed to enhance the center positioning, etc., so that the regulated sound effect processing strategy template is used as the sound effect processing dynamic strategy; finally, the audio signal received by the digital speaker can be enhanced according to the sound effect processing dynamic strategy, that is, based on the strategy parameters dynamically generated in the sound effect processing dynamic strategy, real-time enhancement operations are performed on the audio signal received by the digital speaker, including but not limited to multi-band dynamic equalization, noise suppression and voice clarity enhancement, 3D spatial positioning or stereo width control, rhythm and layering enhancement, and gain adaptive control and distortion protection.

[0081] It should be noted that dynamic sound enhancement of the audio signal received by the digital speaker according to the audio content type can achieve intelligent adaptation and precise control of different audio scenes. By first identifying and classifying the audio signal content (such as voice, music, ambient sound, etc.), and then calling the matching sound enhancement strategy template based on the classification result, and dynamically adjusting the strategy parameters in combination with the playback environment and real-time audio characteristics, the sound effect processing process is no longer static and fixed, but has adaptive capabilities, which significantly improves the pertinence and responsiveness of the sound effect enhancement, and achieves the purpose of continuously optimizing the auditory experience in complex environments, thereby significantly improving the digital speaker's sound effect adjustment responsiveness and intelligence level for the input audio.

[0082] It can be seen that in this application, first, the audio signal is converted into an audio frequency domain correction signal, and its distortion is suppressed to obtain an audio distortion trade-off signal, and the energy dynamic trend vector is further extracted, which can effectively remove the non-ideal factors introduced by equipment characteristics, signal interference or environmental noise, and extract more expressive energy change laws. Through fine and multi-dimensional content modeling, a clearer and more recognizable state input is provided for the reinforcement learning model, which helps the model to accurately judge the content type or scene state of the current audio; then, the obtained energy dynamic trend vector is input into the pre-trained audio content classification model based on reinforcement learning for content recognition, which can give full play to the autonomous strategy optimization of the reinforcement learning model in complex environments. It takes advantage of the advantages of digitalization and dynamic decision-making, thereby achieving high-precision and real-time discrimination of audio content types; finally, it dynamically enhances the audio signal received by the digital speaker according to the audio content type, which can realize intelligent adaptation and precise control of different audio scenes. By first identifying and classifying the audio signal content, and then calling the matching sound enhancement strategy template based on the classification result, and dynamically adjusting the strategy parameters in combination with the playback environment and real-time audio characteristics, the sound effect processing process is no longer static and fixed, but has adaptive capabilities, which significantly improves the pertinence and responsiveness of the sound effect enhancement, and achieves the purpose of continuously optimizing the auditory experience in complex environments, thereby significantly improving the digital speaker's sound effect adjustment responsiveness and intelligence level for the input audio.

[0083] In summary, the technical solution adopted in this application can make full use of the audio content recognition results to dynamically drive the switching and adjustment of the enhancement strategy, so as to improve the digital speaker's ability to adjust the sound effects of the input audio.

[0084] In the second embodiment, the present application provides an audio processing system based on reinforcement learning, referring to Figure 4 As shown in FIG, this figure is a module structure diagram of an audio processing system based on reinforcement learning according to this embodiment of the present application, and the audio processing system includes:

[0085] The audio acquisition module 100 is used to acquire the audio signal received by the digital speaker;

[0086] a feature extraction module 200 for converting the audio signal into an audio frequency domain correction signal, performing distortion suppression on the audio frequency domain correction signal to obtain an audio distortion trade-off signal, and further extracting an energy dynamic trend vector from the audio distortion trade-off signal;

[0087] a content classification module 300 for obtaining a pre-trained audio content classification model based on reinforcement learning, performing content classification on the audio signal based on the energy dynamic trend vector and the audio content classification model based on reinforcement learning, and thereby obtaining the audio content type of the audio signal received by the digital speaker;

[0088] The sound effect enhancement module 400 is configured to dynamically enhance the sound effect of the audio signal received by the digital speaker according to the audio content type.

[0089] In a third embodiment, the present application further provides a digital speaker, which includes the above-mentioned audio processing system based on reinforcement learning, and is used to execute the audio processing method based on reinforcement learning in the present application.

[0090] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0091] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0092] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

Claims

1. An audio processing method based on reinforcement learning, characterized in that: The audio processing method comprises the following steps: Collecting audio signals received by digital speakers; converting the audio signal into an audio frequency domain correction signal, performing distortion suppression on the audio frequency domain correction signal to obtain an audio distortion trade-off signal, and then extracting an energy dynamic trend vector from the audio distortion trade-off signal; Obtaining a pre-trained audio content classification model based on reinforcement learning, and performing content classification on the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning, thereby obtaining an audio content type of the audio signal received by the digital speaker; Dynamic sound effect enhancement is performed on the audio signal received by the digital speaker according to the audio content type.

2. The audio processing method based on reinforcement learning according to claim 1, wherein: Use a digital signal monitor to capture the audio signal received by the digital speaker.

3. The audio processing method based on reinforcement learning according to claim 1, wherein: Converting the audio signal into an audio frequency domain correction signal specifically includes: Performing baseline drift correction on the audio signal to obtain an audio correction signal; Perform frequency domain conversion on the audio correction signal to obtain an audio frequency domain correction signal.

4. The audio processing method based on reinforcement learning according to claim 1, wherein: Performing distortion suppression on the audio frequency domain correction signal to obtain an audio distortion trade-off signal specifically includes: Performing convolution on the audio frequency domain correction signal to obtain a convolution-combined audio frequency domain correction signal; Determining the noise estimation spectrum in the audio frequency domain corrected signal after convolution; Determining a noise suppression spectrum based on the spectrum of the audio frequency domain correction signal after convolution, the noise estimation spectrum, and a preset distortion equalization factor; The noise suppression spectrum is converted into an audio distortion trade-off signal.

5. The audio processing method based on reinforcement learning according to claim 1, wherein: Extracting the energy dynamic trend vector from the audio distortion tradeoff signal specifically includes: generating an energy dynamic trend curve corresponding to the audio distortion trade-off signal; Performing multi-dimensional feature extraction on the energy dynamic trend curve to obtain all energy dynamic trend features; An energy dynamic trend vector in the audio distortion trade-off signal is determined according to all energy dynamic trend features.

6. The audio processing method based on reinforcement learning according to claim 1, characterized in that: The audio content classification model based on reinforcement learning is a self-learning model.

7. The audio processing method based on reinforcement learning according to claim 1, characterized in that: Content classification of the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning is to input the energy dynamic trend vector as input data into the audio content classification model based on reinforcement learning to perform content classification on the audio signal.

8. The audio processing method based on reinforcement learning according to claim 1, wherein: Dynamically enhancing the audio signal received by the digital speaker according to the audio content type specifically includes: Acquire a corresponding sound effect processing strategy template based on the audio content type; Acquiring content characteristics of an audio signal received by the digital speaker and environmental data of an environment in which the digital speaker is located; Dynamically regulating the sound effect processing strategy template using the content features and the environmental data, thereby obtaining a sound effect processing dynamic strategy; The sound effect is enhanced on the audio signal received by the digital speaker according to the dynamic sound effect processing strategy.

9. An audio processing system based on reinforcement learning, configured to execute the audio processing method based on reinforcement learning according to any one of claims 1 to 8, characterized in that: The audio processing system comprises: An audio acquisition module, used to collect audio signals received by the digital speaker; a feature extraction module, configured to convert the audio signal into an audio frequency domain correction signal, perform distortion suppression on the audio frequency domain correction signal to obtain an audio distortion trade-off signal, and further extract an energy dynamic trend vector from the audio distortion trade-off signal; a content classification module, configured to obtain a pre-trained audio content classification model based on reinforcement learning, perform content classification on the audio signal according to the energy dynamic trend vector and the audio content classification model based on reinforcement learning, and thereby obtain the audio content type of the audio signal received by the digital speaker; The sound effect enhancement module is used to dynamically enhance the sound effect of the audio signal received by the digital speaker according to the audio content type.

10. A digital speaker, characterized in that: The invention comprises the audio processing system based on reinforcement learning as claimed in claim 9.