An audio processing system for smart homes

CN122578352APending Publication Date: 2026-08-14GUANGZHOU TCL IND RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,现有的语音识别技术通常直接对获取的音频数据进行语音识别,由于获取的音频数据包含各种环境噪声,导致语音识别的准确率低

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122578352A_ABST
    Figure CN122578352A_ABST
Patent Text Reader

Abstract

This application discloses an audio processing system for smart homes, which determines target feature information based on the audio data to be processed; and determines target data information based on the target feature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically to an audio processing system. Background Technology

[0002] With the rapid development of artificial intelligence and speech recognition technology, voice interaction has been widely used in the smart home field. However, existing speech recognition technologies typically perform speech recognition directly on the acquired audio data. Because the acquired audio data contains various environmental noises, the accuracy of speech recognition is low. Summary of the Invention

[0003] This application provides an audio processing system.

[0004] In a first aspect, this application provides a method comprising:

[0005] Obtain the audio data to be processed;

[0006] Based on the audio data to be processed, determine the target feature information;

[0007] Based on the target feature information, the target data information is determined.

[0008] Secondly, this application provides a system comprising:

[0009] The data acquisition module is used to acquire audio data information to be processed;

[0010] The first determining module is used to determine target feature information based on the audio data to be processed;

[0011] The second determination module is used to determine target data information based on target feature information.

[0012] Thirdly, this application also provides an apparatus, the apparatus comprising:

[0013] One or more processors;

[0014] Memory; and

[0015] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the methods of any one of the first aspects.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the method in any of the first aspects. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a scene of the audio processing system provided in an embodiment of the present invention;

[0019] Figure 2 This is a flowchart of an embodiment of the audio processing method provided by the present invention;

[0020] Figure 3 This is a flowchart illustrating a specific embodiment of determining target feature information provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the structure of the audio enhancement model provided in an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram of the structure of the second feature extraction module provided in an embodiment of the present invention;

[0023] Figure 6 This is a schematic diagram of the structure of the third feature extraction module provided in an embodiment of the present invention;

[0024] Figure 7 This is a schematic diagram of the structure of the second feature extraction model provided in an embodiment of the present invention;

[0025] Figure 8 This is a schematic diagram of the structure of the third feature extraction unit provided in an embodiment of the present invention;

[0026] Figure 9 This is a flowchart illustrating a specific embodiment of determining target data information provided in this invention.

[0027] Figure 10 This is a schematic diagram of the feature fusion model provided in an embodiment of the present invention;

[0028] Figure 11 This is a schematic block diagram of the audio processing system provided in the embodiments of the present invention;

[0029] Figure 12 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0031] In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features. In the description of this application, "several" means one or more, unless otherwise stated.

[0032] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.

[0033] This application provides an audio processing method and system. Furthermore, the above-mentioned audio processing method and system can be used for smart home audio processing, which will be described in detail below.

[0034] Please see Figure 1 , Figure 1 This is a schematic diagram of a scenario for an audio processing system provided in an embodiment of this application. The audio processing system may include a computer device 100, such as... Figure 1 Computer equipment in the country.

[0035] In this embodiment, the computer device 100 is mainly used to acquire audio data information to be processed; determine target feature information based on the audio data information to be processed; and determine target data information based on the target feature information, which can improve the accuracy of speech recognition.

[0036] In this embodiment, the computer device 100 can be a standalone server, a server network, or a server cluster. For example, the computer device 100 described in this embodiment includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud server composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0037] It is understood that the computer device 100 used in the embodiments of this application can be a device that includes both receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such a device may include: cellular or other communication devices having a single-line display, a multi-line display, or a cellular or other communication device without a multi-line display. Specifically, the computer device 100 may be a desktop terminal or a mobile terminal, and may also be one of the following: an augmented reality (AR) device, a mobile phone, a tablet computer, a laptop computer, etc.

[0038] Those skilled in the art will understand that Figure 1 The application environment shown is merely one application scenario of the solution in this application and does not constitute a limitation on the application scenario of the solution in this application. Other application environments may include those that are more specific to this application. Figure 1 The number of computer devices shown is more or less, for example Figure 1 Only one computer device is shown in the document. It is understood that the audio processing system may also include one or more other services, which are not limited here.

[0039] In addition, such as Figure 1 As shown, the audio processing system may also include a memory 200 for storing data, such as feature information, such as first feature information, second feature information, etc., and data information, such as audio data information to be processed, target audio data information, etc.

[0040] It should be noted that, Figure 1 The schematic diagram of the audio processing system shown is merely an example. The audio processing system and scenario described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of audio processing systems and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0041] First, this application provides an audio processing method applied to a computer device. The audio processing method includes: acquiring audio data information to be processed; determining target feature information based on the audio data information to be processed; and determining target data information based on the target feature information.

[0042] like Figure 2 The diagram shown is a flowchart of an embodiment of the audio processing method in this application. The audio processing method may include the following steps S201 to S203, as detailed below:

[0043] S201. Obtain the audio data information to be processed.

[0044] In this embodiment, the audio data to be processed is audio data that requires audio recognition. This audio data can be directly acquired by a computer device from other computer devices via a network, Bluetooth, or infrared, or it can be acquired by the computer device through its own configured microphone module. This embodiment does not impose any limitations. For example, when the audio processing method of this application is applied to a smart TV, the smart TV can acquire the audio data to be processed through its own configured microphone; when the audio processing method of this application is applied to a server, the server can acquire the audio data to be processed from other computer devices (e.g., smart TVs) via a network, Bluetooth, or infrared.

[0045] S202. Based on the audio data to be processed, determine the target feature information.

[0046] In this embodiment, the target feature information is a feature representation extracted from the audio data to be processed. Optionally, the target feature information includes first feature information and / or second feature information. The first feature information is the noise-removed audio feature extracted from the audio data to be processed, and the second feature information is the noise feature extracted from the audio data to be processed. Audio processing based on the target feature information allows for the learning of complementary information between the audio features and the noise features, thereby improving the accuracy of speech recognition.

[0047] In some embodiments, refer to Figure 3 As shown, the step S202 above, which determines the target feature information based on the audio data to be processed, may include steps S301 to S302, as follows:

[0048] S301. Enhance the audio data information to be processed to obtain the target audio data information.

[0049] In this embodiment, the target audio data information is the denoised audio representation corresponding to the audio data information to be processed. Compared with the audio data information to be processed, the target audio data information contains less noise. By enhancing the audio data information to be processed, and then performing speech recognition based on the target audio data information obtained from the enhancement process, the audio signal can be enhanced in a targeted manner, better preserving the features and information of the original audio and improving the accuracy of audio recognition.

[0050] In some embodiments, the target audio data information is obtained by enhancing the audio data information to be processed using an audio enhancement model, referring to... Figure 4As shown, the audio enhancement model includes a first feature extraction module, a second feature extraction module, and a first fusion module. The input of the first feature extraction module is configured to receive audio data to be processed. The output of the first feature extraction module is connected to the input of the second feature extraction module, and the outputs of both modules are connected to the input of the first fusion module. The first feature extraction module is configured to extract features from the audio data to be processed; the second feature extraction module is configured to extract features from the feature representation output by the first feature extraction module; and the first fusion module is configured to fuse the feature representations output by the first and second feature extraction modules. This embodiment enhances the audio data to be processed using the audio enhancement model, simultaneously enhancing both the amplitude and phase information. Furthermore, the data enhancement process does not require waveform reconstruction of the audio data, avoiding the phase distortion problem caused by waveform reconstruction algorithms such as Short-Time Fourier Transform (STFT).

[0051] Optionally, continue to refer to Figure 4 As shown, the first feature extraction module includes a first feature extraction unit, a second feature extraction unit, and a first convolution unit. The input of the first feature extraction unit is configured to receive audio data to be processed; the output of the first feature extraction unit is connected to the input of the second feature extraction unit; and the output of the second feature extraction unit is connected to the input of the first convolution unit. The first feature extraction unit is configured to extract features from the audio data to be processed; the second feature extraction unit is configured to extract features from the feature representation output by the first feature extraction unit; and the first convolution unit is configured to perform a convolution operation on the feature representation output by the second feature extraction unit.

[0052] In some embodiments, continue to refer to Figure 4 As shown, the first feature extraction unit includes a first convolutional layer, a first activation layer, and a first normalization layer. The input of the first convolutional layer is configured to receive audio data to be processed, the output of the first convolutional layer is connected to the input of the first activation layer, and the output of the first activation layer is connected to the input of the first normalization layer.

[0053] Furthermore, continue to refer to Figure 4 As shown, the second feature extraction unit includes a second convolutional layer, a second activation layer, and a second normalization layer. The input of the second convolutional layer is configured to receive the feature representation output by the first feature extraction unit. The output of the second convolutional layer is connected to the input of the second activation layer, and the output of the second activation layer is connected to the input of the second normalization layer.

[0054] Optionally, the first convolutional unit, the first convolutional layer, and the second convolutional layer can be constructed based on a 1×1 convolutional kernel, or they can be constructed based on a 3×3 convolutional kernel, or they can be constructed based on a 5×5 convolutional kernel. This embodiment does not impose any limitations.

[0055] In some embodiments, the first convolutional layer can be constructed based on a 3×3 1D convolutional kernel. This first convolutional layer increases the number of channels in the audio data to be processed, thereby capturing the temporal features of the audio data. Simultaneously, padding can be used to maintain the sequence length of the extracted feature representation. Furthermore, the second convolutional layer can be constructed based on a 1D dilated convolutional kernel with a dilation rate of 'd' and a kernel size of 3×3. Feature extraction using the second convolutional layer expands the receptive field of the features, thereby capturing richer temporal contextual information. The first convolutional unit can be constructed based on a 1×1 convolutional kernel. The first convolutional unit projects the feature representation output by the second feature extraction unit onto the latent representation dimension, thereby obtaining the latent representation of the audio data to be processed.

[0056] In some embodiments, the first and second activation layers can be constructed based on the ReLU activation function, the Sigmoid activation function, or the softmax activation function; this embodiment does not impose any limitations. In a specific embodiment, the first and second activation layers can be constructed based on the ReLU activation function.

[0057] In some embodiments, the first and second normalization layers can be constructed based on a batch normalization (BN) layer, a group normalization (GN) layer, or a layer normalization (LN) layer; this embodiment does not impose any limitations. In a specific embodiment, the first and second normalization layers can be constructed based on a layer normalization (LN) layer.

[0058] In some embodiments, refer to Figure 5As shown, the second feature extraction module includes: a first attention unit, a first fusion unit, a first feature processing unit, a first linear unit, and a first activation unit. The input of the first attention unit is configured to receive the feature representation output by the first feature extraction module. The output of the first attention unit is connected to the input of the first fusion unit, the output of the first fusion unit is connected to the input of the first feature processing unit, the output of the first feature processing unit is connected to the input of the first linear unit, and the output of the first linear unit is connected to the input of the first activation unit. The first attention unit is configured to perform attention calculation processing on the feature representation output by the first feature extraction module; the first fusion unit is configured to perform fusion processing on the feature representation output by the first attention unit; the first feature processing unit is configured to perform weighted calculation processing on the feature representation output by the first fusion unit; the first linear unit is configured to perform linear transformation processing on the feature representation output by the first feature processing unit; and the first activation unit is configured to perform nonlinear transformation processing on the feature representation output by the first linear unit.

[0059] In some embodiments, the step of the first attention unit performing attention calculation processing on the feature representation output by the first feature extraction module specifically includes: performing linear mapping processing on the feature representation output by the first feature extraction module to obtain third feature information, fourth feature information, and fifth feature information; the third feature information includes several first feature information to be processed, and the fifth feature information includes several second feature information to be processed; based on the third feature information and the fourth feature information, determining the first weight information corresponding to each first feature information to be processed; and performing weighted summation processing on several second feature information to be processed based on the first weight information to obtain several sixth feature information.

[0060] In this embodiment, linear mapping of the feature representation output by the first feature extraction module refers to the process of linearly mapping the feature representation output by the first feature extraction module through query vector, key vector, and value vector. In a specific embodiment, the process of linear mapping of the feature representation output by the first feature extraction module can be expressed as: Q = MW Q K = MW K V = MW V Where Q represents the third feature information, K represents the fourth feature information, V represents the fifth feature information, M represents the feature representation output by the first feature extraction module, and W represents the fifth feature information. Q W represents the query vector. K W represents the key vector. V Represents a value vector.

[0061] In some embodiments, the fourth feature information includes a plurality of third feature information to be processed. The step of determining the first weight information corresponding to each first feature information to be processed based on the third feature information and the fourth feature information specifically includes: performing calculation processing on each first feature information to be processed and the plurality of third feature information to be processed to obtain a plurality of fourth feature information to be processed corresponding to each first feature information to be processed; performing calculation processing on the plurality of fourth feature information to be processed to obtain fifth feature information to be processed; and performing calculation processing on each fourth feature information to be processed and the fifth feature information to be processed to obtain the first weight information corresponding to each first feature information to be processed.

[0062] Optionally, the process of calculating and processing each first feature information to be processed and several third feature information to be processed can be represented as follows: Among them, e ij q represents the j-th fourth feature information corresponding to the i-th first feature information to be processed. i Let k represent the i-th first feature information to be processed. j d represents the j-th third feature information to be processed. k This represents the feature dimension corresponding to the j-th third feature information to be processed.

[0063] Further, the step of calculating and processing several fourth feature information to obtain fifth feature information specifically includes: performing exponential operations on several fourth feature information respectively; and performing weighted summation on the exponentially processed fourth feature information to obtain fifth feature information. Optionally, the process of calculating and processing several fourth feature information can be represented as follows: Where S1 represents the fifth feature information to be processed, e ij This represents the j-th fourth feature information corresponding to the i-th first feature information to be processed, where n represents the number of fourth feature information to be processed, and λ represents the number of such fourth feature information. j This indicates the first parameter.

[0064] In this embodiment of the application, the first parameter λ j The settings can be adjusted according to actual needs; in some embodiments, 0 < λ. j ≤1, optionally, λ j If = 1, then the process of calculating and processing several fourth feature information to be processed can be represented as:

[0065] In some embodiments, the first weight information includes several weight values, each corresponding to a number of fourth features to be processed. The step of calculating and processing each fourth feature to be processed and a fifth feature to be processed to obtain the first weight information corresponding to each first feature to be processed specifically includes: performing exponential operations on the several fourth features to be processed to obtain several sixth features to be processed; dividing each sixth feature to be processed and a fifth feature to be processed to obtain the weight value corresponding to each fourth feature to be processed; and fusing the weight values ​​corresponding to the several fourth features to be processed to obtain the first weight information corresponding to each first feature to be processed.

[0066] Optionally, the process of determining the weight value corresponding to each fourth feature information to be processed can be expressed as: e ij a represents the j-th fourth feature information corresponding to the i-th first feature information to be processed. ij Let S1 represent the weight value corresponding to the j-th fourth feature information to be processed, and S1 represent the fifth feature information to be processed. Optionally, λ j If = 1, then the process of determining the weight value corresponding to each fourth feature information to be processed can be represented as:

[0067] In some embodiments, the first weight information includes the weight value corresponding to each fourth feature information to be processed, and the process of performing weighted summation on a plurality of second feature information to be processed based on the first weight information can be represented as follows: Among them, P i Let a represent the i-th sixth feature information. ij v represents the weight value corresponding to the j-th fourth feature information to be processed. j Let j represent the j-th second feature information to be processed, and n represent the number of fourth feature information to be processed.

[0068] Furthermore, the fusion processing of the feature representation output by the first attention unit refers to the concatenation of several sixth feature information output by the first attention unit. Multi-head attention calculation is performed on the feature representation output by the first feature extraction module through the first attention unit, and the feature representation output by the first attention unit is fused through the first fusion unit. This can capture information from different representation subspaces and improve the accuracy of speech recognition.

[0069] In some embodiments, the first feature processing unit may be built based on a feedforward neural network (FFN). Through its hierarchical structure, the feedforward neural network can learn the complex relationships and patterns between input data. Each layer transforms and abstracts the input data, gradually extracting higher-level features. This enables the network to capture details and patterns that are difficult for traditional computing models to discover, which can further improve the accuracy of speech recognition.

[0070] Optionally, the first activation unit can be constructed based on the ReLU activation function, the Sigmoid activation function, or the softmax activation function; this embodiment does not impose any limitations. In a specific embodiment, the first activation unit can be constructed based on the softmax activation function. In this embodiment, the first linear unit performs a linear transformation on the feature representation output by the first feature processing unit, converting the feature representation output by the first feature processing unit into an output with the same size as the input. The first activation unit performs a non-linear transformation on the feature representation output by the first linear unit, converting the feature values ​​corresponding to the feature representation output by the first linear unit to between 0 and 1, facilitating subsequent fusion processing of the feature representation output by the first feature extraction module and the feature representation output by the second feature extraction module.

[0071] In some embodiments, the feature representation output by the second feature extraction module is used to distinguish between speech features and noise features in the audio data to be processed. The step of the first fusion module fusing the feature representations output by the first and second feature extraction modules specifically includes: performing element-wise multiplication on the feature representations output by the first and second feature extraction modules. This embodiment, by fusing the feature representations output by the first and second feature extraction modules, can effectively remove noise information while preserving speech features, thereby providing a cleaner speech representation for subsequent speech recognition.

[0072] S302. Based on the audio data information to be processed and the target audio data information, determine the target feature information.

[0073] In some embodiments, the target feature information includes first feature information and / or second feature information. The first feature information is an audio feature extracted from the target audio data information, and the second feature information is a noise feature extracted from the audio data information to be processed. The step of determining the target feature information based on the audio data information to be processed and the target audio data information specifically includes: performing feature extraction on the target audio data information to obtain the first feature information; and / or, performing feature extraction on the audio data information to be processed to obtain the second feature information. This embodiment enhances the audio data information to be processed to obtain the target audio data information, and determines the target feature information based on the target audio data information. This can reduce the interference of noise in the audio data information to be processed and improve the accuracy of the determined target feature information.

[0074] In some embodiments, the first feature information is obtained by extracting features from the target audio data information using a first feature extraction model. The first feature extraction model includes several serially connected third feature extraction modules, as described above. Figure 6 As shown, each third feature extraction module includes: a second convolutional unit, a second activation unit, a first normalization unit, and a first pooling unit. The output of the second convolutional unit is connected to the input of the second activation unit, the output of the second activation unit is connected to the input of the first normalization unit, and the output of the first normalization unit is connected to the input of the first pooling unit. The second convolutional unit is configured to extract features from the feature representation input to the third feature extraction module; the second activation unit is configured to perform nonlinear transformation processing on the feature representation output by the second convolutional unit; the first normalization unit is configured to normalize the feature representation output by the second activation unit; and the first pooling unit is configured to pool the feature representation output by the first normalization unit. In this embodiment, the second convolutional units in several cascaded third feature extraction modules can extract temporal features of the target audio data. Compared to extracting frequency domain features, this avoids the phase distortion problem caused by waveform reconstruction. Furthermore, the first pooling unit can reduce the dimensionality of the features while retaining important information, thereby reducing the computational load during feature extraction.

[0075] Optionally, the second convolutional unit can be constructed based on a 1×1 1D convolutional kernel, or it can be constructed based on a 3×3 1D convolutional kernel, or it can be constructed based on a 5×5 1D convolutional kernel. This embodiment does not impose any limitations.

[0076] Furthermore, the second activation unit can be constructed based on the ReLU activation function, the sigmoid activation function, or the softmax activation function; this embodiment does not impose any limitations. In one specific embodiment, the second activation unit can be constructed based on the ReLU activation function.

[0077] Optionally, the first normalization unit can be constructed based on a batch normalization (BN) layer, a group normalization (GN) layer, or a layer normalization (LN) layer; this embodiment does not impose any limitations. In a specific embodiment, the first normalization unit can be constructed based on a layer normalization (LN) layer.

[0078] Furthermore, the first pooling unit can be constructed based on the max pooling layer, or it can be constructed based on the average pooling layer; this embodiment does not impose any limitations.

[0079] In some embodiments, the second feature information is obtained by feature extraction of the audio data information to be processed using a second feature extraction model, referring to... Figure 7 As shown, the second feature extraction model includes a fourth feature extraction module and a first normalization module. The input of the fourth feature extraction module is configured to receive the audio data to be processed, and the output of the fourth feature extraction module is connected to the input of the first normalization module. The fourth feature extraction module is configured to extract features from the audio data to be processed; the first normalization module is configured to normalize the feature representation output by the fourth feature extraction module.

[0080] Optionally, refer to Figure 7 and Figure 8As shown, the fourth feature extraction module includes several cascaded third feature extraction units. Each third feature extraction unit includes a third convolutional unit, a third activation unit, a second normalization unit, and a second pooling unit. The output of the third convolutional unit is connected to the input of the third activation unit, the output of the third activation unit is connected to the input of the second normalization unit, and the output of the second normalization unit is connected to the input of the second pooling unit. The third convolutional unit is configured to extract features from the feature representation input to the third feature extraction unit; the third activation unit is configured to perform nonlinear transformation processing on the feature representation output by the third convolutional unit; the second normalization unit is configured to normalize the feature representation output by the third activation unit; and the second pooling unit is configured to pool the feature representation output by the second normalization unit. In this embodiment, the third convolutional unit in the cascaded third feature extraction units can extract the temporal features of the audio data to be processed. The second pooling unit can reduce the dimensionality of the features while retaining important information, thereby reducing the computational load during feature extraction.

[0081] Optionally, the third convolutional unit can be constructed based on a 1×1 1D convolutional kernel, or it can be constructed based on a 3×3 1D convolutional kernel, or it can be constructed based on a 5×5 1D convolutional kernel. This embodiment does not impose any limitations.

[0082] Furthermore, the third activation unit can be constructed based on the ReLU activation function, the sigmoid activation function, or the softmax activation function; this embodiment does not impose any limitations. In one specific embodiment, the third activation unit can be constructed based on the ReLU activation function.

[0083] Optionally, the second normalization unit can be constructed based on a batch normalization (BN) layer, a group normalization (GN) layer, or a layer normalization (LN) layer; this embodiment does not impose any limitations. In a specific embodiment, the second normalization unit can be constructed based on a layer normalization (LN) layer.

[0084] Furthermore, the second pooling unit can be constructed based on the max pooling layer, or it can be constructed based on the average pooling layer; this embodiment does not impose any limitations.

[0085] In some embodiments, the feature representation output by the fourth feature extraction module includes a plurality of first feature representations. The step of the first normalization module normalizing the feature representation output by the fourth feature extraction module specifically includes: for any one of the plurality of first feature representations, performing calculation processing on the first feature representation to obtain a first feature value and a second feature value corresponding to the first feature representation; subtracting the first feature representation from the first feature value to obtain a second feature representation corresponding to the first feature representation; adding the second feature value to the second parameter to obtain a third feature value; and dividing the second feature representation from the third feature value to obtain the normalized feature representation corresponding to the first feature representation.

[0086] In this embodiment, the first eigenvalue is used to characterize the mean of the eigenvector corresponding to the first eigenvalue, and the second eigenvalue is used to characterize the standard deviation of the eigenvector corresponding to the first eigenvalue. The process of normalizing the first eigenvalue can be expressed as follows: in, Let n represent the t-th first feature after normalization. t Let μ represent the t-th first feature. t Let σ represent the first feature value corresponding to the t-th first feature. t Let t represent the first feature and t represent the corresponding second feature value. ε is the second parameter, a preset positive constant, which can prevent σ from being induced. t The case where it is zero.

[0087] S203. Determine target data information based on target feature information.

[0088] In this embodiment, the target data information is the speech recognition result information corresponding to the audio data information to be processed determined based on the target feature information. This embodiment determines the target feature information based on the audio data information to be processed and determines the target data information based on the target feature information, which can improve the accuracy of speech recognition.

[0089] In some embodiments, the target feature information includes first feature information and / or second feature information, referring to... Figure 9 As shown, the step S203 above, which determines the target data information based on the target feature information, may include steps S401 to S402, as follows:

[0090] S401. The first feature information and the second feature information are fused to obtain the target feature information.

[0091] In this embodiment, the target feature information is a feature representation obtained by fusing the first feature information and the second feature information. By fusing the first feature information and the second feature information, complementary information can be learned from the first feature information and the second feature information, thereby reducing speech distortion and improving the accuracy of speech recognition.

[0092] In some embodiments, the target feature information is obtained by fusing the first feature information and the second feature information through a feature fusion model, referring to... Figure 10 As shown, the feature fusion model includes a second fusion module, a third fusion module, and a fourth fusion module. The input of the second fusion module is configured to receive first feature information and second feature information; the input of the third fusion module is configured to receive first feature information and second feature information; and the outputs of the second and third fusion modules are respectively connected to the input of the fourth fusion module. The second fusion module is configured to perform fusion processing on the first and second feature information; the third fusion module is configured to perform fusion processing on the first and second feature information; and the fourth fusion module is configured to perform fusion processing on the feature representations output by the second and third fusion modules.

[0093] Optionally, continue to refer to Figure 10 As shown, the second fusion module includes: a second linear unit, a third linear unit, a fourth linear unit, a second attention unit, a second fusion unit, and a fifth linear unit. The input of the second linear unit is configured to receive first feature information; the inputs of the third and fourth linear units are respectively configured to receive second feature information; the outputs of the second, third, and fourth linear units are respectively connected to the input of the second attention unit; the output of the second attention unit is connected to the input of the second fusion unit; and the output of the second fusion unit is connected to the input of the fifth linear unit.

[0094] Furthermore, the second linear unit is configured to perform linear transformation processing on the first feature information; the third and fourth linear units are respectively configured to perform linear transformation processing on the second feature information; the second attention unit is configured to perform attention calculation processing on the feature representations output by the second linear unit, the feature representations output by the third linear unit, and the feature representations output by the fourth linear unit; the second fusion unit is configured to perform fusion processing on the feature representations output by the second attention unit; and the fifth linear unit is configured to perform linear transformation processing on the feature representations output by the second fusion unit.

[0095] In this embodiment, linear transformation of the first feature information refers to the process of linearly mapping the first feature information through a query vector, and linear transformation of the second feature information refers to the process of linearly mapping the second feature information through a key vector and a value vector. In a specific embodiment, the process of performing attention calculation on the feature representations output by the second linear unit, the feature representations output by the third linear unit, and the feature representations output by the fourth linear unit can be represented as follows: in, Z represents the i-th feature representation output by the second attention unit. en Z represents the first feature information. nosiy Indicates the second feature information. Represents the query vector. Represents the key vector. This represents a value vector, and Attention(·) represents attention calculation.

[0096] Furthermore, performing a linear transformation on the feature representation output by the second fusion unit refers to the process of linearly mapping the feature representation output by the second fusion unit using a preset first vector. Optionally, the process of performing a linear transformation on the feature representation output by the second fusion unit can be expressed as: W represents the i-th feature representation output by the second attention unit. 1 This represents the first vector, and Concat(·) represents the fusion process.

[0097] In some embodiments, continue to refer to Figure 10 As shown, the third fusion module includes: a sixth linear unit, a seventh linear unit, an eighth linear unit, a third attention unit, a third fusion unit, and a ninth linear unit. The input of the sixth linear unit is configured to receive second feature information; the inputs of the seventh and eighth linear units are respectively configured to receive first feature information; the outputs of the sixth, seventh, and eighth linear units are respectively connected to the input of the third attention unit; the output of the third attention unit is connected to the input of the third fusion unit; and the output of the third fusion unit is connected to the input of the ninth linear unit.

[0098] Furthermore, the sixth linear unit is configured to perform linear transformation processing on the second feature information; the seventh and eighth linear units are respectively configured to perform linear transformation processing on the first feature information; the third attention unit is configured to perform attention calculation processing on the feature representations output by the sixth, seventh, and eighth linear units; the third fusion unit is configured to perform fusion processing on the feature representations output by the third attention unit; and the ninth linear unit is configured to perform linear transformation processing on the feature representations output by the third fusion unit.

[0099] In this embodiment, linear transformation of the second feature information refers to the process of linearly mapping the second feature information through the query vector, and linear transformation of the first feature information refers to the process of linearly mapping the first feature information through the key vector and the value vector. In a specific embodiment, the process of performing attention calculation on the feature representations output by the sixth linear unit, the feature representations output by the seventh linear unit, and the feature representations output by the eighth linear unit can be represented as follows: in, Z represents the i-th feature representation output by the third attention unit. en Z represents the first feature information. nosiy Indicates the second feature information. Represents the query vector. Represents the key vector. This represents a value vector, and Attention(·) represents attention calculation.

[0100] Furthermore, performing a linear transformation on the feature representation output by the third fusion unit refers to the process of linearly mapping the feature representation output by the third fusion unit using a preset second vector. Optionally, the process of performing a linear transformation on the feature representation output by the third fusion unit can be expressed as:

[0101] W represents the i-th feature representation output by the third attention unit. 2 This represents the second vector, and Concat(·) indicates the fusion process.

[0102] In some embodiments, the process by which the fourth fusion module fuses the feature representations output by the second fusion module and the feature representations output by the third fusion module can be represented as: Z fusion =Linear(Multihead(Z) en Z nosiy Z nosiy ))+Linear(Multihead(Z nosiy Z en Zen Z fusion The feature representation output by the fourth fusion module is represented by Linear(Multihead(Z). en Z nosiy Z nosiy Linear(Multihead(Z)) represents the feature representation output by the second fusion module. nosiy Z en Z en )) represents the feature representation output by the third fusion module.

[0103] S402. Use an audio recognition model to identify and process the target feature information to obtain target data information.

[0104] In this embodiment, the audio recognition model is a network model used for audio recognition. The audio recognition model can be built based on a Conformer model, a Transformer-Transducer model, or even a Transformer model; this embodiment does not limit the specific model. In one particular embodiment, the audio recognition model can be built based on a Conformer model. The Conformer model is an attention mechanism model. By adding convolutional layers, the Conformer model can increase its ability to model local information, thereby automatically capturing the contextual information of the audio signal and matching the contextual information with the corresponding content in the text, thus improving the accuracy of speech recognition.

[0105] In some embodiments, the audio enhancement model and the speech recognition model can share parameters, and the audio enhancement model, the first feature extraction model, the second feature extraction model, the feature fusion model, and the speech recognition model can be jointly trained using the same training data. By jointly training the audio enhancement model, the first feature extraction model, the second feature extraction model, the feature fusion model, and the speech recognition model, the speech distortion problem caused by speech enhancement can be avoided, and the accuracy of speech recognition can be further improved.

[0106] In some embodiments, before determining the target feature information based on the audio data to be processed in step S202, the process includes: acquiring a training sample set, which includes several first sample audio data, label audio data corresponding to each first sample audio data, and real recognition result information corresponding to each first sample audio data; inputting several first sample audio data into an audio enhancement model, and outputting second sample audio data corresponding to each first sample audio data through the audio enhancement model; inputting the second sample audio data and the first sample audio data into a first feature extraction model and a second feature extraction model respectively for feature extraction to obtain first sample feature information and second sample feature information; and inputting the first sample feature information and the second sample audio data into a first feature extraction model and a second feature extraction model respectively for feature extraction to obtain first sample feature information and second sample feature information; and inputting the first sample feature information and the second sample audio data into a first feature extraction model and a second feature extraction model respectively for feature extraction. The first feature information is input into the feature fusion model for fusion processing to obtain the third sample feature information. This third sample feature information is then input into the speech recognition model, which outputs the predicted recognition result information corresponding to each first sample audio data information. Based on the labeled audio data information, the second sample audio data information, the actual recognition result information, and the predicted recognition result information, a target loss value is determined. If the target loss value does not meet the first condition, the model parameters of the audio enhancement model, the first feature extraction model, the second feature extraction model, the feature fusion model, and the speech recognition model are adjusted based on a preset learning rate. The process of inputting several first sample audio data information into the audio enhancement model continues until the target loss value meets the first condition or the number of model updates reaches a threshold. Meeting the first condition can be defined as the target loss value being less than the first threshold or the difference between two consecutive target loss values ​​being less than the second threshold; this embodiment does not impose any limitations on this.

[0107] In some embodiments, the step of determining the target loss value based on tag audio data information, second sample audio data information, actual recognition result information, and predicted recognition result information specifically includes: determining a first loss value based on the actual recognition result information and the predicted recognition result information; determining a second loss value based on the tag audio data information and the second sample audio data information; and performing a weighted summation of the first loss value and the second loss value to obtain the target loss value. Optionally, the process of determining the target loss value can be expressed as: L = β·L ASR +α·L Enh Where L represents the target loss value, L ASR L represents the first loss value. Enh Let α represent the first parameter and β represent the second parameter. The first parameter α and the second parameter β can be set according to actual needs. In some embodiments, 0 < β ≤ 1; optionally, β = 1. Then, the process of determining the target loss value can be expressed as: L = L ASR +α·L Enh .

[0108] Furthermore, L ASR =-InP(S * |Z fusion ), where InP(·) represents the conditional probability distribution, S * Z represents the actual recognition result information. fusion L represents the feature information of the third sample. ASR The calculation is based on the speech recognition model Z. fusion The difference between the predicted recognition result and the actual recognition result. Optionally, L Enh The mean squared error loss function can be used, then M represents the number of audio data points in the first sample, y i This represents the tag audio data information corresponding to the i-th first sample audio data information. This represents the second sample audio data information corresponding to the i-th first sample audio data information.

[0109] Optionally, the training sample set can utilize existing publicly available datasets, such as the LibriSpeech and CHiME-4 datasets, which contain clean speech as well as speech with different noise types and signal-to-noise ratios. Furthermore, the audio enhancement model, first feature extraction model, second feature extraction model, feature fusion model, and speech recognition model can be jointly trained on a computer device configured with an NVIDIA A100 GPU and an Intel Core i9-12900K CPU. This significantly accelerates the training speed and improves the model's convergence speed and accuracy. Additionally, this computer device can utilize 128GB of DDR4 memory and a solid-state drive (SSD). DDR4 memory ensures that more data and tasks can be processed simultaneously during training, thereby improving training efficiency, while the SSD offers faster read and write speeds, significantly reducing data read and write time.

[0110] In some embodiments, the threshold for the number of training iterations can be set to 150 to 250, optionally with a threshold of 200. The initial learning rate during model training is 0.001, and the learning rate can gradually decrease according to the cosine function as the number of training iterations increases.

[0111] In some embodiments, after determining the target data information based on the target feature information in step S203 above, the process may include: determining a target control instruction based on the target data information; and executing the control operation corresponding to the target control instruction. For example, if the audio processing method is applied to a smart TV and the target control instruction is to increase the volume, the smart TV will automatically increase the volume of the currently playing media data.

[0112] In summary, the audio processing method provided in this implementation scheme acquires the audio data to be processed, determines target feature information based on the audio data, and then determines target data information based on the target feature information. This scheme reduces noise interference in the audio data and improves the accuracy of speech recognition. Furthermore, by using an audio enhancement model to enhance the audio data, both amplitude and phase information can be enhanced simultaneously. Moreover, waveform reconstruction is unnecessary during the data enhancement process, avoiding phase distortion problems caused by waveform reconstruction algorithms. Even further, feature extraction is performed on the audio data to be processed and the target audio data to obtain first and second feature information respectively. The first and second feature information are then fused to obtain the target feature information. Complementary information can be learned from noise and audio information, reducing speech distortion and further improving the accuracy of speech recognition.

[0113] To better implement the audio processing method in the embodiments of this application, an audio processing system is also provided in the embodiments of this application, such as... Figure 11 As shown, the audio processing system 600 includes:

[0114] Data acquisition module 610 is used to acquire audio data information to be processed;

[0115] The first determining module 620 is used to determine target feature information based on the audio data to be processed;

[0116] The second determining module 630 is used to determine target data information based on target feature information.

[0117] In this embodiment, the target feature information is determined based on the audio data information to be processed, and then the target data information is determined based on the target feature information. This can reduce noise interference in the audio data information to be processed and improve the accuracy of speech recognition.

[0118] In some embodiments of this application, the first determining module 620 determines target feature information based on the audio data to be processed, including:

[0119] The audio data to be processed is enhanced to obtain the target audio data.

[0120] Based on the audio data to be processed and the target audio data, the target feature information is determined.

[0121] In some embodiments of this application, the target audio data information is obtained by enhancing the audio data information to be processed through an audio enhancement model, which includes: a first feature extraction module, a second feature extraction module, and a first fusion module;

[0122] The input of the first feature extraction module is configured to receive audio data information to be processed. The output of the first feature extraction module is connected to the input of the second feature extraction module. The outputs of the first feature extraction module and the second feature extraction module are connected to the input of the first fusion module.

[0123] The first feature extraction module is configured to extract features from the audio data to be processed.

[0124] The second feature extraction module is configured to extract features from the feature representation output by the first feature extraction module.

[0125] The first fusion module is configured to fuse the feature representations output by the first feature extraction module and the feature representations output by the second feature extraction module.

[0126] In some embodiments of this application, the first feature extraction module includes: a first feature extraction unit, a second feature extraction unit, and a first convolution unit;

[0127] The input of the first feature extraction unit is configured to receive audio data information to be processed, the output of the first feature extraction unit is connected to the input of the second feature extraction unit, and the output of the second feature extraction unit is connected to the input of the first convolution unit.

[0128] The first feature extraction unit is configured to extract features from the audio data to be processed;

[0129] The second feature extraction unit is configured to extract features from the feature representation output by the first feature extraction unit;

[0130] The first convolutional unit is configured to perform a convolution operation on the feature representation output by the second feature extraction unit.

[0131] In some embodiments of this application, the first feature extraction unit includes: a first convolutional layer, a first activation layer, and a first normalization layer;

[0132] The input of the first convolutional layer is configured to receive audio data to be processed; the output of the first convolutional layer is connected to the input of the first activation layer; and the output of the first activation layer is connected to the input of the first normalization layer; and / or,

[0133] The second feature extraction unit includes: a second convolutional layer, a second activation layer, and a second normalization layer;

[0134] The input of the second convolutional layer is configured to receive the feature representation output by the first feature extraction unit, the output of the second convolutional layer is connected to the input of the second activation layer, and the output of the second activation layer is connected to the input of the second normalization layer.

[0135] In some embodiments of this application, the second feature extraction module includes: a first attention unit, a first fusion unit, a first feature processing unit, a first linear unit, and a first activation unit;

[0136] The input of the first attention unit is configured to receive the feature representation output by the first feature extraction module. The output of the first attention unit is connected to the input of the first fusion unit. The output of the first fusion unit is connected to the input of the first feature processing unit. The output of the first feature processing unit is connected to the input of the first linear unit. The output of the first linear unit is connected to the input of the first activation unit.

[0137] The first attention unit is configured to perform attention calculation processing on the feature representation output by the first feature extraction module;

[0138] The first fusion unit is configured to perform fusion processing on the feature representation output by the first attention unit;

[0139] The first feature processing unit is configured to perform weighted calculation processing on the feature representation output by the first fusion unit;

[0140] The first linear unit is configured to perform a linear transformation on the feature representation output by the first feature processing unit;

[0141] The first activation unit is configured to perform a nonlinear transformation on the feature representation output by the first linear unit.

[0142] In some embodiments of this application, the target feature information includes first feature information and / or second feature information. The first determining module 620 determines the target feature information based on the audio data information to be processed and the target audio data information, including:

[0143] Feature extraction is performed on the target audio data to obtain the first feature information; and / or,

[0144] Feature extraction is performed on the audio data to be processed to obtain the second feature information.

[0145] In some embodiments of this application, the first feature information is obtained by extracting features from the target audio data information through a first feature extraction model. The first feature extraction model includes several third feature extraction modules connected in series. Each third feature extraction module includes: a second convolution unit, a second activation unit, a first normalization unit, and a first pooling unit.

[0146] The output of the second convolutional unit is connected to the input of the second activation unit, the output of the second activation unit is connected to the input of the first normalization unit, and the output of the first normalization unit is connected to the input of the first pooling unit.

[0147] The second convolutional unit is configured to extract features from the feature representation input to the third feature extraction module;

[0148] The second activation unit is configured to perform a nonlinear transformation on the feature representation output by the second convolutional unit;

[0149] The first normalization unit is configured to normalize the feature representation output by the second activation unit;

[0150] The first pooling unit is configured to perform pooling processing on the feature representation output by the first normalization unit.

[0151] In some embodiments of this application, the second feature information is obtained by extracting features from the audio data information to be processed using a second feature extraction model. The second feature extraction model includes: a fourth feature extraction module and a first normalization module.

[0152] The input of the fourth feature extraction module is configured to receive audio data information to be processed, and the output of the fourth feature extraction module is connected to the input of the first normalization module.

[0153] The fourth feature extraction module is configured to extract features from the audio data to be processed;

[0154] The first normalization module is configured to normalize the feature representation output by the fourth feature extraction module.

[0155] In some embodiments of this application, the fourth feature extraction module includes a plurality of third feature extraction units connected in series, each third feature extraction unit including: a third convolution unit, a third activation unit, a second normalization unit, and a second pooling unit;

[0156] The output of the third convolutional unit is connected to the input of the third activation unit, the output of the third activation unit is connected to the input of the second normalization unit, and the output of the second normalization unit is connected to the input of the second pooling unit.

[0157] The third convolutional unit is configured to extract features from the feature representation input to the third feature extraction unit;

[0158] The third activation unit is configured to perform a nonlinear transformation on the feature representation output by the third convolutional unit;

[0159] The second normalization unit is configured to normalize the feature representation output by the third activation unit;

[0160] The second pooling unit is configured to perform pooling processing on the feature representation output by the second normalization unit.

[0161] In some embodiments of this application, the target feature information includes first feature information and / or second feature information. The second determining module 630 determines target data information based on the target feature information, including:

[0162] The first and second feature information are fused to obtain the target feature information;

[0163] The target feature information is identified and processed using an audio recognition model to obtain target data information.

[0164] In some embodiments of this application, target feature information is obtained by fusing first feature information and second feature information through a feature fusion model. The feature fusion model includes: a second fusion module, a third fusion module, and a fourth fusion module.

[0165] The input terminal of the second fusion module is configured to receive the first feature information and the second feature information, the input terminal of the third fusion module is configured to receive the first feature information and the second feature information, and the output terminals of the second fusion module and the third fusion module are respectively connected to the input terminal of the fourth fusion module.

[0166] The second fusion module is configured to fuse the first feature information and the second feature information.

[0167] The third fusion module is configured to fuse the first feature information and the second feature information;

[0168] The fourth fusion module is configured to perform fusion processing on the feature representations output by the second fusion module and the feature representations output by the third fusion module.

[0169] In some embodiments of this application, the second fusion module includes: a second linear unit, a third linear unit, a fourth linear unit, a second attention unit, a second fusion unit, and a fifth linear unit;

[0170] The input of the second linear unit is configured to receive first feature information, the inputs of the third linear unit and the fourth linear unit are respectively configured to receive second feature information, the outputs of the second linear unit, the third linear unit and the fourth linear unit are respectively connected to the input of the second attention unit, the output of the second attention unit is connected to the input of the second fusion unit, and the output of the second fusion unit is connected to the input of the fifth linear unit.

[0171] The second linear unit is configured to perform a linear transformation on the first feature information;

[0172] The third and fourth linear units are respectively configured to perform linear transformation processing on the second feature information;

[0173] The second attention unit is configured to perform attention computation processing on the feature representations output by the second linear unit, the feature representations output by the third linear unit, and the feature representations output by the fourth linear unit.

[0174] The second fusion unit is configured to perform fusion processing on the feature representation output by the second attention unit;

[0175] The fifth linear unit is configured to perform a linear transformation on the feature representation output by the second fusion unit.

[0176] In some embodiments of this application, the third fusion module includes: a sixth linear unit, a seventh linear unit, an eighth linear unit, a third attention unit, a third fusion unit, and a ninth linear unit;

[0177] The input of the sixth linear unit is configured to receive second feature information, the inputs of the seventh and eighth linear units are configured to receive first feature information, the outputs of the sixth, seventh and eighth linear units are connected to the input of the third attention unit, the output of the third attention unit is connected to the input of the third fusion unit, and the output of the third fusion unit is connected to the input of the ninth linear unit.

[0178] The sixth linear unit is configured to perform a linear transformation on the second feature information;

[0179] The seventh and eighth linear units are configured to perform linear transformation processing on the first feature information;

[0180] The third attention unit is configured to perform attention computation processing on the feature representations output by the sixth linear unit, the seventh linear unit, and the eighth linear unit.

[0181] The third fusion unit is configured to perform fusion processing on the feature representation output by the third attention unit;

[0182] The ninth linear unit is configured to perform a linear transformation on the feature representation output by the third fusion unit.

[0183] In some embodiments of this application, after the second determining module 630 determines the target data information based on the target feature information, the second determining module 630 is further configured to:

[0184] Based on the target data information, determine the target control commands;

[0185] Execute the control operation corresponding to the target control command.

[0186] This application embodiment also provides a computer device, the computer device including:

[0187] One or more processors;

[0188] Memory; and

[0189] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the audio processing method in any of the embodiments described above.

[0190] This application also provides a computer device, such as... Figure 12 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:

[0191] The computer device may include components such as a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, and an input unit 804. Those skilled in the art will understand that... Figure 12 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0192] The processor 801 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 802, and by calling data stored in the memory 802, thereby providing overall monitoring of the computer device. Optionally, the processor 801 may include one or more processing cores; optionally, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 801.

[0193] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0194] The computer device also includes a power supply 803 that supplies power to the various components. Optionally, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 803 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0195] The computer device may also include an input unit 804, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0196] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 runs the application programs stored in the memory 802 to realize various functions, as follows:

[0197] Obtain the audio data to be processed;

[0198] Based on the audio data to be processed, determine the target feature information;

[0199] Based on the target feature information, the target data information is determined.

[0200] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0201] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the audio processing methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps:

[0202] Obtain the audio data to be processed;

[0203] Based on the audio data to be processed, determine the target feature information;

[0204] Based on the target feature information, the target data information is determined.

[0205] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0206] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0207] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0208] The above provides a detailed description of an audio processing method and system provided by the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method, characterized in that, include: Obtain the audio data to be processed; Based on the audio data to be processed, the target feature information is determined; Based on the target feature information, the target data information is determined.

2. The method according to claim 1, characterized in that, The step of determining target feature information based on the audio data to be processed includes: The audio data to be processed is enhanced to obtain the target audio data. Based on the audio data to be processed and the target audio data, target feature information is determined.

3. The method according to claim 2, characterized in that, The target audio data information is obtained by enhancing the audio data information to be processed through an audio enhancement model, which includes: a first feature extraction module, a second feature extraction module, and a first fusion module. The input terminal of the first feature extraction module is configured to receive the audio data information to be processed, the output terminal of the first feature extraction module is connected to the input terminal of the second feature extraction module, and the output terminals of the first feature extraction module and the second feature extraction module are connected to the input terminal of the first fusion module. The first feature extraction module is configured to extract features from the audio data information to be processed; The second feature extraction module is configured to extract features from the feature representation output by the first feature extraction module; The first fusion module is configured to perform fusion processing on the feature representation output by the first feature extraction module and the feature representation output by the second feature extraction module.

4. The method according to claim 3, characterized in that, The first feature extraction module includes: a first feature extraction unit, a second feature extraction unit, and a first convolution unit; Wherein, the input end of the first feature extraction unit is configured to receive the audio data information to be processed, the output end of the first feature extraction unit is connected to the input end of the second feature extraction unit, and the output end of the second feature extraction unit is connected to the input end of the first convolution unit; The first feature extraction unit is configured to extract features from the audio data information to be processed; The second feature extraction unit is configured to extract features from the feature representation output by the first feature extraction unit; The first convolutional unit is configured to perform a convolution operation on the feature representation output by the second feature extraction unit.

5. The method according to claim 4, characterized in that, The first feature extraction unit includes: a first convolutional layer, a first activation layer, and a first normalization layer; Wherein, the input of the first convolutional layer is configured to receive the audio data to be processed, the output of the first convolutional layer is connected to the input of the first activation layer, and the output of the first activation layer is connected to the input of the first normalization layer; and / or, The second feature extraction unit includes: a second convolutional layer, a second activation layer, and a second normalization layer; The input of the second convolutional layer is configured to receive the feature representation output by the first feature extraction unit, the output of the second convolutional layer is connected to the input of the second activation layer, and the output of the second activation layer is connected to the input of the second normalization layer.

6. The method according to claim 3, characterized in that, The second feature extraction module includes: a first attention unit, a first fusion unit, a first feature processing unit, a first linear unit, and a first activation unit; Wherein, the input end of the first attention unit is configured to receive the feature representation output by the first feature extraction module, the output end of the first attention unit is connected to the input end of the first fusion unit, the output end of the first fusion unit is connected to the input end of the first feature processing unit, the output end of the first feature processing unit is connected to the input end of the first linear unit, and the output end of the first linear unit is connected to the input end of the first activation unit. The first attention unit is configured to perform attention calculation processing on the feature representation output by the first feature extraction module; The first fusion unit is configured to perform fusion processing on the feature representation output by the first attention unit; The first feature processing unit is configured to perform weighted calculation processing on the feature representation output by the first fusion unit; The first linear unit is configured to perform a linear transformation on the feature representation output by the first feature processing unit; The first activation unit is configured to perform nonlinear transformation processing on the feature representation output by the first linear unit.

7. The method according to claim 2, characterized in that, The target feature information includes first feature information and / or, second feature information; The step of determining target feature information based on the audio data to be processed and the target audio data includes: Feature extraction is performed on the target audio data to obtain first feature information; And / or, Feature extraction is performed on the audio data to be processed to obtain second feature information.

8. The method according to claim 7, characterized in that, The first feature information is obtained by extracting features from the target audio data information through a first feature extraction model. The first feature extraction model includes several serially connected third feature extraction modules. Each third feature extraction module includes: a second convolution unit, a second activation unit, a first normalization unit, and a first pooling unit. Wherein, the output of the second convolutional unit is connected to the input of the second activation unit, the output of the second activation unit is connected to the input of the first normalization unit, and the output of the first normalization unit is connected to the input of the first pooling unit. The second convolutional unit is configured to extract features from the feature representation input to the third feature extraction module; The second activation unit is configured to perform a nonlinear transformation on the feature representation output by the second convolutional unit; The first normalization unit is configured to normalize the feature representation output by the second activation unit; The first pooling unit is configured to perform pooling processing on the feature representation output by the first normalization unit.

9. The method according to claim 7, characterized in that, The second feature information is obtained by extracting features from the audio data information to be processed through a second feature extraction model. The second feature extraction model includes a fourth feature extraction module and a first normalization module. The input terminal of the fourth feature extraction module is configured to receive the audio data information to be processed, and the output terminal of the fourth feature extraction module is connected to the input terminal of the first normalization module. The fourth feature extraction module is configured to extract features from the audio data information to be processed. The first normalization module is configured to normalize the feature representation output by the fourth feature extraction module.

10. The method according to claim 9, characterized in that, The fourth feature extraction module includes several cascaded third feature extraction units, each of which includes: a third convolution unit, a third activation unit, a second normalization unit, and a second pooling unit. Wherein, the output of the third convolutional unit is connected to the input of the third activation unit, the output of the third activation unit is connected to the input of the second normalization unit, and the output of the second normalization unit is connected to the input of the second pooling unit. The third convolutional unit is configured to extract features from the feature representation input to the third feature extraction unit; The third activation unit is configured to perform a nonlinear transformation on the feature representation output by the third convolution unit; The second normalization unit is configured to normalize the feature representation output by the third activation unit; The second pooling unit is configured to perform pooling processing on the feature representation output by the second normalization unit.

11. The method according to claim 1, characterized in that, The target feature information includes first feature information and / or, second feature information; The step of determining target data information based on the target feature information includes: The first feature information and the second feature information are fused to obtain the target feature information; The target feature information is identified and processed using an audio recognition model to obtain target data information.

12. The method according to claim 11, characterized in that, The target feature information is obtained by fusing the first feature information and the second feature information through a feature fusion model, which includes a second fusion module, a third fusion module and a fourth fusion module. The input terminal of the second fusion module is configured to receive the first feature information and the second feature information, the input terminal of the third fusion module is configured to receive the first feature information and the second feature information, and the output terminals of the second fusion module and the third fusion module are respectively connected to the input terminal of the fourth fusion module. The second fusion module is configured to perform fusion processing on the first feature information and the second feature information; The third fusion module is configured to perform fusion processing on the first feature information and the second feature information; The fourth fusion module is configured to perform fusion processing on the feature representations output by the second fusion module and the feature representations output by the third fusion module.

13. The method according to claim 12, characterized in that, The second fusion module includes: a second linear unit, a third linear unit, a fourth linear unit, a second attention unit, a second fusion unit, and a fifth linear unit; Wherein, the input terminal of the second linear unit is configured to receive the first feature information, the input terminals of the third linear unit and the fourth linear unit are respectively configured to receive the second feature information, the output terminals of the second linear unit, the third linear unit and the fourth linear unit are respectively connected to the input terminal of the second attention unit, the output terminal of the second attention unit is connected to the input terminal of the second fusion unit, and the output terminal of the second fusion unit is connected to the input terminal of the fifth linear unit; The second linear unit is configured to perform a linear transformation on the first feature information; The third linear unit and the fourth linear unit are respectively configured to perform linear transformation processing on the second feature information; The second attention unit is configured to perform attention calculation processing on the feature representations output by the second linear unit, the feature representations output by the third linear unit, and the feature representations output by the fourth linear unit; The second fusion unit is configured to perform fusion processing on the feature representation output by the second attention unit; The fifth linear unit is configured to perform a linear transformation on the feature representation output by the second fusion unit.

14. The method according to claim 12, characterized in that, The third fusion module includes: a sixth linear unit, a seventh linear unit, an eighth linear unit, a third attention unit, a third fusion unit, and a ninth linear unit; Wherein, the input terminal of the sixth linear unit is configured to receive the second feature information, the input terminals of the seventh linear unit and the eighth linear unit are respectively configured to receive the first feature information, the output terminals of the sixth linear unit, the seventh linear unit and the eighth linear unit are respectively connected to the input terminal of the third attention unit, the output terminal of the third attention unit is connected to the input terminal of the third fusion unit, and the output terminal of the third fusion unit is connected to the input terminal of the ninth linear unit; The sixth linear unit is configured to perform a linear transformation on the second feature information; The seventh linear unit and the eighth linear unit are respectively configured to perform linear transformation processing on the first feature information; The third attention unit is configured to perform attention calculation processing on the feature representations output by the sixth linear unit, the seventh linear unit, and the eighth linear unit; The third fusion unit is configured to perform fusion processing on the feature representation output by the third attention unit; The ninth linear unit is configured to perform a linear transformation on the feature representation output by the third fusion unit.

15. The method according to any one of claims 1 to 14, characterized in that, After determining the target data information based on the target feature information, the process includes: Based on the target data information, the target control command is determined; Execute the control operation corresponding to the target control command.

16. A system, characterized in that, include: The data acquisition module is used to acquire audio data information to be processed; The first determining module is used to determine target feature information based on the audio data to be processed; The second determining module is used to determine target data information based on the target feature information; Optionally, the first determining module determines target feature information based on the audio data to be processed, including: The audio data to be processed is enhanced to obtain the target audio data. Based on the audio data to be processed and the target audio data, target feature information is determined; Optionally, the target audio data information is obtained by enhancing the audio data information to be processed through an audio enhancement model, wherein the audio enhancement model includes: a first feature extraction module, a second feature extraction module, and a first fusion module; The input terminal of the first feature extraction module is configured to receive the audio data information to be processed, the output terminal of the first feature extraction module is connected to the input terminal of the second feature extraction module, and the output terminals of the first feature extraction module and the second feature extraction module are connected to the input terminal of the first fusion module. The first feature extraction module is configured to extract features from the audio data information to be processed; The second feature extraction module is configured to extract features from the feature representation output by the first feature extraction module; The first fusion module is configured to perform fusion processing on the feature representation output by the first feature extraction module and the feature representation output by the second feature extraction module; Optionally, the first feature extraction module includes: a first feature extraction unit, a second feature extraction unit, and a first convolution unit; Wherein, the input end of the first feature extraction unit is configured to receive the audio data information to be processed, the output end of the first feature extraction unit is connected to the input end of the second feature extraction unit, and the output end of the second feature extraction unit is connected to the input end of the first convolution unit; The first feature extraction unit is configured to extract features from the audio data information to be processed; The second feature extraction unit is configured to extract features from the feature representation output by the first feature extraction unit; The first convolutional unit is configured to perform a convolution operation on the feature representation output by the second feature extraction unit; Optionally, the first feature extraction unit includes: a first convolutional layer, a first activation layer, and a first normalization layer; Wherein, the input of the first convolutional layer is configured to receive the audio data to be processed, the output of the first convolutional layer is connected to the input of the first activation layer, and the output of the first activation layer is connected to the input of the first normalization layer; and / or, The second feature extraction unit includes: a second convolutional layer, a second activation layer, and a second normalization layer; The input of the second convolutional layer is configured to receive the feature representation output by the first feature extraction unit, the output of the second convolutional layer is connected to the input of the second activation layer, and the output of the second activation layer is connected to the input of the second normalization layer. Optionally, the second feature extraction module includes: a first attention unit, a first fusion unit, a first feature processing unit, a first linear unit, and a first activation unit; Wherein, the input end of the first attention unit is configured to receive the feature representation output by the first feature extraction module, the output end of the first attention unit is connected to the input end of the first fusion unit, the output end of the first fusion unit is connected to the input end of the first feature processing unit, the output end of the first feature processing unit is connected to the input end of the first linear unit, and the output end of the first linear unit is connected to the input end of the first activation unit. The first attention unit is configured to perform attention calculation processing on the feature representation output by the first feature extraction module; The first fusion unit is configured to perform fusion processing on the feature representation output by the first attention unit; The first feature processing unit is configured to perform weighted calculation processing on the feature representation output by the first fusion unit; The first linear unit is configured to perform a linear transformation on the feature representation output by the first feature processing unit; The first activation unit is configured to perform a nonlinear transformation on the feature representation output by the first linear unit; Optionally, the target feature information includes first feature information and / or second feature information. The first determining module determines the target feature information based on the audio data information to be processed and the target audio data information, including: Feature extraction is performed on the target audio data to obtain first feature information; and / or, Feature extraction is performed on the audio data to be processed to obtain second feature information; Optionally, the first feature information is obtained by extracting features from the target audio data information through a first feature extraction model. The first feature extraction model includes several third feature extraction modules connected in series. Each third feature extraction module includes: a second convolution unit, a second activation unit, a first normalization unit, and a first pooling unit. Wherein, the output of the second convolutional unit is connected to the input of the second activation unit, the output of the second activation unit is connected to the input of the first normalization unit, and the output of the first normalization unit is connected to the input of the first pooling unit. The second convolutional unit is configured to extract features from the feature representation input to the third feature extraction module; The second activation unit is configured to perform a nonlinear transformation on the feature representation output by the second convolutional unit; The first normalization unit is configured to normalize the feature representation output by the second activation unit; The first pooling unit is configured to perform pooling processing on the feature representation output by the first normalization unit; Optionally, the second feature information is obtained by extracting features from the audio data information to be processed through a second feature extraction model, wherein the second feature extraction model includes: a fourth feature extraction module and a first normalization module; The input terminal of the fourth feature extraction module is configured to receive the audio data information to be processed, and the output terminal of the fourth feature extraction module is connected to the input terminal of the first normalization module. The fourth feature extraction module is configured to extract features from the audio data information to be processed. The first normalization module is configured to normalize the feature representation output by the fourth feature extraction module; Optionally, the fourth feature extraction module includes a plurality of third feature extraction units connected in series, each of the third feature extraction units including: a third convolution unit, a third activation unit, a second normalization unit, and a second pooling unit; Wherein, the output of the third convolutional unit is connected to the input of the third activation unit, the output of the third activation unit is connected to the input of the second normalization unit, and the output of the second normalization unit is connected to the input of the second pooling unit. The third convolutional unit is configured to extract features from the feature representation input to the third feature extraction unit; The third activation unit is configured to perform a nonlinear transformation on the feature representation output by the third convolution unit; The second normalization unit is configured to normalize the feature representation output by the third activation unit; The second pooling unit is configured to perform pooling processing on the feature representation output by the second normalization unit; Optionally, the target feature information includes first feature information and / or second feature information, and the second determining module determines target data information based on the target feature information, including: The first feature information and the second feature information are fused to obtain the target feature information; The target feature information is identified and processed using an audio recognition model to obtain target data information; Optionally, the target feature information is obtained by fusing the first feature information and the second feature information through a feature fusion model, wherein the feature fusion model includes: a second fusion module, a third fusion module and a fourth fusion module; The input terminal of the second fusion module is configured to receive the first feature information and the second feature information, the input terminal of the third fusion module is configured to receive the first feature information and the second feature information, and the output terminals of the second fusion module and the third fusion module are respectively connected to the input terminal of the fourth fusion module. The second fusion module is configured to perform fusion processing on the first feature information and the second feature information; The third fusion module is configured to perform fusion processing on the first feature information and the second feature information; The fourth fusion module is configured to perform fusion processing on the feature representations output by the second fusion module and the feature representations output by the third fusion module; Optionally, the second fusion module includes: a second linear unit, a third linear unit, a fourth linear unit, a second attention unit, a second fusion unit, and a fifth linear unit; Wherein, the input terminal of the second linear unit is configured to receive the first feature information, the input terminals of the third linear unit and the fourth linear unit are respectively configured to receive the second feature information, the output terminals of the second linear unit, the third linear unit and the fourth linear unit are respectively connected to the input terminal of the second attention unit, the output terminal of the second attention unit is connected to the input terminal of the second fusion unit, and the output terminal of the second fusion unit is connected to the input terminal of the fifth linear unit; The second linear unit is configured to perform a linear transformation on the first feature information; The third linear unit and the fourth linear unit are respectively configured to perform linear transformation processing on the second feature information; The second attention unit is configured to perform attention calculation processing on the feature representations output by the second linear unit, the feature representations output by the third linear unit, and the feature representations output by the fourth linear unit; The second fusion unit is configured to perform fusion processing on the feature representation output by the second attention unit; The fifth linear unit is configured to perform a linear transformation on the feature representation output by the second fusion unit; Optionally, the third fusion module includes: a sixth linear unit, a seventh linear unit, an eighth linear unit, a third attention unit, a third fusion unit, and a ninth linear unit; Wherein, the input terminal of the sixth linear unit is configured to receive the second feature information, the input terminals of the seventh linear unit and the eighth linear unit are respectively configured to receive the first feature information, the output terminals of the sixth linear unit, the seventh linear unit and the eighth linear unit are respectively connected to the input terminal of the third attention unit, the output terminal of the third attention unit is connected to the input terminal of the third fusion unit, and the output terminal of the third fusion unit is connected to the input terminal of the ninth linear unit; The sixth linear unit is configured to perform a linear transformation on the second feature information; The seventh linear unit and the eighth linear unit are respectively configured to perform linear transformation processing on the first feature information; The third attention unit is configured to perform attention calculation processing on the feature representations output by the sixth linear unit, the seventh linear unit, and the eighth linear unit; The third fusion unit is configured to perform fusion processing on the feature representation output by the third attention unit; The ninth linear unit is configured to perform a linear transformation on the feature representation output by the third fusion unit; Optionally, after the second determining module determines the target data information based on the target feature information, the second determining module is further configured to: Based on the target data information, the target control command is determined; Execute the control operation corresponding to the target control command.

17. A device, characterized in that, The device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that, It contains a computer program that is loaded by a processor to perform the steps of the method according to any one of claims 1 to 15.