Method, apparatus, electronic device, and storage medium for audio processing
By switching the decoding path diagonally in audio processing, limiting the number of output positions and merging the decoding path, the problem of large amount of calculation and high memory usage of recurrent neural network converter technology is solved, and efficient audio processing is achieved.
Patent Information
- Application Number
- CN202210616304.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The existing recurrent neural network converter technology has a large amount of computation and a lot of memory in audio processing, which makes it impossible to achieve parallel rapid computing, reducing practicality.
By obtaining the audio encoding results and adding 1 to the audio frame number dimension and text label sequence dimension respectively in the decoding path, the output position of the next frame is determined, and the output position of the next frame is switched diagonally to the output position of the next frame, limiting the number of output positions per frame, retaining the decoding path with the highest probability value, combining the same path, reducing the calculation amount and memory usage.
It improves the decoding efficiency of audio processing, reduces the computational volume and memory usage, realizes parallel rapid computing, and improves the practicality of the technology.
Smart Images

Figure CN115035902B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio processing, and particularly relates to a method, an apparatus, an electronic device, and a storage medium for audio processing. Background Art
[0002] In recent years, audio processing technologies such as audio recognition have been gradually developed, and their accuracy has become higher and higher, playing an important role in many fields. Currently, in the field of audio processing, there are technologies such as Connectionist Temporal Classification (CTC), Attention-based model technology, and Recurrent Neural Network Transducer (RNN-T) technology. Among them, the Recurrent Neural Network Transducer technology has the best effect in practice. However, in related technologies, when using the Recurrent Neural Network Transducer technology for audio processing, the amount of calculation is large and the memory occupation is large, resulting in the inability to achieve parallel and fast calculation, greatly reducing the practicability of this technology. Summary of the Invention
[0003] To overcome the problems existing in related technologies, embodiments of the present disclosure provide a method, an apparatus, an electronic device, and a storage medium for audio processing to solve the defects in related technologies.
[0004] According to a first aspect of the embodiments of the present disclosure, there is provided a method for audio processing, including:
[0005] Obtaining an audio encoding result, where each element in the audio encoding result has coordinates in the dimension of the number of audio frames and coordinates in the dimension of the text label sequence;
[0006] In response to the output result of the i-th frame in the decoding path being a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in both the dimension of the number of audio frames and the dimension of the text label sequence to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1;
[0007] Determining the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio encoding result.
[0008] In one embodiment, it further includes:
[0009] In response to the output result of the i-th frame in the decoding path being an empty character, adding 1 to the coordinate of the output position of the i-th frame in the dimension of the number of audio frames to obtain the output position of the (i + 1)-th frame in the decoding path.
[0010] In one embodiment, when the output result of the i-th frame in the decoding path is a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively to obtain the output position of the (i + 1)-th frame in the decoding path includes:
[0011] When the number of output positions of the i-th frame in the decoding path is 1, when the output result of the i-th frame in the decoding path is a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively to obtain the output position of the (i + 1)-th frame in the decoding path; and / or,
[0012] When the output result of the i-th frame in the decoding path is an empty character, adding 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path includes:
[0013] When the number of output positions of the i-th frame in the decoding path is 1, when the output result of the i-th frame in the decoding path is an empty character, adding 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path.
[0014] In one embodiment, when the output result of the i-th frame in the decoding path is a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively to obtain the output position of the (i + 1)-th frame in the decoding path includes:
[0015] When the output result of the i-th frame in the n-th decoding path among N decoding paths is a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively to obtain the output position of the (i + 1)-th frame in the n-th decoding path, where N is an integer greater than 1, and n is an integer not less than 1 and not greater than N;
[0016] When the output result of the i-th frame in the decoding path is an empty character, adding 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path includes:
[0017] When the output result of the i-th frame in the n-th decoding path among N decoding paths is an empty character, adding 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the n-th decoding path.
[0018] In one embodiment, determining the output result of the (i + 1)-th frame in the decoding path based on the output result of the i-th frame in the decoding path and the elements of the (i + 1)-th frame in the audio encoding result includes:
[0019] Determining the output result of the (i + 1)-th frame in the n-th decoding path based on the output result of the i-th frame in the n-th decoding path and the elements of the (i + 1)-th frame in the audio encoding result.
[0020] In one embodiment, the output results of the first i frames in the decoding path are the character types and the corresponding probability values, and the output result of the (i + 1)-th frame in the decoding path includes the probability value of each character type;
[0021] After determining the output result of the (i + 1)-th frame in the n-th decoding path based on the output result of the i-th frame in the n-th decoding path and the elements of the (i + 1)-th frame in the audio encoding result, the method further includes:
[0022] Determining the probability values of multiple candidate decoding paths formed by the n-th decoding path based on the output results of the first i frames and the output result of the (i + 1)-th frame of the n-th decoding path, where each character type in the output result of the (i + 1)-th frame corresponds to a candidate decoding path;
[0023] Among the multiple candidate decoding paths formed by each of the N decoding paths, retaining the top N candidate decoding paths with the highest probability values as the decoding paths.
[0024] In one embodiment, before retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, the method further includes:
[0025] Deleting the empty characters in each candidate decoding path among the multiple candidate decoding paths formed by each of the N decoding paths to obtain the string corresponding to each candidate decoding path;
[0026] Merging at least two candidate decoding paths with the same corresponding strings, and determining the sum of the probability values of the at least two candidate decoding paths as the probability value of the merged candidate decoding path.
[0027] In one embodiment, the (i + 1)-th frame is the last frame in the dimension of the number of audio frames;
[0028] After retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, the method further includes:
[0029] Determine the target text according to the decoding path with the highest probability value.
[0030] In one embodiment, the obtaining of the audio coding result includes:
[0031] Encoding the audio to be processed through the encoding sub-network of the neural network model to obtain the audio coding result; and / or,
[0032] The determining of the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the elements of the (i + 1)-th frame in the audio coding result includes:
[0033] The joint sub-network of the neural network model joints the output result of the i-th frame and the elements of the (i + 1)-th frame in the audio coding result to obtain a first joint result;
[0034] The decoding sub-network of the neural network model decodes the first joint result to obtain the output result of the (i + 1)-th frame.
[0035] In one embodiment, the method further includes:
[0036] Input the training audio and the corresponding text labels into the neural network model, and the neural network model outputs a second joint result, where each element in the second joint result has coordinates in the audio frame number dimension and the text label sequence dimension, and probability values of each character type;
[0037] Determine all first training paths within the second joint result, where in each first training path, each frame corresponds to a serial number, and when the output result of the i-th frame is a non-empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinates in the audio frame number dimension and the text label sequence dimension of the output position of the i-th frame, and when the output result of the i-th frame is an empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinate in the audio frame number dimension of the output position of the i-th frame;
[0038] Adjust the network parameters of the neural network model according to the sum of the probability values of all first training paths.
[0039] In one embodiment, the determining of all first training paths within the second joint result includes:
[0040] For a preset proportion of the training audio among all the training audio, determine all first training paths within the second joint result of the training audio;
[0041] The adjusting of the network parameters of the neural network model according to the sum of the probability values of all first training paths includes:
[0042] Adjust the network parameters of the neural network model according to all the first training paths determined within the second joint result of each training audio in the training audio according to the preset ratio.
[0043] It further includes:
[0044] For all the training audios, determine all the first training paths and all the second training paths within the second joint result of the training audio. Among them, in each second training path, each frame corresponds to a serial number. And when the output result of the i-th frame is a non-empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the text label sequence dimension to obtain the output position of the (i + 1)-th frame. When the output result of the i-th frame is an empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the audio frame number dimension to obtain the output position of the (i + 1)-th frame.
[0045] Adjust the network parameters of the neural network model according to all the first training paths and all the second training paths determined within the second joint result of each training audio in all the training audios.
[0046] According to the second aspect of the embodiments of the present disclosure, there is provided a device for audio processing, including:
[0047] An acquisition module, configured to acquire an audio coding result, where each element in the audio coding result has a coordinate in the audio frame number dimension and a coordinate in the text label sequence dimension;
[0048] A first position module, configured to, in response to the output result of the i-th frame in the decoding path being a non-empty character, increment the output position of the i-th frame by 1 respectively in the coordinates of the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1;
[0049] A decoding module, configured to determine the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio coding result.
[0050] In one embodiment, it further includes a second position module, configured to:
[0051] In response to the output result of the i-th frame in the decoding path being an empty character, increment the output position of the i-th frame by 1 in the coordinate of the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path.
[0052] In one embodiment, the first position module is specifically configured to:
[0053] When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively to obtain the output position of the (i + 1)-th frame in the decoding path; and / or,
[0054] The second position module is specifically configured to:
[0055] When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path.
[0056] In one embodiment, the first position module is specifically configured to:
[0057] In response to the output result of the i-th frame in the n-th decoding path among N decoding paths being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively to obtain the output position of the (i + 1)-th frame in the n-th decoding path, where N is an integer greater than 1, and n is an integer not less than 1 and not greater than N;
[0058] The second position module is specifically configured to:
[0059] In response to the output result of the i-th frame in the n-th decoding path among N decoding paths being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the n-th decoding path.
[0060] In one embodiment, the decoding module is specifically configured to:
[0061] Determine the output result of the (i + 1)-th frame in the n-th decoding path according to the output result of the i-th frame in the n-th decoding path and the element of the (i + 1)-th frame in the audio coding result.
[0062] In one embodiment, the output results of the first i frames in the decoding path are the character types and the corresponding probability values, and the output result of the (i + 1)-th frame in the decoding path includes the probability values of each character type;
[0063] The apparatus for audio processing further includes a screening module, configured to:
[0064] After determining the output result of the (i + 1)-th frame in the n-th decoding path based on the output result of the i-th frame in the n-th decoding path and the elements of the (i + 1)-th frame in the audio coding result, determine the probability values of multiple candidate decoding paths formed by the n-th decoding path according to the output results of the first i frames and the output result of the (i + 1)-th frame in the n-th decoding path, where each character type in the output result of the (i + 1)-th frame corresponds to a candidate decoding path;
[0065] Among the multiple candidate decoding paths formed by each of the N decoding paths, retain the top N candidate decoding paths with the highest probability values as the decoding paths.
[0066] In one embodiment, the apparatus for audio processing further includes a merging module, configured to:
[0067] Before retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, delete the null characters in each candidate decoding path among the multiple candidate decoding paths formed by each of the N decoding paths to obtain the string corresponding to each candidate decoding path;
[0068] Merge at least two candidate decoding paths with the same corresponding strings, and determine the sum of the probability values of the at least two candidate decoding paths as the probability value of the merged candidate decoding path.
[0069] In one embodiment, the (i + 1)-th frame is the last frame in the dimension of the number of audio frames;
[0070] The apparatus for audio processing further includes a target module, configured to:
[0071] After retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, determine the target text according to the decoding path with the highest probability value.
[0072] In one embodiment, the obtaining module is specifically configured to:
[0073] The encoding sub-network of the neural network model encodes the audio to be processed to obtain the audio coding result; and / or,
[0074] The decoding module is specifically configured to:
[0075] The joint sub-network of the neural network model jointly processes the output result of the i-th frame and the elements of the (i + 1)-th frame in the audio coding result to obtain a first joint result;
[0076] The decoding sub-network of the neural network model decodes the first joint result to obtain the output result of the (i + 1)-th frame.
[0077] In one embodiment, the apparatus for audio processing further includes a first training module for:
[0078] Inputting the training audio and the corresponding text labels into the neural network model, and the neural network model outputs a second joint result, where each element in the second joint result has coordinates in the audio frame number dimension and the text label sequence dimension, and probability values of various character types;
[0079] Determining all first training paths within the second joint result, where in each first training path, each frame corresponds to a sequence number, and when the output result of the i-th frame is a non-empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinates in the audio frame number dimension and the text label sequence dimension of the output position of the i-th frame, and when the output result of the i-th frame is an empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinate in the audio frame number dimension of the output position of the i-th frame;
[0080] Adjusting the network parameters of the neural network model according to the sum of the probability values of all first training paths.
[0081] In one embodiment, when the training module is used to determine all first training paths within the second joint result, it is specifically used for:
[0082] For a preset proportion of the training audio among all the training audio, determining all first training paths within the second joint result of the training audio;
[0083] When the first training module is used to adjust the network parameters of the neural network model according to the sum of the probability values of all first training paths, it is specifically used for:
[0084] Adjusting the network parameters of the neural network model according to all the first training paths determined within the second joint result of each training audio in the training audio of the preset proportion;
[0085] The apparatus for audio processing further includes a second training module for:
[0086] For all the training audio, determine all the first training paths and all the second training paths in the second joint result of the training audio. Wherein, in each second training path, each frame corresponds to a serial number, and when the output result of the i-th frame is a non-empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the text label sequence dimension to obtain the output position of the (i + 1)-th frame. When the output result of the i-th frame is an empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the audio frame number dimension to obtain the output position of the (i + 1)-th frame;
[0087] According to all the first training paths and all the second training paths determined in the second joint result of each training audio in all the training audio, adjust the network parameters of the neural network model.
[0088] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes a memory and a processor. The memory is used to store computer instructions that can be run on the processor, and the processor is used to perform the method according to the first aspect when executing the computer instructions.
[0089] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect is implemented.
[0090] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0091] The method for audio processing provided by the present disclosure can, by obtaining the audio coding result, in response to the output result of the i-th frame in the decoding path being a non-empty character, increment the output position of the i-th frame by 1 in the coordinates of both the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1. Finally, according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio coding result, determine the output result of the (i + 1)-th frame in the decoding path. Since the output position of the next frame is switched diagonally in the decoding path, compared with the method of switching to the output position of the next frame through two steps horizontally and vertically in the related art, it can reduce the amount of computation in the decoding process, improve the decoding efficiency, reduce the computational amount and memory occupancy of audio processing, achieve parallel and fast calculation, and improve the practicability of this technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0093] Figure 1It is a flowchart of a method for audio processing shown in an exemplary embodiment of the present disclosure;
[0094] Figure 2 It is a schematic diagram of the output position switching between adjacent frames shown in an exemplary embodiment of the present disclosure;
[0095] Figure 3 It is a schematic diagram of the output position switching between adjacent frames in the related art;
[0096] Figure 4 It is a flowchart of a training method for a neural network model shown in an exemplary embodiment of the present disclosure;
[0097] Figure 5 It is a schematic structural diagram of a device for audio processing shown in an exemplary embodiment of the present disclosure;
[0098] Figure 6 It is a block diagram of the structure of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0099] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0100] The terms used in the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0101] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0102] The field of speech recognition has experienced decades of development, from the initial sequence similarity matching, to the modeling based on Gaussian mixture models and hidden Markov models, and then to the hybrid systems based on neural networks developed later. For decades, building a speech recognition system has been a rather complex task, and a complicated data alignment process is required before building the model. In recent years, the end-to-end model has entered a stage of rapid development. It can not only greatly simplify the modeling process of speech recognition, that is, remove the complex alignment process, but also achieve better recognition results.
[0103] Generally speaking, there are three main categories of current implementation methods for end-to-end models, namely Connectionist Temporal Classification (CTC), Attention-based model, and Recurrent Neural Network Transducer (RNN-T). Among these three categories of models, the CTC model and the RNN-T model are naturally streaming models and can perform frame-synchronous decoding. However, for the Attention-based model to achieve streaming decoding, some additional corrections are needed, which are not only troublesome but also result in a loss of some recognition accuracy. In the CTC model and the RNN-T model, the CTC model has a prerequisite assumption that frames are probabilistically independent of each other, which makes it difficult for the CTC model to achieve excellent recognition rates without adding an external language model. Therefore, comprehensively comparing various aspects, the RNN-T model has a greater scope of application in production practice.
[0104] During the application and training of the RNN-T model, it is necessary to determine the decoding path on the plane composed of the audio frame number dimension and the text label sequence dimension. During the process of determining the decoding path, the switching between the frame number and the text label serial number is rather arbitrary and comprehensive, lacking necessary restrictions. Therefore, the computational complexity is large and the memory occupancy is high, resulting in the inability to achieve parallel and fast computing, and greatly reducing the practicality of this technology.
[0105] Based on this, on the one hand, at least one embodiment of the present disclosure provides a method for audio processing. Please refer to the appendix Figure 1 , which shows the flow of this method, including step S101 and step S103.
[0106] Among them, the method can use a neural network to process the audio to be processed, such as the RNN-T model, etc. The RNN-T model usually includes three parts, namely an encoder network, a prediction network, and a joiner network. This method can be applied to scenarios such as speech recognition, lip reading recognition, and machine translation for generating text label sequences.
[0107] In step S101, an audio encoding result is obtained, where each element in the audio encoding result has coordinates in the audio frame number dimension and coordinates in the text label sequence dimension.
[0108] Among them, the encoder network of the neural network can encode the audio to be processed to obtain the audio encoding result. The audio encoding result can be in the form of a feature vector. Exemplarily, if the audio frame number dimension of the audio encoding result is T and the text label sequence dimension is U, then the audio encoding result can be a feature vector of dimension (T, U). The element with coordinates (t, u) in the audio encoding result represents the feature vector when the t-th frame of audio is the u-th text label.
[0109] The audio encoding result is used for decoding path search during the decoding process. By searching the decoding path, the output position of each frame can be determined within the plane composed of the audio frame number dimension and the text label sequence dimension, and the output result of each frame can be determined. The output position can be identified by the serial number in the text label sequence dimension, and the output result can be represented by the character type of the text label at this output position. The character type can be a specific character or an empty character. After obtaining the decoding path, the target text, that is, the speech recognition result, can be generated according to the output position and output result of each frame on the decoding path.
[0110] In step S102, in response to the output result of the i-th frame in the decoding path being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in both the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1.
[0111] Among them, the i-th frame can be each frame in the decoding path. The output result of the i-th frame in the decoding path being a non-empty character means that the output result of the i-th frame in the decoding path is a specific character other than the empty character. Please refer to the appendix Figure 2, which shows the process of switching the output position of the i-th frame to the output position of the (i + 1)-th frame when the output result of the i-th frame is a non-empty character, that is, switching diagonally from (t, u) to (t + 1, u + 1). It should be understood that in the related art, when the output result of the i-th frame is a non-empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the text label sequence dimension to obtain the output position of the (i + 1)-th frame (please refer to the appendix Figure 3 , which shows the process of switching the output position of the i-th frame to the output position of the (i + 1)-th frame in the related art when the output result of the i-th frame is a non-empty character, that is, switching upward from (t, u) to (t, u + 1)), which can reduce the computational amount in the decoding process and improve the decoding efficiency.
[0112] In addition, in response to the output result of the i-th frame in the decoding path being an empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path. Please refer to the appendix Figure 2 , which shows the process of switching the output position of the i-th frame to the output position of the (i + 1)-th frame when the output result of the i-th frame is an empty character, that is, switching to the right from (t, u) to (t + 1, u).
[0113] It is also possible to limit the number of output positions of each frame during the process of searching the decoding path, that is, to limit the number of text labels of each frame. For example, the output position of each frame is limited to 1, so that different samples between batches can be decoded in parallel. Exemplarily, when the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being a non-empty character, the output position of the i-th frame is incremented by 1 in the coordinates of both the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the decoding path; when the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being an empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path. Limiting the number of output positions of each frame is to constrain the maximum number of labels that can be emitted at each time step. For example, limiting the number of output positions of each frame to 1 is to constrain the maximum number of labels that can be emitted at each time step to 1.
[0114] In step S103, according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio coding result, the output result of the (i + 1)-th frame in the decoding path is determined.
[0115] In a possible embodiment, first, the joint sub-network of the neural network model combines the output result of the i-th frame and the elements of the (i + 1)-th frame in the audio coding result to obtain a first joint result; next, the decoding sub-network of the neural network model decodes the first joint result to obtain the output result of the (i + 1)-th frame.
[0116] Among them, the output results of the first i frames in the decoding path are the character types (a specific character or an empty character) and the corresponding probability values. The output result of the (i + 1)-th frame in the decoding path includes the character type (a specific character or an empty character) and the corresponding probability value, or the probability values of each character type in the character library.
[0117] It can be understood that in the case where the (i + 1)-th frame is the last frame in the audio frame number dimension, the target text can be determined according to the decoding path, that is, the audio recognition is completed to obtain the audio recognition result.
[0118] The method for audio processing provided by the present disclosure can, by obtaining the audio coding result, in response to the output result of the i-th frame in the decoding path being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in both the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1. Finally, according to the output result of the i-th frame in the decoding path and the elements of the (i + 1)-th frame in the audio coding result, the output result of the (i + 1)-th frame in the decoding path can be determined. Since the output position of the next frame is switched in a diagonal manner in the decoding path, compared with the method of switching to the output position of the next frame in two steps, horizontally and vertically, in the related art, it can reduce the amount of computation in the decoding process, improve the decoding efficiency, reduce the computational amount and memory occupancy of audio processing, achieve parallel and fast calculation, and improve the practicability of this technology.
[0119] In some embodiments of the present disclosure, a reserved quantity N (for example, 4) can be set during the process of searching for the decoding path, that is, N decoding paths can be reserved simultaneously. The N reserved decoding paths are the decoding paths with the highest fractional probability values. Moreover, after determining the output position and output result of each frame, the N decoding paths to be reserved can be determined.
[0120] That is to say, in response to the output result of the i-th frame in the decoding path being a non-empty character, when obtaining the output position of the (i + 1)-th frame in the decoding path by adding 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively, it is possible to, in response to the output result of the i-th frame in the n-th decoding path among N decoding paths being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively, to obtain the output position of the (i + 1)-th frame in the n-th decoding path, where N is an integer greater than 1, and n is an integer not less than 1 and not greater than N.
[0121] That is to say, in response to the output result of the i-th frame in the decoding path being an empty character, when obtaining the output position of the (i + 1)-th frame in the decoding path by adding 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension, it is possible to, in response to the output result of the i-th frame in the n-th decoding path among N decoding paths being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension, to obtain the output position of the (i + 1)-th frame in the n-th decoding path.
[0122] That is to say, when determining the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio coding result, it is possible to determine the output result of the (i + 1)-th frame in the n-th decoding path according to the output result of the i-th frame in the n-th decoding path and the element of the (i + 1)-th frame in the audio coding result. The output result of the (i + 1)-th frame in the decoding path includes the probability value of each character type.
[0123] Based on this, after determining the output result of the (i + 1)-th frame in the n-th decoding path (that is, after determining the output result of the (i + 1)-th frame in each of the N decoding paths), it is possible to determine the probability value of the multiple candidate decoding paths formed by the n-th decoding path according to the output result of the first i frames and the output result of the (i + 1)-th frame of the n-th decoding path, where each character type in the output result of the (i + 1)-th frame corresponds to a candidate decoding path; then among the multiple candidate decoding paths formed by each of the N decoding paths, retain the top N candidate decoding paths with the highest probability values as the decoding paths. Thus, after the output result of each frame is output, it is possible to accurately determine the decoding path with the highest probability value for retention, so that the text information corresponding to the retained decoding path is closest to the true text of the audio to be processed.
[0124] In addition, before retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed in each of the N decoding paths, the null characters in each of the multiple candidate decoding paths formed in each of the N decoding paths may be removed to obtain a string corresponding to each candidate decoding path; then, at least two candidate decoding paths with the same corresponding strings are merged, and the sum of the probability values of the at least two candidate decoding paths is determined as the probability value of the merged candidate decoding path. That is, the substantially identical decoding paths are merged by removing the null characters, so that the decoding paths are more accurate and the probability values are more accurate. For example, if a candidate decoding path includes A, null, B, C, D, and another candidate decoding path includes A, B, null, C, D, then these two candidate decoding paths can be merged.
[0125] It can be understood that in the case where the (i + 1)-th frame is the last frame in the dimension of the number of audio frames, the target text can be determined according to the decoding path with the highest probability value among the N retained decoding paths.
[0126] In this embodiment, by retaining multiple decoding paths in real time and merging and updating the retained multiple decoding paths in real time, the decoding paths can be made more accurate and reliable. Moreover, by restricting the number of retained decoding paths, the computing load and memory occupancy can also be reduced.
[0127] In some embodiments of the present disclosure, the neural network may be trained in the manner as Figure 4 shown, including steps S401 to S403.
[0128] In step S401, the training audio and the corresponding text label are input into the neural network model, and the neural network model outputs a second joint result, where each element in the second joint result has coordinates in the dimension of the number of audio frames and coordinates in the dimension of the text label sequence, as well as probability values of various character types.
[0129] Among them, the encoding sub-network of the neural network model can encode the training audio to obtain a training encoding result, the decoding sub-network of the neural network model can decode the text label to obtain a training decoding result, and the joint sub-network of the neural network model can jointly process the training encoding result and the training decoding result to obtain a second joint result.
[0130] In step S402, all first training paths are determined within the second joint result. In each of the first training paths, each frame corresponds to a sequence number. When the output result of the i-th frame is a non-empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinates in the audio frame number dimension and the text label sequence dimension. When the output result of the i-th frame is an empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinate in the audio frame number dimension.
[0131] That is, according to the Figure 1 decoding path search method in the method for audio processing shown in the appendix, the first training paths are searched. That is, by restricting the determination method of the output position of the (i + 1)-th frame when the output result of the i-th frame in the first training path is a non-empty character, and the number of text label sequence numbers (i.e., the number of text labels) per frame, the training process and the decoding process have the same constraint conditions, so that the model has better performance.
[0132] In step S403, the network parameters of the neural network model are adjusted according to the sum of the probability values of all the first training paths.
[0133] First, the first network loss value (such as modified RNN-T loss) can be determined according to the sum of the probability values of all the first training paths, and then the network parameters of the neural network model are adjusted by using the first network loss value, that is, the network parameters of the encoding sub-network, the decoding sub-network, and the joint sub-network are adjusted.
[0134] In the training method provided in this embodiment, the same constraint conditions as those in the Figure 1 decoding process shown in the appendix are added during the search of the first training paths, so that the decoding process and the training process are more closely combined, and the performance of the trained neural network model is significantly improved.
[0135] Furthermore, the following training method can be provided on the basis of the Figure 4 training method shown in the appendix:
[0136] First, step S401 is executed to obtain the second joint result.
[0137] Then, for a preset proportion (such as 25%) of the training audio in all the training audio, step S402 is executed, that is, all the first training paths are determined within the second joint result of the training audio.
[0138] Then, perform step S403, and adjust the network parameters of the neural network model according to all the first training paths determined in the second joint result of each training audio in the preset ratio of training audios.
[0139] Then, for all the training audios, determine all the first training paths and all the second training paths in the second joint result of the training audios. Among them, in each second training path, each frame corresponds to a serial number. And when the output result of the i-th frame is a non-empty character, the output position of the i-th frame is incremented by 1 in the coordinates of the text label sequence dimension to obtain the output position of the (i + 1)-th frame. When the output result of the i-th frame is an empty character, the output position of the i-th frame is incremented by 1 in the coordinates of the audio frame number dimension to obtain the output position of the (i + 1)-th frame.
[0140] Finally, adjust the network parameters of the neural network model according to all the first training paths and all the second training paths determined in the second joint result of each training audio in all the training audios. Exemplarily, a second network loss value (such as modified RNN-T loss) can be determined according to all the first training paths, a third network loss value (such as ordinary RNN-T loss) can be determined according to all the second training paths, and finally, the second network loss value and the third network loss value are weighted and summed to obtain a comprehensive network loss value, and the network parameters of the neural network model are adjusted according to the comprehensive network loss value.
[0141] In this embodiment, the Figure 4 shown neural network training method is combined with the training method in the related art, so that the training method can improve the accuracy of network loss calculation while conforming to the decoding process, and further improve the training accuracy.
[0142] According to the second aspect of the embodiments of the present disclosure, there is provided a device for audio processing. Please refer to the attached Figure 5 , including:
[0143] An acquisition module 501, configured to acquire an audio coding result, where each element in the audio coding result has coordinates in the audio frame number dimension and coordinates in the text label sequence dimension;
[0144] A first position module 502, configured to, in response to the output result of the i-th frame in the decoding path being a non-empty character, increment the output position of the i-th frame by 1 in the coordinates of both the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1;
[0145] A decoding module 503, configured to determine an output result of the (i + 1)-th frame in the decoding path according to an output result of the i-th frame in the decoding path and an element of the (i + 1)-th frame in the audio coding result.
[0146] In some embodiments of the present disclosure, it further includes a second position module, configured to:
[0147] In response to the output result of the i-th frame in the decoding path being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path.
[0148] In some embodiments of the present disclosure, the first position module is specifically configured to:
[0149] When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in both the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the decoding path; and / or,
[0150] The second position module is specifically configured to:
[0151] When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the decoding path.
[0152] In some embodiments of the present disclosure, the first position module is specifically configured to:
[0153] In response to the output result of the i-th frame in the n-th decoding path among N decoding paths being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in both the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame in the n-th decoding path, where N is an integer greater than 1, and n is an integer not less than 1 and not greater than N;
[0154] The second position module is specifically configured to:
[0155] In response to the output result of the i-th frame in the n-th decoding path among N decoding paths being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame in the n-th decoding path.
[0156] In some embodiments of the present disclosure, the decoding module is specifically configured to:
[0157] Determine the output result of the (i + 1)-th frame in the n-th decoding path based on the output result of the i-th frame in the n-th decoding path and the elements of the (i + 1)-th frame in the audio encoding result.
[0158] In some embodiments of the present disclosure, the output results of the first i frames in the decoding path are character types and corresponding probability values, and the output result of the (i + 1)-th frame in the decoding path includes probability values for each character type;
[0159] The apparatus for audio processing further includes a screening module, configured to:
[0160] After determining the output result of the (i + 1)-th frame in the n-th decoding path based on the output result of the i-th frame in the n-th decoding path and the elements of the (i + 1)-th frame in the audio encoding result, determine the probability values of a plurality of candidate decoding paths formed by the n-th decoding path according to the output results of the first i frames and the output result of the (i + 1)-th frame in the n-th decoding path, where each character type in the output result of the (i + 1)-th frame corresponds to a candidate decoding path;
[0161] Among the plurality of candidate decoding paths formed by each decoding path among the N decoding paths, retain the top N candidate decoding paths with the highest probability values as the decoding paths.
[0162] In some embodiments of the present disclosure, the apparatus for audio processing further includes a merging module, configured to:
[0163] Before retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the plurality of candidate decoding paths formed by each decoding path among the N decoding paths, delete the null characters in each candidate decoding path among the plurality of candidate decoding paths formed by each decoding path among the N decoding paths to obtain a string corresponding to each candidate decoding path;
[0164] Merge at least two candidate decoding paths with the same corresponding strings, and determine the sum of the probability values of the at least two candidate decoding paths as the probability value of the merged candidate decoding path.
[0165] In some embodiments of the present disclosure, the (i + 1)-th frame is the last frame in the dimension of the number of audio frames;
[0166] The apparatus for audio processing further includes a target module, configured to:
[0167] After retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the plurality of candidate decoding paths formed by each decoding path among the N decoding paths, determine the target text according to the decoding path with the highest probability value.
[0168] In some embodiments of the present disclosure, the obtaining module is specifically configured to:
[0169] The encoding sub-network of the neural network model encodes the audio to be processed to obtain the audio encoding result; and / or,
[0170] The decoding module is specifically configured to:
[0171] The joint sub-network of the neural network model combines the output result of the i-th frame and the elements of the (i + 1)-th frame in the audio encoding result to obtain a first joint result;
[0172] The decoding sub-network of the neural network model decodes the first joint result to obtain the output result of the (i + 1)-th frame.
[0173] In some embodiments of the present disclosure, the apparatus for audio processing further includes a first training module, which is configured to:
[0174] Input the training audio and the corresponding text labels into the neural network model, and the neural network model outputs a second joint result, where each element in the second joint result has coordinates in the audio frame number dimension and the text label sequence dimension, as well as probability values of various character types;
[0175] Determine all first training paths within the second joint result, where in each first training path, each frame corresponds to a serial number, and when the output result of the i-th frame is a non-empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinates in the audio frame number dimension and the text label sequence dimension of the output position of the i-th frame, and when the output result of the i-th frame is an empty character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinate in the audio frame number dimension of the output position of the i-th frame;
[0176] Adjust the network parameters of the neural network model according to the sum of the probability values of all first training paths.
[0177] In some embodiments of the present disclosure, when the training module is configured to determine all first training paths within the second joint result, it is specifically configured to:
[0178] For a preset proportion of the training audio among all the training audio, determine all first training paths within the second joint result of the training audio;
[0179] When the first training module is configured to adjust the network parameters of the neural network model according to the sum of the probability values of all first training paths, it is specifically configured to:
[0180] Adjust the network parameters of the neural network model according to all the first training paths determined within the second joint result of each training audio in the training audio according to the preset ratio;
[0181] The apparatus for audio processing further includes a second training module for:
[0182] For all the training audios, determine all the first training paths and all the second training paths within the second joint result of the training audio. Wherein, in each second training path, each frame corresponds to a serial number, and when the output result of the i-th frame is a non-empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the text label sequence dimension to obtain the output position of the (i + 1)-th frame. When the output result of the i-th frame is an empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the audio frame number dimension to obtain the output position of the (i + 1)-th frame;
[0183] Adjust the network parameters of the neural network model according to all the first training paths and all the second training paths determined within the second joint result of each training audio in all the training audio.
[0184] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments of the method in the first aspect, and will not be elaborated here.
[0185] According to the third aspect of the embodiments of the present disclosure, please refer to the attached Figure 6 , which exemplarily shows a block diagram of an electronic device. For example, the apparatus 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0186] Refer to Figure 6 , the apparatus 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0187] The processing component 602 generally controls the overall operation of the apparatus 600, such as operations associated with display, telephone calls, data communication, camera program operations, and recording operations. The processing element 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.
[0188] The memory 604 is configured to store various types of data to support the operation of the device 600. Examples of such data include instructions for any application or method operating on the device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0189] The power component 606 provides power to the various components of the device 600. The power component 606 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 600.
[0190] The multimedia component 608 includes a screen that provides an output interface between the device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0191] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC) that is configured to receive external audio signals when the device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.
[0192] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface module, which can be a keyboard, click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, volume buttons, a power-on button, and a lock button.
[0193] The sensor assembly 614 includes one or more sensors for providing a status assessment of various aspects of the device 600. For example, the sensor assembly 614 can detect the on / off state of the device 600, the relative positioning of components, such as the display and keypad of the device 600. The sensor assembly 614 can also detect a change in the position of the device 600 or a component of the device 600, the presence or absence of user contact with the device 600, the orientation or acceleration / deceleration of the device 600, and a change in the temperature of the device 600. The sensor assembly 614 can also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0194] The communication component 616 is configured to facilitate communication between the device 600 and other devices in a wired or wireless manner. The device 600 can access a wireless network based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0195] In an exemplary embodiment, the device 600 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the power supply method of the above electronic device.
[0196] In a fourth aspect, in an exemplary embodiment, the present disclosure also provides a non-transitory computer-readable storage medium including instructions, such as a memory 604 including instructions, which can be executed by a processor 620 of the device 600 to complete the power supply method of the above electronic device. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0197] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0198] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for audio processing, characterized in that, Including: Obtaining an audio encoding result, wherein each element in the audio encoding result has coordinates in the dimension of audio frame number and coordinates in the dimension of text label sequence; In response to the output result of the i-th frame in the decoding path being a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in both the dimension of audio frame number and the dimension of text label sequence to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1; Determining the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio encoding result; The step of, in response to the output result of the i-th frame in the decoding path being a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in both the dimension of audio frame number and the dimension of text label sequence to obtain the output position of the (i + 1)-th frame in the decoding path includes: In response to the output result of the i-th frame in the n-th decoding path among N decoding paths being a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in both the dimension of audio frame number and the dimension of text label sequence to obtain the output position of the (i + 1)-th frame in the n-th decoding path, where N is an integer greater than 1, and n is an integer not less than 1 and not greater than N.
2. The method according to claim 1, characterized in that, The method further includes: In response to the output result of the i-th frame in the decoding path being an empty character, adding 1 to the coordinate of the output position of the i-th frame in the dimension of audio frame number to obtain the output position of the (i + 1)-th frame in the decoding path.
3. The method according to claim 2, wherein The step of, in response to the output result of the i-th frame in the decoding path being a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in both the dimension of audio frame number and the dimension of text label sequence to obtain the output position of the (i + 1)-th frame in the decoding path includes: When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in both the dimension of audio frame number and the dimension of text label sequence to obtain the output position of the (i + 1)-th frame in the decoding path; and / or The step of, in response to the output result of the i-th frame in the decoding path being an empty character, adding 1 to the coordinate of the output position of the i-th frame in the dimension of audio frame number to obtain the output position of the (i + 1)-th frame in the decoding path includes: When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being an empty character, adding 1 to the coordinate of the output position of the i-th frame in the dimension of audio frame number to obtain the output position of the (i + 1)-th frame in the decoding path.
4. The method according to claim 2, wherein The step of, in response to the output result of the i-th frame in the decoding path being an empty character, adding 1 to the coordinate of the output position of the i-th frame in the dimension of audio frame number to obtain the output position of the (i + 1)-th frame in the decoding path includes: In response to the output result of the i-th frame in the n-th decoding path among the N decoding paths being a null character, increment the coordinate of the output position of the i-th frame in the dimension of the number of audio frames by 1 to obtain the output position of the (i + 1)-th frame in the n-th decoding path.
5. The method according to claim 4, characterized in that, The determining of the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio coding result includes: Determine the output result of the (i + 1)-th frame in the n-th decoding path according to the output result of the i-th frame in the n-th decoding path and the element of the (i + 1)-th frame in the audio coding result.
6. The method according to claim 5, wherein The output results of the first i frames in the decoding path are the character types and the corresponding probability values, and the output result of the (i + 1)-th frame in the decoding path includes the probability values of each character type; After determining the output result of the (i + 1)-th frame in the n-th decoding path according to the output result of the i-th frame in the n-th decoding path and the element of the (i + 1)-th frame in the audio coding result, the method further includes: Determine the probability values of the multiple candidate decoding paths formed by the n-th decoding path according to the output results of the first i frames and the output result of the (i + 1)-th frame in the n-th decoding path, where each character type in the output result of the (i + 1)-th frame corresponds to a candidate decoding path; Among the multiple candidate decoding paths formed by each of the N decoding paths, retain the top N candidate decoding paths with the highest probability values as the decoding paths.
7. The method according to claim 6, wherein Before retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, the method further includes: Delete the null characters in each candidate decoding path among the multiple candidate decoding paths formed by each of the N decoding paths to obtain the string corresponding to each candidate decoding path; Merge at least two candidate decoding paths with the same corresponding strings, and determine the sum of the probability values of the at least two candidate decoding paths as the probability value of the merged candidate decoding path.
8. The method according to claim 6, wherein The (i + 1)-th frame is the last frame in the dimension of the number of audio frames; After retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, the method further includes: Determine the target text according to the decoding path with the highest probability value.
9. The method according to claim 1, characterized in that, The obtaining of the audio coding result includes: Encoding the audio to be processed through the encoding sub-network of the neural network model to obtain the audio coding result; and / or, The determining of the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio coding result includes: The joint sub-network of the neural network model joints the output result of the i-th frame and the element of the (i + 1)-th frame in the audio coding result to obtain a first joint result; The decoding sub-network of the neural network model decodes the first joint result to obtain the output result of the (i + 1)-th frame.
10. The method according to claim 9, characterized in that The method further includes: Input the training audio and the corresponding text labels into the neural network model, and the neural network model outputs a second joint result, where each element in the second joint result has coordinates in the audio frame number dimension and the text label sequence dimension, as well as probability values for each character type; Determine all first training paths within the second joint result. In each first training path, each frame corresponds to a serial number. When the output result of the i-th frame is a non-empty character, add 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension to obtain the output position of the (i + 1)-th frame. When the output result of the i-th frame is an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame; Adjust the network parameters of the neural network model according to the sum of the probability values of all first training paths.
11. The method according to claim 10, wherein The determining all first training paths within the second joint result includes: For a preset proportion of the training audio among all the training audio, determine all first training paths within the second joint result of the training audio; The adjusting the network parameters of the neural network model according to the sum of the probability values of all first training paths includes: Adjust the network parameters of the neural network model according to all the first training paths determined within the second joint result of each training audio in the preset proportion of the training audio; The method further includes: For all the training audio, determine all first training paths and all second training paths within the second joint result of the training audio. In each second training path, each frame corresponds to a serial number. When the output result of the i-th frame is a non-empty character, add 1 to the coordinate of the output position of the i-th frame in the text label sequence dimension to obtain the output position of the (i + 1)-th frame. When the output result of the i-th frame is an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension to obtain the output position of the (i + 1)-th frame; Adjust the network parameters of the neural network model according to all the first training paths and all the second training paths determined within the second joint result of each training audio among all the training audio.
12. An apparatus for audio processing, characterized in that, Includes: An acquisition module for acquiring an audio encoding result, where each element in the audio encoding result has coordinates in the audio frame number dimension and the text label sequence dimension; A first position module for, in response to the output result of the i-th frame in the decoding path being a non-empty character, adding 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively to obtain the output position of the (i + 1)-th frame in the decoding path, where i is an integer not less than 1; A decoding module for determining the output result of the (i + 1)-th frame in the decoding path according to the output result of the i-th frame in the decoding path and the element of the (i + 1)-th frame in the audio encoding result; The first position module is specifically used for: In response to the output result of the i-th frame in the n-th decoding path among the N decoding paths being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively, to obtain the output position of the (i + 1)-th frame in the n-th decoding path, where N is an integer greater than 1, and n is an integer not less than 1 and not greater than N.
13. The device according to claim 12, characterized in that, It further includes a second position module for: In response to the output result of the i-th frame in the decoding path being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension, to obtain the output position of the (i + 1)-th frame in the decoding path.
14. The device according to claim 13, characterized in that, The first position module is specifically used for: When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being a non-empty character, add 1 to the coordinates of the output position of the i-th frame in the audio frame number dimension and the text label sequence dimension respectively, to obtain the output position of the (i + 1)-th frame in the decoding path; and / or, The second position module is specifically used for: When the number of output positions of the i-th frame in the decoding path is 1, in response to the output result of the i-th frame in the decoding path being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension, to obtain the output position of the (i + 1)-th frame in the decoding path.
15. The device according to claim 13, characterized in that, The second position module is specifically used for: In response to the output result of the i-th frame in the n-th decoding path among the N decoding paths being an empty character, add 1 to the coordinate of the output position of the i-th frame in the audio frame number dimension, to obtain the output position of the (i + 1)-th frame in the n-th decoding path.
16. The device according to claim 15, characterized in that, The decoding module is specifically used for: According to the output result of the i-th frame in the n-th decoding path and the element of the (i + 1)-th frame in the audio coding result, determine the output result of the (i + 1)-th frame in the n-th decoding path.
17. The device according to claim 16, characterized in that, The output results of the first i frames in the decoding path are the character types and the corresponding probability values, and the output result of the (i + 1)-th frame in the decoding path includes the probability values of each character type; And the apparatus for audio processing further includes a screening module for: After determining the output result of the (i + 1)-th frame in the n-th decoding path according to the output result of the i-th frame in the n-th decoding path and the element of the (i + 1)-th frame in the audio coding result, determine the probability values of the multiple candidate decoding paths formed by the n-th decoding path according to the output results of the first i frames and the (i + 1)-th frame of the n-th decoding path, where each character type in the output result of the (i + 1)-th frame corresponds to a candidate decoding path; Among the multiple candidate decoding paths formed by each decoding path among the N decoding paths, retain the top N candidate decoding paths with the highest probability values as the decoding paths.
18. The device according to claim 17, characterized in that, The apparatus for audio processing further includes a merging module for: Before retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, delete the null characters in each candidate decoding path among the multiple candidate decoding paths formed by each of the N decoding paths to obtain the string corresponding to each candidate decoding path; Merge at least two candidate decoding paths with the same corresponding strings, and determine the sum of the probability values of the at least two candidate decoding paths as the probability value of the merged candidate decoding path.
19. The device according to claim 17, characterized in that, The (i + 1)-th frame is the last frame in the audio frame number dimension; The apparatus for audio processing further includes a target module for: After retaining the top N candidate decoding paths with the highest probability values as the decoding paths among the multiple candidate decoding paths formed by each of the N decoding paths, determine the target text according to the decoding path with the highest probability value.
20. The device according to claim 12, characterized in that, The obtaining module is specifically configured to: The encoding sub-network of the neural network model encodes the audio to be processed to obtain the audio encoding result; and / or, The decoding module is specifically configured to: The joint sub-network of the neural network model joints the output result of the i-th frame and the elements of the (i + 1)-th frame in the audio encoding result to obtain a first joint result; The decoding sub-network of the neural network model decodes the first joint result to obtain the output result of the (i + 1)-th frame.
21. The device according to claim 20, characterized in that, It further includes a first training module for: Input the training audio and the corresponding text labels into the neural network model, and the neural network model outputs a second joint result, where each element in the second joint result has coordinates in the audio frame number dimension and the text label sequence dimension, and probability values of each character type; Determine all first training paths within the second joint result, where in each first training path, each frame corresponds to a serial number, and when the output result of the i-th frame is a non-null character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinates in the audio frame number dimension and the text label sequence dimension of the output position of the i-th frame, and when the output result of the i-th frame is a null character, the output position of the (i + 1)-th frame is obtained by adding 1 to the coordinate in the audio frame number dimension of the output position of the i-th frame; Adjust the network parameters of the neural network model according to the sum of the probability values of all the first training paths.
22. The device according to claim 21, wherein, When the training module is used to determine all the first training paths within the second joint result, it is specifically configured to: For a preset proportion of the training audio among all the training audio, determine all the first training paths within the second joint result of the training audio; When the first training module is used to adjust the network parameters of the neural network model according to the sum of the probability values of all the first training paths, it is specifically configured to: Adjust the network parameters of the neural network model according to all the first training paths determined within the second joint result of each training audio in the training audio of the preset proportion; The apparatus for audio processing further includes a second training module, configured to: For all training audio, determine all first training paths and all second training paths in the second joint result of the training audio. Wherein, in each second training path, each frame corresponds to a serial number, and when the output result of the i-th frame is a non-empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the text label sequence dimension to obtain the output position of the (i + 1)-th frame; when the output result of the i-th frame is an empty character, the output position of the i-th frame is incremented by 1 in the coordinate of the audio frame number dimension to obtain the output position of the (i + 1)-th frame; Adjust the network parameters of the neural network model according to all the first training paths and all the second training paths determined in the second joint result of each training audio in all the training audio.
23. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory is used to store computer instructions that can be run on the processor, and the processor is used to perform the method according to any one of claims 1 to 11 when executing the computer instructions.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Coding method, decoding method, coding apparatus, and decoding apparatus
US20120127002A1