Audio Processing Method, Apparatus and Computer-Readable Storage Medium
By segmenting audio frames and identifying non-speech audio frames, combined with speech recognition and synthesis technology, the problem of low speech accuracy in audio style conversion is solved, and higher audio quality and cross-language labeling accuracy are achieved.
Patent Information
- Application Number
- CN202110872240.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-07-30
AI Technical Summary
When existing audio style conversion technology deals with abnormal voice, it causes speech accuracy to decrease and the content cannot be expressed correctly.
By segmenting audio frames, non-speech audio frames are identified and eliminated, and style conversion is used to use speech recognition technology and speech synthesis technology to improve audio quality and accuracy.
Effectively eliminate non-voice audio, improve the accuracy of audio style conversion, and ensure that the labeler can accurately realize cross-language audio annotation.
Smart Images

Figure CN113823287B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio processing method, device and computer-readable storage medium. Background Art
[0002] Audio is an important medium in multimedia. The speech in the audio is the sound produced by humans through their vocal organs, which has certain meaning and is used for social communication. To convert the style of audio, we convert the language type of the speech in the audio. For example, if the language type of the speech in the audio is the Weizang dialect, we can convert it to the Kham dialect.
[0003] There are usually some abnormal speech in the audio, such as humming, hesitation, laughter, shouting and other noises, which lead to poor accuracy in audio style conversion. That is, before and after the style conversion, the text information corresponding to the speech in the audio changes. For example, the text corresponding to the speech in the original audio originally asked "Where are you going", but after the style conversion, the text corresponding to the speech in the audio becomes "Are you going to eat?" At this time, although the style conversion can solve the problem of language barriers, it cannot express the content correctly. Therefore, it is very necessary to improve the accuracy of the speech involved in audio style conversion. Summary of the invention
[0004] The embodiments of the present application provide an audio processing method, device, and computer-readable storage medium, which can improve the accuracy of speech involved in audio style conversion.
[0005] On the one hand, an embodiment of the present application provides an audio processing method, the method comprising:
[0006] Acquire audio to be processed, where the audio to be processed includes one or more audio frames;
[0007] For any audio frame among the one or more audio frames, segment the audio frame to obtain a plurality of audio segments, determine an audio category of each audio segment among the plurality of audio segments, and determine a speech recognition result of the audio frame according to the audio category of each audio segment;
[0008] According to the speech recognition results of each audio frame, the audio frames whose speech recognition results are target recognition results in the audio to be processed are removed to obtain processed audio;
[0009] Performing style conversion processing on the processed audio to obtain target audio.
[0010] On the other hand, an embodiment of the present application provides an audio processing device, the device comprising:
[0011] An acquisition module, configured to acquire an audio to be processed, where the audio to be processed includes one or more audio frames;
[0012] A processing module, configured to perform segmentation processing on any one of the one or more audio frames to obtain a plurality of audio segments, determine the audio category of each audio segment in the plurality of audio segments, and determine the speech recognition result of the any one audio frame according to the audio category of each audio segment;
[0013] The processing module is further configured to remove the audio frames with the speech recognition result being the target recognition result in the audio to be processed according to the speech recognition results of the respective audio frames, so as to obtain a processed audio;
[0014] The processing module is further configured to perform style conversion processing on the processed audio to obtain a target audio.
[0015] Correspondingly, an embodiment of the present application provides a computer device, which includes a processor, a communication interface, and a memory. The processor, the communication interface, and the memory are interconnected. Wherein, the memory stores a computer program, and the processor is configured to call the computer program to execute the audio processing method according to any one of the above possible implementation manners.
[0016] Correspondingly, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the processor executes the computer program involved in the audio processing method according to any one of the above possible implementation manners.
[0017] Correspondingly, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio processing method according to any one of the above possible implementation manners.
[0018] In the embodiments of the present application, first, for any one of the one or more audio frames included in the audio to be processed, the any one audio frame is segmented to obtain a plurality of audio segments, the audio category of each audio segment in the plurality of audio segments is determined, the speech recognition result of the any one audio frame is determined according to the audio category of each audio segment, and then, according to the speech recognition results of the respective audio frames, the audio frames in the audio to be processed with the speech recognition result being the target recognition result are removed to obtain the processed audio. Finally, the processed audio is subjected to a style conversion process to obtain the target audio. The above audio processing method can remove the non-speech audio in the audio to be processed, thereby reducing external interference and improving the audio quality of the audio, which is beneficial to improving the accuracy of the speech involved in the audio style conversion. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0020] Figure 1 Schematic diagram of the architecture of an audio processing system provided by an embodiment of the present application;
[0021] Figure 2 Schematic diagram of the flow of an audio processing method provided by an embodiment of the present application;
[0022] Figure 3 Schematic diagram of the flow of another audio processing method provided by an embodiment of the present application;
[0023] Figure 4 Schematic diagram of the flow for determining the audio category provided by an embodiment of the present application;
[0024] Figure 5 Schematic diagram of the model structure of the x-vector model provided by an embodiment of the present application;
[0025] Figure 6 Schematic diagram of the flow of another audio processing method provided by an embodiment of the present application;
[0026] Figure 7 Schematic diagram of the principle of the speech recognition technology provided by an embodiment of the present application;
[0027] Figure 8 Schematic diagram of the principle of the speech synthesis technology provided by an embodiment of the present application;
[0028] Figure 9 Schematic diagram of the processing of the speech recognition technology provided by an embodiment of the present application;
[0029] Figure 10 It is a schematic flowchart of another audio processing method provided by an embodiment of the present application;
[0030] Figure 11 It is a schematic structural diagram of an audio processing device provided by an embodiment of the present application;
[0031] Figure 12 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific embodiments
[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0033] In order to improve the accuracy of audio style conversion, an audio processing method is proposed in the embodiments of the present application based on cloud technology and artificial intelligence skills.
[0034] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model. It can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0035] Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services according to needs. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0036] Artificial Intelligence (AI) technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic AI technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, cloud storage, big data processing technology, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0037] The key technologies of Speech Technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0038] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstrations.
[0039] With the research and progress of cloud technology and artificial intelligence technology, cloud technology and artificial intelligence technology are being studied and applied in multiple fields. The embodiments of this application involve cloud computing technology of cloud technology, as well as speech technology and machine learning of artificial intelligence in the process of realizing audio style conversion, and will be specifically described through the following embodiments.
[0040] Audio annotation is to convert the speech in the audio into text, which is one of the tasks of annotators. Even when using the same writing system, due to different regions, there will be language differences. For example, in the Tibetan writing system, it includes three major types of languages: Weizang dialect, Kangba dialect, and Amdo dialect. Before annotators convert the speech in the audio into text, due to the limitations of the languages usually mastered by annotators, such as not being proficient in Weizang dialect, Kangba dialect, and Amdo dialect at the same time, before converting the speech in the audio into text, it is necessary to determine the language type of the speech in the audio and let the corresponding annotators perform the audio annotation work. For example, if the language type of the speech in the audio is Weizang dialect, let the annotators proficient in Weizang dialect perform the audio annotation work. It can be seen that due to the inability of annotators to perform cross-language annotation, the selectivity of annotators in audio annotation is small, and the output quantity level of audio annotation is unbalanced. At the same time, it is also necessary to consider how annotators are invested. This application can perform style conversion on the audio, so that regardless of the language type of the audio, annotators can convert the language type of the speech in the audio into the language type they are good at, so as to perform audio annotation, which can solve problems such as difficult cross-language annotation, unbalanced output quantity level, and unreasonable investment of annotators caused by language barriers, and can effectively control the progress and delivery output of audio annotation, reducing progress management and investment costs.
[0041] Please refer to Figure 1 , Figure 1 FIG. is a schematic diagram of an audio processing system provided by an embodiment of the present application. The audio processing system may specifically include a terminal device 101 and a server 102, and the terminal device 101 and the server 102 are connected through a network, for example, through a wireless network connection or the like.
[0042] The terminal device 101 is also referred to as a terminal (Termina), a user equipment (UE), an access terminal, a user unit, a mobile device, a user terminal, a wireless communication device, a user agent, or a user device. The terminal device may be a smart TV, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC), a vehicle-mounted device, a wearable device, or other intelligent devices), but is not limited thereto.
[0043] The server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0044] In one embodiment, the terminal device 102 can send the audio to be processed to the server 102. The server 102 obtains the audio to be processed. For any one of the one or more audio frames included in the audio to be processed, the server 102 performs segmentation processing on the audio frame to obtain multiple audio segments, and obtains the audio category of each audio segment among the multiple audio segments. The server 102 determines the speech recognition result of any one audio frame according to the audio category of each audio segment, and according to the speech recognition results of each audio frame, the server 102 removes the audio frames in the audio to be processed whose speech recognition results are the target recognition results, to obtain the processed audio, and performs style conversion processing on the processed audio to obtain the target audio. Through this embodiment, the speech recognition result of any one audio frame can be determined according to the audio categories of the multiple audio segments included in the audio frame, and the non-speech audio in the audio to be processed can be screened according to the speech recognition results of each audio frame, which can improve the audio quality of the audio, so that the style conversion of the audio can be accurately achieved. When the annotator performs cross-language audio annotation, the annotator can also hear the audio that is identical to the text content of the original audio, which can solve the problem of difficult cross-language annotation.
[0045] It can be understood that the schematic diagram of the system architecture described in the embodiments of the present application is for more clearly explaining the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0046] As Figure 2 shown, it is an audio processing method provided by the audio processing system based on Figure 1 in the embodiments of the present application. Taking the server 102 mentioned in Figure 1 as an example. The following describes the method of the embodiments of the present application in conjunction with Figure 2 .
[0047] S201. Obtain the audio to be processed, where the audio to be processed includes one or more audio frames.
[0048] The audio to be processed refers to the audio that needs to undergo style conversion. Style conversion means converting the language type of the speech in the audio. For example, if the language type of the speech in the audio is the U-Tsang dialect, the language type of the speech in the audio can be converted to the Khampa dialect. An audio frame is the audio obtained after splitting the audio to be processed.
[0049] In one embodiment, the server can split the audio to be processed in a timed splitting manner to obtain one or more audio frames. For example, the audio to be processed is split once every 5 seconds of playback; it can also split the audio to be processed in an equal division manner to obtain one or more audio frames. For example, if the audio to be processed is split into 4 audio frames and the playback duration of the audio to be processed is 20 seconds, then each of the 4 obtained audio frames plays for 5 seconds; it can also obtain the audio waveform data of the audio to be processed (such as a time-domain graph, a frequency-domain graph, a spectrogram), and split the audio to be processed according to the audio waveform data to obtain one or more audio frames. For example, the audio to be processed is split at a frequency lower than the frequency threshold (such as the frequency threshold is 150 Hz) in the audio waveform data to obtain one or more audio frames.
[0050] S202. For any one of the one or more audio frames, split the any one audio frame to obtain multiple audio segments, determine the audio category of each audio segment among the multiple audio segments, and determine the speech recognition result of the any one audio frame according to the audio category of each audio segment.
[0051] An audio segment is the audio obtained after splitting an audio frame. The audio category refers to the nature of the audio. For example, the nature of the audio can be divided into speech, laughter, song, pure music, noise (such as background noise). Then the audio category can be speech, laughter, song, pure music, noise. The audio category can also be normal speech and abnormal speech. Normal speech is the sound that annotators usually need to annotate when performing audio annotation. For example, when the nature of the audio segment is speech (such as a conversation, a reading), since speech usually needs to be audio-annotated, the audio segment is normal speech. Abnormal speech is the sound that annotators usually do not need to annotate when performing audio annotation. For example, when the nature of the audio segment is laughter, song, pure music, noise, it usually does not need to be audio-annotated, so the audio segment is abnormal speech.
[0052] The speech recognition results include two types: one is speech audio, and the other is non-speech audio. Speech audio usually refers to the audio that includes normal speaking voices (such as conversations, readings), which is the audio that annotators usually need to annotate during audio annotation. Non-speech audio refers to the audio that includes abnormal speaking voices (such as laughter, songs, pure music, noise), which is the audio that annotators usually do not need to annotate during audio annotation. Therefore, the speech recognition results can also reflect the nature of the audio. For example, when the recognition result of an audio frame is speech audio, it means that the audio frame is the audio that annotators need to perform audio annotation on. When the recognition result of an audio frame is non-speech audio, it means that the audio frame is the audio that annotators do not need to perform audio annotation on.
[0053] In one embodiment, when the server obtains the speech recognition result of any one of the one or more audio frames included in the audio to be processed, it is necessary to perform a segmentation process on any one of the audio frames to obtain multiple audio segments included in any one of the audio frames. The server can obtain the audio waveform data of any one of the audio frames (such as a time-domain graph, a frequency-domain graph, a spectrogram), and use the audio waveform data of any one of the audio frames to perform a segmentation process on any one of the audio frames to obtain one or more audio segments.
[0054] Optionally, the server can use the audio silent interval (the time duration segment corresponding to an amplitude of zero or near zero) in the audio waveform data of any one of the audio frames to perform a segmentation process on any one of the video frames. For example, if the audio waveform data is a time-domain graph, the time-domain graph is a two-dimensional graph of the audio amplitude changing with time. That is, in the time-domain graph, the abscissa is the playing duration, and the ordinate is the amplitude (i.e., the sound intensity). Then, the playing duration interval corresponding to an amplitude of zero or near zero in the time-domain graph can be determined as the audio silent interval, and the audio silent interval can be used as a segmentation point to perform a segmentation process on any one of the audio frames to obtain multiple audio segments. Among them, the audio silent interval is also an audio segment.
[0055] Optionally, the server can also use the low-frequency (frequency lower than the frequency threshold) in the audio waveform data of any one of the audio frames to perform a segmentation process on any one of the video frames to obtain multiple audio segments.
[0056] In one embodiment, when the server determines the speech recognition result of any audio frame by using the audio category of each audio segment included in the any audio frame, the server may obtain the proportion of the audio segments with the audio category being the target category (such as normal speech or voice) in the any audio frame, and judge the speech recognition result of the any audio frame according to the proportion. For example, when the proportion of the audio segments with the audio category being normal speech in the any audio frame is greater than or equal to 50%, the speech recognition result of the any audio frame is speech audio; when the proportion of the audio segments with the audio category being normal speech is less than 50%, the speech recognition result of the any audio frame is non-speech audio. Suppose an audio frame includes 5 audio segments, and 3 of the 5 audio segments have the audio category of normal speech, then when the proportion of the audio segments with the audio category being normal speech is 60%, the speech recognition result of this any audio frame is speech audio; if 2 of the 5 audio segments have the audio category of normal speech, then when the proportion of the audio segments with the audio category being normal speech is 40%, the speech recognition result of this any audio frame is non-speech audio.
[0057] S203. According to the speech recognition results of each audio frame, remove the audio frames in the to-be-processed audio whose speech recognition results are the target recognition results, so as to obtain the processed audio.
[0058] The target recognition result refers to that the speech recognition result is non-speech audio. When the speech recognition result of an audio frame is the target recognition result, that is, when it is non-speech audio, it indicates that the audio frame includes abnormal speaking sounds (such as abnormal speech, laughter, songs, pure music, noise), and the annotator does not need to perform audio annotation on this audio frame.
[0059] In one embodiment, the server obtains the speech recognition results of each audio frame, and when it determines that the speech recognition result of an audio frame is the target recognition result, removes this audio frame from the to-be-processed audio. This embodiment can remove the non-speech audio in the to-be-processed audio, so that the audio quality of the processed audio is higher than that of the to-be-processed audio. When performing style conversion later, external interference can be reduced and the accuracy can be improved.
[0060] S204. Perform style conversion processing on the processed audio to obtain the target audio.
[0061] The style conversion processing of the processed audio refers to converting the language type of the speech in the audio. For example, if the language type of the speech in the audio is Weizang dialect, the language type of the speech in the audio can be converted to Khampa dialect.
[0062] In one embodiment, the server may perform a style conversion process on the processed audio to obtain the target audio. For example, the text information of the processed audio is obtained by using speech recognition technology, and then the text information of the processed audio is processed by using text-to-speech technology to obtain the target audio. Since the non-speech audio in the audio is removed, the text information of the obtained audio is more accurate, and thus the audio synthesized by speech is also more accurate.
[0063] Among them, speech recognition technology is the process of recognizing the speech of others to obtain text information, and text-to-speech (TTS) technology is the process of converting any text information (such as help files or web pages) into standard and fluent speech in real time.
[0064] In the embodiment of the present application, the server obtains the audio to be processed, performs segmentation processing on any one of the one or more audio frames included in the audio to be processed to obtain a plurality of audio segments, determines the audio category of each audio segment in the plurality of audio segments, determines the speech recognition result of any one audio frame according to the audio category of each audio segment, and removes the audio frames with the speech recognition result as the target recognition result in the audio to be processed according to the speech recognition results of each audio frame to obtain the processed audio, and performs a style conversion process on the processed audio to obtain the target audio. This embodiment can use the speech recognition results of each audio frame included in the audio to be processed to remove the non-speech audio in the audio to be processed, which can reduce external interference, improve the audio quality of the audio, and is beneficial to improving the accuracy of the speech involved in the audio style conversion. At the same time, since the audio after style conversion has high accuracy, the annotator can also more accurately perform cross-language audio annotation.
[0065] As Figure 3 shown, it is another audio processing method provided by the audio processing system based on Figure 1 of the embodiment of the present application. Taking the server 102 mentioned in Figure 1 as an example. The following describes the method of the embodiment of the present application in conjunction with Figure 3 .
[0066] S301. Obtain the audio to be processed, where the audio to be processed includes one or more audio frames.
[0067] The specific implementation of S301 may refer to the relevant description of S201 in the foregoing embodiment, and will not be elaborated here.
[0068] S302. For any one of the one or more audio frames, perform segmentation processing on the any one audio frame to obtain a plurality of audio segments.
[0069] In one embodiment, the server performs segmentation processing on any audio frame to obtain multiple audio segments, including the following steps:
[0070] (1) Obtain the audio waveform data of any audio frame.
[0071] (2) According to the audio waveform data, determine the audio silent interval in any audio frame as the segmentation point.
[0072] (3) Perform segmentation processing on any audio frame according to the determined segmentation points to obtain multiple audio segments.
[0073] Optionally, the server can obtain the audio waveform data of any audio frame (such as time-domain graph, frequency-domain graph, spectrogram), determine the audio silent interval (the duration segment corresponding to zero amplitude or near-zero amplitude) in any audio frame according to the audio waveform data, and use the audio silent interval as the segmentation point to perform segmentation processing on any audio frame to obtain multiple audio segments, where the audio silent interval is also an audio segment.
[0074] Optionally, the server can also perform segmentation processing on any video frame at the low frequency (frequency lower than the frequency threshold) in the audio waveform data of any audio frame to obtain multiple audio segments.
[0075] S303. Determine the audio category of each audio segment in the multiple audio segments.
[0076] As Figure 4 shown, when the server determines the audio category of each audio segment in the multiple audio segments included in the audio frame, first, for any audio segment in the multiple audio segments, perform feature extraction on any audio segment to obtain the speech feature of any audio segment.
[0077] Among them, the speech feature is an acoustic feature, and the acoustic feature refers to a physical quantity representing the acoustic characteristics of the audio. For example, the speech feature can be Mel Frequency Cepstrum Coefficient (MFCC), Fbank (Filter Bank) feature, Linear Prediction Coefficient (LPC). Any audio segment is a digital audio, that is, the any audio segment is a digital audio signal composed of binary numbers 1 or 0, which is convenient for machine processing.
[0078] Then the server uses the feature processing module of the audio classification model to process the speech feature of any audio segment to obtain the speech feature vector of any audio segment.
[0079] Among them, the feature processing module is used to extract speech feature vectors, which can be machine learning models, such as Gaussian Mixture Model (GMM), Time-Delay Neural Network (TDNN), Universal Background Model (UBM), x-vector model, and so on.
[0080] Optionally, this application uses the x-vector model in the field of speaker recognition (i.e., identifying who the speaker is) to process the speech features of any audio segment, and obtains the speech feature vector of any audio segment. Among them, the x-vector model is as Figure 5 shown, including a frame processing layer, a statistical pooling layer, and a segment processing layer. The frame processing layer is composed of 5 layers of TDNN. After the statistical pooling layer calculates the mean and standard deviation of the output of the frame processing layer respectively, it concatenates the output mean and standard deviation. The segment processing layer extracts segment-level vectors through 2 layers of forward DNN (Deep Neural Network) to represent the speaker. This application can use any one of the 2 layers of DNN as the speech feature vector of any audio segment. For example Figure 5 the speech feature vector 1 and the speech feature vector 2 in can both be used as the speech feature vector of any audio segment.
[0081] Finally, the server uses the classification processing module of the audio classification model to process the speech feature vector of any audio segment, and obtains the audio category of any audio segment.
[0082] Among them, the classification processing module is used to perform classification tasks and can be a classification model, such as Probabilistic Linear Discriminant Analysis (LDAP), Logistic Regression (LR), Support Vector Machine (SVM), and so on.
[0083] In one embodiment, the server can use the x-vector model to obtain the speech feature vectors of multiple audios, use the speech feature vectors of multiple audios to train the parameters of the initial classification model, and after completion, use the trained classification model as the classification processing module in the audio classification model.
[0084] In one embodiment, the server processes the speech feature vector of any audio clip using the classification processing module of the audio classification model, and the probability that the audio clip belongs to each audio category can be obtained. For example, the audio categories include speech, laughter, songs, pure music, and noise. The probabilities that the audio clip belongs to speech, laughter, songs, pure music, and noise are 0.5, 0.1, 0.2, 0, and 0.2 respectively. The server can use the audio category corresponding to the maximum probability as the audio category of the audio clip, that is, the audio category of the audio clip is speech.
[0085] S304. Determine the proportion of audio clips with the audio category of the target category in any audio frame according to the audio category of each audio clip, and determine the speech recognition result of any audio frame according to the proportion.
[0086] The target category is the audio category that needs to be labeled by the annotator when annotating audio. For example, the audio categories include normal speech and abnormal speech, and the target category can be normal speech; or, the audio categories include speech, laughter, songs, pure music, and noise, and the target category can be speech.
[0087] In one embodiment, the server obtains the proportion of audio clips with the audio category of the target category in any audio frame according to the audio category of each audio clip included in any audio frame, and determines the speech recognition result of any audio frame according to the proportion. For example, the audio categories include speech, laughter, songs, pure music, and noise, and the target category is speech. Assume that any audio frame includes 5 audio clips, and 3 of the 5 audio clips have the audio category of speech. Then the proportion of audio clips with the audio category of the target category in any audio frame is 60%; or the audio categories include normal speech and abnormal speech, and the target category is normal speech. Any audio frame includes 5 audio clips, and 3 of the 5 audio clips have the audio category of normal speech. Then the proportion of audio clips with the audio category of the target category in any audio frame is 60%. The server can set a proportion threshold (which can be set manually). This proportion threshold is a percentage value. When the obtained proportion is less than the proportion threshold, for example, the proportion threshold is 70% and the proportion is 60%, at this time, it can be determined that the speech recognition result of any audio frame is the target recognition result, that is, it is determined that any audio frame is a non-speech audio, which is an audio that the annotator does not need to annotate.
[0088] In one embodiment, when the server determines the speech recognition result of any audio frame according to the proportion, it can also obtain the prediction value that any audio frame is speech, and use the proportion and the prediction value to obtain the reference probability that any audio frame is a speech audio. When the reference probability is less than the probability threshold, the speech recognition result of any audio frame is determined as the target recognition result, that is, it is determined that any audio frame is a non-speech audio, which is an audio that the annotator does not need to annotate.
[0089] The predicted value of an audio frame being speech refers to the predicted probability that the audio category of the audio frame is speech. When the server obtains the predicted value of the speech of an audio frame, the audio frames can be classified into two categories: one is speech and the other is non-speech. It should be noted that the speech here is essentially different from the speech, laughter, songs, pure music, and noise included in the audio category in the foregoing embodiments. When the audio category of an audio frame is speech, the speech includes the speech, laughter, songs, pure music, and noise in the foregoing audio category, and the sounds other than the speech, laughter, songs, pure music, and noise in the foregoing audio category are regarded as non-speech. In addition, when the audio category of an audio frame is speech, the speech also includes the normal speech and abnormal speech in the foregoing audio category, and the sounds other than the normal speech and abnormal speech in the foregoing audio category are regarded as non-speech.
[0090] In one embodiment, the predicted value of any audio frame being speech can be used as a weight value to process the proportion, and the reference probability of the audio frame being speech audio can be obtained. For example, if the predicted value of any audio frame being speech is 0.8 and the proportion is 0.6, the predicted value and the proportion value can be multiplied as the reference probability. For example, multiplying the predicted value 0.8 by the proportion 0.6 gives the reference probability 0.48. When the reference probability is less than the probability threshold (which can be set artificially), for example, the probability threshold is set to 0.5, and at this time the reference probability is less than the probability threshold, the server determines the speech recognition result of any audio frame as the target recognition result, that is, determines that any audio frame is non-speech audio, which is the audio that the annotator does not need to annotate.
[0091] S305. According to the speech recognition results of each audio frame, remove the audio frames in the audio to be processed whose speech recognition results are the target recognition results, and obtain the processed audio.
[0092] Among them, for the specific implementation of S305, reference can be made to the relevant description of S203 in the foregoing embodiments, which will not be elaborated here.
[0093] S306. Perform style conversion processing on the processed audio to obtain the target audio.
[0094] The target language is the language type of the speech in the audio after style conversion. For example, if the language type of the target audio is the Khampa dialect, the target language is the Khampa dialect.
[0095] In one embodiment, as Figure 6 shown, the server can call the speech recognition interface to obtain the speech recognition model to perform speech recognition processing on the processed audio, obtain the text information of the processed audio, and call the speech synthesis interface to obtain the speech synthesis model corresponding to the target language (for example, the Khampa speech synthesis model corresponding to the Khampa dialect) to perform speech synthesis processing on the text information of the processed audio to obtain the target audio.
[0096] The speech recognition model can be obtained based on speech recognition technology. For example, Figure 7 As shown, it is the schematic diagram of speech recognition technology. The speech recognition technology first preprocesses the processed audio, such as filtering, framing, extracting speech features, etc. Then, the decoder uses the acoustic model, language model, speech data, text data, and pronunciation dictionary to decode the audio to obtain the most likely word sequence, thereby obtaining the text information of the processed audio. The speech recognition model can also be a model obtained through end-to-end training based on deep learning.
[0097] The speech synthesis model can be obtained based on speech synthesis technology. For example, Figure 8 As shown, it is the schematic diagram of speech synthesis technology. Its core idea is to store natural speech waveforms to form a large-scale sound library, and then perform text analysis on the text information of the processed audio, so as to select appropriate waveforms from the large-scale sound library and splice them together. The speech synthesis model can also be a model obtained through end-to-end training based on deep learning.
[0098] In an embodiment, when the user clicks the "Online Recognition" function key on the display interface of the terminal device, the intelligent terminal obtains the processed audio and calls the speech recognition interface, thereby using the speech recognition model to perform speech recognition processing on the processed audio to obtain the text information of the processed audio. For example, Figure 9 As shown, after the intelligent terminal calls the Tibetan speech recognition interface to perform speech recognition processing on the Tibetan audio, it can add its text information to the text box area below the audio band. When the user clicks the "Text-to-Speech" function key again, "Please select the target language" can be displayed on the display interface. For example, "Tibetan - U-Tsang", "Tibetan - Amdo", and "Tibetan - Khams" can be displayed as three sub-items. The user selects one of them as the target language. The intelligent terminal accesses the corresponding speech synthesis interface according to the target language selected by the user, thereby calling the speech synthesis model corresponding to the target language to perform speech synthesis processing on the text information of the processed audio to obtain the target audio. If the target language is the Khams dialect, the speech synthesis model can convert the text information into an audio in the Khams dialect. The server can play the target audio on the display interface of the intelligent terminal and store the target audio at the same time, facilitating the user to listen to it repeatedly.
[0099] In one embodiment, the server may obtain a basic pronunciation vocabulary, which includes a mapping relationship between text and one or more language pronunciations. For example, as shown in Table 1 below, it is a partial example of a basic pronunciation vocabulary for Tibetan, including mapping relationships between common Tibetan and U-Tsang pronunciation, Amdo pronunciation, and Kham pronunciation. The server may perform speech recognition processing on the processed audio to obtain the language pronunciation of each word segment in the processed audio. The language pronunciation of the word segment obtained by the server may be the language pronunciation of a character, such as the Mandarin pronunciation "xi" of "west", or the language pronunciation of a word, such as the Mandarin pronunciation "xi an" of "Xi'an", or the language pronunciation of a sentence, such as the Mandarin pronunciation "xi an shi gu du" of "Xi'an is an ancient capital". After obtaining the language pronunciation of each word segment in the processed audio, the server may use the basic pronunciation vocabulary to obtain text information of the processed audio. For example, if the language pronunciation of the word segment included in the processed audio is "gafkaf", the text information of the processed audio is
[0100] Table 1
[0101]
[0102] In one embodiment, after the server obtains the text information of the processed audio using the basic pronunciation vocabulary, it can obtain the target audio according to the target language, the text information of the processed audio and the basic pronunciation vocabulary. For example, the text information of the processed audio is If the target language is Kham dialect, the Kham pronunciation corresponding to the text information can be obtained from the basic pronunciation vocabulary, for example The corresponding Kangba pronunciation is "gàkǎ", and the target audio can be generated according to the Kangba pronunciation.
[0103] In one embodiment, considering that there may be different pronunciations for the same text even in the same language type, the present application can establish a special pronunciation vocabulary based on this. In the special pronunciation vocabulary, the text has a mapping relationship with multiple voice pronunciations in a language type, that is, there can be multiple pronunciations for the same text in the same language type, which can improve the accuracy of speech recognition and speech synthesis.
[0104] like Figure 10As shown, the server can first preprocess the Tibetan audio. The preprocessing methods are the aforementioned steps S201, S202, S203, or S301, S302, S303, S304, which will not be elaborated in this embodiment. Then, the server can use Tibetan language recognition technology to obtain the text information in the Tibetan audio, thereby obtaining the speech recognition result. The Tibetan audio can be the audio included in the video. Finally, the server can use Tibetan language synthesis technology to generate the target audio based on the text information in the Tibetan audio. In this embodiment, the Tibetan audio can be converted into multiple language types for playback, realizing the autonomous selection of audio. The annotator does not need to judge the language type of the audio, and can perform audio annotation, reducing the difficulty of cross-language annotation and achieving non-differentiation of language types.
[0105] In this application, the server performs segmentation processing on any one of the one or more audio frames included in the audio to be processed to obtain multiple audio segments, determines the audio category of each audio segment among the multiple audio segments, determines the speech recognition result of any one audio frame according to the audio category of each audio segment, and according to the speech recognition results of each audio frame, eliminates the audio frames in the audio to be processed whose speech recognition results are the target recognition results to obtain the processed audio. Finally, style conversion processing is performed on the processed audio to obtain the target audio, which can eliminate the non-speech audio in the audio to be processed, improve the audio quality, and make this application more accurate in style conversion compared to directly performing style conversion on the audio to be processed. At the same time, because the audio after style conversion has high accuracy, the annotator can also more accurately perform cross-language audio annotation.
[0106] The method of the embodiment of the present application is elaborated in detail above. To facilitate better implementation of the above solution of the embodiment of the present application, correspondingly, the device of the embodiment of the present application is provided below. Please refer to Figure 11 , Figure 11 is a schematic structural diagram of an audio processing device provided by an exemplary embodiment of the present application. The device 110 may include:
[0107] An acquisition module 1101, configured to acquire the audio to be processed, where the audio to be processed includes one or more audio frames;
[0108] A processing module 1102, configured to perform segmentation processing on any one of the one or more audio frames to obtain multiple audio segments, determine the audio category of each audio segment among the multiple audio segments, and determine the speech recognition result of the any one audio frame according to the audio category of each audio segment;
[0109] The processing module 1102 is further configured to remove the audio frames in the to-be-processed audio whose speech recognition results are the target recognition results according to the speech recognition results of each audio frame, so as to obtain the processed audio;
[0110] The processing module 1102 is further configured to perform a style conversion process on the processed audio to obtain the target audio.
[0111] In one embodiment, the processing module 1102 is specifically configured to:
[0112] For any one of the multiple audio segments, perform feature extraction on the any one of the audio segments to obtain the speech features of the any one of the audio segments;
[0113] Use the feature processing module of the audio classification model to process the speech features of the any one of the audio segments to obtain the speech feature vector of the any one of the audio segments;
[0114] Use the classification processing module of the audio classification model to process the speech feature vector of the any one of the audio segments to obtain the audio category of the any one of the audio segments.
[0115] In one embodiment, the processing module 1102 is specifically configured to:
[0116] According to the audio category of each audio segment, determine the proportion of the audio segments in the any one of the audio frames whose audio category is the target category;
[0117] Determine the speech recognition result of the any one of the audio frames according to the proportion.
[0118] In one embodiment, the processing module 1102 is specifically configured to:
[0119] When the proportion is less than the proportion threshold, determine that the speech recognition result of the any one of the audio frames is the target recognition result, and the target recognition result is used to indicate that the any one of the audio frames is non-speech audio.
[0120] In one embodiment, the processing module 1102 is specifically configured to:
[0121] For the any one of the audio frames, determine the prediction value that the any one of the audio frames is speech;
[0122] Wherein, the determining the speech recognition result of the any one of the audio frames according to the proportion includes:
[0123] Determine the reference probability that the any one of the audio frames is speech audio according to the proportion and the prediction value;
[0124] When the reference probability is less than the probability threshold, determine the speech recognition result of any one of the audio frames as the target recognition result, where the target recognition result is used to indicate that any one of the audio frames is non-speech audio.
[0125] In one embodiment, the processing module 1102 is specifically configured to:
[0126] Perform speech recognition processing on the processed audio to obtain the text information of the processed audio;
[0127] Determine the target language, and perform speech synthesis processing on the text information of the processed audio according to the target language to obtain the target audio.
[0128] In one embodiment, the processing module 1102 is specifically configured to:
[0129] Obtain a basic pronunciation vocabulary, where the basic pronunciation vocabulary includes the mapping relationship between text and one or more language pronunciations;
[0130] Perform speech recognition processing on the processed audio to determine the language pronunciation of each word segment in the processed audio;
[0131] Determine the text information of the processed audio according to the language pronunciation of each word segment and the basic pronunciation vocabulary.
[0132] In one embodiment, the processing module 1102 is specifically configured to:
[0133] Obtain the audio waveform data of any one of the audio frames;
[0134] According to the audio waveform data, determine the audio silent interval in any one of the audio frames as the segmentation point;
[0135] Perform segmentation processing on any one of the audio frames according to the determined segmentation points to obtain a plurality of audio segments.
[0136] In the embodiment of the present application, the server obtains the audio to be processed, performs segmentation processing on any one of the one or more audio frames included in the audio to be processed to obtain a plurality of audio segments, determines the audio category of each audio segment in the plurality of audio segments, determines the speech recognition result of any one of the audio frames according to the audio category of each audio segment, and according to the speech recognition results of each audio frame, removes the audio frames with the speech recognition result as the target recognition result in the audio to be processed to obtain the processed audio, and performs style conversion processing on the processed audio to obtain the target audio. This embodiment can use the speech recognition results of each audio frame included in the audio to be processed to remove the non-speech audio in the audio to be processed, reduce external interference, improve the audio quality of the audio, and is beneficial to improving the accuracy of the speech involved in the audio style conversion.
[0137] As Figure 12 shown Figure 12 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The internal structure of the computer device 120 is as Figure 12 shown, including: one or more processors 1201, a memory 1202, and a communication interface 1203. The above-mentioned processors 1201, memory 1202, and communication interface 1203 can be connected through a bus 1204 or other means. In the embodiment of the present application, taking the connection through the bus 1204 as an example.
[0138] Among them, the processor 1201 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device 120. It can parse various instructions in the computer device 120 and process various data of the computer device 120. For example: The CPU can be used to parse the power-on and power-off instructions sent by the user to the computer device 120 and control the computer device 120 to perform power-on and power-off operations; Another example: The CPU can transmit various interactive data between the internal structures of the computer device 120, and so on. The communication interface 1203 may optionally include a standard wired interface, a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), and is controlled by the processor 1201 to receive and send data. The memory 1202 (Memory) is a memory device in the computer device 120, used to store programs and data. It can be understood that the memory 1202 here can include both the built-in memory of the computer device 120, and of course can also include the extended memory supported by the computer device 120. The memory 1202 provides a storage space, and this storage space stores the operating system of the computer device 120, which may include but is not limited to: Windows system, Linux system, etc. The present application does not make any limitations on this.
[0139] In one embodiment, the processor 1201 is specifically used for:
[0140] Obtain the audio to be processed, where the audio to be processed includes one or more audio frames;
[0141] For any one of the one or more audio frames, perform segmentation processing on the any one audio frame to obtain a plurality of audio segments, determine the audio category of each audio segment in the plurality of audio segments, and determine the speech recognition result of the any one audio frame according to the audio category of each audio segment;
[0142] According to the speech recognition results of each audio frame, remove the audio frames in the audio to be processed whose speech recognition results are the target recognition results to obtain the processed audio;
[0143] Perform a style conversion process on the processed audio to obtain the target audio.
[0144] In one embodiment, the processor 1201 is specifically configured to:
[0145] For any one of the multiple audio segments, extract features from the any one audio segment to obtain the speech features of the any one audio segment;
[0146] Use the feature processing module of the audio classification model to process the speech features of the any one audio segment to obtain the speech feature vector of the any one audio segment;
[0147] Use the classification processing module of the audio classification model to process the speech feature vector of the any one audio segment to obtain the audio category of the any one audio segment.
[0148] In one embodiment, the processor 1201 is specifically configured to:
[0149] According to the audio category of each audio segment, determine the proportion of the audio segments with the audio category being the target category in the any one audio frame;
[0150] Determine the speech recognition result of the any one audio frame according to the proportion.
[0151] In one embodiment, the processor 1201 is specifically configured to:
[0152] When the proportion is less than the proportion threshold, determine the speech recognition result of the any one audio frame as the target recognition result, and the target recognition result is used to indicate that the any one audio frame is non-speech audio.
[0153] In one embodiment, the processor 1201 is specifically configured to:
[0154] For the any one audio frame, determine the prediction value that the any one audio frame is speech;
[0155] Wherein, the determining the speech recognition result of the any one audio frame according to the proportion includes:
[0156] Determine the reference probability that the any one audio frame is speech audio according to the proportion and the prediction value;
[0157] When the reference probability is less than the probability threshold, determine the speech recognition result of the any one audio frame as the target recognition result, and the target recognition result is used to indicate that the any one audio frame is non-speech audio.
[0158] In one embodiment, the processor 1201 is specifically configured to:
[0159] Perform speech recognition processing on the processed audio to obtain the text information of the processed audio;
[0160] Determine the target language, and perform speech synthesis processing on the text information of the processed audio according to the target language to obtain the target audio.
[0161] In one embodiment, the processor 1201 is specifically configured to:
[0162] Obtain a basic pronunciation word list, where the basic pronunciation word list includes the mapping relationship between characters and one or more language pronunciations;
[0163] Perform speech recognition processing on the processed audio to determine the language pronunciation of each segmented word in the processed audio;
[0164] Determine the text information of the processed audio according to the language pronunciation of each segmented word and the basic pronunciation word list.
[0165] In one embodiment, the processor 1201 is specifically configured to:
[0166] Obtain the audio waveform data of any audio frame;
[0167] According to the audio waveform data, determine the audio silent interval in any audio frame as the segmentation point;
[0168] Perform segmentation processing on any audio frame according to the determined segmentation points to obtain multiple audio segments.
[0169] In the embodiment of the present application, the server obtains the audio to be processed, performs segmentation processing on any audio frame in one or more audio frames included in the audio to be processed to obtain multiple audio segments, determines the audio category of each audio segment in the multiple audio segments, determines the speech recognition result of any audio frame according to the audio category of each audio segment, and according to the speech recognition results of each audio frame, eliminates the audio frames in the audio to be processed whose speech recognition results are the target recognition results to obtain the processed audio, and performs style conversion processing on the processed audio to obtain the target audio. This embodiment can eliminate the non-speech audio in the audio to be processed by using the speech recognition results of each audio frame included in the audio to be processed, reduce external interference, improve the audio quality of the audio, and is beneficial to improving the accuracy of the speech involved in the audio style conversion.
[0170] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above audio processing method. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0171] One or more embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps performed in the embodiments of the above various methods.
[0172] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An audio processing method, characterized in that, The method includes: Obtaining the audio to be processed, where the audio to be processed includes one or more audio frames; For any one of the one or more audio frames, determining a predicted value that the any one audio frame is speech, and performing segmentation processing on the any one audio frame to obtain a plurality of audio segments, determining the audio category of each audio segment in the plurality of audio segments, according to the audio category of each audio segment, determining the proportion of the audio segments with the audio category being the target category in the any one audio frame, processing the proportion with the predicted value that the any one audio frame is speech as a weight value to obtain a reference probability that the any one audio frame is speech audio, and determining the speech recognition result of the any one audio frame according to the reference probability; According to the speech recognition results of each audio frame, removing the audio frames in the audio to be processed whose speech recognition results are the target recognition results to obtain the processed audio; Performing style conversion processing on the processed audio to obtain the target audio.
2. The method according to claim 1, characterized in that, The determining the audio category of each audio segment in the plurality of audio segments includes: For any one of the plurality of audio segments, performing feature extraction on the any one audio segment to obtain the speech feature of the any one audio segment; Using the feature processing module of the audio classification model to process the speech feature of the any one audio segment to obtain the speech feature vector of the any one audio segment; Using the classification processing module of the audio classification model to process the speech feature vector of the any one audio segment to obtain the audio category of the any one audio segment.
3. The method according to claim 1, wherein The determining the speech recognition result of the any one audio frame according to the reference probability includes: When the reference probability is less than the probability threshold, determining the speech recognition result of the any one audio frame as the target recognition result, where the target recognition result is used to indicate that the any one audio frame is non-speech audio.
4. The method according to claim 1, characterized in that, The performing style conversion processing on the processed audio to obtain the target audio includes: Performing speech recognition processing on the processed audio to obtain the text information of the processed audio; Determining the target language, and performing speech synthesis processing on the text information of the processed audio according to the target language to obtain the target audio.
5. The method according to claim 4, characterized in that, The performing speech recognition processing on the processed audio to obtain the text information of the processed audio includes: Obtaining a basic pronunciation word list, where the basic pronunciation word list includes the mapping relationship between characters and one or more language pronunciations; Performing speech recognition processing on the processed audio to determine the language pronunciation of each word segment in the processed audio; According to the language pronunciation of each word segment and the basic pronunciation word list, determining the text information of the processed audio.
6. The method according to claim 1, characterized in that, The performing segmentation processing on the any one audio frame to obtain a plurality of audio segments includes: Obtaining the audio waveform data of the any one audio frame; According to the audio waveform data, determining the audio silent intervals in the any one audio frame as segmentation points; Performing segmentation processing on the any one audio frame according to the determined segmentation points to obtain a plurality of audio segments.
7. An audio processing device, characterized in that, The device includes: An acquisition module, configured to acquire an audio to be processed, where the audio to be processed includes one or more audio frames; A processing module, configured to, for any one of the one or more audio frames, determine a predicted value that the any one audio frame is speech, and perform segmentation processing on the any one audio frame to obtain a plurality of audio segments, determine the audio category of each audio segment in the plurality of audio segments, according to the audio category of each audio segment, determine the proportion of the audio segments with the audio category being the target category in the any one audio frame, use the predicted value that the any one audio frame is speech as a weight value to process the proportion, obtain a reference probability that the any one audio frame is a speech audio, and determine the speech recognition result of the any one audio frame according to the reference probability; The processing module is further configured to, according to the speech recognition results of each audio frame, remove the audio frames with the speech recognition result being the target recognition result in the audio to be processed, to obtain a processed audio; The processing module is further configured to perform style conversion processing on the processed audio to obtain a target audio.
8. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a processor, the audio processing method according to any one of claims 1 to 6 is implemented.
9. A computer device, characterized in that, The computer device includes a processor, a communication interface, and a memory, the processor, the communication interface, and the memory are connected to each other, wherein the memory stores a computer program, and the processor is configured to call the computer program to implement the audio processing method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program or instruction, the computer program or instruction is stored in a computer-readable storage medium, the processor reads the computer program or instruction from the computer-readable storage medium, and the processor executes the computer program or instruction to implement the audio processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Audio classification model training and junk audio recognition method and device
CN111816170A
Video translation method, system and device and storage medium
CN112562721A