Audio Processing Method, Apparatus, Device, Program Product and Storage Medium

The proposed method uses an attention-based neural network to fuse audio and phoneme features, enhancing phoneme alignment accuracy in speech processing systems by accurately aligning audio frames with phonemes.

CN114360504BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111421900.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-07-15
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

In the prior art, phonemes cannot be accurately aligned with text, resulting in inaccurate phoneme alignment.

Method used

Using an audio processing method based on artificial intelligence, the characteristics of phonemes and audio frames are acquired, and the attention mechanism is used to perform fusion processing, and the start and end time of phonemes is determined, and the training and loss function optimization are combined with phoneme classification network and loudness classification network.

Benefits of technology

The accuracy of phoneme alignment is improved, and the fusion characteristics are obtained through attention mechanism calculation, which effectively characterizes the relationship between audio frames and phonemes, and improves the accuracy of phoneme classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360504B_ABST
    Figure CN114360504B_ABST
Patent Text Reader

Abstract

The present application provides an audio processing method, apparatus, electronic device, computer program product, and computer-readable storage medium based on artificial intelligence; the method includes: obtaining at least one phoneme of a given text and determining the phoneme features of each of the phonemes; obtaining at least one audio frame of audio data corresponding to the given text and determining the audio features of each of the audio frames; performing the following processing for each of the audio frames: performing an attention mechanism-based fusion process on the audio features of the audio frame and the phoneme features of at least one of the phonemes to obtain a fusion feature corresponding to each of the audio frames; determining the phoneme corresponding to each of the audio frames based on the fusion feature of each of the audio frames, and determining the start and end times of each of the phonemes in the audio data based on the phoneme corresponding to each of the audio frames. Through the present application, the accuracy of phoneme alignment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology, and in particular, to an audio processing method, apparatus, electronic device, computer program product, and computer-readable storage medium based on artificial intelligence. Background Art

[0002] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0003] More and more artificial intelligence products have the function of voice interaction. Voice interaction can be applied to various voice scoring systems. For example, language test systems for language education applications, oral examination systems, etc. In the process of using the voice interaction function, it is necessary to align phonemes with text. However, in related technologies, it is impossible to accurately align phonemes with text. Summary of the Invention

[0004] Embodiments of the present application provide an audio processing method, apparatus, electronic device, computer program product, and computer-readable storage medium based on artificial intelligence, which can improve the accuracy of phoneme alignment.

[0005] The technical solution of the embodiments of the present application is implemented as follows:

[0006] Embodiments of the present application provide an audio processing method based on artificial intelligence, including:

[0007] Obtain at least one phoneme of a given text, and determine the phoneme feature of each said phoneme;

[0008] Obtain at least one audio frame of audio data corresponding to the given text, and determine the audio feature of each said audio frame;

[0009] For each said audio frame, perform the following processing: perform fusion processing based on an attention mechanism on the audio feature of the audio frame and the phoneme features of at least one said phoneme to obtain a fusion feature corresponding to each said audio frame;

[0010] Based on the fusion feature of each said audio frame, determine the phoneme corresponding to each said audio frame, and based on the phoneme corresponding to each said audio frame, determine the start and end times of each said phoneme.

[0011] Embodiments of the present application provide an audio processing apparatus based on artificial intelligence, including:

[0012] A phoneme module, configured to obtain at least one phoneme of a given text, and determine the phoneme feature of each said phoneme;

[0013] An audio module, configured to obtain at least one audio frame of audio data corresponding to the given text, and determine the audio feature of each audio frame;

[0014] A fusion module, configured to perform the following processing for each audio frame: perform an attention mechanism-based fusion processing on the audio feature of the audio frame and the phoneme features of at least one phoneme, to obtain a fusion feature corresponding to each audio frame;

[0015] An alignment module, configured to determine the phoneme corresponding to each audio frame based on the fusion feature of each audio frame, and determine the start and end times of each phoneme based on the phoneme corresponding to each audio frame.

[0016] In the above solution, the determination of the audio feature of each audio frame is implemented by calling an audio encoder, the audio encoder includes a plurality of convolutional networks and a normalization network, and the audio module is further configured to: perform feature extraction processing on the at least one audio frame through the plurality of cascaded convolutional networks, to obtain a convolutional feature extraction result corresponding to each audio frame; perform normalization processing on the convolutional feature extraction result of each audio frame through the normalization network, to obtain the audio feature of each audio frame.

[0017] In the above solution, the determination of the phoneme feature of each phoneme is implemented by calling a phoneme encoder, the phoneme encoder includes a phoneme characteristic representation network and a phoneme position representation network, and the phoneme module is further configured to: perform the following processing for each phoneme: determine the characteristic representation feature of the phoneme through the phoneme characteristic representation network, where the characteristic representation feature is used to characterize the characteristics of the phoneme; determine the position representation feature of the phoneme through the phoneme position representation network, where the position representation feature is used to characterize the position of the phoneme in the corresponding text unit; perform an addition process on the position representation feature and the characteristic representation feature, to obtain the phoneme feature of the phoneme.

[0018] In the above solution, the attention mechanism-based fusion processing is implemented by calling an attention fusion network, the attention fusion network includes an attention layer and a fusion layer, and the fusion module is further configured to: perform attention processing on the audio feature of the audio frame and the phoneme features of at least one phoneme through the attention layer, to obtain an attention result; perform a fusion process on the attention result and the audio feature through the fusion layer, to obtain a fusion feature corresponding to the audio frame.

[0019] In the above solution, the fusion module is further configured to perform the following processing for each of the phonemes: based on the audio features of the audio frame and the phoneme features of the phoneme, determine the attention score corresponding to the phoneme; perform a value vector transformation process on the phoneme features of the phoneme to obtain a value vector; multiply the attention score corresponding to the phoneme by the value vector to obtain the attention result corresponding to the phoneme.

[0020] In the above solution, the fusion module is further configured to: perform a query vector transformation process on the audio features to obtain a query vector; perform a key vector transformation process on the phoneme features to obtain a key vector; multiply the query vector and the transpose of the key vector to obtain a multiplication result; determine the ratio of the multiplication result to the square root of the dimension of the key vector as the attention feature; perform a maximum likelihood process on the attention feature to obtain the attention score corresponding to the phoneme.

[0021] In the above solution, determining the phoneme corresponding to each audio frame is implemented by calling a phoneme classification network, and the phoneme classification network includes at least one cascaded phoneme fully connected layer. The alignment module is further configured to perform the following processing for each audio frame: when the number of phoneme fully connected layers is one, perform a first fully connected process on the fusion features through the phoneme fully connected layer to obtain the first probability that the audio frame belongs to each candidate phoneme; when the number of phoneme fully connected layers is multiple, perform a first fully connected process on the input of the n-th phoneme fully connected layer through the n-th phoneme fully connected layer in N cascaded phoneme fully connected layers, and transmit the n-th phoneme fully connected result output by the n-th phoneme fully connected layer to the (n + 1)-th phoneme fully connected layer to continue the first fully connected process to obtain the (n + 1)-th phoneme fully connected result corresponding to the (n + 1)-th phoneme fully connected layer; where N is an integer greater than or equal to 2, n is an integer variable starting from 1 and increasing, and the value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the n-th phoneme fully connected layer is the fusion feature. When n takes the value of 2 ≤ n < N, the input of the n-th phoneme fully connected layer is the (n - 1)-th phoneme fully connected result output by the (n - 1)-th phoneme fully connected layer. When n takes the value of N - 1, the (n + 1)-th phoneme fully connected result is the first probability that the audio frame belongs to each candidate phoneme; determine the candidate phoneme with the largest first probability as the phoneme corresponding to the audio frame.

[0022] In the above solution, the alignment module is further configured to: determine at least one audio frame corresponding to each phoneme based on the phonemes corresponding to each audio frame; perform the following processing for each phoneme: when the phoneme corresponds to multiple consecutive audio frames, determine the start and end times of the consecutive audio frames corresponding to the phoneme as the start and end times of the phoneme; when the phoneme corresponds to one audio frame, determine the time of the audio frame corresponding to the phoneme as the start and end times of the phoneme.

[0023] In the above solution, the fusion processing based on the attention mechanism is implemented by invoking an attention fusion network. Determining the phoneme corresponding to each audio frame is implemented by invoking a phoneme classification network. The phoneme classification network shares the attention fusion network with the loudness classification network. The input of the attention fusion network is the outputs of an audio encoder and a phoneme encoder. The apparatus further includes: a training module, configured to: obtain an audio data sample and a given text sample; obtain at least one phoneme sample of the given text sample, and determine the phoneme features of each phoneme sample through the phoneme encoder; obtain at least one audio frame sample of the audio data sample corresponding to the given text sample, and determine the audio features of each audio frame sample through the audio encoder; perform the following processing for each audio frame sample: perform forward propagation of the audio features of the audio frame sample and the phoneme features of at least one phoneme sample in the attention fusion network and the phoneme classification network to obtain a first forward propagation result; perform the following processing for each audio frame sample: perform forward propagation of the audio features of the audio frame sample and the phoneme features of at least one phoneme sample in the attention fusion network and the loudness classification network to obtain a second forward propagation result; determine a joint loss according to the first forward propagation result and the second forward propagation result; update the parameters of the attention fusion network, the loudness classification network, the audio encoder, and the phoneme encoder according to the joint loss.

[0024] In the above solution, for the audio features of the audio frame sample and the phoneme features of at least one phoneme sample, the training module is further configured to: perform attention mechanism-based fusion processing on the audio features of the audio frame sample and the phoneme features of at least one phoneme sample through the attention fusion network to obtain fusion features corresponding to each audio frame sample; perform second fully connected processing on the fusion features of each audio frame sample through the loudness classification network to obtain the second probabilities of each audio frame sample belonging to each loudness category, which constitute the second forward propagation result.

[0025] In the above solution, the training module is further configured to perform the following processing on each phoneme sample through the attention layer of the attention fusion network: perform the following processing on each phoneme sample through the attention layer of the attention fusion network: based on the audio features of the audio frame sample and the phoneme features of the phoneme sample, determine the attention score of the corresponding phoneme sample; perform a value vector transformation process on the phoneme features of the phoneme sample, multiply the attention score of the corresponding phoneme sample by the value vector transformation result to obtain the attention result of the corresponding phoneme sample; through the fusion layer of the attention fusion network, fuse the attention results of each corresponding phoneme sample and the audio features of the audio frame sample to obtain the fusion features of the corresponding audio frame sample; perform a first fully connected process on the fusion features of the audio frame sample through the phoneme classification network to obtain the third probability that the audio frame sample belongs to each candidate phoneme; and form the first forward propagation result with the third probability and the attention score.

[0026] In the above solution, the training module is further configured to: determine the first phoneme category loss based on the third probabilities of each audio frame sample corresponding to multiple candidate phonemes and the pre-labeled candidate phonemes of each audio frame sample; determine the second loudness category loss based on the second probabilities of each audio frame sample corresponding to multiple loudness categories and the pre-labeled loudness categories of each audio frame sample; determine the third alignment loss based on the attention scores of each audio frame sample corresponding to each phoneme sample and the pre-labeled alignment identifiers of each audio frame sample corresponding to each phoneme sample; and perform a fusion process on the first phoneme category loss, the second loudness category loss, and the third alignment loss to obtain the joint loss.

[0027] An embodiment of the present application provides an electronic device, including:

[0028] A memory for storing executable instructions;

[0029] A processor, when executing the executable instructions stored in the memory, implements the audio processing method based on artificial intelligence provided by the embodiments of the present application.

[0030] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, which when executed by a processor, implement the audio processing method based on artificial intelligence provided by the embodiments of the present application.

[0031] An embodiment of the present application provides a computer program product including a computer program or instruction, which when executed by a processor, implements the audio processing method based on artificial intelligence provided by the embodiments of the present application.

[0032] The embodiments of the present application have the following beneficial effects:

[0033] Through the embodiments of the present application, the audio features and the text sequence are calculated by the attention mechanism to obtain the fusion features. Therefore, the fusion features can effectively represent the relationship between the audio frames and the phonemes. Then, based on the fusion features, phoneme classification is performed on each audio frame in the audio, which can effectively improve the classification accuracy, thereby improving the phoneme alignment accuracy. Description of the Drawings

[0034] Figure 1 It is a schematic structural diagram of an audio processing system based on artificial intelligence provided by an embodiment of the present application;

[0035] Figure 2 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0036] Figures 3A - 3C It is a schematic flowchart of an audio processing method based on artificial intelligence provided by an embodiment of the present application;

[0037] Figures 4A - 4D It is a schematic interface diagram of an audio processing method based on artificial intelligence provided by an embodiment of the present application;

[0038] Figure 5 It is a schematic flowchart of an audio processing method based on artificial intelligence provided by an embodiment of the present application;

[0039] Figure 6 It is a schematic structural diagram of a phoneme alignment model of an audio processing method based on artificial intelligence provided by an embodiment of the present application;

[0040] Figure 7 It is a schematic data flow diagram of an audio processing method based on artificial intelligence provided by an embodiment of the present application;

[0041] Figures 8A - 8C It is an alignment time matrix of an audio processing method based on artificial intelligence provided by an embodiment of the present application. Detailed Embodiments

[0042] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0043] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0044] In the following description, the terms "first", "second", and "third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0046] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0047] 1) Speech recognition technology: Automatic Speech Recognition (ASR), whose goal is to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0048] 2) Hidden Markov Model (HMM): A statistical model used to describe a Markov process with hidden unknown parameters.

[0049] 3) Maximum Likelihood Estimation (MLE), also known as maximum likelihood estimation, is a method used to estimate the parameters of a probability model.

[0050] 4) Discriminative model: In the field of machine learning, a discriminative model is a method for modeling the relationship between unknown data y and known data x. A discriminative model is a method based on probability theory. Given the input variable x, the discriminative model predicts y by constructing the conditional probability distribution P(y|x).

[0051] 5) Full Connection (FC): Each neuron in the fully connected layer is fully connected to all neurons in the previous layer. The fully connected layer can integrate the local information with class discrimination in the convolutional layer or pooling layer.

[0052] 6) Pearson correlation coefficient: In statistics, the Pearson correlation coefficient is used to measure the linear correlation between two variables X and Y, and its value ranges from -1 to 1.

[0053] 7) Support Vector Machine (SVM): Often abbreviated as support vector network in machine learning, it is a supervised learning model for analyzing data in classification and regression analysis.

[0054] 8) Phoneme: A phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing based on the pronunciation actions in a syllable, one action constitutes one phoneme. Phonemes are divided into two major categories: vowels and consonants. In the embodiments of this application, phonemes also include silent phonemes. For example, if a certain audio frame is silent, that is, the audio frame corresponds to a silent phoneme.

[0055] In related technologies, there are two phoneme alignment methods. One is independent of the given text, and the other is text-dependent. The method independent of the text usually classifies phoneme boundaries to determine whether the time of a certain frame in the audio is a phoneme boundary. For example, the Viterbi algorithm is used to distinguish the pronunciation segment and the non-pronunciation segment, or a recurrent neural network is used to classify phoneme boundaries. The method dependent on the text usually uses HMM based on maximum likelihood to obtain the most likely sequence, or uses a discriminant model, or designs an alignment function and uses a support vector machine for phoneme alignment.

[0056] In related technologies, the alignment method based on HMM mainly uses phoneme boundary judgment as the hidden state and optimizes it using maximum likelihood, without directly and explicitly optimizing phoneme alignment. Other phoneme alignment methods in related technologies require manually designing alignment functions and performing manual feature engineering. The embodiments of this application propose an audio processing method based on artificial intelligence, which can automatically learn the mapping relationship between the phoneme sequence and audio data based on a neural network including an attention mechanism without relying on manually designed alignment functions, and explicitly optimize the loss function during the training phase, jointly train with multiple tasks, and perform constrained learning through the loss function during the attention processing phase, effectively improving the accuracy of phoneme alignment.

[0057] In view of the above problems in related technologies, the embodiments of this application provide an audio processing method, device, electronic device, computer program product, and computer-readable storage medium based on artificial intelligence, which can perform attention mechanism calculation on audio features and text sequences to obtain a fusion feature, thereby classifying phonemes for each frame in the audio based on the fusion feature, effectively improving the classification accuracy, and thus improving the phoneme alignment accuracy.

[0058] The following describes the exemplary applications of the electronic device provided in the embodiments of this application. The electronic device provided in the embodiments of this application can be implemented as a server. Next, the exemplary applications will be described when the electronic device is implemented as a server.

[0059] See Figure 1 , Figure 1It is a schematic structural diagram of an audio processing system based on artificial intelligence provided by an embodiment of the present application. The audio processing system can be used in an oral examination scenario. In the audio processing system, the terminal 400 is connected to the server 200 through a network. The network can be a wide area network or a local area network, or a combination of the two.

[0060] In some embodiments, the functions of the audio processing system are implemented based on various modules in the server 200. During the user's use of the terminal 400, the terminal 400 receives audio data of the user for a given text, and the terminal 400 sends the audio data and the given text to the server 200. The server 200 determines the phoneme features of each phoneme in the given text and the audio features of each audio frame in the audio data, and performs the following processing for each audio frame: performing an attention mechanism-based fusion process on the audio features of the audio frame and the phoneme features of multiple phonemes to obtain a fusion feature corresponding to each audio frame, determining the phoneme corresponding to each audio frame based on the fusion feature of each audio frame, and determining the start and end times of each phoneme based on the phoneme corresponding to each audio frame, and sending the start and end times of each phoneme to the terminal 400, so that the terminal 400 directly presents the start and end times of each phoneme, thereby completing the phoneme alignment process.

[0061] In some embodiments, when the audio processing system is applied to an oral examination scenario, for example, the oral examination question requires the user to read the given text aloud in English. The terminal 400 receives the audio data of the user corresponding to the given text, and the terminal 400 sends the audio data to the server 200. The server 200 performs an attention mechanism-based fusion process on the audio features of the audio frame and the phoneme features of multiple phonemes to obtain a fusion feature corresponding to each audio frame, determines the phoneme corresponding to each audio frame based on the fusion feature of each audio frame, and determines the start and end times of each phoneme based on the phoneme corresponding to each audio frame, and sends them to the terminal 400, so that the terminal 400 directly presents the start and end times of each phoneme. In response to the user's scoring operation, the terminal 400 can display the scoring results for each phoneme. The user participating in the reading aloud and the user performing the scoring can be the same or different users.

[0062] In some embodiments, when the audio processing system is applied to a spoken language practice scenario, for example, the spoken language practice question requires the user to follow and read aloud a given text in English, the terminal 400 receives the audio data corresponding to the given text from the user, and the terminal 400 sends the audio data to the server 200. The server 200 performs fusion processing based on the attention mechanism on the audio features of the audio frames and the phoneme features of multiple phonemes to obtain the fusion features corresponding to each audio frame. Based on the fusion features of each audio frame, the phoneme corresponding to each audio frame is determined, and based on the phoneme corresponding to each audio frame, the start and end times of each phoneme are determined and sent to the terminal 400, so that the terminal 400 can directly present the start and end times of each phoneme, thereby being used for the user's playback operation for each phoneme. The terminal 400 can play the audio frames corresponding to each phoneme separately.

[0063] In other embodiments, the terminal can also perform fusion processing based on the attention mechanism on the audio features of the audio frames and the phoneme features of multiple phonemes to obtain the fusion features corresponding to each audio frame. Based on the fusion features of each audio frame, the phoneme corresponding to each audio frame is determined, and based on the phoneme corresponding to each audio frame, the start and end times of each phoneme are determined and the start and end times of each phoneme are directly presented.

[0064] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0065] In some embodiments, the terminal or the server can implement the audio processing method provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a spoken language examination APP or a spoken language learning APP; it can also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded into any APP. In short, the above computer program can be any form of application program, module or plug-in.

[0066] Next, the structure of the electronic device provided in the embodiments of the present application for implementing the audio processing method based on artificial intelligence will be described. As before, the electronic device provided in the embodiments of the present application can be Figure 1 the server 200 in Figure 2 . Refer to Figure 2 which is a schematic structural diagram of the server 200 provided in the embodiments of the present application. The server 200 shown in Figure 2 includes: at least one processor 210, a memory 250, and at least one network interface 220. Each component in the server 200 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 240.

[0067] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0068] The memory 250 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. Optionally, the memory 250 includes one or more storage devices that are physically remote from the processor 210.

[0069] The memory 250 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0070] In some embodiments, the memory 250 is capable of storing data to support various operations. Examples of these data include programs, modules, and data structures, or subsets or supersets thereof, which will be described below by way of example.

[0071] The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks; the network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.

[0072] In some embodiments, the audio processing device based on artificial intelligence provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is the audio processing device 255 based on artificial intelligence stored in the memory 250, which can be software in the form of programs and plugins, etc., including the following software modules: the phoneme module 2551, the audio module 2552, the fusion module 2553, the alignment module 2554, and the training module 2555. These modules are logical, so they can be combined arbitrarily or further split according to the implemented functions. The functions of each module will be described below.

[0073] The audio processing method based on artificial intelligence provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server 200 provided by the embodiments of the present application.

[0074] See Figure 6 , Figure 6 is a schematic structural diagram of the phoneme alignment model of the audio processing method based on artificial intelligence provided by the embodiments of the present application. The phoneme alignment model includes an attention fusion network, a phoneme classification network (corresponding to the second task), and a loudness classification network (corresponding to the first task). The attention fusion network is used to perform attention mechanism-based fusion processing on phoneme features and audio features, so that the fusion features output by the attention fusion network are shared by the loudness classification network corresponding to the first task and the phoneme classification network corresponding to the second task. The input of the attention fusion network is the audio features obtained based on audio data and the phoneme features obtained based on the given text. The output of the attention fusion network is the fusion features of audio features and phoneme features. Then, full connection processing is performed through the fully connected layers corresponding to the loudness classification network and the phoneme classification network respectively to obtain the loudness classification result and the phoneme classification result. Among them, the first task is to identify the phoneme of a certain audio frame from multiple candidate phonemes, and the second task is to judge whether a certain audio frame is a silent audio frame.

[0075] See Figure 7 , Figure 7It is a schematic diagram of the data flow of the audio processing method based on artificial intelligence provided by an embodiment of the present application. The phoneme alignment model includes an attention fusion network, a phoneme classification network (corresponding to the second task), and a loudness classification network (corresponding to the first task). The input of the audio encoder is audio data, and the output of the audio encoder is the audio feature (in vector form) of each audio frame. The input of the phoneme encoder is a phoneme sequence (given text), and the output of the phoneme encoder is the phoneme feature (in vector form) of each phoneme. The input of the attention fusion network is the output of the audio encoder and the output of the phoneme encoder. The output of the attention fusion network is the fusion feature of the phoneme feature and the audio feature. The fusion feature is classified through two parallel phoneme classification networks and a loudness classification network respectively. The phoneme classification network outputs the probability that each audio frame belongs to each candidate phoneme, and the loudness classification network outputs the probability that each audio frame belongs to each given loudness category. The loudness categories include silence and non-silence. For example, the identifier of non-silence is 1, and the identifier of silence is 0. The candidate phonemes are W, IH, L, etc.

[0076] Take the server 200 in Figure 1 executing the audio processing method based on artificial intelligence provided by an embodiment of the present application as an example to illustrate the audio processing method based on artificial intelligence provided by an embodiment of the present application.

[0077] See Figure 3A , Figure 3A It is a schematic diagram of the process of the audio processing method based on artificial intelligence provided by an embodiment of the present application, and will be described in combination with Figure 3A the steps 101-104 shown.

[0078] In step 101, at least one phoneme of the given text is obtained, and the phoneme feature of each phoneme is determined.

[0079] In some embodiments, the phoneme feature of each phoneme determined in step 101 can be implemented through the following technical solutions: the following processing is performed for each phoneme: the characteristic representation feature of the phoneme is determined through the phoneme characteristic representation network, where the characteristic representation feature is used to characterize the characteristics of the phoneme; the position representation feature of the phoneme is determined through the phoneme position representation network, where the position representation feature is used to characterize the position of the phoneme in the corresponding text unit; the position representation feature and the characteristic representation feature are added to obtain the phoneme feature of the phoneme.

[0080] As an example, determining the phoneme features of each phoneme is achieved by invoking a phoneme encoder, which includes a phoneme feature representation network and a phoneme position representation network. Different languages contain different phonemes. Taking English as an example, when the given text is "ever forget", the phonemes of the given text include EH1, V, ER, sp, F, R, G, EH, T, where EH1, V, ER, F, R, G, EH, T are different phonemes, and sp represents a silent phoneme. Silence is also one of the candidate phonemes. Each phoneme is encoded by the phoneme feature representation network to obtain the feature representation of each phoneme. The feature representations of different phonemes are different, and the features include pronunciation features, semantic features, etc. The feature representation is used to distinguish different phonemes. Each phoneme has four position possibilities in the corresponding text unit. In English, the text unit is a word. When a word contains multiple phonemes, there are the start position (B), middle position (I), and end position (E) of the word. When a word contains one phoneme, the position of the phoneme is represented by S. The position of the phoneme in the corresponding text unit is encoded by the phoneme position representation network to obtain the position representation of each phoneme. Finally, the unique encoding representation (pronunciation vector) of each phoneme is added to the position encoding representation (position vector) to obtain the final phoneme feature. Through this phoneme encoding method, the differences between each phoneme can be effectively represented, and the differences between the same phoneme in different positions can also be effectively represented.

[0081] In step 102, at least one audio frame of the audio data corresponding to the given text is obtained, and the audio feature of each audio frame is determined.

[0082] In some embodiments, the audio feature of each audio frame determined in step 102 can be implemented by the following technical solution: The feature extraction process is performed on at least one audio frame through multiple cascaded convolutional networks to obtain the convolutional feature extraction result corresponding to each audio frame; the convolutional feature extraction result of each audio frame is normalized through a normalization network to obtain the audio feature of each audio frame.

[0083] As an example, determining the audio features of each audio frame is achieved by calling an audio encoder. The audio encoder includes multiple convolutional networks and normalization networks. Based on the audio encoder, audio features are obtained. At least one audio frame is used as a whole for feature extraction processing through multiple cascaded convolutional networks. When there are multiple audio frames, the outputs of the multiple convolutional networks are low-frequency feature representations. For example, it encodes audio data of 16 kHz for approximately 30 milliseconds, and a low-frequency feature representation is generated every set time step, thereby obtaining the convolutional feature extraction result of each audio frame. Then, the convolutional feature extraction result of each audio frame is normalized through the normalization network to obtain the audio feature of each audio frame. The structure of the audio encoder can be the network structure of wav2vec, and the parameters of the audio encoder are trained based on the network structure of wav2vec.

[0084] In step 103, the following processing is performed for each audio frame: perform attention mechanism-based fusion processing on the audio features of the audio frame and the phoneme features of at least one phoneme to obtain the fusion feature corresponding to each audio frame.

[0085] In some embodiments, refer to Figure 3B , Figure 3B is a schematic flowchart of the audio processing method based on artificial intelligence provided by the embodiments of the present application. In step 103, the attention mechanism-based fusion processing on the audio features of the audio frame and the phoneme features of at least one phoneme to obtain the fusion feature corresponding to each audio frame can be described through Figure 3B the steps 1031-1032 shown.

[0086] In step 1031, perform attention processing on the audio features of the audio frame and the phoneme features of at least one phoneme through the attention layer to obtain an attention result.

[0087] In some embodiments, the attention processing on the audio features of the audio frame and the phoneme features of at least one phoneme in step 1031 to obtain an attention result can be achieved through the following technical solution: perform the following processing for each phoneme: based on the audio features of the audio frame and the phoneme features of the phoneme, determine the attention score corresponding to the phoneme; perform value vector transformation processing on the phoneme features of the phoneme to obtain a value vector; multiply the attention score corresponding to the phoneme by the value vector to obtain the attention result corresponding to the phoneme.

[0088] In some embodiments, the attention score corresponding to a phoneme can be determined based on the audio features of audio frames and the phoneme features of phonemes through the following technical solutions: performing query vector transformation processing on the audio features to obtain a query vector; performing key vector transformation processing on the phoneme features to obtain a key vector; multiplying the query vector and the transpose of the key vector to obtain a multiplication result; determining the ratio of the multiplication result to the square root of the dimension of the key vector as the attention feature; and performing maximum likelihood processing on the attention feature to obtain the attention score corresponding to the phoneme.

[0089] As an example, as an example, the attention mechanism is used to fuse the phoneme features and the audio features. The attention mechanism is used to model the relationship between the query vector Q, the key vector K, and the value vector V. Refer to formulas (1) and (2):

[0090]

[0091] Attention(Q, K, V) = AttentionScore(Q, K) * V (2);

[0092] Among them, based on the audio features of each audio frame the query vector Q is obtained, and based on the phoneme features H of all phonemes in the given text phone the key vector K and the value vector V are obtained. It is also possible to directly use the audio features of each audio frame as the query vector, or directly use the phoneme features H of all phonemes in the given text phone as the key vector K and the value vector V. AttentionScore(Q, K) is the attention score, and Attention(Q, K, V) is the attention result of each audio frame corresponding to all phonemes. d k is the dimension of the key vector K.

[0093] As an example, performing query vector transformation processing on the audio features of each audio frame to obtain the query vector Q, performing key vector transformation processing on the phoneme features H of all phonemes in the given text phone to obtain the key vector K and performing value vector transformation processing to obtain the value vector V. The parameters involved in these transformation processes can be obtained through overall training of the phoneme alignment model. It is also possible to directly use the audio features of each audio frame as the query vector, or directly use the phoneme features H of all phonemes in the given text phone as the key vector K and the value vector V.

[0094] In step 1032, the attention result and the audio features are fused through a fusion layer to obtain the fusion feature corresponding to the audio frame.

[0095] As an example, the fusion processing based on the attention mechanism is implemented by calling an attention fusion network. The attention fusion network includes an attention layer and a fusion layer. The fusion processing is actually a feature splicing process, which splices the attention result of a certain audio frame with the audio feature of the audio frame to obtain the fusion feature corresponding to the audio frame. See formula (3):

[0096]

[0097] Among them, is the attention result of audio frame i obtained based on the attention mechanism. The attention result of audio frame i is a matrix, and each column in the matrix represents the attention result of each phoneme among all phonemes with audio frame i. is the audio feature of audio frame i, and H phone is the phoneme feature of all phonemes of the given text. is the fusion feature corresponding to each audio frame.

[0098] In step 104, based on the fusion feature of each audio frame, the phoneme corresponding to each audio frame is determined, and based on the phoneme corresponding to each audio frame, the start and end times of each phoneme are determined.

[0099] In some embodiments, determining the phoneme corresponding to each audio frame is implemented by calling a phoneme classification network. The phoneme classification network includes at least one cascaded phoneme fully connected layer. Based on the fusion feature of each audio frame in step 104, determining the phoneme corresponding to each audio frame can be achieved through the following technical solution: the following processing is performed for each audio frame: when the number of phoneme fully connected layers is one, the fusion feature is subjected to a first fully connected process through the phoneme fully connected layer to obtain the first probability that the audio frame belongs to each candidate phoneme; when the number of phoneme fully connected layers is multiple, through the nth phoneme fully connected layer in the N cascaded phoneme fully connected layers, the input of the nth phoneme fully connected layer is subjected to a first fully connected process, and the nth phoneme fully connected result output by the nth phoneme fully connected layer is transmitted to the (n + 1)th phoneme fully connected layer to continue the first fully connected process to obtain the (n + 1)th phoneme fully connected result corresponding to the (n + 1)th phoneme fully connected layer; where N is an integer greater than or equal to 2, n is an integer variable starting from 1 and increasing, and the value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the nth phoneme fully connected layer is the fusion feature. When n takes the value of 2 ≤ n < N, the input of the nth phoneme fully connected layer is the (n - 1)th phoneme fully connected result output by the (n - 1)th phoneme fully connected layer. When n takes the value of N - 1, the (n + 1)th phoneme fully connected result is the first probability that the audio frame belongs to each candidate phoneme; the candidate phoneme with the largest first probability is determined as the phoneme corresponding to the audio frame.

[0100] As an example, refer to Figure 6 , after the attention fusion network, an external phoneme classification network (phoneme fully connected layer) is connected. The phoneme classification network performs phoneme classification for each audio frame. The candidate phonemes total 40 phonemes (including 39 phonemes in the phoneme dictionary and the silence phoneme). When there is only one phoneme fully connected layer, the first probability that a certain audio frame belongs to each candidate phoneme is output through the phoneme fully connected layer, that is, 40 first probabilities are output for audio frame A. The candidate phoneme corresponding to the highest first probability is determined as the phoneme of audio frame A. When there are multiple phoneme fully connected layers, due to the cascaded relationship, more in-depth features can be learned through multiple cascaded fully connected layers, thereby effectively improving the subsequent phoneme recognition accuracy.

[0101] In some embodiments, in step 104, based on the phoneme corresponding to each audio frame, the start and end times of each phoneme are determined, which can be achieved through the following technical solutions: based on the phoneme corresponding to each audio frame, at least one audio frame corresponding to each phoneme is determined; the following processing is performed for each phoneme: when the phoneme corresponds to multiple consecutive audio frames, the start and end times of the consecutive audio frames corresponding to the phoneme are determined as the start and end times of the phoneme; when the phoneme corresponds to one audio frame, the time of the audio frame corresponding to the phoneme is determined as the start and end times of the phoneme.

[0102] As an example, the start and end times include the start time and the end time of the phoneme. Taking 10 audio frames as an example for illustration, based on the phoneme corresponding to each audio frame, at least one audio frame corresponding to each phoneme is determined. The following processing is performed for each phoneme: when the phoneme corresponds to multiple consecutive audio frames, the start and end times of the consecutive audio frames corresponding to the phoneme are determined as the start and end times of the phoneme. For example, the 1st to 3rd audio frames all correspond to the phoneme W, then the phoneme W corresponds to the 1st to 3rd audio frames, and the start and end times of the 1st to 3rd audio frames are determined as the start and end times of the phoneme W, that is, the time of the 1st audio frame is determined as the start time in the start and end times, and the time of the 3rd audio frame is determined as the end time in the start and end times. When the phoneme corresponds to one audio frame, the time of the audio frame corresponding to the phoneme is determined as the start and end times of the phoneme. For example, the 1st audio frame corresponds to the phoneme W, and the 2nd audio frame corresponds to the silence audio frame, then the phoneme W corresponds to the 1st audio frame, and the start and end times of the 1st audio frame are determined as the start and end times of the phoneme W, that is, the time of the 1st audio frame is determined as the start time in the start and end times, and at the same time, the time of the 1st audio frame is also determined as the end time in the start and end times.

[0103] In some embodiments, refer to Figure 3C , Figure 3CIt is a schematic flowchart of an audio processing method based on artificial intelligence provided by an embodiment of the present application. Before obtaining at least one phoneme of a given text and determining the phoneme features of each phoneme in step 101, or before obtaining at least one audio frame of audio data corresponding to the given text and determining the audio features of each audio frame in step 102, the steps 105-111 shown in Figure 3C can be executed.

[0104] In step 105, an audio data sample and a given text sample are obtained.

[0105] As an example, the given text sample corresponds to the audio data sample. For example, the audio data sample is obtained by a user reading the given text aloud.

[0106] In step 106, at least one phoneme sample of the given text sample is obtained, and the phoneme features of each phoneme sample are determined through a phoneme encoder.

[0107] In step 107, at least one audio frame sample of the audio data sample corresponding to the given text sample is obtained, and the audio features of each audio frame sample are determined through an audio encoder.

[0108] As an example, the audio encoder and the phoneme encoder participating in the training can be pre-trained network structures. In an embodiment of the present application, a pre-trained acoustic model is used for audio feature extraction, such as a voice vector model. The voice vector model is composed of a multi-layer convolutional network. The voice vector model is pre-trained based on a contrast loss using a large number of unlabeled tasks. When training a phoneme alignment model, audio data (audio waveform features) is input into the pre-trained network structure.

[0109] As an example, the phoneme alignment model includes a phoneme classification network, a loudness classification network, a shared attention fusion network, an audio encoder, and a phoneme encoder. The fusion processing based on the attention mechanism is implemented by calling the attention fusion network. Determining the phoneme corresponding to each audio frame is implemented by calling the phoneme classification network. The phoneme classification network and the loudness classification network share the attention fusion network. The input of the attention fusion network is the output of the audio encoder and the output of the phoneme encoder.

[0110] In step 108, the following processing is performed for each audio frame sample: The audio features of the audio frame sample and the phoneme features of at least one phoneme sample are propagated forward in the attention fusion network and the phoneme classification network to obtain a first forward propagation result.

[0111] In some embodiments, the forward propagation of the audio features of the audio frame samples and the phoneme features of at least one phoneme sample in the attention fusion network and the phoneme classification network to obtain the first forward propagation result can be achieved through the following technical solutions: The attention layer of the attention fusion network performs the following processing for each phoneme sample: Based on the audio features of the audio frame samples and the phoneme features of the phoneme sample, determine the attention score of the corresponding phoneme sample; perform a value vector transformation process on the phoneme features of the phoneme sample, and multiply the attention score of the corresponding phoneme sample by the value vector transformation result to obtain the attention result of the corresponding phoneme sample; through the fusion layer of the attention fusion network, fuse the attention results of each phoneme sample and the audio features of the audio frame samples to obtain the fusion features of the corresponding audio frame samples; perform a first fully connected process on the fusion features of the audio frame samples through the phoneme classification network to obtain the third probability that the audio frame samples belong to each candidate phoneme; and form the first forward propagation result with the third probability and the attention score.

[0112] As an example, in order to better fuse the phoneme features and audio feature representations, it is necessary to constrain the attention score matrix in the embodiments of the present application, that is, perform attention weight constraint, where each row in the attention score matrix represents an audio frame, and each column represents the probability distribution of each phoneme corresponding to the audio frame.

[0113] In step 109, the following processing is performed for each audio frame sample: The audio features of the audio frame sample and the phoneme features of at least one phoneme sample are forward propagated in the attention fusion network and the loudness classification network to obtain the second forward propagation result.

[0114] In some embodiments, the forward propagation of the audio features of the audio frame samples and the phoneme features of at least one phoneme sample in the attention fusion network and the loudness classification network to obtain the second forward propagation result can be achieved through the following technical solutions: The attention fusion network performs an attention mechanism-based fusion process on the audio features of the audio frame samples and the phoneme features of at least one phoneme sample to obtain the fusion features of each audio frame sample; the loudness classification network performs a second fully connected process on the fusion features of each audio frame sample to obtain the second probability that each audio frame sample belongs to each loudness category, and forms the second forward propagation result.

[0115] As an example, during the forward propagation of data, the input of the loudness classification network is the same as the input of the phoneme classification network.

[0116] In some embodiments, the above-mentioned forward propagation of the audio features of the audio frame samples and the phoneme features of at least one phoneme sample in the attention fusion network and the loudness classification network to obtain the second forward propagation result can be achieved through the following technical solutions: The attention layer of the attention fusion network performs the following processing for each phoneme sample: Based on the audio features of the audio frame samples and the phoneme features of the phoneme sample, determine the attention score of the corresponding phoneme sample; perform a value vector transformation process on the phoneme features of the phoneme sample, and multiply the attention score of the corresponding phoneme sample by the value vector transformation result to obtain the attention result of the corresponding phoneme sample; through the fusion layer of the attention fusion network, fuse the attention results corresponding to each phoneme sample and the audio features of the audio frame samples to obtain the fusion features corresponding to the audio frame samples; perform a second fully connected process on the fusion features of the audio frame samples through the loudness classification network to obtain the second probability that the audio frame sample belongs to each candidate phoneme; and form the second forward propagation result with the second probability and the attention score.

[0117] As an example, the phoneme alignment model includes an attention fusion network, a phoneme classification network, and a loudness classification network. The input of the audio encoder is an audio data sample, and the output of the audio encoder is the audio features (in vector form) of each audio frame sample. The input of the phoneme encoder is a phoneme sequence sample (a given text sample), and the output of the phoneme encoder is the phoneme features (in vector form) of each phoneme sample. The input of the attention fusion network is the output of the audio encoder and the output of the phoneme encoder. The output of the attention fusion network is the fusion features of the phoneme features and the audio features. The audio features of each audio frame are calculated with all phonemes through an attention mechanism to obtain the fusion features, determine the representation of the audio frame corresponding to the candidate phoneme and the representation of whether it is corresponding to silence or not. Through two parallel phoneme classification networks and a loudness classification network, classify the fusion features respectively. The phoneme classification network outputs the third probability that each audio frame belongs to each candidate phoneme, and the loudness classification network outputs the second probability that each audio frame belongs to each loudness category. The loudness categories include silence and non-silence. For example, the identifier of non-silence is 1, and the identifier of silence is 0. The loudness categories can also be more fine-grained divisions, such as silence, 10 decibels, 20 decibels, 30 decibels, etc. The candidate phonemes are W, IH, L, etc.

[0118] In step 110, determine the joint loss according to the first forward propagation result and the second forward propagation result.

[0119] In some embodiments, the joint loss can be determined based on the first forward propagation result and the second forward propagation result through the following technical solution: determining a first phoneme category loss based on the third probability corresponding to multiple candidate phonemes for each audio frame sample and the pre-labeled candidate phoneme for each audio frame sample; determining a second loudness category loss based on the second probability corresponding to multiple loudness categories for each audio frame sample and the pre-labeled loudness category for each audio frame sample; determining a third alignment loss based on the attention score corresponding to each phoneme sample for each audio frame sample and the pre-labeled alignment identifier corresponding to each phoneme sample for each audio frame sample; and performing a fusion process on the first phoneme category loss, the second loudness category loss, and the third alignment loss to obtain the joint loss.

[0120] As an example, during the training process of the phoneme alignment model, the cross loss is used to calculate the losses of the two classifications. See Formulas (4) and (5):

[0121]

[0122]

[0123] where L phone is the phoneme classification loss (the first phoneme category loss), L sil is the loudness classification loss (the second loudness category loss), m is the number of audio frames, c is the number of candidate phonemes, is the true identification result of the j-th phoneme corresponding to the i-th audio frame, is the first probability of the j-th phoneme corresponding to the i-th audio frame, is the pre-labeled alignment identifier of the i-th audio frame, 1 for non-silent and 0 for silent, is the probability that the i-th audio frame is a non-silent audio frame.

[0124] In some embodiments, in order to better fuse the phoneme features and audio feature representations, the attention score matrix in the embodiments of the present application is constrained, that is, attention weight constraint is performed. Among them, each row in the matrix represents an audio frame, and each column represents the probability distribution of each phoneme in the audio frame. The probability distribution of the phonemes in each audio frame is calculated with the actual phoneme corresponding to the audio frame to obtain the attention mechanism loss. See Formula (6):

[0125]

[0126] where L align is the attention mechanism loss, m is the number of audio frames, N p is the number of phonemes in the given text, is 1 or 0, where 1 indicates that the i-th audio frame is aligned with the j-th phoneme, and 0 indicates that the i-th audio frame is not aligned with the j-th phoneme. is the attention score between the i-th audio frame and the j-th phoneme.

[0127] In some embodiments, the joint loss of the entire phoneme alignment network consists of three parts, including phoneme classification loss (the first phoneme class loss), loudness classification loss (the second loudness class loss), and alignment loss (the third alignment loss). The three losses are weighted and summed using different weights, and the final joint loss is shown in Equation (7):

[0128] L total = λL phone + βL sil + γL align (7);

[0129] where the weights (λ, β, and γ) of each loss are preset weights, and the sum of the three is equal to 1. L phone is the phoneme classification loss (the first phoneme class loss), L sil is the loudness classification loss (the second loudness class loss), and L align is the alignment loss (the third alignment loss), and L total is the joint loss.

[0130] In step 111, the parameters of the attention fusion network, loudness classification network, phoneme encoder, and audio encoder are updated according to the joint loss.

[0131] As an example, when updating the parameters of the attention fusion network, loudness classification network, and loudness classification network according to the joint loss, the gradient is determined according to the joint loss, and then the parameters of each network are updated through a descent algorithm, so as to minimize the joint loss as much as possible.

[0132] Through the embodiments of the present application, the audio features and the text sequence are calculated by the attention mechanism to obtain the fusion features. Therefore, the fusion features can effectively represent the relationship between the audio frames and the phonemes. Then, based on the fusion features, phoneme classification is performed on each audio frame in the audio, which can effectively improve the classification accuracy and thus improve the phoneme alignment accuracy.

[0133] Next, an exemplary application of the embodiments of the present application in a practical application scenario will be described.

[0134] In some embodiments, when the audio processing system is applied to a spoken test scenario, for example, the spoken test question requires the candidate user to follow and read aloud a given text in English. The candidate terminal receives the audio data corresponding to the given text from the user, and the candidate terminal sends the audio data to the server. The server performs an attention mechanism-based fusion process on the audio features of the audio frames and the phoneme features of multiple phonemes to obtain the fusion features corresponding to each audio frame. Based on the fusion features of each audio frame, the phoneme corresponding to each audio frame is determined, and based on the phoneme corresponding to each audio frame, the start and end times of each phoneme are determined and sent to the judge terminal, so that the judge terminal directly presents the start and end times of each phoneme. In response to the scoring operation of the judge user, the judge terminal can display the scoring results for each phoneme. That is, the embodiments of the present application mainly provide an automated tool for phoneme annotation, which annotates the corresponding positions of each phoneme in the given text in the audio data, and on this basis, can further annotate whether the pronunciation of phonemes and words is incorrect, thereby effectively reducing the manual annotation cost and providing a more convenient scoring environment for subsequent judge scoring.

[0135] In some embodiments, when the audio processing system is applied to a spoken practice scenario, for example, the spoken practice question requires the student user to follow and read aloud a given text in English. The student terminal receives the audio data corresponding to the given text from the user, and the student terminal sends the audio data to the server. The server performs an attention mechanism-based fusion process on the audio features of the audio frames and the phoneme features of multiple phonemes to obtain the fusion features corresponding to each audio frame. Based on the fusion features of each audio frame, the phoneme corresponding to each audio frame is determined, and based on the phoneme corresponding to each audio frame, the start and end times of each phoneme are determined and sent to the candidate terminal, so that the candidate terminal directly presents the start and end times of each phoneme. In response to the scoring operation of the candidate user, the candidate terminal can display the scoring results for each phoneme, and the scoring results can be annotations on whether the pronunciation of the phoneme is correct. That is, the embodiments of the present application mainly provide an automated tool for phoneme annotation, which annotates the corresponding positions of each phoneme in the given text in the audio data, and on this basis, can further annotate whether the pronunciation of phonemes and words is incorrect, thereby effectively reducing the manual annotation cost and providing a more convenient self-checking environment for subsequent candidate scoring self-check.

[0136] Phoneme forced alignment refers to aligning a given phoneme sequence text with the corresponding audio to obtain the time positions of each phoneme in the text in the audio. Phoneme alignment has different applications in speech processing, such as speech recognition, speech keyword detection, etc. In the embodiments of the present application, the audio features and the text sequence are subjected to an attention mechanism calculation to obtain the fused audio and text features, and phoneme classification is performed on each frame in the audio. In order to make the alignment more accurate, auxiliary tasks are added, such as judging whether each frame in the audio is silent. At the same time, the obtained attention score matrix is constrained to achieve a more accurate alignment.

[0137] In some embodiments, referring to Figure 4A , Figure 4A is a schematic diagram of the interface of the audio processing method based on artificial intelligence provided by the embodiments of the present application. In the human-computer interaction interface 401A, a reading button 402A and an end reading button 403A are displayed. The given text "What are you doing?" is also displayed in the human-computer interaction interface 401A. In response to the triggering operation of the candidate user on the reading button 402A, the candidate terminal receives the audio data corresponding to the given text. In response to the triggering operation of the candidate user on the end reading button 403A, the candidate terminal stops receiving the audio data corresponding to the given text.

[0138] In some embodiments, referring to Figure 4B , Figure 4B is a schematic diagram of the interface of the audio processing method based on artificial intelligence provided by the embodiments of the present application. The phoneme annotation function can be embedded in a web page or in a client. The process for the user to perform phoneme-level annotation of pronunciation is as follows. The given text 403B and an annotation button 402B are displayed in the human-computer interaction interface 401B. In response to the triggering operation on the annotation button 402B, an annotation page for the given text 403B is displayed in the human-computer interaction interface 401B.

[0139] In some embodiments, referring to Figure 4C , Figure 4C is a schematic diagram of the interface of the audio processing method based on artificial intelligence provided by the embodiments of the present application. In the human-computer interaction interface 401C, an annotation page 403C is displayed. In the annotation page 403C, the start and end times of the phoneme 402C in the audio and the start and end times of the word 404C in the audio are displayed. The start and end times of the word 404C in the audio are determined by the start and end times of the phoneme 402C in the audio.

[0140] In some embodiments, referring to Figure 4D , Figure 4D is a schematic diagram of the interface of the audio processing method based on artificial intelligence provided by the embodiments of the present application. In the human-computer interaction interface 401D, an annotation page 403D is displayed. In the annotation page 403D, the start and end times of the phoneme 402D in the audio and the start and end times of the word 404D in the audio are displayed. The start and end times of the word 404D in the audio are determined by the start and end times of the phoneme 402D in the audio. Therefore, the divided phonemes are displayed in the human-computer interaction interface 401D. In response to the user's annotation operation on the phoneme, the pronunciation annotation 405D for the phoneme is displayed in the last layer of the annotation page. For example, whether a certain phoneme is incorrect.

[0141] In some embodiments, referring to Figure 5It is a schematic flowchart of an audio processing method based on artificial intelligence provided by an embodiment of the present application. The overall business flowchart based on phoneme forced alignment is as Figure 5 shown. The steps are as follows: After the webpage of the phoneme annotation tool is opened, the user can select the audio to be annotated and the corresponding shadowing text; in response to the user's selection operation, the audio to be annotated and the corresponding phoneme text sequence (derived from the shadowing text of the title) are determined and the annotation starts; the webpage sends the audio data and the phoneme text sequence (derived from the shadowing text of the title) to the server; the server sends the audio data and the phoneme text sequence (derived from the shadowing text of the title) to the phoneme forced alignment module; the phoneme forced alignment module returns the start and end times of each phoneme in the audio data (phoneme boundary information) to the server; the server returns the audio segmented based on the phoneme boundary information to the user; in response to the user's annotation operation, pronunciation annotation is performed at the phoneme level based on each segmented phoneme pronunciation segment.

[0142] In some embodiments, referring to Figure 6 , the phoneme alignment model provided by the embodiment of the present application includes a phoneme encoder, an audio encoder, an attention fusion network, a phoneme classification network, and a loudness classification network. The phoneme encoder is used to extract phoneme features, and the audio encoder is used to extract audio features. The two encoders are fused based on the attention mechanism to obtain a fused feature, which contains the information of the audio feature and the phoneme feature. After the attention fusion network, a phoneme classification network (fully connected layer) and a loudness classification network (fully connected layer) are externally connected. Phoneme classification is performed on each audio frame through the phoneme classification network. The phoneme classification includes a total of 40 phonemes (including 39 phonemes in the phoneme dictionary and a silent phoneme), and whether each audio frame is a silent audio frame is classified through the loudness classification network (including silent or non-silent).

[0143] In some embodiments, an audio feature representation is obtained based on an audio encoder. In the embodiments of the present application, a pre-trained acoustic model is used for audio feature extraction, such as a voice vector model. The voice vector model is composed of a multi-layer convolutional network. The voice vector model is pre-trained based on a contrastive loss using a large number of unlabeled tasks. When training a phoneme alignment model, audio data (audio waveform features) are input into the pre-trained network structure, and the audio features of each audio frame in the audio data are output. A phoneme feature is obtained based on a phoneme encoder. In the embodiments of the present application, a phoneme encoding method is used for phoneme feature extraction, and the characteristics of each phoneme are represented by a unique vector (characteristic representation feature). The characteristic vector (characteristic representation feature) of each phoneme is initialized in a randomly initialized manner. At the same time, in order to make the representations of phonemes in different positions in a word different, the position vector (position representation feature) of each phoneme is randomly initialized, including four positions. When a word contains multiple phonemes, it represents the start position (B), middle position (I), and end position (E) of the word. When a word contains one phoneme, it is represented by S. These positions are encoded to obtain the position vector of each phoneme. Finally, the unique encoded representation (pronunciation vector) of each phoneme and the position encoded representation (position vector) are added together to obtain the final phoneme feature. After the phonemes of a given text are input into the phoneme encoder, a depth feature representation (phoneme feature) of each phoneme is obtained.

[0144] In some embodiments, the phoneme feature and the audio feature are fused based on an attention mechanism. In the embodiments of the present application, an attention mechanism is used to fuse the phoneme feature and the audio feature. The attention mechanism is used to model the relationship between a query vector Q, a key vector K, and a value vector V. See formulas (8) and (9):

[0145]

[0146] Attention(Q,K,V)=AttentionScore(Q,K)*V (9);

[0147] Among them, the audio feature of each audio frame is used as the query vector Q, and the phoneme features H of all phonemes of a given text phone are used as the key vector K and the value vector V. AttentionScore(Q,K) is the attention score, and Attention(Q,K,V) is the attention result of each audio frame corresponding to all phonemes. d k is the dimension of the key vector K.

[0148] In some embodiments, the matrix obtained based on the attention mechanism is concatenated with the audio feature, and finally a fused feature is obtained. See formula (10):

[0149]

[0150] Among them, is the attention result of audio frame i obtained based on the attention mechanism. The attention result of audio frame i is a matrix, and each column in the matrix represents the attention result of each phoneme among all phonemes to audio frame i. is the audio feature of audio frame i, H phone is the phoneme feature of all phonemes of the given text. is the fused feature corresponding to each audio frame.

[0151] In some embodiments, during the training process of the phoneme alignment model, the cross loss is used to calculate the losses of the two classifications. Refer to Formula (11) and Formula (12):

[0152]

[0153]

[0154] Among them, L phone is the phoneme classification loss (the first phoneme category loss), L sil is the loudness classification loss (the second loudness category loss), m is the number of audio frames, c is the number of candidate phonemes, is the true identification result of the j-th phoneme corresponding to the i-th audio frame. is the first probability of the j-th phoneme corresponding to the i-th audio frame. is the pre-labeled alignment identification of the i-th audio frame, 1 for non-silent and 0 for silent. is the probability that the i-th audio frame is a non-silent audio frame.

[0155] In some embodiments, in order to better fuse the phoneme feature and the audio feature representation, the attention score matrix in the embodiments of the present application is constrained, that is, attention weight constraint is performed. Among them, each row in the matrix represents an audio frame, and each column represents the probability distribution of each phoneme in the audio frame. The probability distribution of the phonemes of each audio frame is calculated with the actual phonemes corresponding to the audio frame to obtain the attention mechanism loss. Refer to Formula (13):

[0156]

[0157] Among them, L align is the attention mechanism loss, m is the number of audio frames, N p is the number of phonemes in the given text. is 1 or 0. 1 indicates that the i-th audio frame is aligned with the j-th phoneme, and 0 indicates that the i-th audio frame is not aligned with the j-th phoneme. is the attention score between the i-th audio frame and the j-th phoneme.

[0158] In some embodiments, the joint loss of the entire phoneme alignment network consists of three parts, including the phoneme classification loss (the first phoneme class loss), the loudness classification loss (the second loudness class loss), and the alignment loss (the third alignment loss). The three losses are weighted and summed using different weights, and the final joint loss is shown in Equation (14):

[0159] L total = λL phone + βL sil + γL align (14);

[0160] where the weights (λ, β, and γ) of each loss are preset weights, and the sum of the three is equal to 1. L phone is the phoneme classification loss (the first phoneme class loss), L sil is the loudness classification loss (the second loudness class loss), and L align is the alignment loss (the third alignment loss), and L total is the joint loss.

[0161] In some embodiments, referring to Figure 7 , Figure 7 is a schematic data flow diagram of an AI-based audio processing method provided by an embodiment of the present application. The phoneme alignment model includes an attention fusion network, a phoneme classification network (corresponding to the second task), and a loudness classification network (corresponding to the first task). The input of the audio encoder is audio data, and the output of the audio encoder is the audio feature (in vector form) of each audio frame. The input of the phoneme encoder is a phoneme sequence (given text), and the output of the phoneme encoder is the phoneme feature (in vector form) of each phoneme. The input of the attention fusion network is the output of the audio encoder and the output of the phoneme encoder, and the output of the attention fusion network is the fusion feature of the phoneme feature and the audio feature. The audio feature of each audio frame is calculated with all phonemes through an attention mechanism to obtain the fusion feature, and the representation of the candidate phoneme corresponding to the audio frame and the representation of whether it is silent or not are determined. The fusion feature is classified through two parallel phoneme classification networks and a loudness classification network. The phoneme classification network outputs the probability that each audio frame belongs to each candidate phoneme, and the loudness classification network outputs the probability that each audio frame belongs to each loudness category. The loudness categories include silent and non-silent. For example, the identifier of non-silent is 1, the identifier of silent is 0, and the candidate phonemes are W, IH, L, etc.

[0162] In some embodiments, the embodiments of the present application conduct experiments on two public datasets, including the TIMIT dataset and the Buckeye dataset. These two datasets mark the time of each phoneme in the audio, and finally calculate metrics, including at least one of the following: the precision P, recall rate R, and F1 score of the phoneme boundaries predicted by the phoneme alignment model. In addition, to solve the problem that the F1 score is relatively high when the recall rate is high and the precision rate is low, R-value is introduced for evaluation. See Formulas (8)-(10):

[0163]

[0164]

[0165]

[0166] Among them, P is the precision rate, R is the recall rate, and OS is R / P - 1.

[0167] For the final results, see Table 1. Discrimi, Montreal, and SEGFEAT are all models in the related technologies. It can be seen from Table 1 that the embodiments of the present application have greatly improved the accuracy of phoneme boundaries on different public datasets.

[0168] Table 1 Scoring of each model in the embodiments of the present application and related technologies on each dataset

[0169] Corpora Model P R F1 R - value TIMIT Ours 93.42 95.96 94.67 95.18 TIMIT Discrimi 90 82.2 85.9 79.51 TIMIT Montreal 83.9 81.6 82.7 85.16 TIMIT SEGFEAT 92.67 93.03 92.85 93.91 Buckeye Ours 88.49 90.33 89.40 90.90 Buckeye SEGFEAT 85.40 89.12 87.23 88.76

[0170] See Figures 8A - 8C , Figures 8A - 8C is the alignment time matrix of the audio processing method based on artificial intelligence provided by the embodiments of the present application. To verify the effectiveness of the attention mechanism constraint, a phoneme alignment matrix is drawn. Among them, the vertical axis is the audio frames divided by time, and the horizontal axis is each phoneme. Figure 8A shows the alignment time matrix without adding the attention weight constraint. Figure 8B shows the alignment time matrix with the added constraint. Figure 8C shows the true alignment time matrix. It can be seen that the matrix with the attention mechanism constraint is overall more in line with the actual alignment time of phonemes and audio.

[0171] It can be understood that in the embodiments of the present application, data related to user information, etc. is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0172] Next, the exemplary structure of the software module implementation of the artificial intelligence-based audio processing device 255 provided in the embodiments of the present application will be further described. In some embodiments, as Figure 2 shown, the software module stored in the artificial intelligence-based audio processing device 255 in the memory 250 may include: a phoneme module 2551, configured to obtain at least one phoneme of a given text and determine the phoneme features of each phoneme; an audio module 2552, configured to obtain at least one audio frame of the audio data corresponding to the given text and determine the audio features of each audio frame; a fusion module 2553, configured to perform the following processing for each audio frame: perform an attention mechanism-based fusion process on the audio features of the audio frame and the phoneme features of at least one phoneme to obtain the fusion features corresponding to each audio frame; an alignment module 2554, configured to determine the phoneme corresponding to each audio frame based on the fusion features of each audio frame, and determine the start and end times of each phoneme based on the phoneme corresponding to each audio frame.

[0173] In some embodiments, determining the audio features of each audio frame is implemented by invoking an audio encoder. The audio encoder includes a plurality of convolutional networks and a normalization network. The audio module 2552 is further configured to: perform feature extraction processing on at least one audio frame through a plurality of cascaded convolutional networks to obtain the convolutional feature extraction result corresponding to each audio frame; perform normalization processing on the convolutional feature extraction result of each audio frame through the normalization network to obtain the audio features of each audio frame.

[0174] In some embodiments, determining the phoneme features of each phoneme is implemented by invoking a phoneme encoder. The phoneme encoder includes a phoneme characteristic representation network and a phoneme position representation network. The phoneme module 2551 is further configured to: perform the following processing for each phoneme: determine the characteristic representation feature of the phoneme through the phoneme characteristic representation network, where the characteristic representation feature is used to characterize the characteristics of the phoneme; determine the position representation feature of the phoneme through the phoneme position representation network, where the position representation feature is used to characterize the position of the phoneme in the corresponding text unit; perform an addition process on the position representation feature and the characteristic representation feature to obtain the phoneme features of the phoneme.

[0175] In some embodiments, the attention mechanism-based fusion process is implemented by invoking an attention fusion network. The attention fusion network includes an attention layer and a fusion layer. The fusion module 2553 is further configured to: perform an attention process on the audio features of the audio frame and the phoneme features of at least one phoneme through the attention layer to obtain an attention result; perform a fusion process on the attention result and the audio features through the fusion layer to obtain the fusion features corresponding to the audio frame.

[0176] In some embodiments, the fusion module 2553 is further configured to perform the following processing for each phoneme: based on the audio features of the audio frames and the phoneme features of the phoneme, determine the attention score of the corresponding phoneme; perform a value vector transformation process on the phoneme features of the phoneme to obtain a value vector; multiply the attention score of the corresponding phoneme by the value vector to obtain the attention result of the corresponding phoneme.

[0177] In some embodiments, the fusion module 2553 is further configured to: perform a query vector transformation process on the audio features to obtain a query vector; perform a key vector transformation process on the phoneme features to obtain a key vector; multiply the query vector by the transpose of the key vector to obtain a multiplication result; determine the ratio of the multiplication result to the square root of the dimension of the key vector as the attention feature; perform a maximum likelihood process on the attention feature to obtain the attention score of the corresponding phoneme.

[0178] In some embodiments, determining the phoneme corresponding to each audio frame is implemented by calling a phoneme classification network, and the phoneme classification network includes at least one cascaded phoneme fully connected layer. The alignment module 2554 is further configured to perform the following processing for each audio frame: when the number of phoneme fully connected layers is one, perform a first fully connected process on the fusion features through the phoneme fully connected layer to obtain the first probability that the audio frame belongs to each candidate phoneme; when the number of phoneme fully connected layers is multiple, perform a first fully connected process on the input of the n-th phoneme fully connected layer through the n-th phoneme fully connected layer in the N cascaded phoneme fully connected layers, and transmit the n-th phoneme fully connected result output by the n-th phoneme fully connected layer to the (n + 1)-th phoneme fully connected layer to continue the first fully connected process to obtain the (n + 1)-th phoneme fully connected result corresponding to the (n + 1)-th phoneme fully connected layer; where N is an integer greater than or equal to 2, n is an integer variable starting from 1 and increasing, and the value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the n-th phoneme fully connected layer is the fusion feature. When 2 ≤ n < N, the input of the n-th phoneme fully connected layer is the (n - 1)-th phoneme fully connected result output by the (n - 1)-th phoneme fully connected layer. When n takes the value of N - 1, the (n + 1)-th phoneme fully connected result is the first probability that the audio frame belongs to each candidate phoneme; determine the candidate phoneme with the largest first probability as the phoneme corresponding to the audio frame.

[0179] In some embodiments, the alignment module 2554 is further configured to: based on the phoneme corresponding to each audio frame, determine at least one audio frame corresponding to each phoneme; perform the following processing for each phoneme: when the phoneme corresponds to multiple consecutive audio frames, determine the start and end times of the consecutive audio frames corresponding to the phoneme as the start and end times of the phoneme; when the phoneme corresponds to one audio frame, determine the time of the audio frame corresponding to the phoneme as the start and end times of the phoneme.

[0180] In some embodiments, the fusion processing based on the attention mechanism is implemented by invoking an attention fusion network, and determining the phoneme corresponding to each audio frame is implemented by invoking a phoneme classification network. The phoneme classification network shares the attention fusion network with the loudness classification network. The apparatus further includes: a training module 2555, configured to: obtain an audio data sample and a given text sample; obtain at least one phoneme sample of the given text sample, and determine the phoneme feature of each phoneme sample through a phoneme encoder; obtain at least one audio frame sample of the audio data sample corresponding to the given text sample, and determine the audio feature of each audio frame sample through an audio encoder; perform the following processing for each audio frame sample: perform forward propagation on the audio feature of the audio frame sample and the phoneme features of at least one phoneme sample in the attention fusion network and the phoneme classification network to obtain a first forward propagation result; perform the following processing for each audio frame sample: perform forward propagation on the audio feature of the audio frame sample and the phoneme features of at least one phoneme sample in the attention fusion network and the loudness classification network to obtain a second forward propagation result; determine a joint loss according to the first forward propagation result and the second forward propagation result; update the parameters of the attention fusion network, the loudness classification network, the loudness classification network, the audio encoder, and the phoneme encoder according to the joint loss.

[0181] In some embodiments, the training module 2555 is further configured to: perform attention mechanism-based fusion processing on the audio feature of the audio frame sample and the phoneme features of at least one phoneme sample through the attention fusion network to obtain a fusion feature corresponding to each audio frame sample; perform a second fully connected processing on the fusion feature of each audio frame sample through the loudness classification network to obtain a second probability that each audio frame sample belongs to each loudness category, and form the second forward propagation result.

[0182] In some embodiments, the training module 2555 is further configured to: for each phoneme sample through the attention layer of the attention fusion network: for each of the phoneme samples through the attention layer of the attention fusion network: determine an attention score corresponding to the phoneme sample based on the audio feature of the audio frame sample and the phoneme feature of the phoneme sample; perform a value vector transformation process on the phoneme feature of the phoneme sample, and multiply the attention score corresponding to the phoneme sample by the value vector transformation result to obtain an attention result corresponding to the phoneme sample; perform fusion processing on the attention result corresponding to each phoneme sample and the audio feature of the audio frame sample through the fusion layer of the attention fusion network to obtain a fusion feature corresponding to the audio frame sample; perform a first fully connected processing on the fusion feature of the audio frame sample through the phoneme classification network to obtain a third probability that the audio frame sample belongs to each candidate phoneme; and form the first forward propagation result by combining the third probability and the attention score.

[0183] In some embodiments, the training module 2555 is further configured to: determine a first phoneme class loss based on the third probabilities corresponding to multiple candidate phonemes for each audio frame sample and the pre-labeled candidate phonemes for each audio frame sample; determine a second loudness class loss based on the second probabilities corresponding to multiple loudness classes for each audio frame sample and the pre-labeled loudness classes for each audio frame sample; determine a third alignment loss based on the attention scores corresponding to each phoneme sample for each audio frame sample and the pre-labeled alignment identifiers corresponding to each phoneme sample for each audio frame sample; and perform a fusion process on the first phoneme class loss, the second loudness class loss, and the third alignment loss to obtain a joint loss.

[0184] An embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the method described above in the embodiments of the present application.

[0185] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions, when executed by a processor, will cause the processor to execute the audio processing method based on artificial intelligence provided by the embodiments of the present application. For example, Figures 3A - 3C the audio processing method based on artificial intelligence shown.

[0186] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0187] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0188] As an example, the executable instructions may or may not correspond to files in a file system, and may be stored as part of a file that holds other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program under discussion, or in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).

[0189] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected via a communication network.

[0190] In summary, through the embodiments of the present application, an attention mechanism calculation is performed on the audio feature and the text sequence to obtain a fusion feature. Therefore, the fusion feature can effectively represent the relationship between the audio frame and the phoneme. Then, based on the fusion feature, phoneme classification is performed on each audio frame in the audio, which can effectively improve the classification accuracy, thereby improving the phoneme alignment accuracy.

[0191] The above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. An audio processing method based on artificial intelligence, characterized in that The method includes: Obtaining at least one phoneme of the given text and determining the phoneme features of each of the phonemes; Obtaining at least one audio frame of the audio data corresponding to the given text and determining the audio features of each of the audio frames; Performing the following processing for each of the audio frames: performing attention processing on the audio features of the audio frame and the phoneme features of at least one of the phonemes through an attention layer included in an attention fusion network to obtain an attention result of the audio frame; performing fusion processing on the attention result of the audio frame and the audio features of the audio frame through a fusion layer included in the attention fusion network to obtain a fusion feature corresponding to the audio frame; Based on the fusion features of each of the audio frames, determining the phoneme corresponding to each of the audio frames, and based on the phoneme corresponding to each of the audio frames, determining the start and end times of each of the phonemes in the audio data.

2. The method according to claim 1, wherein The determining of the phoneme features of each of the phonemes is implemented by invoking a phoneme encoder, and the phoneme encoder includes a phoneme characteristic representation network and a phoneme position representation network. The determining of the phoneme features of each of the phonemes includes: Performing the following processing for each of the phonemes: Determining a characteristic representation feature of the phoneme through the phoneme characteristic representation network, where the characteristic representation feature is used to characterize the characteristics of the phoneme; Determining a position representation feature of the phoneme through the phoneme position representation network, where the position representation feature is used to characterize the position of the phoneme in the corresponding text unit; Performing an addition process on the position representation feature and the characteristic representation feature to obtain the phoneme feature of the phoneme.

3. The method according to claim 1, characterized in that, The performing of attention processing on the audio features of the audio frame and the phoneme features of at least one of the phonemes to obtain an attention result of the audio frame includes: Performing the following processing for each of the phonemes: Based on the audio features of the audio frame and the phoneme features of the phoneme, determining an attention score corresponding to the phoneme; Performing a value vector transformation process on the phoneme features of the phoneme to obtain a value vector; Performing a multiplication process on the attention score corresponding to the phoneme and the value vector to obtain an attention result corresponding to the phoneme.

4. The method according to claim 3, characterized in that The determining of the attention score corresponding to the phoneme based on the audio features of the audio frame and the phoneme features of the phoneme includes: Performing a query vector transformation process on the audio features to obtain a query vector; Performing a key vector transformation process on the phoneme features to obtain a key vector; Performing a multiplication process on the query vector and the transpose of the key vector to obtain a multiplication result; Determining the ratio of the multiplication result to the square root of the dimension of the key vector as the attention feature; Performing a maximum likelihood process on the attention feature to obtain an attention score corresponding to the phoneme.

5. The method according to claim 1, characterized in that, The determining of the phoneme corresponding to each of the audio frames is implemented by invoking a phoneme classification network, and the phoneme classification network includes at least one cascaded phoneme fully connected layer. The determining of the phoneme corresponding to each of the audio frames based on the fusion features of each of the audio frames includes: Performing the following processing for each of the audio frames: When the number of the phoneme fully-connected layers is one, perform a first fully-connected process on the fusion feature through the phoneme fully-connected layer to obtain a first probability that the audio frame belongs to each candidate phoneme; When the number of the phoneme fully-connected layers is multiple, perform a first fully-connected process on the input of the n-th phoneme fully-connected layer through the n-th phoneme fully-connected layer in N cascaded phoneme fully-connected layers, and transmit the n-th phoneme fully-connected result output by the n-th phoneme fully-connected layer to the (n + 1)-th phoneme fully-connected layer to continue the first fully-connected process to obtain the (n + 1)-th phoneme fully-connected result corresponding to the (n + 1)-th phoneme fully-connected layer; where N is an integer greater than or equal to 2, n is an integer variable starting from 1 and increasing, the value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the n-th phoneme fully-connected layer is the fusion feature. When n takes the value of 2 ≤ n < N, the input of the n-th phoneme fully-connected layer is the (n - 1)-th phoneme fully-connected result output by the (n - 1)-th phoneme fully-connected layer. When n takes the value of N - 1, the (n + 1)-th phoneme fully-connected result is the first probability that the audio frame belongs to each candidate phoneme; Determine the candidate phoneme with the largest first probability as the phoneme corresponding to the audio frame.

6. The method according to claim 1, wherein Determining the start and end times of each phoneme in the audio data based on the phoneme corresponding to each audio frame includes: Based on the phoneme corresponding to each audio frame, determine at least one audio frame corresponding to each phoneme; Perform the following processing for each phoneme: When the phoneme corresponds to multiple consecutive audio frames, determine the start and end times of the consecutive audio frames corresponding to the phoneme as the start and end times of the phoneme; When the phoneme corresponds to one audio frame, determine the time of the audio frame corresponding to the phoneme as the start and end times of the phoneme in the audio data.

7. The method according to claim 1, wherein The fusion processing based on the attention mechanism is implemented by calling an attention fusion network. Determining the phoneme corresponding to each audio frame is implemented by calling a phoneme classification network. The phoneme classification network shares the attention fusion network with the loudness classification network. The input of the attention fusion network is the output of an audio encoder and a phoneme encoder. The method further includes: Obtain a given text sample and an audio data sample corresponding to the given text sample; Obtain at least one phoneme sample of the given text sample, and determine the phoneme feature of each phoneme sample through the phoneme encoder; Obtain at least one audio frame sample of the audio data sample, and determine the audio feature of each audio frame sample through the audio encoder; Perform the following processing for each audio frame sample: Propagate the audio feature of the audio frame sample and the phoneme features of at least one phoneme sample forward in the attention fusion network and the phoneme classification network to obtain a first forward propagation result; Perform the following processing for each of the audio frame samples: Propagate the audio features of the audio frame sample and the phoneme features of at least one of the phoneme samples forward in the attention fusion network and the loudness classification network to obtain a second forward propagation result; Determine a joint loss according to the first forward propagation result and the second forward propagation result; Update the parameters of the attention fusion network, the phoneme classification network, the loudness classification network, the audio encoder, and the phoneme encoder according to the joint loss.

8. The method according to claim 7, wherein The step of propagating the audio features of the audio frame sample and the phoneme features of at least one of the phoneme samples forward in the attention fusion network and the loudness classification network to obtain a second forward propagation result includes: Perform attention mechanism-based fusion processing on the audio features of the audio frame sample and the phoneme features of at least one of the phoneme samples through the attention fusion network to obtain fusion features corresponding to each audio frame sample; Perform a second fully connected process on the fusion features of each audio frame sample through the loudness classification network to obtain a second probability of each audio frame sample belonging to each loudness category, which constitutes the second forward propagation result.

9. The method according to claim 7, wherein The step of propagating the audio features of the audio frame sample and the phoneme features of at least one of the phoneme samples forward in the attention fusion network and the phoneme classification network to obtain a first forward propagation result includes: Perform the following processing for each phoneme sample through the attention layer of the attention fusion network: Determine an attention score corresponding to the phoneme sample based on the audio features of the audio frame sample and the phoneme features of the phoneme sample; Perform a value vector transformation process on the phoneme features of the phoneme sample, and multiply the attention score corresponding to the phoneme sample by the value vector transformation result to obtain an attention result corresponding to the phoneme sample; Perform fusion processing on the attention results corresponding to each phoneme sample and the audio features of the audio frame sample through the fusion layer of the attention fusion network to obtain fusion features corresponding to the audio frame sample; Perform a first fully connected process on the fusion features of the audio frame sample through the phoneme classification network to obtain a third probability of the audio frame sample belonging to each candidate phoneme; The third probability and the attention score constitute the first forward propagation result.

10. The method according to claim 9, wherein The step of determining a joint loss according to the first forward propagation result and the second forward propagation result includes: Determine a first phoneme category loss based on the third probabilities of each audio frame sample corresponding to multiple candidate phonemes and the pre-labeled candidate phonemes of each audio frame sample; Determine a second loudness category loss based on the second probabilities of each audio frame sample corresponding to multiple loudness categories and the pre-labeled loudness categories of each audio frame sample; Determine a third alignment loss based on the attention scores of each audio frame sample corresponding to each phoneme sample and the pre-labeled alignment identifiers of each audio frame sample corresponding to each phoneme sample; Fuse the first phoneme class loss, the second loudness class loss, and the third alignment loss to obtain the joint loss.

11. An audio processing device based on artificial intelligence, characterized in that, The apparatus includes: A phoneme module, configured to obtain at least one phoneme of a given text and determine the phoneme features of each of the phonemes; An audio module, configured to obtain at least one audio frame of audio data corresponding to the given text and determine the audio features of each of the audio frames; A fusion module, configured to perform the following processing for each of the audio frames: perform attention processing on the audio features of the audio frame and the phoneme features of at least one of the phonemes through an attention layer included in an attention fusion network to obtain an attention result of the audio frame; fuse the attention result of the audio frame and the audio features of the audio frame through a fusion layer included in the attention fusion network to obtain a fusion feature corresponding to the audio frame; An alignment module, configured to determine the phoneme corresponding to each of the audio frames based on the fusion feature of each of the audio frames, and determine the start and end times of each of the phonemes based on the phoneme corresponding to each of the audio frames.

12. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable instructions; A processor, configured to implement the artificial intelligence-based audio processing method according to any one of claims 1 to 10 when executing the executable instructions stored in the memory.

13. A computer-readable storage medium storing executable instructions, characterized in that, The executable instructions, when executed by the processor, implement the artificial intelligence-based audio processing method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program or instructions, characterized in that, The computer program or instructions, when executed by the processor, implement the artificial intelligence-based audio processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Combining auditory attention cues with phoneme posterior scores for phone / vowel / syllable boundary detection

    CN104756182A

  • Pronunciation mistake detection method and device based on depth learning

    CN106297828A