Audio processing methods, apparatus, devices, and computer programs
By extracting audio content and paralinguistic vectors and fusing them with suggested words, the voice processing method enhances speech recognition accuracy by considering hidden information like emotions and intonation, improving the conversion to text.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-07-11
- Publication Date
- 2026-05-26
AI Technical Summary
Current voice recognition technologies fail to accurately convert voice data into text information due to the lack of consideration for auxiliary information such as paralinguistic features, leading to reduced accuracy in speech understanding.
A voice processing method that extracts both audio content and paralinguistic vectors from voice data, fuses these with suggested words to obtain audio fusion features, and performs speech conversion processing to enhance speech recognition accuracy.
Improves the accuracy of speech recognition by incorporating paralinguistic information, enabling advanced speech understanding and conversion into more accurate text information.
Smart Images

Figure 2026516854000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the priority of a Chinese patent application filed with the Chinese Patent Office on September 12, 2023, with an application number of 2023111711595 and an invention title of "Voice Processing Method, Apparatus, Device, and Storage Medium", and the entire content thereof is incorporated herein by reference.
[0002] This application relates to the field of artificial intelligence technology, and particularly to voice processing methods, apparatuses, devices, and storage media.
Background Art
[0003] Voice recognition technology is utilized in various scenarios. For example, in an intelligent conversation scenario, by recognizing and understanding the voice data of the conversation partner, it is possible to understand the intention that the conversation partner wants to express, select appropriate response data, and accurately respond. However, the voice data of the conversation partner usually contains auxiliary information for assisting in the recognition of the voice data of the conversation partner in addition to the text content. However, the current voice recognition and understanding technology only converts the voice data of the conversation partner into text information, and cannot reflect the auxiliary information contained in the voice data in the text information, so the accuracy of voice recognition is low.
Summary of the Invention
[0004] Embodiments of this application provide a voice processing method, apparatus, device, and storage medium that can improve the accuracy of voice recognition.
[0005] In a first aspect, this application provides a voice processing method. This method includes: Performing feature extraction on the processed voice data to obtain target voice expression information of the processed voice data, where the target voice expression information includes a voice content vector and a para-language vector corresponding to the processed voice data, and the para-language vector is used to assist in the recognition of the text information corresponding to the processed voice data; The process involves obtaining suggested words related to the audio data to be processed, fusing the audio content vector, paralinguistic vector, and suggested words to obtain audio fusion features, and The process includes the steps of performing speech conversion processing on speech fusion features and obtaining text information corresponding to the processed speech data.
[0006] In a second embodiment, the present application provides an audio processing device. This device is A feature extraction unit for extracting features from audio data to be processed and obtaining target audio representation information of the audio data to be processed, wherein the target audio representation information includes an audio content vector and a paralinguistic vector corresponding to the audio data to be processed, and the paralinguistic vector is used to assist in the recognition of text information corresponding to the audio data to be processed, and the feature extraction unit An information fusion unit for obtaining suggested words related to the audio data to be processed, and for fusing the audio content vector, paralinguistic vector, and suggested words to obtain audio fusion features, The system includes a speech conversion unit that performs speech conversion processing on speech fusion features and obtains text information corresponding to the processed speech data.
[0007] In a third embodiment, the present invention provides a computer device comprising a processor, memory, and a network interface. The processor is connected to the memory and the network interface, the network interface is used to provide data communication functions, and the memory is used to store computer programs, including program instructions. The processor is configured to call program instructions to cause the computer device comprising the processor to execute the above-described voice processing method.
[0008] In a fourth embodiment, the present invention provides a computer-readable storage medium in which a computer program is stored. The computer program is suitable for causing a computer device equipped with a processor to perform the above-described speech processing method by being loaded and executed by a processor.
[0009] In a fifth embodiment, the present application provides a computer program product or computer program including computer instructions. When the computer instructions are executed by a processor, they can realize the above-described speech processing method.
[0010] In the embodiments of this invention, feature extraction is performed on the audio data to be processed to obtain target audio representation information of the audio data to be processed, suggested words related to the audio data to be processed are obtained, the target audio representation information and suggested words are fused to obtain audio fusion features, and audio conversion processing is performed on the audio fusion features to obtain text information corresponding to the audio data to be processed. The target audio representation information includes an audio content vector and a paralanguage vector corresponding to the audio data to be processed, and the paralanguage vector is used to assist in the recognition of text information corresponding to the audio data to be processed. Therefore, when performing speech recognition on the audio data to be processed, speech recognition can be performed by combining information about the audio content of the audio data to be processed, information about the paralanguage of the audio data to be processed, and text content corresponding to suggested words. Because speech recognition processing is performed using more comprehensive and richer audio representation information, this invention enables advanced speech recognition and understanding of the audio data to be processed, and improves the accuracy of speech recognition. [Brief explanation of the drawing]
[0011] [Figure 1] This is a schematic diagram of the network architecture of the voice processing system according to an embodiment of the present invention. [Figure 2] This is a schematic diagram of an application scenario for the speech processing method according to an embodiment of the present invention. [Figure 3] This is a flowchart of the audio processing method according to an embodiment of the present invention. [Figure 4] This is a schematic diagram of the architecture of the speech feature extraction model according to an embodiment of the present invention. [Figure 5] This is a flowchart of the training method for the speech feature extraction model according to the embodiment of the present invention. [Figure 6]This is a flowchart of the training method for the speech conversion model according to the embodiment of the present invention. [Figure 7] This is a schematic diagram of the parameter adjustment of the speech conversion model according to the embodiment of the present invention. [Figure 8] This is a schematic diagram of the configuration of the audio processing device according to an embodiment of the present invention. [Figure 9] This is a schematic diagram of the configuration of a computer device according to an embodiment of the present invention. [Modes for carrying out the invention]
[0012] In some speech processing tasks, speech recognition and speech understanding are performed by first performing speech recognition on audio data to obtain text data, and then processing the text data to achieve speech understanding. Speech understanding is performed based on the text data obtained by speech recognition, and since the text data contains only text content, speech understanding can be performed based solely on the text content, and paralinguistic information contained in the audio data cannot be used. Paralinguistic information may be information hidden in the audio data, such as emotions, facial expressions, and intonation. Therefore, if the speech recognition effect is poor in a complex scenario, and the speech recognition result is used as a prerequisite for speech understanding, errors in speech understanding will accumulate, making error correction impossible and reducing the accuracy of speech recognition.
[0013] Therefore, this application provides a speech processing method that can be applied to any speech interaction scenario, such as meeting scenarios, interview scenarios, intelligent dialogue scenarios, and game scenarios, and that can improve the accuracy of speech recognition by converting speech data into text information in the above speech interaction scenarios, thereby facilitating the execution of related business processing operations in the speech interaction scenarios. Specifically, this speech processing method does not directly recognize and understand speech data as text data, but rather extracts speech representation information including speech content vectors and paralinguistic vectors from the speech data, and obtains final text information by further processing the speech content vectors, paralinguistic vectors, and presented words. Since speech data contains paralinguistic information, the accuracy of speech understanding can be improved by combining the paralinguistic information with the speech data. Specifically, the principle of this speech processing method is generally as follows. (1) Acquire the audio data to be processed in the business scenario. Here, the business scenario includes, but is not limited to, any audio interaction scenario such as a meeting scenario, an interview scenario, an intelligent dialogue scenario, or a game scenario. The audio data to be processed may be generated by any object in the business scenario, such as a meeting host in a meeting scenario, or a game-controlled object or game character in a game scenario. (2) Feature extraction is performed on the audio data to be processed to obtain target audio representation information for the audio data to be processed. Here, the target audio representation information includes an audio content vector and a paralinguistic vector corresponding to the audio data to be processed, and the paralinguistic vector is used to assist in the recognition of text information corresponding to the audio data to be processed. (3) The presented words related to the audio data to be processed are obtained, and the audio content vector, paralinguistic vector, and presented words are fused to obtain an audio fusion feature. Audio conversion processing is performed on the audio fusion feature to obtain text information corresponding to the audio data to be processed.
[0014] As can be seen from the above, in the embodiment of the present invention, the target speech representation information includes a speech content vector and a paralanguage vector corresponding to the speech data to be processed. The paralanguage vector is used to assist in the recognition of text information corresponding to the speech data to be processed. Therefore, when performing speech recognition on the speech data to be processed, speech recognition can be performed by combining information about the speech content of the speech data to be processed, information about the paralanguage of the speech data to be processed, and text content corresponding to the presented words. Because more comprehensive and richer speech representation information is used, the present invention enables advanced speech recognition and understanding of the speech data to be processed, and improves the accuracy of speech recognition.
[0015] The technical means of this application can be applied to scenarios that convert audio data into text information through speech recognition. For example, it can be used in scenarios such as speech-to-text conversion in online meetings, voice input in social applications, speech-to-text conversion from recording devices in interview scenarios, and voice conversations in intelligent dialogue. For example, in an online meeting scenario, the efficiency of obtaining important meeting content can be improved by recording the audio during the meeting to obtain audio data, and then performing speech-to-text conversion on the audio data to obtain corresponding meeting minutes. In a social application, the efficiency of text information input can be improved by obtaining user audio data, performing speech recognition, and obtaining text information. Furthermore, in an interview scenario, the efficiency of obtaining interview text can be improved by recognizing and converting audio data recorded by a recording device into text. Furthermore, in an intelligent dialogue scenario, accurate text dialogue can be achieved by performing speech recognition on the voice data of the dialogue partner and obtaining text information. Optionally, the technical means of this application can also be used in a variety of scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and driving assistance.
[0016] Note that the embodiments of the present application include data related to target information (such as processed voice data, sample voice data, prompting words, etc.). When the embodiments of the present application are used in specific products or technologies, permission or consent of the target is required, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards in the relevant region. For example, the target may be a user of a terminal device or a computer device.
[0017] Hereinafter, the system architecture according to the embodiments of the present application will be described in detail.
[0018] Please refer to FIG. 1. FIG. 1 is a schematic diagram of the network architecture of the voice processing system according to the embodiments of the present application. As shown in FIG. 1, the computer device can exchange data with the terminal device, and the number of terminal devices may be one or at least two. For example, when the number of terminal devices is plural, the terminal devices may include terminal device 101a, terminal device 101b, terminal device 101c, etc. in FIG. 1. Here, taking terminal device 101a as an example, computer device 102 can perform feature extraction on the processed voice data and obtain the target voice expression information of the processed voice data. Furthermore, computer device 102 can obtain the prompting words related to the processed voice data, and fuse-process the voice content vector, the paralanguage vector, and the prompting words to obtain the voice fusion feature. Furthermore, computer device 102 can perform voice conversion processing on the voice fusion feature and obtain the text information corresponding to the processed voice data. Optionally, computer device 102 can also send the text information corresponding to the processed voice data to terminal device 101a so that terminal device 101a displays the text data, or computer device 102 can determine the response text information based on the text information corresponding to the processed voice data and send the response text information to terminal device 101a.
[0019] To ensure understanding, the computer equipment referred to in the embodiments of this application includes, but is not limited to, terminal equipment or servers. In other words, computer equipment may be a server or terminal equipment, or a system consisting of a server and terminal equipment. Here, terminal equipment may be electronic devices and may include, but is not limited to, mobile phones, tablet computers, desktop computers, laptop computers, personal digital assistants, in-vehicle devices, smart voice interaction devices, augmented reality / virtual reality (AR / VR) devices, helmet displays, wearable devices, smart speakers, smart home appliances, aircraft, digital cameras, cameras, and other mobile internet devices (MIDs) with network access capabilities. The servers may be independent physical servers, server clusters or distributed systems consisting of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-to-infrastructure communication, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0020] Please refer to FIG. 2. FIG. 2 is a schematic diagram of an application scenario of the voice processing method according to an embodiment of the present application. As shown in FIG. 2, the processed voice data 21 is input into the voice feature extraction model 22, and the voice feature extraction model 22 is used to perform feature encoding on the processed voice data 21 to obtain a voice vector matrix of the processed voice data 21. For example, the processed voice data 21 is "Yes, yes, yes, so it is, yes". The voice expression fully connected layer of the voice feature extraction model 22 is used to perform feature transformation on the voice vector matrix of the processed voice data 21 to obtain the target voice expression information 23 of the processed voice data 21. Here, the target voice expression information includes a voice content vector and a paralanguage vector corresponding to the processed voice data 21. Further, a prompt word 24 (such as a non-colloquial prompt word) related to the processed voice data 21 is obtained, and the target voice expression information 23 (voice content vector and paralanguage vector) and the prompt word 24 are input into the voice conversion model 25. The voice conversion model 25 is used to perform feature processing such as feature encoding on the prompt word 24 to obtain a feature vector matrix 26 corresponding to the prompt word 24. The voice conversion model 25 is used to perform fusion processing on the feature vector matrix 26 corresponding to the prompt word and the target voice expression information 23 of the processed voice data 21 to obtain a voice fusion feature. Thus, the voice conversion model 25 can convert the voice fusion feature into text information 27 corresponding to the processed voice data 21 and finally output the text information 27. For example, the non-colloquial text information 27 is "Yes, so it is".
[0021] Therefore, in the above application scenario, the target voice expression information includes a voice content vector and a paralanguage vector corresponding to the processed voice data. Since the paralanguage vector is used to assist in the recognition of the text information corresponding to the processed voice data, when performing voice recognition on the processed voice data, information about the voice content of the processed voice data, information about the paralanguage of the processed voice data, and the text content corresponding to the prompt word can be combined to perform voice recognition, enabling advanced voice recognition and understanding of the processed voice data and improving the accuracy of voice recognition.
[0022] The following describes specific embodiments relating to this application in detail.
[0023] Please refer to Figure 3. Figure 3 is a flowchart of an audio processing method according to an embodiment of the present invention. As shown in Figure 3, this audio processing method can be used in computer equipment and includes, but is not limited to, the following steps S101 to S103.
[0024] S101: Feature extraction is performed on the audio data to be processed, and target audio representation information of the audio data to be processed is obtained.
[0025] Specifically, the following methods can be used to acquire the audio data to be processed. Method 1: Collect audio data from business scenarios in real time and use it as the audio data to be processed. In possible implementations, audio data can be collected in real time from business scenarios (e.g., any business scenario such as a meeting scenario, interview scenario, intelligent dialogue scenario, or game scenario). For example, game audio data can be acquired from a game scenario using a recording device (such as a mobile phone microphone), or meeting audio data can be acquired from a meeting scenario using a recording device (such as a mobile phone microphone). The collected audio data can then be used as the audio data to be processed. Optionally, after collecting audio data in real time, the collected audio data can be preprocessed, and the preprocessed audio data can then be used as the audio data to be processed. Here, preprocessing includes, but is not limited to, processes such as audio noise reduction, audio alignment, and data normalization. Preprocessed audio data becomes more accurate. In this implementation, audio data can be collected in real time from business scenarios and used as the audio data to be processed. Real-time data collection enhances real-time performance and enables timely data processing. Method 2: The collected audio data is retrieved from a pre-configured database and used as the audio data to be processed. In a possible implementation, the pre-configured database stores audio data collected from various business scenarios, with different business scenarios associated with corresponding audio data. For example, game scenarios store game audio data, and meeting scenarios store meeting audio data. In this way, according to business needs, audio data collected in any business scenario can be directly retrieved from the pre-configured database and used as the audio data to be processed. In this implementation, the data acquisition efficiency is improved because the audio data to be processed can be retrieved directly from the pre-configured database without having to collect it in real time.
[0026] Here, the target speech representation information includes a speech content vector and a paralinguistic vector corresponding to the speech data to be processed. Specifically, (1) the speech content vector can reflect the speech content of the speech data to be processed. For example, the speech content may be the specific content contained in the speech frame corresponding to each character of the speech data to be processed. For example, if the speech data to be processed is "Have you eaten rice?", the speech data to be processed contains the characters "you", "rice", "eat", "did", and "are", so the speech content is the speech obtained after pronouncing the characters "you", "rice", "eat", "did", and "are". (2) The paralinguistic vector is used to assist in the recognition of text information corresponding to the speech data to be processed. The paralinguistic vector can reflect the paralinguistic information of the speech data to be processed, and so-called paralinguistic information can be used to assist in the recognition of the speech data to be processed. Paralinguistic information may include, for example, information such as the volume, timbre, and speed of the pronunciation of each character in the speech data to be processed, or it may include corresponding emotional information at the time of pronunciation. Therefore, paralinguistic information is data that reflects hidden information such as emotion and intonation in the audio data being processed, and so-called hidden information is information that cannot be directly discovered from the surface of the data.
[0027] Furthermore, since both the paralinguistic vector and the audio content vector of the audio data to be processed are vectors corresponding to information of the speech modality, it is difficult to separate the information of the two speech modalities, paralinguistic information and audio content. If the audio data to be processed is directly converted to text data, only the audio content can be converted to text data, but the paralinguistic information cannot. Moreover, since the converted text data contains only characters, information such as the emotion the speaker feels when pronouncing the characters and the speaker's volume cannot be reflected in the text data. In the embodiment of the present invention, since the paralinguistic information of the audio data reflects this information of the speaker in the form of speech, it is possible to adequately convey information such as the content of the speaker's utterance using the audio content and paralinguistic information.
[0028] The following describes in detail the specific process for extracting target speech representation information from the audio data to be processed.
[0029] Method 1: Extract target speech representation information from the audio data to be processed through presented words.
[0030] In one embodiment, the process for extracting target speech representation information from the audio data to be processed includes the following steps (1) to (3).
[0031] (1) Obtain feature transformation parameters for performing feature transformation on the presented words. Here, the feature transformation parameters can be used to perform feature transformation on the presented words and obtain a feature vector matrix corresponding to the presented words.
[0032] (2) Feature coding is performed on the audio data to be processed to obtain the audio vector matrix of the audio data to be processed. Here, by performing feature coding on the audio data to be processed, each character of the audio data to be processed can be coded into an audio vector. Feature coding of the audio data to be processed involves embedding the audio signal into a fixed-dimensional vector space and obtaining an audio vector. Since the audio data to be processed contains multiple characters, the audio vector matrix corresponding to the audio data to be processed can be obtained by coding the audio vectors corresponding to multiple characters and obtaining an audio vector matrix. For example, the dimension of the audio vector obtained by feature coding one character is m dimensions, and a matrix of m*n dimensions is obtained by feature coding audio data to be processed containing n characters. That is, the audio vector matrix may also be a matrix of m*n dimensions.
[0033] In one embodiment, the process of obtaining the audio vector matrix of the audio data to be processed may include the following steps 1-2.
[0034] 1. Divide the audio data to be processed to obtain N (where N is a positive integer) audio frames. For example, when obtaining audio data to be processed, frame division processing can be performed on the audio data to obtain N audio frames. Frame division processing involves dividing the audio data to be processed by frame length to obtain N audio frames. For example, the frame length can be any value, such as 10 milliseconds to 30 milliseconds. Generally, one character of the audio content of the audio signal to be processed can correspond to N audio frames. For example, if the audio signal corresponding to one character of the audio content is 1 second and the frame length is 30 milliseconds, the number of audio frames corresponding to that character is approximately 33.
[0035] 2. Feature coding is performed on each audio frame to obtain the audio vector matrix for each audio frame. Here, feature coding is the process of converting audio frames from audio data to feature vectors. Converting audio data to an audio vector matrix simplifies subsequent calculations. In selectable implementations, the position features of each audio frame can be combined to perform feature coding and obtain the audio vector matrix for each audio frame. For example, the position features of each audio frame can be determined based on the partitioning order of N audio frames. The position features indicate the position of the corresponding audio frame in the audio data being processed. Feature coding is performed on each audio frame to obtain the coded features of each audio frame, and for audio frame i among the N audio frames, feature concatenation is performed on the position features and coded features of audio frame i to obtain the audio vector matrix for audio frame i. Here, the position features of each audio frame can indicate the position of each audio frame in the audio data being processed. When splitting audio data to be processed, it is generally done according to the pronunciation order of the audio data. Therefore, the audio frames corresponding to audio data with earlier pronunciation are split first, and the audio frames corresponding to audio data with later pronunciation are split later. Thus, the positional features of each audio frame can be determined based on the splitting order of the N audio frames. Then, when concatenating the encoded features of the N audio frames, the positional features of each audio frame are combined to perform feature concatenation, and an audio vector matrix of the N audio frames can be obtained. By introducing positional features during feature encoding, the positional features of the audio frames can be combined and concatenated when concatenating the encoded features of the audio frames in the subsequent process, ensuring accuracy of character order in text information and improving the accuracy of speech recognition results.
[0036] In the embodiments of this application, if the probability that the currently traversing audio frame maps to a specific character in the audio content is greater than the probability threshold, it means that the currently traversing audio frame can be mapped to that character in the audio content. If the maximum probability among the probabilities that the currently traversing audio frame maps to each character in the audio content is less than the probability threshold, it means that the currently traversing audio frame cannot be mapped to each character in the audio content. That is, since there is little audio information in that audio frame, the candidate audio representation information of the currently traversing audio frame can be removed from the candidate audio representation information of the N audio frames. By removing the candidate audio representation information of audio frames whose probability of mapping to each character in the audio content of the audio data to be processed is less than the probability threshold from the N audio frames, after the traverse is complete, the target audio representation information of the audio data to be processed can be obtained based on the remaining candidate audio representation information. For example, the target audio representation information of the audio data to be processed can be obtained by concatenating or combining the remaining candidate audio representation information.
[0037] In the embodiments of this invention, the probability may be, for example, a posterior probability output from a speech feature extraction model. If the probability corresponding to a speech frame is smaller than the probability threshold, the corresponding speech frame may be omitted. By dividing the speech data to be processed into N speech frames, each character's speech data contains N speech frames, so among the N speech frames corresponding to one character, there are speech frames with little speech information. Even if these speech frames with little speech information are deleted, the overall speech recognition result will have little effect. Therefore, by deleting speech frames with little information in the speech data to be processed, the computational load can be reduced and computational efficiency can be improved.
[0038] (3) Feature transformation is performed on the audio vector matrix of the audio data to be processed using feature transformation parameters to obtain target audio representation information of the audio data to be processed. The dimension of the feature vector matrix represented by the target audio representation information is the same as the dimension of the feature vector matrix corresponding to the presented word. The purpose of performing feature transformation on the audio vector matrix of the audio data to be processed using feature transformation parameters is to make the dimension of the audio vector matrix of the audio data to be processed the same as the dimension of the feature vector matrix corresponding to the presented word, so that the audio vector matrix of the audio data to be processed and the feature vector matrix corresponding to the presented word can be concatenated in the subsequent steps.
[0039] In one embodiment, when performing a feature transformation on the audio vector matrix of the audio data to be processed and obtaining target audio representation information for the audio data to be processed, the computational complexity can be reduced by removing meaningless audio frames from the audio data to be processed. Specifically, a feature transformation is performed on the audio vector matrix of each audio frame using feature transformation parameters, candidate audio representation information for each audio frame is obtained, N audio frames are traversed, and the probability that the currently traversed audio frame is mapped to each character in the audio content can be predicted based on the candidate audio representation information of the currently traversed audio frame. The audio content is the content indicated by the audio content vector. If the maximum probability of the currently traversed audio frame being mapped to each character in the audio content is smaller than a probability threshold, the candidate audio representation information of the currently traversed audio frame is removed from the candidate audio representation information of each audio frame. After the traverse is complete, the target audio representation information of the audio data to be processed is obtained based on the remaining candidate audio representation information. In this embodiment, by traversing N audio frames of the audio data to be processed, the probability that each of the N audio frames is mapped to each character in the audio content of the audio data can be predicted based on candidate audio representation information of the N audio frames. The probability that each audio frame is mapped to each character in the audio content of the audio data can be used to reflect the likelihood that each audio frame is mapped to each character in the audio content of the audio data. That is, a higher probability indicates a higher likelihood that the character corresponding to that audio frame is one of the characters in the audio content of the audio data, and a lower probability indicates a lower likelihood that the audio frame is one of the characters in the audio content of the audio data. The character corresponding to an audio frame may be an audio frame obtained by pronouncing that character. If the probability that an audio frame is mapped to each character in the audio content of the audio data is lower than a probability threshold, it means that the audio frame is meaningless. That is, since the audio information contained in the audio frame is small, meaningless audio frames can be omitted to improve computational efficiency.
[0040] As can be seen from steps (1) to (3) above, when extracting target speech expression information, the dimension of the feature vector matrix represented by the target speech expression information is the same as the dimension of the feature vector matrix corresponding to the presented word, and the target speech expression information includes speech content vectors and paralinguistic vectors. Therefore, in this application, performing a fusion process on the target speech expression information and the feature vector matrix corresponding to the presented word actually means performing a fusion process on the speech content vector, paralinguistic vector, and presented word. By performing a feature transformation on the speech vector matrix of the audio data to be processed using feature transformation parameters and obtaining the target speech expression information of the audio data to be processed, the fusion of the subsequent speech expression information and the text information corresponding to the presented word becomes better, and the accuracy of speech recognition is improved.
[0041] Method 2: Use a pre-trained speech feature extraction model to extract target speech representation information from the audio data to be processed.
[0042] Specifically, the efficiency of feature extraction can be improved by using a pre-trained speech feature extraction model. Here, the speech feature extraction model may include, but is not limited to, automatic speech recognition (ASR) models, Transformer models based on self-attention mechanisms, convolutional reinforcement Transformer models, and connectionist temporal classification (CTC) models based on neural networks. Optionally, the speech feature extraction model may include a speech vector matrix extraction layer and a fully connected speech representation layer, the speech vector matrix extraction layer being used to extract the speech vector matrix of the audio data to be processed, and the fully connected speech representation layer being used to convert the speech vector matrix of the audio data to be processed into target speech representation information of the audio data to be processed.
[0043] For example, please refer to Figure 4. Figure 4 is a schematic diagram of the architecture of an audio feature extraction model according to an embodiment of the present invention. Here, the audio feature extraction model may include an audio vector matrix extraction layer and an audio representation fully connected layer. The audio vector matrix extraction layer is used to output an audio vector matrix of the audio data to be processed, and the audio representation fully connected layer is used to output target audio representation information of the audio data to be processed. Furthermore, optionally, the audio vector matrix extraction layer may include an encoding layer, a multi-head attention layer, and a normalization layer. Specifically, the audio data to be processed is input to the audio feature extraction model, and the audio data to be processed is processed by the audio vector matrix extraction layer of the audio feature extraction model. For example, each audio frame of the audio data to be processed can be encoded by the encoding layer of the audio vector matrix extraction layer, and encoded features can be obtained. The encoding layer obtains the positional features of each audio frame of the audio data to be processed, and the encoded features and positional features of each audio frame are concatenated to obtain concatenated encoded features of each audio frame. The concatenated encoded features are audio vectors that have positional information. The multi-head attention layer calculates the similarity between the concatenated coding features of each audio frame, determining the similarity score between pairs of audio frames, and the normalization layer normalizes the similarity score to a range of 0 to 1. For pairs of audio frames, the higher the similarity score between the two audio frames, the larger the weight between them; and the lower the similarity score, the smaller the weight between them. By combining the weights between audio frames and weighting each audio frame with other frames, the acquired audio frame includes both the audio information of that audio frame and the audio information of other audio frames; in other words, contextual audio information of the audio data being processed is introduced, so that each audio frame contains information of the entire audio data being processed. The normalization layer outputs a matrix with the same dimensions as the concatenated coding features, i.e., the audio vector matrix of the audio data being processed. Furthermore, the feature transformation parameters in the fully connected audio representation layer can be fixed in advance.Therefore, by transforming the audio vector matrix of the audio data to be processed using the feature transformation parameters in the fully connected audio representation layer, it is possible to output the target audio representation information of the audio data to be processed.
[0044] Optionally, the speech feature extraction model may further include a text output layer. By inputting the target speech representation information of each audio frame of the audio data to be processed into the text output layer, the text output layer can predict the probability that the target speech representation information of each audio frame of the audio data is mapped to each character in the audio content of the audio data, i.e., the probability that the target speech representation information of each audio frame consists of multiple characters, determine the text data corresponding to the audio data, and output text data such as "Have you eaten?".
[0045] For example, if the speech feature extraction model is an ASR model, the fully connected layer before the CTC fully connected layer can be determined as the speech representation fully connected layer. For instance, feature transfer parameters can be obtained, and the original parameters of the fully connected layer can be replaced with these feature transfer parameters. The fully connected layer with replaced parameters is called the speech representation fully connected layer. The role of the speech representation fully connected layer is to achieve information fusion between two different modalities, speech and text, by sharing parameters with the feature transfer parameters used in the transformation process of presented words, thereby aligning the latent spatial representation distributions of speech and text.
[0046] S102: The presented words related to the audio data to be processed are obtained, and the audio content vector, paralinguistic vector, and presented words are fused to obtain audio fusion features.
[0047] In the embodiments of this application, by obtaining suggested words related to the audio data to be processed, and fusing the audio content vector, paralinguistic vector, and suggested words to obtain an audio fusing feature, the audio fusing feature includes not only the suggested words of the audio data to be processed, but also the audio content and paralinguistic information. Therefore, the text information obtained by subsequently processing the audio fusing feature can reflect not only the audio content of the audio data to be processed, but also the paralinguistic information of the audio data to be processed, and the text content corresponding to the suggested words. This enables a high level of speech understanding of the audio data and improves the accuracy of speech recognition. Since the target audio representation information includes the audio content vector and paralinguistic vector, fusing the audio content vector, paralinguistic vector, and suggested words is, in practice, fusing the target audio representation information and the suggested words.
[0048] In one embodiment, suggested words related to the audio data to be processed can be obtained in the following manner. That is, based on the display interface, a plurality of pre-set suggested words are output, and in response to a selection operation involving suggested words on the display interface, suggested words related to the audio data to be processed are selected from the plurality of pre-set suggested words output to the display interface based on the selection operation. Here, the suggested words are used to reflect the speech understanding method of the audio data to be processed. Suggested words may include, but are not limited to, recognition suggested words, emotion recognition suggested words, de-colloquialization suggested words, punctuation addition suggested words, text smoothing suggested words, error correction suggested words, etc. The types of emotions may include, but are not limited to, joy, sadness, fear, anger, surprise, disgust, etc. Error correction suggested words may include specialized terminology from various fields. By selecting suggested words related to the audio data to be processed, the suggested words can be combined to output corresponding text information.
[0049] For example, if the selected presented word is a re-recognition presented word, the audio data to be processed can be re-recognized and text information of the re-recognized audio data can be output. Alternatively, if the selected presented word is an emotion recognition presented word, the output text information may include the type of emotion corresponding to the audio data to be processed, i.e., the type of emotion the speaker felt at the time of utterance, or text information corresponding to the audio data to be processed may be output simultaneously. Alternatively, if the selected presented word is a de-colloquialization presented word, the output text information may be text information of the audio data to be processed in a de-colloquialized form. Alternatively, if the selected presented word is a punctuation addition presented word, the output text information may be text information of the text content corresponding to the audio data to be processed with punctuation added. Alternatively, if the selected presented word is a text smoothing presented word, the output text information may be text information of the audio data to be processed that has undergone text smoothing, making the text information more fluent. Alternatively, if the selected presented word is an error correction presented word, the output text information may be text information of the text content corresponding to the audio data to be processed that has been corrected.
[0050] In selectable implementations, a presentation word corresponding to the audio data to be processed can be selected, and text information corresponding to the selected presentation word can be output based on the target speech expression information of the acquired audio data to be processed. For example, if the selected presentation word is a recognition presentation word, text information can be output based on the target speech expression information of the acquired audio data to be processed, and the output text information will be as fluent and consistent as possible. If the selected presentation word is an emotion recognition presentation word, the type of emotion of the speaker can be determined based on the target speech expression information of the acquired audio data to be processed, and based on the type of emotion of the speaker, the type of emotion corresponding to the audio data to be processed can be selected from multiple types of emotions (joy, sadness, fear, anger, surprise, disgust, etc.), and the selected type of emotion can be output. If the selected presentation word is a non-colloquial presentation word, text information can be obtained based on the target speech expression information of the acquired audio data to be processed, and colloquial words can be removed from the text information to make the text information as fluent and easy to read as possible, and the text information with the colloquial words removed can be output. If the selected suggested word is a word that requires punctuation, text information can be obtained based on the target speech representation information of the acquired audio data to be processed, punctuation can be added to the text information, and the punctuated text information can be output.
[0051] In the embodiments of this invention, by selecting suggested words according to the requirements corresponding to the audio data to be processed, it is possible to recognize the audio data to be processed and perform speech comprehension simultaneously, thereby improving the accuracy of speech recognition and obtaining more accurate text information.
[0052] In one embodiment, the presented word and target speech representation information can be fused using the following method. Specifically, the presented word is transformed using feature transformation parameters to obtain a feature vector matrix corresponding to the presented word, and feature concatenation is performed on the speech content vector, paralinguistic vector, and feature vector matrix corresponding to the presented word to obtain a speech fusion feature.
[0053] Here, performing feature concatenation on the audio content vector, paralinguistic vector, and feature vector matrix corresponding to the presented words is actually performing feature concatenation on the target audio representation information and the feature vector matrix corresponding to the presented words, and the speech fusion features obtained by the two concatenation methods are identical. Since the dimension of the feature vector matrix corresponding to the presented words is the same as the dimension of the feature vector matrix represented by the target audio representation information, it is possible to obtain speech fusion features by performing feature concatenation on the target audio representation information and the feature vector matrix corresponding to the presented words. Speech fusion features can reflect not only the audio content of the audio data being processed, but also the paralinguistic information of the audio data being processed, and the text content corresponding to the presented words. Therefore, the text information obtained by subsequently processing the speech fusion features can not only include the audio content of the audio data being processed, but also reflect the paralinguistic information of the audio data being processed, as well as the text content corresponding to the presented words, thereby improving the accuracy of speech recognition.
[0054] Optionally, the feature transformation parameters may also be the parameters of the word embedding layer. The word embedding layer can perform feature transformations on input text data, such as presented words, and convert the text data into a feature vector matrix. Performing feature transformations on presented words using the parameters of the word embedding layer is, in practice, equivalent to performing feature coding, that is, encoding text-dimensional data into feature vectors. Word embedding refers to the process of encoding segmented words into dense vectors, that is, mapping words into a mathematical space. For example, by pre-setting the parameters of the word embedding layer, i.e., the feature transformation parameters, and inputting presented words into the word embedding layer, it is possible to convert the presented words into a feature vector matrix using the parameters of the word embedding layer, i.e., the feature transformation parameters. Performing feature transformations on presented words using the word embedding layer facilitates feature fusion and speech understanding using feature vector matrices, thereby improving the accuracy of speech recognition.
[0055] In the embodiments of the present invention, the audio data to be processed is converted into audio representation information of the same dimension as the feature vector matrix corresponding to the presented words, thereby aligning the audio representation information with the text meaning. That is, the latent space representation of the audio in the audio data to be processed matches the latent space representation of the text corresponding to the audio data to be processed, so that the audio representation information and the text meaning information can be merged, improving the accuracy of speech recognition.
[0056] S103: Perform speech conversion processing on the speech fusion features to obtain text information corresponding to the processed speech data.
[0057] In the embodiments of this invention, the speech fusion function includes text content corresponding to the presented words, audio content of the audio data to be processed, and paralinguistic information. By performing speech conversion processing on the speech fusion function and obtaining text information corresponding to the audio data to be processed, the text information corresponding to the audio data to be processed not only includes the audio content of the audio data to be processed, but also reflects the paralinguistic information of the audio data to be processed, as well as the text content corresponding to the presented words. This enables advanced speech understanding and improves the accuracy of speech recognition.
[0058] Optionally, a pre-trained speech-to-text model can be used to perform speech-to-text conversion on speech fusion features, thereby obtaining text information corresponding to the processed speech data. Here, the speech-to-text model can include, but is not limited to, large-scale language models (LLMs), generative dialogue models (chat General Language Models, chatGLMs), open-source dialogue language models (MOSSs), and generative pre-training models (GPTs).
[0059] For example, the process of performing speech conversion on speech fusion features using a trained speech conversion model to obtain text information corresponding to the processed speech data may involve dividing the speech fusion features into multiple feature units, predicting the next feature unit based on the feature unit sequence input to the speech conversion model, adding the input feature unit and the predicted feature unit to the feature unit sequence, and continuing to predict the next feature unit until multiple feature units corresponding to the speech fusion features are predicted. Each time a feature unit is predicted, that feature unit and the feature unit preceding it are added to the feature unit sequence, and the next feature unit for that feature unit is predicted. Here, the feature unit may be a basic character unit such as a single character or a single word. Optionally, the Byte-pair encoding (BPE) method can be used to divide words into smaller units such as substrings or characters as basic units. In selectable implementations, the basic components of text can be trained as feature units based on a text corpus.
[0060] In concrete implementation, the feature vector matrix representing the target speech representation information is a multidimensional matrix, the feature vector matrix corresponding to the presented word is a multidimensional matrix, and since the dimensions of these two matrices are the same, the speech fusion feature obtained by fusion is also a multidimensional matrix, and the dimensions of these three matrices are the same. When inputting the multidimensional matrix corresponding to the speech fusion feature into a trained speech conversion model, one column of the multidimensional matrix corresponding to the speech fusion feature can be input as a feature unit. Since the first column of the matrix contains the speech corresponding to one character of the processed speech data, the feature unit of the next column can be predicted. When predicting the feature unit of the next column, the previously predicted feature unit is input into the speech conversion model as a feature unit sequence, thereby realizing the prediction of text information corresponding to the speech fusion feature.
[0061] In the embodiments of this application, since speech recognition belongs to a perceptual task, combining it with an LLM model can improve recognition capabilities for speech data, and combining modal information of speech and text can improve comprehension capabilities for speech data, thereby improving performance in a wider range of speech and semantic-related tasks. Since the LLM model can process text tasks in any format, the scope of application for speech and semantic-related tasks can be expanded based on the technical means of this application. For example, in addition to speech and text, more modalities such as visual information can be fused. For example, by converting visual information into text representation information, and then combining the text information and speech representation information and inputting it into the LLM model for processing, the speech comprehension content is enriched and the accuracy of speech recognition is improved.
[0062] In the embodiments of this invention, feature extraction is performed on the audio data to be processed to obtain target audio representation information of the audio data to be processed, suggested words related to the audio data to be processed are obtained, the target audio representation information and suggested words are fused to obtain audio fusion features, and audio conversion processing is performed on the audio fusion features to obtain text information corresponding to the audio data to be processed. The target audio representation information includes an audio content vector and a paralanguage vector corresponding to the audio data to be processed, and the paralanguage vector is used to assist in the recognition of text information corresponding to the audio data to be processed. Therefore, when performing speech recognition on the audio data to be processed, speech recognition can be performed by combining information about the audio content of the audio data to be processed, information about the paralanguage of the audio data to be processed, and text content corresponding to suggested words. Because speech recognition processing is performed using more comprehensive and richer audio representation information, this invention enables advanced speech recognition and understanding of the audio data to be processed, and improves the accuracy of speech recognition.
[0063] Furthermore, please refer to Figure 5. Figure 5 is a flowchart of the training method for a speech feature extraction model according to an embodiment of the present invention. This method can be used with computer equipment and includes, but is not limited to, the following steps, as shown in Figure 5.
[0064] S201: Acquire sample audio data, perform feature extraction on the sample audio data using an audio feature extraction model, and obtain sample audio representation information of the sample audio data.
[0065] In the embodiments of this invention, the sample audio data may be obtained in advance, for example, by downloading from an audio data storage website, uploading from a terminal device, or obtaining from locally stored audio data. To increase the number of training data, the sample audio data can be further processed by cutting, rotating, tuning, adding noise, etc. The accuracy of the audio feature extraction model can be improved by using a large amount of sample audio data as training data for the audio feature extraction model.
[0066] In embodiments of the present invention, for example, a speech feature extraction model can be used to perform feature encoding on sample speech data, obtain a speech vector matrix of the processed speech data, perform a feature transformation on the speech vector matrix of the processed speech data using the feature transformation parameters of the speech feature extraction model, and obtain sample speech representation information of the sample speech data. The sample speech representation information of the sample speech data may include a sample speech content vector and a sample paralinguistic vector corresponding to the sample speech data.
[0067] In one embodiment, the speech feature extraction model may include a speech vector matrix extraction layer and a fully connected speech representation layer, and the sample speech representation information of the sample speech data can be determined by combining the speech vector matrix extraction layer and the fully connected speech representation layer. For example, sample speech data is input to the speech vector matrix extraction layer, the speech vector matrix extraction layer performs feature encoding on the sample speech data, obtains the speech vector matrix of the sample speech data, and inputs the speech vector matrix of the sample speech data to the fully connected speech representation layer. A feature transformation is performed on the speech vector matrix of the sample speech data using the feature transformation parameters of the fully connected speech representation layer to obtain the sample speech representation information of the sample speech data.
[0068] Furthermore, optionally, the speech vector matrix extraction layer may also include an encoding layer, a multi-head attention layer, and a normalization layer. Specifically, sample speech data is input to the speech feature extraction model, the encoding layer encodes each speech frame of the sample speech data to obtain encoded features, the position features of each speech frame of the sample speech data are obtained, and the encoded features and position features of each speech frame are concatenated to obtain a concatenated encoded feature of each speech frame. The concatenated encoded feature is a speech vector that has position information. The multi-head attention layer calculates the similarity between the concatenated encoded features of each speech frame of the sample speech data to determine the similarity score between pairs of speech frames, and the normalization layer normalizes the similarity score to a range of 0 to 1. For pairs of speech frames of the sample speech data, the higher the similarity score between the two speech frames, the larger the weight between the two speech frames, and the lower the similarity score between the two speech frames, the smaller the weight between the two speech frames. By combining the weights of each audio frame in the sample audio data with those of other audio frames and weighting them together, the acquired audio frame includes both the audio information of that audio frame and the audio information of other audio frames in the sample audio data. Contextual audio information from the sample audio data is also introduced, so that each audio frame in the sample audio data contains information from the entire sample audio data.
[0069] S202: Retrieve the sample speech representation label for the sample speech data.
[0070] Here, the sample speech representation label may be a pre-obtained speech representation label that reflects the true value of the sample speech data. By obtaining the sample speech representation label of the sample speech data, the speech feature extraction model can be adjusted by combining the sample speech representation label with the sample speech representation information output from the speech feature extraction model when training the speech feature extraction model later.
[0071] S203: Train a speech feature extraction model based on sample speech representation labels and sample speech representation information, and obtain the trained speech feature extraction model.
[0072] Here, the sample speech representation information is the model output value of the speech feature extraction model, and the sample speech representation label is the sample true value. The purpose of training the speech feature extraction model is to make the model output value and the sample true value match as closely as possible. If the model output value and the sample true value do not match, the model parameters of the speech feature extraction model are further adjusted until they match. If the model output value and the sample true value match, the speech feature extraction model at this point is considered the trained speech feature extraction model.
[0073] Training a speech feature extraction model involves comparing the difference between a sample speech representation label and the sample speech representation information, and determining the loss function of the speech feature extraction model based on this difference. Here, the difference between the sample speech representation label and the sample speech representation information can be calculated using a similarity calculation method. That is, the greater the similarity between the sample speech representation label and the sample speech representation information, the smaller the difference between them. The smaller the similarity between the sample speech representation label and the sample speech representation information, the larger the difference between them. If the difference between the sample speech representation label and the sample speech representation information is greater than the difference threshold, the loss function of the speech feature extraction model becomes greater than the first loss threshold, and the model parameters of the speech feature extraction model are further adjusted to reduce the loss function of the speech feature extraction model. If the difference between the sample speech representation label and the sample speech representation information is less than or equal to the difference threshold, the loss function of the speech feature extraction model is less than or equal to the first loss threshold, and the speech feature extraction model at this point can be saved as a trained speech feature extraction model.
[0074] Optionally, if the number of training iterations for the speech feature extraction model exceeds the iteration threshold, or if the speech feature extraction model reaches the convergence condition, the adjustment of the model parameters of the speech feature extraction model can be stopped, and the trained speech feature extraction model can be obtained.
[0075] In the selectable implementations, the speech feature extraction model can also be trained in the following way: the speech feature extraction model is used to extract features from sample audio data, sample audio representation information is obtained from the sample audio data, the speech feature extraction model is used to predict sample text data corresponding to the sample audio representation information of the sample audio data, sample text labels are obtained from the sample audio data, the speech feature extraction model is trained based on the sample text labels and sample text data, and the trained speech feature extraction model is obtained.
[0076] Predicting the text data of sample speech representation information from sample speech data is equivalent to converting the sample speech representation information into text modality information and training a speech feature extraction model based on the difference between the two texts. The sample text label may be the actual text of the sample speech data, and the text data of the sample speech representation information may be the text predicted by the speech feature extraction model, i.e., the text output by the model. By comparing the difference between the sample text label and the text data of the sample speech representation information, the speech feature extraction model is trained based on the difference between the sample text label and the text data of the sample speech representation information. The difference between the sample text label and the text data of the sample speech representation information can be calculated by a text similarity calculation method, but the embodiments of this application are not limited to this. By converting the sample speech representation information into text modality information and comparing them, a text comparison can be performed to determine the difference between the texts and adjust the speech feature extraction model.
[0077] In one embodiment, if the speech feature extraction model includes a speech vector matrix extraction layer and a fully connected speech representation layer, the speech feature extraction model can be trained as follows. The parameters of the speech vector matrix extraction layer are adjusted based on the sample speech representation labels and sample speech representation information to obtain a trained speech feature extraction model.
[0078] In the embodiment of this invention, when training the speech feature extraction model, the parameters of the fully connected layer of speech representation are fixed; that is, the parameters of the fully connected layer of speech representation are fixed as the feature transformation parameters of the word embedding layer. By fixing the parameters of the fully connected layer of speech representation, speech recognition can be performed using the speech feature extraction model, and since the latent spatial representation matches the speech transformation model, the target speech representation information of the processed speech data output from the speech feature extraction model can be input into the speech transformation model.
[0079] In the selectable implementations, meaningless frames from the sample audio data can be individually removed when training the speech feature extraction model, thereby reducing computational complexity and improving the training efficiency of the speech feature extraction model.
[0080] In the embodiments of this invention, the computational complexity can be reduced by removing meaningless frames from the audio data to be processed and obtaining target speech representation information from the audio data to be processed. By sharing the parameter weights of the fully connected speech representation layer and the word embedding layer of the speech conversion model, the speech representation information output from the fully connected speech representation layer can be matched with the encoded features of the presented words output from the word embedding layer, thereby realizing information fusion between the two modalities. The parameters of the fully connected speech representation layer are derived from the word embedding layer of the speech conversion model.
[0081] In the embodiments of this invention, by training a speech feature extraction model, feature extraction can be performed on the audio data to be processed using the trained speech feature extraction model, target speech representation information of the audio data to be processed can be obtained, and the processing efficiency of the audio data can be improved. Since the speech feature extraction model is trained using a large amount of sample audio data, the accuracy of the speech feature extraction model can be improved.
[0082] Please refer to Figure 6. Figure 6 is a flowchart of a training method for a speech conversion model according to an embodiment of the present invention, which can be used with computer equipment and includes, but is not limited to, the following steps as shown in Figure 6.
[0083] S301: Obtain sample audio expression information and sample suggested words corresponding to the sample audio data.
[0084] In the embodiments of this application, sample speech representation information can be obtained by performing feature extraction on sample speech data. For example, sample speech representation information may be obtained by dividing the sample speech data to obtain N (where N is a positive integer) speech frames, performing feature encoding on the N speech frames of the sample speech data to obtain a speech vector matrix of the N speech frames, performing a feature transformation on the speech vector matrix of the N speech frames using the feature transformation parameters of the speech feature extraction model to obtain candidate speech representation information for the N speech frames, and then removing meaningless speech frames from the candidate speech representation information of the N speech frames of the sample speech data. In the embodiments of this application, multiple sample presentation words corresponding to different scenarios may be included. When creating training data, corresponding sample presentation words can be selected according to actual needs. For example, sample presentation words may include words such as text smoothing, decolluding, error correction, re-recognition, sentiment recognition, and punctuation addition. For example, error correction presentation words may include specialized terminology from various fields. Decolluding presentation words may include several colloquial words. Text smoothing presentation words may include words composed of repeating characters.
[0085] To make it easier to understand, the presented words do not have a fixed form. For example, in scenarios such as de-colloquialization and emotion recognition, the presented words are letters, in the punctuation addition scenario, the presented words are punctuation marks, and in the recognition scenario, the presented words are information indicating recognition. Therefore, in actual use scenarios, it is sufficient to match the selected presented words during model training.
[0086] S302: Using a speech conversion model, sample speech content vectors, sample paralinguistic vectors, and sample presented words are fused to obtain sample speech fusion features.
[0087] Here, the sample speech representation information of the sample speech data includes the sample speech content vector and sample paralinguistic vector corresponding to the sample speech data. Therefore, merging the sample speech content vector, sample paralinguistic vector, and sample presented words is, in practice, merging the sample speech representation information and the sample presented words. Since the speech representation information is represented by a feature vector matrix and the presented words are represented by characters, converting the presented words into a feature vector matrix before merging the two makes feature merging, such as feature concatenation, easier.
[0088] S303: Using a speech conversion model, perform speech conversion processing on sample speech fusion features to obtain text information corresponding to the sample speech data.
[0089] For example, the process of using a speech conversion model to perform speech conversion on sample speech fusion features and obtain text information corresponding to the sample speech data may involve dividing the sample speech fusion features into multiple feature units, predicting the next feature unit based on the feature unit sequence input to the speech conversion model, adding the input feature unit and the predicted feature unit to the feature unit sequence, and continuing to predict the next feature unit until multiple feature units corresponding to the sample speech fusion features are predicted.
[0090] In the selectable implementation, the feature vector matrix representing the sample speech representation information is a multidimensional matrix, the feature vector matrix corresponding to the sample presented words is a multidimensional matrix, and since these two matrices have the same dimensions, the sample speech fusion feature obtained by fusion is also a multidimensional matrix, and the dimensions of these three matrices are the same. When inputting the multidimensional matrix corresponding to the sample speech fusion feature into a trained speech conversion model, one column of the multidimensional matrix corresponding to the sample speech fusion feature can be input as a feature unit. Since one column of the matrix can represent the speech data corresponding to one character of the sample speech data, the feature unit of the next column can be predicted. When predicting the feature unit of the next column, the previously predicted feature unit is input into the speech conversion model as a feature unit sequence, thereby realizing the prediction of text information corresponding to the sample speech fusion feature.
[0091] Optionally, the speech conversion model can use currently open-source large-scale language models, such as models based on the widely used Transformer structure. Autoregression is used to predict the next token based on the input token (feature unit) sequence, and then predict the next token based on the input token and the predicted token, thereby achieving the prediction of text information corresponding to the processed speech data.
[0092] S304: Obtain sample text labels corresponding to sample audio data, train a speech conversion model based on the sample text labels and text information corresponding to the sample audio data, and obtain the trained speech conversion model.
[0093] In the embodiments of this application, the sample text label may be the actual text label of the sample audio data, and the text information corresponding to the sample audio data may be a model output value produced based on the speech conversion model. The purpose of training the speech conversion model is to make the sample text label and the text information corresponding to the sample audio data match as closely as possible. When the sample text label and the text information corresponding to the sample audio data match, the speech conversion model at that time can be determined as a trained speech conversion model. The sample text label and the text information corresponding to the sample audio data can be calculated by a text similarity calculation method.
[0094] Here, training a speech-to-speech model based on the text information corresponding to the sample text labels and sample audio data means determining the loss function of the speech-to-speech model based on the difference between the sample text labels and the text information corresponding to the sample audio data. If the loss function of the speech-to-speech model is greater than the second loss threshold, the model parameters of the speech-to-speech model are further adjusted to reduce the loss function of the speech-to-speech model. If the loss function of the speech-to-speech model is less than or equal to the second loss threshold, the speech-to-speech model at this point is determined to be the trained speech-to-speech model.
[0095] The process of training a speech-to-speech model is, in essence, the process of adjusting the parameters of the speech-to-speech model. Since a speech-to-speech model contains a large number of parameters, adjusting all of them during training would be time-consuming and inefficient. Therefore, by adjusting only some of the parameters of the speech-to-speech model, the goal of improving the model's efficiency can be achieved.
[0096] Please refer to Figure 7. Figure 7 is a schematic diagram of the parameter tuning of the speech-to-translate model according to an embodiment of the present invention. The left side of Figure 7 shows the pre-trained model parameters W (i.e., pre-trained weights) of the speech-to-translate model (such as an LLM model), with a branch added next to the pre-trained model structure. This branch contains two structures, A and B, and the two parameters A and B are initialized to a Gaussian distribution and 0, respectively. At the start of training, the added parameters are 0, and the input dimension of A and the output dimension of B are the same as the input and output dimensions of the original model, respectively, but the output dimension of A and the input dimension of B are much smaller than the input and output dimensions of the original model. This significantly reduces the number of parameters to be trained in the LLM model. When training the LLM model, only the parameters of A and B are updated, and the pre-trained model parameters W remain unchanged. By integrating A and B with the original model parameter matrix W, no additional computation is required during inference, and for different downstream tasks, only A and B need to be retrained based on the pre-trained model. After training with new parameters, integrating them with the old parameters and reparameterizing allows for fine-tuning on new tasks without increasing model inference time, thereby improving the training efficiency of the model. Since the LLM model itself has numerous parameters, only a few training parameters need to be added during fine-tuning, thus improving training efficiency.
[0097] When training a speech conversion model, the key to predicting the output data from the input data lies in preparing the training set. For a specific task (such as text smoothing, de-colloquialization, or error correction), a training set like the following can be prepared. After removing meaningless frames from sample audio data using the above speech feature conversion model, the speech representation information of the sample audio data is obtained and input into the speech conversion model along with the presented words corresponding to the sample audio data, thereby outputting text information for the corresponding task.
[0098] Example 1 is a text smoothing scenario. Input: "Did you eat?" (The input is the audio representation information of the speech modality), and the corresponding presentation word (for example, a repeated word such as "you, you, me, me"). Output: "Did you eat it?"
[0099] Example 2 is a non-colloquial scenario. Input: "Yes, yes, yes, that's right, yes" (the input is speech modality information), and the corresponding presented word (for example, colloquial words such as "Yes, yes yes yes"). Output: "Yes, that's right."
[0100] Example 3 is an error correction scenario. Input: "The field effect of this system is not very good" (the input is the voice representation information of the voice modality), and the corresponding presented word (e.g., a technical term in the corresponding scenario such as "far-field"). Output: "The long-range field effect of this system is not very good."
[0101] Example 4 is a scenario involving the addition of punctuation marks. Input: "This sentence is long, but are punctuation marks necessary? I think they are." (Input is phonetic representation information for the phonetic modality), and corresponding suggested words (e.g., punctuation marks such as "comma, question mark, period"). Output: "This sentence is long, but are punctuation marks necessary? I think they are."
[0102] In the embodiments of this application, a speech feature extraction model, such as an ASR model, reuses the word embedding mechanism of the LLM model and aligns the speech representation information with the text semantic space of the LLM model, thereby allowing the speech representation information to be used directly as input to the LLM model. Since the speech representation information includes text content and paralinguistic information, the LLM model can fully utilize information from the speech modality and further improve its speech recognition and comprehension capabilities. Because the LLM model can fully utilize information other than text content from the speech data, it can improve its speech recognition and comprehension capabilities. By setting different presented words during speech recognition, speech recognition capabilities such as de-colloquialization, hot word substitution, sentiment recognition, and punctuation addition can be directly improved at the speech recognition stage, forming an end-to-end model and outputting corresponding text information. This eliminates the need to input the text into the LLM model for processing after it has been recognized as text at the speech recognition stage. The technical means of this application can be extended to include information from other modalities, such as visual information, as input to the LLM model, thereby further improving the speech recognition and comprehension capabilities of the model.
[0103] In the embodiments of this invention, by training a speech conversion model, speech conversion processing can be performed on speech fusion features using the trained speech conversion model, text information corresponding to the processed speech data can be obtained, and the processing efficiency of the speech data can be improved. Since the speech conversion model is trained using a large amount of sample speech data, the accuracy of the speech conversion model can be improved.
[0104] The above describes the method of the embodiment of the present application; below, the apparatus of the embodiment of the present application will be described.
[0105] Please refer to Figure 8. Figure 8 is a schematic diagram of the configuration of an audio processing device according to an embodiment of the present application. This audio processing device can be placed in a computer device and used to perform the corresponding steps of the audio processing method according to an embodiment of the present application. This audio processing device 80 is A feature extraction unit 801 for extracting features from audio data to be processed and obtaining target audio representation information of the audio data to be processed, wherein the target audio representation information includes an audio content vector and a paralinguistic vector corresponding to the audio data to be processed, and the paralinguistic vector is used to assist in the recognition of text information corresponding to the audio data to be processed, and the feature extraction unit 801, An information fusion unit 802 for obtaining suggested words related to the audio data to be processed, and for fusing the audio content vector, paralinguistic vector, and suggested words to obtain audio fusion features, The system includes a speech conversion unit 803 for performing speech conversion processing on speech fusion features and obtaining text information corresponding to the processed speech data.
[0106] Optionally, the information fusion unit 802 specifically... The process involves performing feature transformation on the presented words using feature transformation parameters and obtaining a feature vector matrix corresponding to the presented words. It is used to obtain speech fusion features by performing feature concatenation on the speech content vector, paralinguistic vector, and feature vector matrix corresponding to the presented words.
[0107] Optionally, the feature extraction unit 801 specifically: Obtaining feature transformation parameters for performing feature transformation on the presented words, This involves performing feature encoding on the audio data to be processed and obtaining the audio vector matrix of the audio data to be processed, This method involves performing a feature transformation on the audio vector matrix of the audio data to be processed using feature transformation parameters to obtain target audio representation information of the audio data to be processed, wherein the dimension of the feature vector matrix represented by the target audio representation information is the same as the dimension of the feature vector matrix corresponding to the presented word.
[0108] Optionally, the feature extraction unit 801 specifically: The process involves dividing the audio data to be processed and obtaining N audio frames, Feature encoding is performed on each audio frame to obtain the audio vector matrix for each audio frame, This involves performing a feature transformation on the audio vector matrix of each audio frame using feature transformation parameters to obtain candidate audio representation information for each audio frame, and The process involves traversing N audio frames and predicting the probability that the currently traversed audio frame is mapped to each character in the audio content, based on the candidate audio representation information of the currently traversed audio frame, wherein the audio content is the content indicated by the audio content vector. If the maximum probability that the currently traversed audio frame is mapped to each character in the audio content is less than the probability threshold, the candidate audio representation information for the currently traversed audio frame is removed from the candidate audio representation information for each audio frame. After the traverse is complete, it is used to obtain the target speech representation information of the audio data to be processed based on the remaining candidate speech representation information.
[0109] Optionally, the feature extraction unit 801 further, specifically, The method involves determining the positional features of each audio frame based on the division order of N audio frames, wherein the positional features are used to indicate the position of the corresponding audio frame in the audio data being processed. The feature extraction unit 801 specifically, The process involves performing feature encoding on each audio frame and obtaining the encoded features of N audio frames. This method is used to obtain the audio vector matrix of audio frame i by performing a feature concatenation operation on the positional features and encoded features of audio frame i among N audio frames, where i is a positive integer and 1 ≤ i ≤ N.
[0110] Optionally, the text information corresponding to the audio data to be processed is obtained by a trained speech conversion model, and the speech processing device 80 further comprises a first training unit 804, the first training unit 804 is The process involves obtaining sample speech expression information and sample presented words corresponding to sample speech data, wherein the sample speech expression information includes a sample speech content vector and a sample paralinguistic vector corresponding to the sample speech data. Using a speech conversion model, sample speech content vectors, sample paralinguistic vectors, and sample presented words are fused to obtain sample speech fusion features, Using a speech conversion model, perform speech conversion processing on sample speech fusion features to obtain text information corresponding to the sample speech data, It is used to obtain sample text labels corresponding to sample audio data, train a speech conversion model based on the sample text labels and text information corresponding to the sample audio data, and obtain the trained speech conversion model.
[0111] Optionally, the target speech representation information of the speech data to be processed is obtained by a trained speech feature extraction model, and the speech processing device 80 further comprises a second training unit 805, the second training unit 805 is This involves acquiring sample audio data, performing feature extraction on the sample audio data using an audio feature extraction model, and obtaining sample audio representation information from the sample audio data. It is used to obtain sample speech representation labels from sample speech data, train a speech feature extraction model based on the sample speech representation labels and sample speech representation information, and obtain the trained speech feature extraction model.
[0112] Optionally, the speech feature extraction model includes a speech vector matrix extraction layer and a fully connected speech representation layer, and the second training unit 805 specifically includes, The process involves using an audio vector matrix extraction layer to perform feature encoding on sample audio data and obtain the audio vector matrix of the sample audio data, and The process involves performing a feature transformation on the audio vector matrix of the sample audio data using the feature transformation parameters in the fully connected layer of the audio representation, thereby obtaining the sample audio representation information of the sample audio data. It is used to adjust the parameters of the speech vector matrix extraction layer based on sample speech representation labels and sample speech representation information, and to obtain a trained speech feature extraction model.
[0113] Regarding the details not mentioned in the embodiment corresponding to Figure 8, please refer to the description of the embodiment of the method, and therefore, the explanation will be omitted here.
[0114] In the embodiments of this invention, feature extraction is performed on the audio data to be processed to obtain target audio representation information of the audio data to be processed, suggested words related to the audio data to be processed are obtained, the target audio representation information and suggested words are fused to obtain audio fusion features, and audio conversion processing is performed on the audio fusion features to obtain text information corresponding to the audio data to be processed. The target audio representation information includes an audio content vector and a paralanguage vector corresponding to the audio data to be processed, and the paralanguage vector is used to assist in the recognition of text information corresponding to the audio data to be processed. Therefore, when performing speech recognition on the audio data to be processed, speech recognition can be performed by combining information about the audio content of the audio data to be processed, information about the paralanguage of the audio data to be processed, and text content corresponding to suggested words. Because speech recognition processing is performed using more comprehensive and richer audio representation information, this invention enables advanced speech recognition and understanding of the audio data to be processed, and improves the accuracy of speech recognition.
[0115] Please refer to Figure 9. Figure 9 is a schematic diagram of the configuration of a computer device according to an embodiment of the present invention. As shown in Figure 9, the computer device 90 may include a processor 901, memory 902, and a network interface 903. The processor 901 is connected to the memory 902 and the network interface 903, and for example, the processor 901 can be connected to the memory 902 and the network interface 903 via a bus. The computer device may be a terminal device or a server.
[0116] The processor 901 is arranged to support the voice processing unit in performing the corresponding functions of the voice processing method described above. The processor 901 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0117] Memory 902 is used to store program instructions, data, and the like. Memory 902 may include volatile memory (VM) such as random access memory (RAM), non-volatile memory (NVM) such as read-only memory (ROM), flash memory, hard disk drive (HDD), and solid-state drive (SSD), and may also include combinations of the above types of memory.
[0118] The network interface 903 is used to provide network communication functionality.
[0119] The processor 901 can call program code to perform the following operations: Feature extraction is performed on the audio data to be processed to obtain target audio representation information for the audio data. The target audio representation information includes an audio content vector and a paralinguistic vector corresponding to the audio data to be processed. The paralinguistic vector is used to assist in the recognition of text information corresponding to the audio data to be processed. The presented words related to the audio data to be processed are obtained, and the audio content vector, paralinguistic vector, and presented words are fused to obtain audio fusion features. Speech conversion processing is performed on the speech fusion features to obtain text information corresponding to the processed speech data.
[0120] Furthermore, the computer device 90 described in the embodiments of this application may be configured according to the methods described in the embodiments corresponding to Figures 3, 5, and 6, or according to the methods described in the embodiment corresponding to Figure 8, and such descriptions are omitted here. The advantageous effects of similar methods are also omitted here.
[0121] In the embodiments of this invention, feature extraction is performed on the audio data to be processed to obtain target audio representation information of the audio data to be processed, suggested words related to the audio data to be processed are obtained, the target audio representation information and suggested words are fused to obtain audio fusion features, and audio conversion processing is performed on the audio fusion features to obtain text information corresponding to the audio data to be processed. The target audio representation information includes an audio content vector and a paralanguage vector corresponding to the audio data to be processed, and the paralanguage vector is used to assist in the recognition of text information corresponding to the audio data to be processed. Therefore, when performing speech recognition on the audio data to be processed, speech recognition can be performed by combining information about the audio content of the audio data to be processed, information about the paralanguage of the audio data to be processed, and text content corresponding to suggested words. Because speech recognition processing is performed using more comprehensive and richer audio representation information, this invention enables advanced speech recognition and understanding of the audio data to be processed, and improves the accuracy of speech recognition.
[0122] Optionally, other steps of the method in the above embodiment can be implemented when program instructions are executed by the processor, but these will not be described here.
[0123] Embodiments of the present invention further provide a computer-readable storage medium in which a computer program is stored. The computer program, when executed by a computer, includes program instructions that cause the computer to perform the method of the above embodiment. The computer may be part of the above computer equipment. For example, the program instructions may be located and executed on one computer device, on multiple computer devices located in the same place, or on multiple computer devices distributed in multiple locations and interconnected via a communication network. Multiple computer devices distributed in multiple locations and interconnected via a communication network can constitute a blockchain network.
[0124] Embodiments of the present application further provide a computer program product or computer program including computer instructions. When executed by a processor, these computer instructions can implement some or all of the steps of the above method. For example, the computer instructions are stored in a computer-readable storage medium, and the processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, thereby causing the computer device to perform the steps performed in each embodiment of the above method.
[0125] As those skilled in the art will understand, all or part of the processes in the above embodiments may be implemented by instructing the relevant hardware with a computer program. The program may be stored on a computer-readable storage medium. When the program is executed, it may include the processes of each embodiment of the above methods. The storage medium may be a magnetic disk, an optical disk, read-only memory (ROM), random access memory (RAM), etc.
[0126] The above disclosures are merely preferred embodiments of the present application and do not limit the scope of the claims. Accordingly, equivalent modifications made based on the claims of the present application remain within the scope of protection.
Claims
1. A method of audio processing, A step of performing feature extraction on audio data to be processed and obtaining target audio representation information for the audio data to be processed, wherein the target audio representation information includes an audio content vector and a paralinguistic vector corresponding to the audio data to be processed, and the paralinguistic vector is used to assist in the recognition of text information corresponding to the audio data to be processed. The steps include obtaining suggested words related to the audio data to be processed, and fusing the audio content vector, the paralinguistic vector, and the suggested words to obtain audio fusion features, The steps include: performing speech conversion processing on the speech fusion features and obtaining text information corresponding to the processed speech data; A method for processing sound, characterized by including [a certain element].
2. The step of fusing the aforementioned audio content vector, the aforementioned paralinguistic vector, and the aforementioned presented words to obtain an audio fusion feature is: The steps include: performing feature transformation on the presented words using feature transformation parameters and obtaining a feature vector matrix corresponding to the presented words; The steps include: performing feature concatenation on the audio content vector, the paralinguistic vector, and the feature vector matrix corresponding to the presented words to obtain the audio fusion features, The method according to claim 1, characterized in that
3. The step of extracting features from the audio data to be processed and obtaining target audio representation information of the audio data to be processed is: The steps include obtaining feature transformation parameters for performing feature transformation on the aforementioned presented words, The steps include: performing feature encoding on the audio data to be processed and obtaining the audio vector matrix of the audio data to be processed; A step of performing a feature transformation on the audio vector matrix of the audio data to be processed using the feature transformation parameters, and obtaining target audio representation information of the audio data to be processed, wherein the dimension of the feature vector matrix represented by the target audio representation information is the same as the dimension of the feature vector matrix corresponding to the presented word, The method according to claim 1 or 2, characterized in that
4. The step of performing feature encoding on the audio data to be processed and obtaining the audio vector matrix of the audio data to be processed is: A step of dividing the audio data to be processed and obtaining N audio frames, where N is a positive integer, and a step of The process includes the steps of: performing feature encoding on N audio frames and obtaining an audio vector matrix of the N audio frames; The step of performing a feature transformation on the audio vector matrix of the audio data to be processed using feature transformation parameters and obtaining target audio representation information of the audio data to be processed is: The steps include: performing a feature transformation on the audio vector matrix of the N audio frames using the feature transformation parameters and obtaining candidate audio representation information for the N audio frames; A step of traversing the N audio frames and predicting the probability that the currently traversed audio frame is mapped to each character in the audio content, based on the candidate audio representation information of the currently traversed audio frame, wherein the audio content is the content indicated by the audio content vector. If the maximum probability among the probabilities that the currently traversing audio frame is mapped to each character in the audio content is less than the probability threshold, the candidate audio representation information of the currently traversing audio frame is removed from the candidate audio representation information of each audio frame. The process includes the step of obtaining target speech representation information for the audio data to be processed based on the remaining candidate speech representation information after the traverse is complete. The method according to any one of claims 1 to 3, characterized in that
5. The aforementioned method, A step of determining the positional features of each audio frame based on the division order of N audio frames, the step of using the positional features to indicate the position of the corresponding audio frame in the audio data to be processed, further comprising: The step of performing feature encoding on each audio frame and obtaining the audio vector matrix for each audio frame is: The steps include: performing feature encoding on each audio frame and obtaining the encoded features of each audio frame; The process includes the step of performing a feature concatenation operation on the positional features and encoded features of audio frame i among N audio frames, thereby obtaining an audio vector matrix of audio frame i, wherein i is a positive integer and 1 ≤ i ≤ N. The method according to any one of claims 1 to 4, characterized in that
6. The text information corresponding to the audio data to be processed is obtained by a trained speech conversion model, and the training method for the trained speech conversion model is: A step of obtaining sample speech expression information and sample presentation words corresponding to sample speech data, wherein the sample speech expression information includes a sample speech content vector and a sample paralinguistic vector corresponding to the sample speech data. The steps include using a speech conversion model to fuse the sample speech content vector, the sample paralinguistic vector, and the sample presented words to obtain sample speech fusion features, The steps include: performing speech conversion processing on the sample speech fusion features using the speech conversion model and obtaining text information corresponding to the sample speech data; The process includes the steps of obtaining sample text labels corresponding to the sample audio data, training the speech conversion model based on the sample text labels and the text information corresponding to the sample audio data, and obtaining the trained speech conversion model. The method according to any one of claims 1 to 5, characterized in that
7. The target speech representation information of the audio data to be processed is obtained by a trained speech feature extraction model, and the training method for the trained speech feature extraction model is: The steps include: acquiring sample audio data, performing feature extraction on the sample audio data using an audio feature extraction model, and obtaining sample audio representation information of the sample audio data; The process includes the steps of obtaining sample speech representation labels for the sample speech data, training the speech feature extraction model based on the sample speech representation labels and the sample speech representation information, and obtaining the trained speech feature extraction model. The method according to any one of claims 1 to 6, characterized in that
8. The speech feature extraction model includes a speech vector matrix extraction layer and a fully connected speech representation layer. The step of performing feature extraction on sample audio data using the aforementioned audio feature extraction model and obtaining sample audio representation information of the sample audio data is: The steps include: performing feature encoding on the sample audio data using the audio vector matrix extraction layer to obtain the audio vector matrix of the sample audio data; The step includes performing a feature transformation on the audio vector matrix of the sample audio data using the feature transformation parameters in the fully connected audio representation layer, and obtaining sample audio representation information of the sample audio data. The steps of training the speech feature extraction model based on the sample speech expression labels and the sample speech expression information, and obtaining the trained speech feature extraction model, are: The steps include adjusting the parameters of the speech vector matrix extraction layer based on the sample speech representation labels and the sample speech representation information, and obtaining the trained speech feature extraction model. The method according to any one of claims 1 to 7, characterized in that
9. A sound processing device, A feature extraction unit for extracting features from audio data to be processed and obtaining target audio representation information of the audio data to be processed, wherein the target audio representation information includes an audio content vector and a paralanguage vector corresponding to the audio data to be processed, and the paralanguage vector is used to assist in the recognition of text information corresponding to the audio data to be processed, and An information fusion unit for obtaining suggested words related to the audio data to be processed, and for fusing the audio content vector, the paralinguistic vector, and the suggested words to obtain audio fusion features, A speech conversion unit for performing speech conversion processing on the aforementioned speech fusion features and obtaining text information corresponding to the processed speech data, A voice processing device characterized by comprising:
10. A computer device comprising a processor, memory, and a network interface, wherein the processor is connected to the memory and the network interface, the network interface is used to provide data communication functions, the memory is used to store computer programs including program instructions, and the processor is configured to call program instructions to cause the computer device to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and the computer program is suitable for causing a computer device equipped with the processor to perform the method described in any one of claims 1 to 8, when loaded and executed by a processor.
12. A computer program product, wherein the computer program product includes a computer program, and the computer program, when executed by a processor, performs the method described in any one of claims 1 to 8.