A voice translation method, device, storage medium and equipment

By aligning the output features of the dual decoders, acoustic feature vectors are extracted using edit distance and multi-head self-attention mechanisms. This solves the problems of low accuracy and error accumulation in dialect translation in existing speech translation methods, achieving higher translation accuracy and consistency.

CN120108399BActive Publication Date: 2025-11-11IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510304971.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-11-11
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Existing speech translation methods suffer from problems such as error propagation, over-translation, and named entity translation bias when dealing with dialects, resulting in low translation accuracy, especially with serious error accumulation in multilingual and multi-dialect scenarios.

Method used

By aligning the output features of the dual decoders, entity word features are aligned using the posterior probabilities of the recognition decoder and the translation decoder. Acoustic feature vectors are extracted by combining the edit distance algorithm and the multi-head self-attention mechanism, resulting in more accurate translations.

Benefits of technology

It improves the accuracy and consistency of voice translation, reduces over-translation and named entity translation bias, and enhances the user's translation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108399B_ABST
    Figure CN120108399B_ABST
Patent Text Reader

Abstract

This application discloses a speech translation method, apparatus, storage medium, and device. The method includes: firstly, extracting the acoustic feature vector of the target speech; then, inputting the target speech into a recognition decoder and a translation decoder respectively to obtain recognized text and translated text; and determining entity word feature vectors based on an edit distance algorithm, according to the recognized text, translated text, and acoustic feature vectors; next, inputting the entity word feature vectors into the recognition decoder and translation decoder respectively to obtain a first posterior probability and a second posterior probability; and then, aligning the entity word features using the first and second posterior probabilities to determine the final translation result corresponding to the target speech. It is evident that, because this application enhances the consistency and transcription accuracy between the two decoders by aligning the output features of the two decoders when translating the target speech, it improves the accuracy of the translation result corresponding to the target speech and enhances the user's translation experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a speech translation method, apparatus, storage medium and device. Background Technology

[0002] With the continuous breakthroughs in artificial intelligence technology and the increasing prevalence of various smart terminal devices, human-computer interaction is occurring more and more frequently in people's daily work and life. Voice, as one of the most convenient and efficient interaction methods, brings great convenience to people's lives. Among the most important aspects is the technology for accurate speech translation. For example, the translation of dialects is not only significant for language preservation and cultural inheritance, but also plays a crucial role in promoting social integration, economic development, and cultural exchange. Furthermore, in today's era of accelerated globalization and technological breakthroughs, speech translation has become a key tool for driving social, economic, cultural, and technological development.

[0003] Currently, commonly used speech translation methods fall into two categories: one is a cascaded decoding method that independently generates two transcriptions. However, because this method separates Automatic Speech Recognition (ASR) and Machine Translation (MT) into two independent steps, it often suffers from recognition errors propagating to the translation stage, resulting in low translation accuracy. The other commonly used speech translation method utilizes an end-to-end trainable direct model for direct translation without transcription. While this method avoids the error propagation problem inherent in the cascaded approach, it still suffers from over-translation and named entity translation bias, leading to inaccurate translation results. Summary of the Invention

[0004] The main objective of this application is to provide a speech translation method, apparatus, storage medium, and device that can effectively improve the accuracy of translation results during speech translation.

[0005] This application provides a speech translation method, including:

[0006] Obtain the target speech to be translated; and extract the acoustic feature vector of the target speech;

[0007] The target speech is input into the recognition decoder and the translation decoder respectively to obtain the recognized text and the translated text; and based on the edit distance algorithm, the entity word feature vector is determined according to the recognized text, the translated text and the acoustic feature vector.

[0008] The entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability.

[0009] The first and second posterior probabilities are used to align entity word features; and after alignment, the final translation result corresponding to the target speech is determined.

[0010] In one possible implementation, extracting the acoustic feature vector of the target speech includes:

[0011] A multi-head self-attention mechanism is used to extract the global contextual information of the target speech, and a convolution module is used to capture the local temporal dependence of the target speech to determine the acoustic feature vector of the target speech.

[0012] In one possible implementation, the determination of entity word feature vectors based on the edit distance algorithm, according to the identified text, translated text, and acoustic feature vectors, includes:

[0013] Based on the edit distance algorithm, the position index of entity words in the recognized and translated texts is calculated.

[0014] After applying the location index to the acoustic feature vector, the entity word feature vector is determined.

[0015] In one possible implementation, the recognition decoder is trained using sample speech with the same representation as the target speech and its corresponding sample recognition text; the translation decoder is trained using the sample speech and its corresponding sample translation text.

[0016] In one possible implementation, the target speech is dialect speech; the sample recognition text is dialect text; the sample translation text is Mandarin text; the step of inputting the target speech into the recognition decoder and the translation decoder respectively to obtain the recognition text and the translation text includes:

[0017] The target speech is input into the recognition decoder and the translation decoder respectively to obtain dialect-recognized text and Mandarin-translated text.

[0018] In one possible implementation, the step of inputting the entity word feature vectors into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability includes:

[0019] The entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain the first posterior probability of the entity word in the dialect dictionary and the second posterior probability in the Mandarin dictionary.

[0020] In one possible implementation, the alignment processing of entity word features using the first posterior probability and the second posterior probability includes:

[0021] The first posterior probability and the second posterior probability are normalized to obtain the normalized first posterior probability and the normalized second posterior probability.

[0022] The bidirectional KL divergence average loss is calculated on the normalized first posterior probability and the normalized second posterior probability. By minimizing the calculation result, the distributions of the first posterior probability and the second posterior probability are brought closer together, thereby achieving alignment processing of entity word features.

[0023] This application also provides a voice translation device, including:

[0024] An extraction unit is used to acquire the target speech to be translated and to extract the acoustic feature vector of the target speech;

[0025] The determining unit is used to input the target speech into the recognition decoder and the translation decoder respectively to obtain the recognized text and the translated text; and to determine the entity word feature vector based on the edit distance algorithm, according to the recognized text, the translated text and the acoustic feature vector;

[0026] The input unit is used to input the entity word feature vector into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability.

[0027] An alignment unit is used to perform entity word feature alignment processing using the first posterior probability and the second posterior probability; and after alignment, to determine the final translation result corresponding to the target speech.

[0028] In one possible implementation, the extraction unit is specifically used for:

[0029] A multi-head self-attention mechanism is used to extract the global contextual information of the target speech, and a convolution module is used to capture the local temporal dependence of the target speech to determine the acoustic feature vector of the target speech.

[0030] In one possible implementation, the determining unit includes:

[0031] The computational subunit is used to calculate the position index of entity words in the recognized and translated texts based on the edit distance algorithm;

[0032] A subunit is defined to determine the entity word feature vector after applying the position index to the acoustic feature vector.

[0033] In one possible implementation, the recognition decoder is trained using sample speech with the same representation as the target speech and its corresponding sample recognition text; the translation decoder is trained using the sample speech and its corresponding sample translation text.

[0034] In one possible implementation, the target speech is dialect speech; the sample recognition text is dialect text; the sample translation text is Mandarin text; and the determining unit is specifically used for:

[0035] The target speech is input into the recognition decoder and the translation decoder respectively to obtain dialect-recognized text and Mandarin-translated text.

[0036] In one possible implementation, the input unit is specifically used for:

[0037] The entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain the first posterior probability of the entity word in the dialect dictionary and the second posterior probability in the Mandarin dictionary.

[0038] In one possible implementation, the alignment unit includes:

[0039] The normalization subunit is used to normalize the first posterior probability and the second posterior probability to obtain the normalized first posterior probability and the normalized second posterior probability.

[0040] The alignment subunit is used to calculate the bidirectional KL divergence average loss of the normalized first posterior probability and the normalized second posterior probability, and to minimize the calculation result to bring the distributions of the first posterior probability and the second posterior probability closer together, thereby realizing the alignment processing of entity word features.

[0041] This application also provides a voice translation device, including: a processor, a memory, and a system bus;

[0042] The processor and the memory are connected via the system bus;

[0043] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the speech translation method.

[0044] This application also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform any of the above-described implementations of the speech translation method.

[0045] This application also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described implementations of the speech translation method.

[0046] This application provides a speech translation method, apparatus, storage medium, and device. First, the target speech to be translated is acquired; then, the acoustic feature vector of the target speech is extracted; the target speech is then input into a recognition decoder and a translation decoder respectively to obtain recognized text and translated text; based on an edit distance algorithm, entity word feature vectors are determined according to the recognized text, translated text, and acoustic feature vectors; next, the entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain a first posterior probability and a second posterior probability; then, the first posterior probability and the second posterior probability can be used for entity word feature alignment; and after alignment, the final translation result corresponding to the target speech is determined.

[0047] As can be seen, this application, when translating target speech, constrains the translation results of named entity words by aligning the output features of the dual decoders (i.e., the recognition decoder and the translation decoder), thereby alleviating phenomena such as over-translation and named entity translation bias, achieving more reasonable translation results. Furthermore, entity feature alignment also improves the consistency and transcription accuracy between the dual decoders, thereby improving the accuracy of the final translation result corresponding to the target speech and enhancing the user's translation experience. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 A flowchart illustrating a speech translation method provided in an embodiment of this application;

[0050] Figure 2 Example diagram of the edit distance matrix provided in the embodiments of this application;

[0051] Figure 3 Example diagram of entity word feature vectors provided in the embodiments of this application;

[0052] Figure 4 Example diagram of dual-decoding output entity word feature alignment provided in the embodiments of this application;

[0053] Figure 5 This is a schematic diagram illustrating the overall implementation process of the speech translation method provided in the embodiments of this application;

[0054] Figure 6 This is a schematic diagram illustrating the composition of a speech translation device provided in an embodiment of this application. Detailed Implementation

[0055] In the field of natural language processing, accurate speech translation is crucial for human-computer interaction. Currently, commonly used speech translation methods include the following two:

[0056] The first method is a translation approach that uses cascaded decoding to independently generate two transcriptions.

[0057] Because existing cascaded speech translation systems typically separate speech recognition (ASR) and machine translation (MT) into two independent steps, they often suffer from recognition errors propagating to the translation stage in practical applications. Any errors in the recognition stage directly affect translation quality, causing the system to deviate from the original meaning when processing long sentences, complex sentence structures, or dialects, and even resulting in overtranslation or named entity translation errors. Furthermore, because existing cascaded models do not fully consider the characteristics of dialects and other speech sounds, their recognition accuracy and translation quality are significantly insufficient when processing audio from different dialects. This error accumulation problem is particularly prominent in complex scenarios involving multiple languages ​​and dialects, making it difficult for the overall speech translation system to provide accurate and consistent results, resulting in low translation accuracy.

[0058] The second method is to use an end-to-end dialect speech translation model to directly generate dialect translation results from the input audio.

[0059] These end-to-end dialect speech translation models can translate without transcription. While reducing error propagation, issues such as inaccurate dialect translation and named entity translation bias still exist. Furthermore, existing end-to-end solutions, while generating reasonable translation results, cannot provide fine-grained control over speech recognition and translation, especially in real-world applications requiring readable speech transcription and high-quality translation, leading to inaccurate translations.

[0060] Taking dialects as an example, as unique variants of language, they possess complex phonetic features and regional differences. Traditional speech translation systems, lacking specialized processing mechanisms for dialect characteristics, often suffer from inaccurate entity recognition and inappropriate vocabulary translation during the speech recognition and translation process. Furthermore, errors in named entity translation are also a common problem, especially when dealing with special words such as personal names and place names. These issues severely impact the accuracy and usability of the translation results, leading to a poor user experience. Traditional end-to-end dialect speech translation models directly generate translation results from the input speech signal. While this simplifies the system architecture and reduces the complexity of processing steps, it also has certain shortcomings. For example, existing translation systems are prone to over-translation or inaccurate named entity translation for certain dialects or accents. This is mainly because existing systems struggle to effectively capture and process subtle differences in dialects, especially in complex or special linguistic contexts where these phenomena are more pronounced.

[0061] To address the aforementioned shortcomings, this application provides a speech translation method. First, the target speech to be translated is acquired; then, the acoustic feature vector of the target speech is extracted. The target speech is then input into a recognition decoder and a translation decoder respectively to obtain recognized text and translated text. Based on an edit distance algorithm, entity word feature vectors are determined according to the recognized text, translated text, and acoustic feature vectors. Next, the entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain a first posterior probability and a second posterior probability. The first and second posterior probabilities are then used for entity word feature alignment. Finally, after alignment, the final translation result corresponding to the target speech is determined.

[0062] As can be seen, this application, when translating target speech, constrains the translation results of named entity words by aligning the output features of the dual decoders (i.e., the recognition decoder and the translation decoder), thereby alleviating phenomena such as over-translation and named entity translation bias, achieving more reasonable translation results. Furthermore, entity feature alignment also improves the consistency and transcription accuracy between the dual decoders, thereby improving the accuracy of the final translation result corresponding to the target speech and enhancing the user's translation experience.

[0063] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0064] First Embodiment

[0065] See Figure 1 This is a flowchart illustrating a speech translation method provided in this embodiment. The method includes the following steps:

[0066] S101: Obtain the target speech to be translated; and extract the acoustic feature vector of the target speech.

[0067] In this embodiment, any speech that needs to be translated is defined as the target speech to be translated. It should be noted that this embodiment does not limit the language type of the target speech. For example, the target speech can be a dialect or English. At the same time, this embodiment does not limit the length of the target speech. For example, the target speech can be a sentence or a paragraph.

[0068] It is understood that the target speech can be obtained through recording or other means as needed. For example, telephone conversations in people's daily lives or recordings of instant messaging software can be used as target speech. At the same time as the target speech is obtained, it can be translated using the solution provided in this embodiment to translate the target speech into text in another language (such as Mandarin).

[0069] Furthermore, after obtaining the target speech to be recognized, in order to improve the translation accuracy of the target speech, it is necessary to use existing or future feature extraction methods to extract the acoustic feature vector of the target speech, and use the acoustic feature vector as the basis for translation, so as to achieve accurate translation of the target speech through subsequent steps S102-S104.

[0070] Specifically, one possible implementation is that when extracting the acoustic features of the target speech, not only can the global context information of the target speech be extracted using a multi-head self-attention mechanism, but the local temporal dependencies of the target speech can also be captured through a convolutional module. Then, the output feature representation is obtained by combining the residual connections of these features and determined as the acoustic feature vector of the target speech.

[0071] In this implementation, it should be noted that this embodiment does not limit the method for extracting the acoustic features of the target speech, nor does it limit the specific extraction process. Appropriate extraction methods and corresponding feature extraction operations can be selected according to the actual situation. For ease of understanding, this embodiment will use the extraction of acoustic features of the target speech using a conformal transform speech recognition model as an example.

[0072] Conformer is a speech recognition model that combines Transformer and Convolutional Neural Networks (CNN). The Conformer model's structure includes a self-attention mechanism and convolutional modules, enabling it to simultaneously capture global semantic information and local temporal features. By combining self-attention and convolutional operations, the Conformer model can effectively handle long-term dependencies and short-term variations in speech signals, thus more accurately representing the acoustic features of speech.

[0073] Specifically, when using the Conformer model to extract features from target speech, after inputting the target speech into the model, the input speech signal is first processed into frames to generate a series of frame-level acoustic features. Then, local temporal information is extracted through a convolutional module, and global contextual information is captured through a multi-head self-attention mechanism, ultimately generating an acoustic feature sequence. Each frame of this sequence corresponds to a feature vector, reflecting the spectral changes and speech components of the speech signal in the time dimension (e.g., high frequencies and low frequencies represent emotional expressions such as screams and interjections, respectively).

[0074] If X represents the target speech input to the Conformer model, then in the Conformer model, it can be processed through a multi-head self-attention module to extract global contextual information, and then through a convolutional module to capture local temporal dependencies. Finally, the residual connections combining these features yield the output feature representation, which serves as the acoustic feature vector of the target speech, and is denoted as H. The specific calculation formula is as follows:

[0075] H=LayerNorm(X+ConvModule(MHSA(X)))

[0076] Here, ConModule represents the function of the local convolutional module; MHSA represents the function of the multi-head self-attention mechanism; LayerNorm represents the function of layer normalization; H, as the output of the Conformer model, can be understood as a matrix of size T*K, where T represents the number of speech frames in the target speech, and K represents the feature dimension of each frame. The specific values ​​are not limited and can be set according to the actual situation and experience. For example, T and K can be set to 5 and 512 respectively. Moreover, this feature matrix H contains multi-level spatiotemporal information of the target speech signal, which can effectively represent the acoustic features of the target speech and is used to execute the subsequent step S102 for the recognition and translation processing of the target speech.

[0077] S102: Input the target speech into the recognition decoder and the translation decoder respectively to obtain the recognition text and the translation text; and determine the entity word feature vector based on the edit distance algorithm, according to the recognition text, the translation text and the acoustic feature vector.

[0078] In this embodiment, after extracting the acoustic feature vector H of the target speech in step S101, in order to improve the translation accuracy of the target speech, the target speech can be further input into the recognition decoder and the translation decoder respectively to obtain the recognition text and the translation text. Then, based on the edit distance algorithm, the position index of the entity word in the recognition text and the translation text is calculated. Then, the position index is applied to the acoustic feature vector H to determine the entity word feature vector (which can be represented by S) for subsequent step S103.

[0079] It should be noted that the specific composition and specific processing process of the recognition decoder and the translation decoder can be set according to the actual situation, and this embodiment does not limit it. An optional implementation is that the recognition decoder can be jointly trained using sample voices with the same form as the target voice and their corresponding sample recognition texts; the translation decoder can be jointly trained using sample voices and their corresponding sample translation texts.

[0080] Specifically, in this implementation, in order to construct the recognition decoder and the translation decoder, a large number of preparatory works need to be carried out in advance. First, a large amount of speech data needs to be collected as sample voices to form model training data. For example, a large amount of recorded data of dialects can be collected in advance. For example, the Cantonese voices obtained after reciting multiple preset texts in Cantonese can be used as sample voices to form model training data, and the corresponding Cantonese texts (such as "点解没声") and Mandarin texts (such as "为什么没声") of these sample voices are manually marked as sample recognition texts and sample translation texts respectively. Then, based on these sample voices and their corresponding sample recognition texts, the initial recognition decoder can be trained to generate the recognition decoder; and based on these sample voices and their corresponding sample translation texts, the initial translation decoder can be trained to generate the translation decoder.

[0081] Among them, the composition structures of the initial recognition decoder and the initial translation decoder in this application are not limited, and their composition structures can be the same or different. An optional implementation is that the initial recognition decoder and the initial translation decoder can be (but not limited to) neural network models based on the Transformer structure.

[0082] Specifically, when performing model training, a sample speech (such as a Cantonese speech) can be sequentially extracted from the training data, and the encoded acoustic vector can be obtained through a common model encoder. Then, the acoustic vector is respectively input into two initial decoders (an initial recognition decoder and an initial translation decoder), and the corresponding sample recognition text (such as the Cantonese text "为什么没声") and sample translation text (such as the Mandarin text "为什么没声") are used as outputs for multiple rounds of model training. The recognition results and translation results obtained in each round of training are respectively compared with the corresponding manually annotated results, and the model parameters are updated according to the differences between the two until the preset conditions are met, such as the value of the loss constraint function is very small and basically unchanged or the preset training times are reached. Then, the update of the model parameters is stopped, the training of the recognition decoder and the translation decoder is completed, and two trained translators (i.e., the recognition decoder and the translation decoder) are generated. Moreover, the model structures of these two decoders can be the same, and the only difference is that the recognition decoder is trained using the acoustic vector of the sample speech and the sample recognition text (such as the Cantonese dialect text), and the translation decoder is trained using the acoustic vector of the sample speech and the sample translation text (such as the Mandarin translation text). That is, a recognition decoder with recognition ability is trained using the acoustic vector of the sample speech and the sample recognition text (such as the Cantonese dialect text), and a translation decoder with translation ability is trained using the acoustic vector of the sample speech and the sample translation text (such as the Mandarin translation text).

[0083] On this basis, after constructing a recognition decoder with recognition ability and a translation decoder with translation ability, further, the target speech can be first input into the recognition decoder and the translation decoder respectively to obtain the recognition text and the translation text, and then the minimum edit operation sequence between the recognition text and the translation text can be calculated based on an existing or future edit distance algorithm (such as the Levenshtein edit distance algorithm). Taking the Levenshtein edit distance algorithm as an example, this algorithm realizes the conversion of strings through three basic operations: insertion, deletion, and replacement. The edit distance, as a metric standard for the difference between two strings, reflects the minimum number of operations required to convert one string to another string. The smaller the distance, the higher the similarity between the two strings, which is particularly important for the alignment of the speech recognition result and the translation text.

[0084] During the calculation process, an edit distance matrix D can be first constructed, where D[i][j] represents the minimum number of operations required to convert the first i characters of the recognition result string to the first j characters in the translation annotation file. This matrix is calculated through a dynamic programming algorithm and follows the following recurrence formula:

[0085]

[0086] Among them, the right - hand formulas respectively represent deletion, insertion, and substitution operations from top to bottom. a[i - 1] and b[j - 1] respectively represent the characters at the corresponding positions in the recognized text and the translated text. The filling of the edit - distance matrix D follows the above formulas, expanding from the initial state to D[len(a)][len(b)] all the way. The final edit - distance value represents the total transformation cost from the recognized text to the translated text.

[0087] In this way, after calculating the edit distance, the feature vector S related to the entity word can be generated according to this matrix and the acoustic feature vector H.

[0088] An optional implementation method is that when the target speech is a dialect speech, correspondingly, the sample recognized text is a dialect text, and the sample translated text is a Mandarin text. Then, the target speech is input into the recognition decoder and the translation decoder respectively, and the recognized text and the translated text obtained are the dialect recognized text and the Mandarin translated text respectively.

[0089] For example: Suppose a Cantonese speech is input into the recognition decoder and the translation decoder respectively as the target speech, and the obtained dialect recognized text and Mandarin translated text are: the Cantonese text "点解没声" (4 characters) and the Mandarin text "为什么没声" (5 characters). Then the process of constructing the edit - distance matrix D is shown in the following three steps:

[0090] (1) Initialize the matrix (5 rows × 6 columns):

[0091] D[0][j] = j (number of insertion operations)

[0092] D[i][0] = i (number of deletion operations)

[0093] (2) Fill the matrix cell by cell: Calculate the minimum edit distance through dynamic programming, and finally obtain the complete matrix as Figure 2 shown. It can be seen that the finally obtained edit distance is: D[4][5]=3.

[0094] (3) Backtrack the edit operation sequence:

[0095] Path: D[4][5] ← D[3][4] (声 → 声, no operation)

[0096] D[3][4] ← D[2][3] (没 → 没, no operation)

[0097] D[2][3] ← D[1][2]+1 (解 → 么, substitution)

[0098] D[1][2] ← D[0][1]+1 (点 → 什, insert "为" and then substitution)

[0099] Thus, according to the above process, the position indexes of the entity word "mòu sīng" in the Cantonese text and the Mandarin text can be calculated through "no operation", that is, the 3rd - 4th characters in the Cantonese text and the 4th - 5th characters in the Mandarin text.

[0100] It should be noted that when matching entity words, to avoid interference from excessive function words (such as "le") and other non - entity words, an optional implementation method is that when generating the entity word feature vector, phrases with a length less than 2 characters can be filtered out to ensure that the extracted entity words are all meaningful. This can not only effectively reduce the interference of noisy data but also improve the accuracy of the translation result.

[0101] Next, after matching the entity word, it can be mapped to a low - dimensional vector representation to characterize the importance and similarity of the entity word in the recognition text and the translation text (for example, performing semantic encoding on "mòu sīng" and mapping it to a low - dimensional vector [0.32, - 0.15, 0.87] for input to the decoder in subsequent steps). Among them, the dimension of the vector is usually the speech frame feature dimension corresponding to the entity word in the acoustic feature vector H, representing the semantic space features of the entity word. These features can not only be used for alignment in the translation process but also provide input for subsequent deep - learning models.

[0102] In this way, through the above steps, the entity words in the recognition text and the translation text can be efficiently matched, and an entity word feature vector S that can reflect its characteristics can be generated, as Figure 3 shown, ensuring a high accuracy in the target speech recognition and translation process.

[0103] S103: Input the entity word feature vector into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability.

[0104] In this embodiment, after determining the entity word feature vector S through step S102, the entity word feature vector S can be further input into the recognition decoder and the translation decoder for decoding (such as further improving the accuracy of speech recognition through the forward decoding process), obtaining the posterior probabilities of the entity word in the respective dictionaries of the two decoders, and respectively defining them as the first posterior probability and the second posterior probability for performing the subsequent step S104. This can ensure the synchronous progress between the speech recognition and translation tasks based on the ability of the dual - decoder structure to parallel - process two types of tasks, namely recognition and translation, and at the same time improve the adaptability to complex language environments.

[0105] Specifically, one possible implementation is that if the target speech is input into the recognition decoder and the translation decoder respectively to obtain dialect recognition text and Mandarin translation text, then the entity word feature vector S can be input into the recognition decoder and the translation decoder respectively to obtain the first posterior probability of the entity word in the dialect dictionary and the second posterior probability of the entity word in the Mandarin dictionary.

[0106] The composition of the recognition decoder and the translation decoder is not limited and can consist of two independent sub-decoders. They are constructed using a neural network model based on the Transformer structure, as described above, to ensure efficient processing of time-series data and text prediction tasks.

[0107] S104: Align entity word features using the first and second posterior probabilities; and after alignment, determine the final translation result corresponding to the target speech.

[0108] In this embodiment, after inputting the entity word feature vector S into the recognition decoder and the translation decoder respectively in step S103 to obtain the first posterior probability and the second posterior probability, since the recognition decoder and the translation decoder may generate output word sequences of different lengths, it is necessary to perform scale unification processing on the probability distributions of their outputs (i.e., the first posterior probability and the second posterior probability) for subsequent fusion operations. Therefore, both the first posterior probability and the second posterior probability can be normalized to obtain the normalized first posterior probability and the normalized second posterior probability, which are respectively denoted as P. recognizer and P translator The specific normalization formula used is as follows:

[0109]

[0110] Among them, S i Let represent the posterior score of the i-th character; Let represent the normalized posterior probability of the i-th character; N represents the total length of the dictionary.

[0111] After normalization, the outputs of the two decoders are fused posteriorly to generate the final recognition result. That is, the word with the highest posterior probability is selected as the recognition output by combining the probability distributions of the two decoders. This process is achieved through the maximum a posteriori probability criterion, which selects the word with the highest posterior probability in each frame as the output, ensuring the accuracy and robustness of the recognition result.

[0112] In the fusion process, to improve the accuracy of the final translation result, the normalized first posterior probability P can be adjusted. recognizer and the normalized second posterior probability P translator The two-way KL divergence average loss can be expressed as The calculation of the KL divergence is as follows:

[0113]

[0114] Based on this, by minimizing the average loss of KL divergence By aligning the distributions of the first and second posterior probabilities using the calculated results, entity word features can be aligned to achieve semantic representation consistency across tasks. This alignment also allows for a more accurate determination of the final translation result corresponding to the target speech. Figure 4 As shown. The specific alignment formula used is as follows:

[0115]

[0116] in, P represents the average loss of the two-way KL divergence. recognizer and P translator Let D represent the normalized first posterior probability and the normalized second posterior probability, respectively; KL (P recognizer ||P translator The expression ) represents the KL divergence calculation from translation to recognition. Here, the translator can be understood as the prior information for translation; that is, the KL divergence of the recognition prior is calculated based on the prior information for translation. Therefore, it is a divergence calculation from translation to recognition. The part after "||" represents the original data distribution, and the part before "||" represents the target data distribution. According to the definition of probability distribution, it is interpreted here as "KL divergence from translation to recognition." Similarly, D... KL (P translator ||P recognizer The ) indicates the KL divergence calculation of the identified translation; the specific reasons will not be elaborated further.

[0117] In this way, feature alignment aims to improve the overall consistency and accuracy of recognition and translation results by minimizing the semantic distribution differences between decoders. Once the feature vectors of the recognition decoder and the translation decoder are successfully aligned, the results can be automatically fused. That is, based on the posterior probability distributions of the recognition decoder and the translation decoder, the optimal entity word is selected for output, thus determining a more accurate entity word translation result.

[0118] Specifically, after receiving the entity word feature vector S, the recognition decoder can decode the possible speech recognition results based on a specific dialect (such as Cantonese) dictionary. This process takes the feature vector of each word as input and performs layer-by-layer forward propagation through a deep neural network model to calculate the posterior probability of each word in the dictionary.

[0119] In parallel with the recognition decoder, the translation decoder processes translation-related feature vectors. Similar to the recognition decoder, based on a specific Mandarin dictionary, it calculates the posterior probability of each translated word using a similar neural network structure, capturing semantically relevant contextual information. Through this process, the translation decoder can infer text highly relevant to the translation result from the speech feature vectors.

[0120] Therefore, by using dual decoders, the parallel processing capability for target speech recognition and translation can be effectively improved, ensuring efficient and accurate output in multilingual speech recognition environments. Compared with existing single-decoder systems, dual decoders can better handle the differentiated needs of speech recognition and translation tasks, and ensure consistency and coherence in the simultaneous output of recognition and translation results, while significantly reducing semantic errors and ambiguities. The final generated recognition and translation results are not only optimized with normalized posterior probabilities, but also possess higher language fluency and contextual coherence, improving the accuracy of speech recognition and translation results and optimizing the user translation experience.

[0121] To facilitate understanding of the speech translation scheme proposed in this application, a schematic diagram of the overall implementation process of the speech translation method is also provided, such as... Figure 5 As shown, specifically: First, the target speech signal to be translated is acquired. Then, acoustic feature extraction is performed on the target speech to generate an acoustic feature vector. This process can use existing or future acoustic models or specific pre-trained speech models to extract spectral features, phoneme information, etc., of the target speech as the acoustic feature vector of the target speech.

[0122] Next, based on the extracted acoustic feature vectors, a similarity matrix is ​​constructed using the edit distance algorithm to match entity words in the speech signal. The edit distance algorithm calculates the degree of matching between the phoneme sequence and the text, generates a specific distance score, selects the most matching entity word based on the score, and generates a corresponding feature vector for it. At this time, the matched entity word feature vectors are sent to two independent decoders for processing: (1) Speech recognition decoder: After receiving the entity word feature vector, the decoder outputs the recognition confidence of the word (which can be reflected as a posterior probability) and generates the final recognition result based on this confidence. (2) Speech translation decoder: The decoder receives the same entity word feature vector and translates the semantics of the entity word according to a specific language model, calculates the confidence of the translation (which can be reflected as a posterior probability), and outputs the translation result. Finally, the results of speech recognition and translation can be presented to the user simultaneously to ensure that the translation result is consistent with the recognition result, or the final translation result can be presented to the user to avoid common problems such as translation deviation caused by recognition errors.

[0123] Thus, when translating target speech using the speech translation method proposed in this application, both the speech transcription result and the translation result can be generated simultaneously during the translation process, avoiding the quality degradation problem caused by error propagation in existing cascaded systems. Furthermore, through the dual-decoder structure, speech recognition and translation tasks can be performed concurrently and mutually constrained, thereby ensuring the consistency of recognition and translation results. In particular, by introducing a feature alignment mechanism, the translation accuracy of named entities can be effectively improved, and over-translation problems in target speech (such as dialect speech) can be reduced. Therefore, this not only solves the error accumulation problem in traditional speech translation modes but also significantly improves the quality and user experience of multi-dialect and multilingual speech translation.

[0124] In summary, the speech translation method provided in this embodiment first acquires the target speech to be translated; then extracts the acoustic feature vector of the target speech; and then inputs the target speech into a recognition decoder and a translation decoder respectively to obtain the recognized text and the translated text; based on the edit distance algorithm, entity word feature vectors are determined according to the recognized text, the translated text, and the acoustic feature vector; next, the entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain a first posterior probability and a second posterior probability; then, the first posterior probability and the second posterior probability can be used to perform entity word feature alignment processing; and after alignment, the final translation result corresponding to the target speech is determined.

[0125] As can be seen, this application, when translating target speech, constrains the translation results of named entity words by aligning the output features of the dual decoders (i.e., the recognition decoder and the translation decoder), thereby alleviating phenomena such as over-translation and named entity translation bias, achieving more reasonable translation results. Furthermore, entity feature alignment also improves the consistency and transcription accuracy between the dual decoders, thereby improving the accuracy of the final translation result corresponding to the target speech and enhancing the user's translation experience.

[0126] Second Embodiment

[0127] This embodiment will introduce a voice translation device; please refer to the above method embodiment for related content.

[0128] See Figure 6 This is a schematic diagram of the composition of a voice translation device provided in this embodiment. The device 600 includes:

[0129] Extraction unit 601 is used to acquire the target speech to be translated and extract the acoustic feature vector of the target speech;

[0130] The determining unit 602 is used to input the target speech into the recognition decoder and the translation decoder respectively to obtain the recognized text and the translated text; and to determine the entity word feature vector based on the edit distance algorithm, according to the recognized text, the translated text and the acoustic feature vector.

[0131] The input unit 603 is used to input the entity word feature vector into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability;

[0132] Alignment unit 604 is used to perform entity word feature alignment processing using the first posterior probability and the second posterior probability; and after alignment, to determine the final translation result corresponding to the target speech.

[0133] In one implementation of this embodiment, the extraction unit 601 is specifically used for:

[0134] A multi-head self-attention mechanism is used to extract the global contextual information of the target speech, and a convolution module is used to capture the local temporal dependence of the target speech to determine the acoustic feature vector of the target speech.

[0135] In one implementation of this embodiment, the determining unit 602 includes:

[0136] The computational subunit is used to calculate the position index of entity words in the recognized and translated texts based on the edit distance algorithm;

[0137] A subunit is defined to determine the entity word feature vector after applying the position index to the acoustic feature vector.

[0138] In one implementation of this embodiment, the recognition decoder is trained using sample speech with the same representation as the target speech and its corresponding sample recognition text; the translation decoder is trained using the sample speech and its corresponding sample translation text.

[0139] In one implementation of this embodiment, the target speech is dialect speech; the sample recognition text is dialect text; the sample translation text is Mandarin text; the determining unit 602 is specifically used for:

[0140] The target speech is input into the recognition decoder and the translation decoder respectively to obtain dialect-recognized text and Mandarin-translated text.

[0141] In one implementation of this embodiment, the input unit 603 is specifically used for:

[0142] The entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain the first posterior probability of the entity word in the dialect dictionary and the second posterior probability in the Mandarin dictionary.

[0143] In one implementation of this embodiment, the alignment unit 604 includes:

[0144] The normalization subunit is used to normalize the first posterior probability and the second posterior probability to obtain the normalized first posterior probability and the normalized second posterior probability.

[0145] The alignment subunit is used to calculate the bidirectional KL divergence average loss of the normalized first posterior probability and the normalized second posterior probability, and to minimize the calculation result to bring the distributions of the first posterior probability and the second posterior probability closer together, thereby realizing the alignment processing of entity word features.

[0146] Furthermore, embodiments of this application also provide a voice translation device, including: a processor, a memory, and a system bus;

[0147] The processor and the memory are connected via the system bus;

[0148] The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform any of the above-described implementations of the speech translation method.

[0149] Furthermore, embodiments of this application also provide a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform any of the above-described implementations of the speech translation method.

[0150] Furthermore, this application also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described implementations of the speech translation method.

[0151] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0152] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0153] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0154] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech translation method, characterized in that, include: Obtain the target speech to be translated; And extract the acoustic feature vector of the target speech; The target speech is input into the recognition decoder and the translation decoder respectively to obtain the recognized text and the translated text; and based on the edit distance algorithm, the entity word feature vector is determined according to the recognized text, the translated text and the acoustic feature vector. The entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability. Alignment of entity word features is performed using the first posterior probability and the second posterior probability. After alignment, the final translation result corresponding to the target speech is determined.

2. The method according to claim 1, characterized in that, The extraction of the acoustic feature vector of the target speech includes: A multi-head self-attention mechanism is used to extract the global contextual information of the target speech, and a convolution module is used to capture the local temporal dependence of the target speech to determine the acoustic feature vector of the target speech.

3. The method according to claim 1, characterized in that, The edit distance-based algorithm determines entity word feature vectors based on the identified text, translated text, and acoustic feature vectors, including: Based on the edit distance algorithm, the position index of entity words in the recognized and translated texts is calculated. After applying the location index to the acoustic feature vector, the entity word feature vector is determined.

4. The method according to claim 1, characterized in that, The recognition decoder is trained using sample speech with the same representation as the target speech and its corresponding sample recognition text; the translation decoder is trained using the sample speech and its corresponding sample translation text.

5. The method according to claim 4, characterized in that, The target speech is a dialect speech; the sample recognition text is a dialect text; the sample translation text is a Mandarin text; The step of inputting the target speech into the recognition decoder and the translation decoder respectively to obtain the recognized text and the translated text includes: The target speech is input into the recognition decoder and the translation decoder respectively to obtain dialect-recognized text and Mandarin-translated text.

6. The method according to claim 5, characterized in that, The step of inputting the entity word feature vectors into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability includes: The entity word feature vectors are input into the recognition decoder and the translation decoder respectively to obtain the first posterior probability of the entity word in the dialect dictionary and the second posterior probability in the Mandarin dictionary.

7. The method according to any one of claims 1-6, characterized in that, The alignment process of entity word features using the first posterior probability and the second posterior probability includes: The first posterior probability and the second posterior probability are normalized to obtain the normalized first posterior probability and the normalized second posterior probability. The bidirectional KL divergence average loss is calculated on the normalized first posterior probability and the normalized second posterior probability. By minimizing the calculation result, the distributions of the first posterior probability and the second posterior probability are brought closer together, thereby achieving alignment processing of entity word features.

8. A voice translation device, characterized in that, include: Extraction unit, used to acquire the target speech to be translated; And extract the acoustic feature vector of the target speech; The determining unit is used to input the target speech into the recognition decoder and the translation decoder respectively to obtain the recognized text and the translated text; and to determine the entity word feature vector based on the edit distance algorithm, according to the recognized text, the translated text and the acoustic feature vector; The input unit is used to input the entity word feature vector into the recognition decoder and the translation decoder respectively to obtain the first posterior probability and the second posterior probability. The alignment unit is used to perform alignment processing of entity word features using the first posterior probability and the second posterior probability. After alignment, the final translation result corresponding to the target speech is determined.

9. A voice translation device, characterized in that, include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Language model translation and training method and apparatus

    US20190163747A1

  • Translation method and related device

    WO2024055707A1