Voice processing method and device, equipment, medium and product

Through the integrated framework's speech processing model, the encoding and decoding strategies are dynamically adjusted, which solves the high cost and real-time limitations caused by independent construction of speech translation and speech simultaneous transmission models, and achieves high-quality and low-latency speech translation and simultaneous transmission effects.

CN120472904APending Publication Date: 2025-08-12XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510739243.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, independent construction of speech translation and speech simultaneous transmission models leads to high deployment costs and limited real-time and quality in multi-translation modes.

Method used

The speech processing model adopts an integrated framework, integrates speech translation and speech simultaneous modes through end-to-end training, dynamically adjusts encoding and decoding strategies, share training data and model parameters, and optimizes the real-time and quality of speech processing.

Benefits of technology

It reduces deployment costs, improves the real-time and quality of voice processing, and realizes the advantages of knowledge sharing and translation under different translation modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472904A_ABST
    Figure CN120472904A_ABST
Patent Text Reader

Abstract

The invention provides a voice processing method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a target voice and a translation mode of the target voice; based on a voice encoder module and a translation mode in the voice processing model, encoding the target voice to obtain voice content features and acoustic features of the target voice; translating the target voice based on a large language model, a translation mode and voice content features in the voice processing model to obtain a target translation text of the target voice; and based on a voice decoder module, the translation mode and the acoustic features in the voice processing model, performing voice synthesis on the target translation text to obtain a target translation voice of the target voice. According to the invention, adaptive processing is carried out on voice coding, translation and voice decoding through a voice processing model of an integrated framework integrating voice translation and voice simultaneous transmission in combination with a translation mode, so that the deployment cost is better reduced, and the real-time performance and quality of voice processing are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a speech processing method, apparatus, device, medium and product. Background Art

[0002] Speech translation and simultaneous interpretation are the two core technologies that support real-time cross-language interaction. Speech translation, also known as non-instantaneous speech translation, generates a translation result in the target language after receiving the user's complete speech input. Simultaneous interpretation, also known as instantaneous speech translation, receives the user's speech in real time and outputs incremental translation results in real time.

[0003] In related technologies, in response to the differences in translation modes for speech processing, independent models (such as simultaneous interpretation models or speech translation models) are constructed to adapt to each translation mode one by one to implement speech processing under the corresponding translation mode. In scenarios with multiple translation modes, this approach not only leads to high deployment costs, but also limits the real-time performance and quality of speech processing. Summary of the Invention

[0004] The present invention provides a speech processing method, device, equipment, medium and product to solve the defects in the prior art.

[0005] The present invention provides a speech processing method, comprising: Acquire a target speech and a translation mode of the target speech; the translation mode includes a speech translation mode or a speech simultaneous interpretation mode; Encoding the target speech based on the speech encoder module in the speech processing model and the translation mode to obtain speech content features and acoustic features of the target speech; Translating the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translation text of the target speech; Performing speech synthesis on the target translation text based on the speech decoder module in the speech processing model, the translation mode, and the acoustic features to obtain a target translation speech of the target speech; The speech processing model is trained based on sample speech, and the sample speech is obtained by training sample translated speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

[0006] According to the present invention, a speech processing method is provided, wherein the target speech is translated based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translation text of the target speech, including: When the translation mode is the speech simultaneous interpretation mode, based on the semantic encoder module in the speech processing model, semantic content prediction and semantic end symbol prediction are performed on the target speech, and the speech content features are divided into a plurality of semantic feature blocks according to the semantic content prediction results and the semantic end symbol prediction results; Based on the large language model, a plurality of the semantic feature blocks are subjected to streaming translation to obtain the target translation text.

[0007] According to the present invention, a speech processing method is provided, wherein the method performs streaming translation on the plurality of semantic feature blocks based on the large language model to obtain the target translation text, comprising: For the current translation, obtaining a current semantic feature block corresponding to the current translation from the plurality of semantic feature blocks; Based on the adapter corresponding to the simultaneous speech interpretation mode in the speech processing model and the preset length information, feature fusion is performed on the historical semantic feature block corresponding to the historical translation, the historical translation text of the historical translation, and the current semantic feature block to obtain a first fused feature; Inputting the first fusion feature and the translation prompt word corresponding to the speech simultaneous interpretation mode into the large language model to obtain the translation text of the current translation; The translation text of the current translation is determined as the target translation text.

[0008] According to the present invention, a speech processing method is provided, wherein based on the speech decoder module in the speech processing model, the translation mode and the acoustic features, speech synthesis is performed on the target translation text to obtain the target translation speech of the target speech, comprising: When the translation mode is the simultaneous voice interpretation mode, for the current translation, based on the voice decoder module, a semantic judgment is performed on the translation text of the current translation. When all text translation data blocks in the translation text of the current translation are determined to be fixed according to the semantic judgment result, speech synthesis is performed on the translation text of the current translation according to the acoustic features to obtain the target translation speech.

[0009] According to the present invention, a speech processing method is provided, the method further comprising: When it is determined according to the semantic judgment result that all text translation data blocks are not to be fixed, determining whether the fixed delay duration corresponding to the current translation is greater than the first delay duration; When it is determined that the fixed delay time is greater than the first delay time, determining whether there is a common prefix between the translation text of the current translation and the historical translation text of the previous translation; When it is determined that the common prefix exists, speech synthesis is performed on the common prefix according to the acoustic features to obtain the target translated speech.

[0010] According to the present invention, a speech processing method is provided, the method further comprising: When it is determined that the common prefix does not exist, speech synthesis is performed on a target number of text translation data blocks ranked top in the translation text of the current translation according to the acoustic features to obtain the target translation speech.

[0011] According to the present invention, a speech processing method is provided, the method further comprising: When the translation mode is the simultaneous voice interpretation mode, for the current translation, when it is determined that the screen-on delay time corresponding to the current translation is greater than the second delay time, the translation text of the current translation is displayed on the display screen.

[0012] According to the present invention, a speech processing method is provided, wherein the target speech is translated based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translation text of the target speech, including: When the translation mode is the speech translation mode, based on the adapter corresponding to the speech translation mode in the speech processing model and preset length information, feature fusion is performed on historical speech content features of historical speech, historical translation text of the historical speech, and speech content features of the target speech to obtain a second fused feature; the historical speech is a speech that was translated before the target speech and is associated with the target speech; The second fusion feature and the translation prompt word corresponding to the speech translation mode are input into the large language model to obtain the target translation text.

[0013] According to the present invention, a speech processing method is provided, wherein the target speech is obtained based on the following steps: Based on the speech detection module in the speech processing model, removing the silent segments in the original speech of the target speech to obtain multiple target speech segments; According to the silent segments, zero padding is performed between every two adjacent target speech segments to obtain the target speech.

[0014] The present invention also provides a speech processing device, comprising: An acquisition unit, configured to acquire a target speech and a translation mode of the target speech; the translation mode includes a speech translation mode or a speech simultaneous interpretation mode; An encoding unit, configured to encode the target speech based on the speech encoder module in the speech processing model and the translation mode to obtain speech content features and acoustic features of the target speech; a first translation unit, configured to translate the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translated text of the target speech; a second translation unit, configured to perform speech synthesis on the target translation text based on the speech decoder module in the speech processing model, the translation mode, and the acoustic features to obtain a target translation speech of the target speech; The speech processing model is trained based on sample speech, and the sample speech is obtained by training sample translated speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described speech processing methods when executing the program.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned speech processing methods when executed by a processor.

[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned speech processing methods.

[0018] The speech processing method, apparatus, device, medium, and product provided by the present invention obtain a speech processing model integrating a speech translation mode or a speech simultaneous interpretation mode through end-to-end training. Under the speech processing model, a speech encoder module is used to dynamically adjust the encoding strategy according to different translation modes to encode the target speech in different translation modes, thereby obtaining speech content features and acoustic features under different translation modes. Furthermore, input information of a large language model under different translation modes is dynamically determined based on the speech content features, so that the large language model can perform corresponding translations based on the input information under different translation modes, thereby obtaining target translated texts under different translation modes. Furthermore, a speech decoder module is used to dynamically adjust the decoding strategy according to different translation modes to perform speech synthesis on the target translated texts under different translation modes based on the acoustic features under different translation modes, thereby obtaining target translated speech of the target speech. Thus, through the speech processing model with an integrated framework including speech translation and speech simultaneous interpretation, the speech encoding, translation, and speech decoding processes are adapted in combination with the translation mode, avoiding the need to independently build models for different translation modes, effectively reducing deployment costs, and simultaneously sharing the internal knowledge and translation advantages of different translation modes, thereby better optimizing the real-time performance and quality of speech processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 This is one of the flow charts of the speech processing method provided by the present invention.

[0021] Figure 2 It is a structural diagram of the semantic encoder module provided by the present invention.

[0022] Figure 3 This is the second flow chart of the speech processing method provided by the present invention.

[0023] Figure 4 This is the third flow chart of the speech processing method provided by the present invention.

[0024] Figure 5 It is a schematic diagram of the flow of streaming decoding provided by the present invention.

[0025] Figure 6 This is the fourth flow chart of the speech processing method provided by the present invention.

[0026] Figure 7It is a structural diagram of the speech processing device provided by the present invention.

[0027] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0028] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0029] With the rapid development of technologies such as artificial intelligence, big data, and cloud computing, significant progress has been made in simultaneous interpretation and translation. The development and application of simultaneous interpretation and translation technologies have broken down language barriers and provided a convenient means of communication for global cross-cultural exchange. In areas such as business negotiations, international conferences, and education and training, these technologies have greatly improved communication efficiency, fostered friendly exchanges between countries, and driven globalization.

[0030] Currently, most simultaneous interpretation and translation technologies utilize a non-real-time cascade architecture for speech translation. The specific process involves first using a speech recognition system to transcribe the user's speech into text in real time. This text is then fed into a text translation system for translation, ensuring semantic integrity. Finally, a speech synthesis system converts the target language text into a speech signal for broadcast. Although speech recognition, text translation, and speech synthesis technologies are relatively mature, and the cascade approach can achieve a stable, reliable, latency-controlled, and effective speech translation system, it has significant drawbacks: first, translation accuracy decreases due to error accumulation; second, the translation process requires waiting for complete clause recognition before machine translation can proceed, resulting in poor real-time translation performance. In contrast, end-to-end translation methods directly translate speech input into the target language, effectively reducing the latency and error accumulation associated with the cascade approach. However, due to limited data resources, data acquisition for end-to-end translation is costly and limited in scale.

[0031] With the rapid development and iteration of large language models, their basic functions have been widely used in fields such as education, healthcare, programming, and science. Users describe questions in natural language and send them to the large language model. The model then infers the questions and generates answers to feedback to the user. Although large language models have helped users analyze and solve problems to a certain extent and promoted social development, in real-world interactive scenarios, interactions are not limited to natural language but also include multiple modalities such as images and video. Therefore, other modal or cross-modal large models based on large language models, such as speech large models, image large models, video large models, and multimodal interaction large models, have gradually become a research hotspot to meet the needs of users to express questions and obtain answers through various means such as speech, images, and video.

[0032] In this context, relevant technologies have begun to explore integrating large speech models into end-to-end translation methods, building end-to-end simultaneous speech interpretation models and end-to-end speech translation models.

[0033] However, speech translation and simultaneous interpretation are two different technologies. Speech translation, or non-instantaneous speech translation, generates a translation of the target language after receiving the entire user's speech input. Simultaneous interpretation, or instantaneous speech translation, receives the user's speech in real time and outputs incremental translation results. Furthermore, speech translation tags differ from simultaneous interpretation tags. Simultaneous interpretation requires human translation capabilities, and already translated content cannot be altered. However, speech translation captures the entire speech content, resulting in a more concise and precise translation.

[0034] Therefore, in related technologies, in response to the differences in translation modes of speech processing, end-to-end models (such as speech simultaneous interpretation models or speech translation models) are usually independently constructed based on large speech models to adapt to the translation modes one by one, so as to realize speech processing in the corresponding translation mode; however, in multi-translation mode scenarios, this approach will not only lead to high deployment costs due to the independent construction of models, but also lead to high deployment costs because the training data cannot be synergistically utilized between different models; in addition, due to the low scalability of independently constructed models, such as the complete distinction between speech simultaneous interpretation scenarios and speech translation scenarios, the knowledge of the speech translation model cannot be smoothly transferred to the speech simultaneous interpretation model, and frequent cross-model switching is required for different scenarios, which will limit the real-time performance and quality of speech processing.

[0035] To reduce deployment costs and improve the real-time performance and quality of speech processing, this application provides a speech processing method. This method can be applied to various translation scenarios, such as meetings and speeches, and this embodiment does not specifically limit this.

[0036] Figure 1This is one of the flow charts of the speech processing method provided by the present invention. Figure 1 As shown, the method includes step 110 , step 120 , step 130 and step 140 .

[0037] Step 110: Acquire a target speech and a translation mode of the target speech; the translation mode includes a speech translation mode or a speech simultaneous interpretation mode.

[0038] The target speech here refers to the speech to be processed. It can be speech obtained by collecting a speaker's speech signal through a sound pickup device, or it can be speech processed by a sound pickup device on a speaker's speech signal, such as speech obtained after noise reduction and / or silence filling. This embodiment does not specifically limit this. The sound pickup device here can be a smartphone, tablet computer, microphone, etc.

[0039] The translation mode here is dynamically determined based on the user's voice processing latency requirements and quality requirements, or is input by the user in the front-end interface, such as information input through a command line interface, a graphical interface, touch input, a drop-down selection input, voice input, gesture input, visual input, brain-computer input, etc. This embodiment does not specifically limit this.

[0040] Step 120: encoding the target speech based on the speech encoder module in the speech processing model and the translation mode to obtain speech content features and acoustic features of the target speech; Step 130: translating the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translation text of the target speech; Step 140: performing speech synthesis on the target translation text based on the speech decoder module in the speech processing model, the translation mode, and the acoustic features to obtain a target translation speech of the target speech; The speech processing model is trained based on sample speech, and the sample speech is obtained by training sample translated speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

[0041] It is understandable that in related technologies, the speech translation model and the speech simultaneous interpretation model are deployed independently, so the speech translation model and the speech simultaneous interpretation model need to be trained and developed independently, resulting in high deployment costs. In addition, the knowledge between the speech translation model and the speech simultaneous interpretation model cannot be effectively shared, and frequent cross-model switching is required for different scenarios, which limits the real-time performance and quality of speech processing.

[0042] To address this issue, the present invention establishes an integrated framework encompassing both speech translation and simultaneous interpretation. This integrated framework dynamically adjusts translation configuration strategies based on different translation modes, avoiding frequent cross-model switching. Furthermore, the integrated framework allows for sharing training data and model parameters across different translation modes, facilitating knowledge transfer and improving the real-time performance and quality of speech processing while reducing deployment costs.

[0043] Optionally, before executing step 120, a speech processing model capable of integrated speech translation and simultaneous speech interpretation tasks may be acquired through end-to-end training. Specifically, an initialization processing model may be constructed first. The initialization processing model herein includes at least an initialization speech encoder module, a pre-trained large language model, and an initialization speech decoder module. In addition, other modules may be included, such as an initialization semantic encoder module, an initialization online adapter corresponding to the simultaneous speech interpretation mode, and an initialization offline adapter corresponding to the speech translation mode. This embodiment does not specifically limit this. Each module in the initialization processing model may be a module after parameter initialization or a pre-trained module. This embodiment does not specifically limit this.

[0044] The pre-trained large language model herein, also referred to as a large model or a general pre-trained large model, refers to a natural language processing (NLP) model with a large number of parameters. The number of model parameters and / or the complexity of the model structure exceed a set threshold. This model is pre-trained on large-scale text data and possesses a high level of semantic understanding and natural language generation capabilities. This pre-trained large language model can be a bidirectional encoder representation transformer (BERT), a generative pre-trained transformer, a large language model meta-AI, etc., which is not specifically limited in this embodiment.

[0045] In addition, a large number of sample speech samples can be collected. These sample speech samples can cover a variety of languages, accents, speaking speeds, and topics to ensure that the trained speech processing model has good generalization capabilities. At the same time, corresponding sample translation speech samples are collected for speech translation mode and speech interpretation mode. For example, for speech translation mode, the sample translation speech samples can be carefully proofread and polished by professional translators, ensuring high translation quality. For speech interpretation mode, the sample translation speech samples can be generated in a simulated real-time translation scenario, focusing on the timeliness of translation.

[0046] Next, when conducting model training, the sample speech can be used as the input of the initialization processing model, and the translation speech prediction results under the speech translation mode and the translation speech prediction results under the speech simultaneous interpretation mode corresponding to the sample speech can be used as output. Multiple rounds of model training are performed, and the translation speech prediction results obtained in each round of training are compared with the corresponding sample translation speech to update the model parameters of the initialization processing model according to the difference between the two. The model parameters are updated until the maximum number of iterations is reached or the model converges, and the updating of the model parameters is stopped to obtain a speech processing model that can perform integrated processing of speech translation and speech simultaneous interpretation with low latency and high quality.

[0047] In practical applications, the speech encoder module in the trained speech processing model can be loaded first. The speech encoder module here is a module that integrates streaming and non-streaming encoding. That is, it can call its own network structure to implement different encoding strategies according to different translation modes. For example, when the translation mode is speech interpretation mode, it can implement a streaming encoding strategy, and when the translation mode is speech translation mode, it can implement a non-streaming encoding strategy. The speech encoder module can be constructed based on a multi-layer single-structure network model or a multi-layer hybrid structure network model. For example, it can be constructed based on a 16-layer convolution-augmented transformer (Conformer) module and a speaker feature extraction module. The speaker feature extraction module is used to take the shallow coding features output by the third-layer Conformer module as input to extract acoustic features based on the shallow coding features, thereby outputting the acoustic features of the target speech. The last-layer Conformer module is used to integrate and encode the coding features processed by the previous multi-layer Conformer modules, thereby outputting the speech content features of the target speech.

[0048] To simplify the description, the method provided in this embodiment is described below by taking an example in which a speech encoder module is constructed based on a 16-layer Conformer module and a speaker feature extraction module.

[0049] After the speech encoder module is loaded into the speech processing model, it can be called to apply the corresponding encoding strategy to the target speech based on the translation mode, thereby encoding the target speech's speech content and acoustic features. Acoustic features can include pitch, intensity, timbre, and emotion. Speech processing based on these features effectively ensures that the generated translated speech has a natural and fluent sound quality.

[0050] For example, when the translation mode is the simultaneous speech interpretation mode, the 16-layer Conformer module and the speaker feature extraction module can be called through the speech encoder module to execute a streaming coding strategy, such as adjusting the window range of the attention layer in each layer of the Conformer module to include only three adjacent data blocks in the target speech, so as to segmentally encode the data blocks within the window range, thereby outputting the speech content features and acoustic features of the target speech; when the translation mode is the speech translation mode, the 16-layer Conformer module and the speaker feature extraction module can be called based on the speech encoder module to execute a non-streaming coding strategy, such as adjusting the window range of the attention layer in each layer of the Conformer module to include all data blocks in the target speech, so as to encode all data blocks as a whole, thereby outputting the speech content features and acoustic features of the target speech.

[0051] After obtaining the speech content features and acoustic features of the target speech, the large language model in the speech processing model can be loaded; the large language model here is obtained by fine-tuning the pre-trained large language model through the sample speech and its sample translated speech in the speech translation mode and the sample translated speech in the speech simultaneous interpretation mode; the fine-tuning training here can be achieved through supervised fine-tuning of low-rank adaptation (LoRA), etc., so as to achieve a large language model that can output translated text more quickly and efficiently while maintaining the fine-tuning quality.

[0052] After the large language model is loaded into the speech processing model, the input information of the large language model can be determined based on the translation mode, and the input information is input into the large language model. The large language model uses its huge knowledge base and powerful semantic understanding ability to translate the target speech based on the input information to obtain the target translation text of the target speech.

[0053] For example, when the translation mode is the speech simultaneous interpretation mode, the speech content features can be divided into blocks, such as semantic block division, so that each block data can be used as input information of the large language model in turn to perform streaming translation on the target speech, so that the translation text output by each streaming translation can be determined as the target translation text of the target speech; when the translation mode is the speech translation mode, the speech content features can be used as the input information of the large language model to perform the overall translation on the target speech, so that the translation text output by the overall translation can be determined as the target translation text of the target speech; for example, when the translation mode is the speech simultaneous interpretation mode, the speech content features can be used as the input information of the large language model to perform the overall translation on the target speech, so that the translation text output by the overall translation can be determined as the target translation text of the target speech; Block processing, such as semantic block division, is performed to fuse each block data with the translation prompt information corresponding to the voice simultaneous interpretation mode and use them as input information of the large language model in sequence to perform streaming translation on the target voice, so that the translation text output by each streaming translation is determined as the target translation text of the target voice; when the translation mode is the voice translation mode, the overall voice content features and the translation prompt information corresponding to the voice translation mode can be fused as input information of the large language model to perform overall translation on the target voice, so that the translation text output by the overall translation is determined as the target translation text of the target voice, etc. This embodiment does not make any specific restrictions on this.

[0054] It should be noted that during streaming translation, the large language model can translate based on the currently received input information to gradually generate the corresponding translated text. In other words, the translated text corresponding to the currently received input information is immediately output as the target translated text to improve the real-time translation. During the integrated translation process, the large language model can translate based on the complete input information received to output the complete translated text of the target speech as the target translated text.

[0055] After obtaining the target translation text output by the target voice, the voice decoder module in the voice processing model can be loaded; the voice decoder module can implement different decoding strategies according to different translation modes. For example, when the translation mode is the voice simultaneous interpretation mode, it can perform translation speech synthesis and broadcast of the target translation text output by the large language model based on the fixed strategy and the on-screen strategy of the translation result. In other words, it can use streaming speech synthesis technology to perform speech synthesis while receiving the translation text output by the large language model to ensure low-latency output, thereby meeting real-time requirements. When the translation mode is the voice translation mode, the complete translation text output by the large language model is directly translated and broadcasted, etc. This embodiment does not specifically limit this.

[0056] After being loaded into the speech decoder module, the speech decoder module can call the decoding strategy of the corresponding mode according to the translation mode to perform speech synthesis on the target translation text output by the target speech according to the decoding strategy of the corresponding mode to obtain the target translation speech of the target speech.

[0057] The method provided in this embodiment obtains a speech processing model that integrates a speech translation mode or a speech simultaneous interpretation mode through end-to-end training. Within this speech processing model, a speech encoder module dynamically adjusts the encoding strategy according to different translation modes to encode the target speech in different translation modes, thereby obtaining speech content features and acoustic features under different translation modes. Furthermore, input information for a large language model under different translation modes is dynamically determined based on the speech content features, so that the large language model can perform corresponding translations based on the input information under different translation modes, thereby obtaining target translated texts under different translation modes. Furthermore, a speech decoder module dynamically adjusts the decoding strategy according to different translation modes to perform speech synthesis on the target translated texts under different translation modes based on the acoustic features under different translation modes, thereby obtaining target translated speech of the target speech. Thus, through a speech processing model that includes an integrated framework for speech translation and speech simultaneous interpretation, the speech encoding, translation, and speech decoding processes are adapted in combination with the translation mode, avoiding the need to build independent models for different translation modes, effectively reducing deployment costs, and simultaneously sharing the internal knowledge and translation advantages of different translation modes, thereby better optimizing the real-time performance and quality of speech processing.

[0058] In some implementations, step 130 specifically includes: Step 131, when the translation mode is the speech simultaneous interpretation mode, based on the semantic encoder module in the speech processing model, performing semantic content prediction and semantic end symbol prediction on the target speech, and dividing the speech content features into a plurality of semantic feature blocks according to the semantic content prediction results and the semantic end symbol prediction results; Step 132: Based on the large language model, perform streaming translation on the plurality of semantic feature blocks to obtain the target translation text.

[0059] The semantic encoder module here is a module that is triggered and called in the simultaneous speech interpretation mode. It can divide the speech content features output by the speech encoder module in the simultaneous speech interpretation mode into semantic blocks; the semantic encoder module can be constructed based on a multi-layer single-structure network model or a multi-layer hybrid-structure network model, and can be set according to actual needs.

[0060] Figure 2 is a structural diagram of the semantic encoder module provided by the present invention; Figure 2As shown, exemplarily, the semantic encoder module may include a shared encoding module shared with the first 12 layers of the Conformer module of the speech encoder module, a convolutional module (Convd for short), a hybrid model integrating a recurrent neural network and a converter network (RWKV for short), a multi-layer perceptron (MLP for short), a connection temporal classification (CTC for short) module and a multi-classification module.

[0061] The following Figure 2 The semantic encoder module shown describes the method provided in the embodiment.

[0062] It is understandable that in the related art, simultaneous interpretation is performed through a wait-k strategy or a fixed window. If the delay of the wait-k strategy or the fixed window is increased, although it can improve the translation quality, it will cause the translation to be more real-time; if the delay of the wait-k strategy or the fixed window is reduced, it will result in less reference information, making it difficult to ensure the quality of translation, that is, it is difficult to take into account both delay and quality at the same time, and each delay setting requires manual adjustment, which is inflexible and affects the user experience.

[0063] To address this issue, an embodiment of the present invention dynamically divides semantic feature blocks through semantic content prediction results and semantic end symbol prediction results when performing simultaneous interpretation, and performs dynamic window translation based on the divided semantic feature blocks, effectively achieving an automatic optimal balance between delay and quality, and flexibly controlling the information reference range according to the context without human intervention, significantly improving the accuracy and real-time performance of simultaneous interpretation.

[0064] Figure 3 This is the second flow chart of the speech processing method provided by the present invention; Figure 3 As shown, when performing text translation, if the translation mode is simultaneous speech interpretation, the semantic encoder module in the speech processing model can be loaded to share the encoded features of the first 12 layers of the Conformer module of the speech encoder module through the shared encoding module in the semantic encoder module. The encoded features are sequentially extracted through Convd, RWKV, and MLP in the semantic encoder module to obtain deep semantic features. Then, the CTC module performs frame-level semantic content prediction on the target speech based on the deep semantic features to obtain semantic content prediction of the target speech. Furthermore, the multi-classification module performs multi-category semantic end symbol prediction on the target speech based on the deep semantic features to obtain semantic end symbol prediction results of the target speech. Then, based on the semantic content prediction results and the semantic end symbol prediction results, the semantic segmentation boundary in the speech content features is determined, and the speech content features are semantically segmented based on the semantic segmentation boundary to obtain multiple segmented semantic feature blocks.

[0065] After obtaining multiple segmented semantic feature blocks, the large language model can be called to perform streaming translation on the multiple semantic feature blocks, so that the translation text output by each streaming translation is determined as the target translation text. It should be noted that during the streaming translation process, the large language model can directly perform real-time translation based on the currently received semantic feature blocks to gradually generate corresponding translation texts to obtain the target translation text; or the large language model can perform real-time translation based on the currently received semantic feature blocks and historical translation records (such as historically received historical semantic feature blocks and / or translation results of historical semantic feature blocks) to gradually generate corresponding translation texts to obtain the target translation text; or the large language model can perform real-time translation based on the currently received semantic feature blocks, historical translation records, and translation prompts corresponding to the simultaneous voice interpretation mode to gradually generate corresponding translation texts to obtain the target translation text, etc. This embodiment does not specifically limit this.

[0066] This embodiment provides a method for performing dynamic window translation by dividing speech content features into semantic blocks using a semantic encoder module in the speech simultaneous interpretation mode, thereby effectively achieving an automatic optimal balance between delay and quality, and significantly improving the accuracy and real-time performance of simultaneous interpretation.

[0067] In one possible implementation, step 132 specifically includes: For the current translation, obtaining a current semantic feature block corresponding to the current translation from the plurality of semantic feature blocks; Based on the adapter corresponding to the simultaneous speech interpretation mode in the speech processing model and the preset length information, feature fusion is performed on the historical semantic feature block corresponding to the historical translation, the historical translation text of the historical translation, and the current semantic feature block to obtain a first fused feature; Inputting the first fusion feature and the translation prompt word corresponding to the speech simultaneous interpretation mode into the large language model to obtain the translation text of the current translation; The translation text of the current translation is determined as the target translation text.

[0068] Figure 4 This is the third flow chart of the speech processing method provided by the present invention; Figure 4 As shown, when obtaining the translated text in the voice interpretation mode, the following steps can be performed: For the current translation, the current semantic feature block required for execution is obtained, and historical translation records, including the historical semantic feature blocks and translated text corresponding to previous translations of the target speech, are dynamically loaded from the long context module. The long context module dynamically removes speech, semantic feature blocks, and translated text that were stored more recently, based on the outputs of the semantic encoder and speech decoder modules, to optimize storage efficiency and maintain real-time translation.

[0069] The preset length information, the historical semantic feature blocks corresponding to the previous translations, the historical translation texts of the previous translations, and the current semantic feature block are input into an adapter (also called an online adapter) corresponding to the simultaneous speech interpretation mode. The online adapter then fuses the historical semantic feature blocks corresponding to the previous translations, the historical translation texts of the previous translations, and the current semantic feature block based on the preset length information to obtain a first fused feature that is compatible with the simultaneous speech interpretation mode and the input dimensions of the large language model. This better constrains the translation length, effectively controls the translation output length, and better meets the real-time and concise requirements of speech translation. Fusion here can include concatenation or feature encoding, etc., which is not specifically limited in this embodiment.

[0070] Then, the first fusion feature and the translation prompt word corresponding to the speech simultaneous interpretation mode are input into the large language model, so that the translation prompt word corresponding to the speech simultaneous interpretation mode guides the large language model to perform instant translation based on the first fusion feature, outputs the translation text of the current translation in real time, and updates the translation text of the current translation to be determined as the target translation text of the target speech. It should be noted that the translation text of the current translation can include the translation text corresponding to the historical semantic feature block and the translation text corresponding to the current semantic feature block. The translation prompt word corresponding to the speech simultaneous interpretation mode here includes the translation instruction corresponding to the speech simultaneous interpretation mode, and the translation instruction includes at least the translation mode, the source language information of the target speech, and the target language of the translation speech corresponding to the target speech. For example, the translation instruction can be "translate the provided English speech into Chinese according to the manual simultaneous interpretation form" or "translate the provided Chinese speech into English according to the manual simultaneous interpretation form", etc. This embodiment does not specifically limit this.

[0071] According to the above steps, each translation is performed traversally to gradually update the target translation text outputted as the target voice.

[0072] The method provided in this embodiment, in the simultaneous voice interpretation mode, assists the large model in utilizing its long-term memory capacity for translation by maintaining historical translation records, thereby improving the effective use of historical information and the long-term consistency of the context of the translation results. At the same time, by explicitly increasing the input voice length information to constrain the length of the translation output by the large language model, the problem of unstable translation result length is alleviated, thereby significantly improving the accuracy and fluency of the translation and enhancing the user experience.

[0073] In some embodiments, step 140 specifically includes: When the translation mode is the simultaneous voice interpretation mode, for the current translation, based on the voice decoder module, a semantic judgment is performed on the translation text of the current translation. When all text translation data blocks in the translation text of the current translation are determined to be fixed according to the semantic judgment result, speech synthesis is performed on the translation text of the current translation according to the acoustic features to obtain the target translation speech.

[0074] Figure 5 This is a flow chart of the streaming decoding process provided by the present invention. Figure 5 As shown, when the translation mode is the voice simultaneous interpretation mode, the following steps can be performed based on the voice decoder module for streaming decoding: For the current translation, a semantic judgment is performed on the translation text of the current translation, and based on the result of the semantic judgment, it is determined whether to fix all the text translation data blocks in the translation text of the current translation. If it is determined that all the text translation data blocks in the translation text of the current translation are fixed, it is determined that the text semantics of the translation text of the current translation are complete and have high stability. All the text translation data blocks in the translation text of the current translation can be used as fixed results, and speech synthesis is performed on all the text translation data blocks in the translation text of the current translation based on acoustic features to obtain the translation speech of the current translation, and the translation speech of the current translation is updated and determined as the target translation speech of the target speech, and is broadcast in real time. In this way, the fixed result is dynamically determined through semantic judgment, and speech synthesis is performed in combination with the fixed result and acoustic features. In the voice simultaneous interpretation mode, it can be ensured that the semantically complete and stable text can generate speech quickly and accurately, thereby improving the accuracy and fluency of the translation, enhancing the user experience, and meeting the user's demand for real-time voice translation.

[0075] In some embodiments, step 140 further includes: When it is determined according to the semantic judgment result that all text translation data blocks are not to be fixed, determining whether the fixed delay duration corresponding to the current translation is greater than the first delay duration; When it is determined that the fixed delay time is greater than the first delay time, determining whether there is a common prefix between the translation text of the current translation and the historical translation text of the previous translation; When it is determined that the common prefix exists, speech synthesis is performed on the common prefix according to the acoustic features to obtain the target translated speech.

[0076] like Figure 5 As shown, when the translation mode is the voice simultaneous interpretation mode, the following steps can also be performed based on the voice decoder module to perform streaming decoding: For the current translation, if it is determined based on the semantic judgment result that all text translation data blocks in the translation text of the current translation are not fixed, then it is further determined whether the fixed delay time corresponding to the current translation is greater than the first delay time m. If it is further determined that the fixed delay time corresponding to the current translation is less than or equal to the first delay time m, then the current semantic feature block and its translation text are cached, and the translation text of the next semantic feature block output by the large language model is waited for until the fixed delay time is greater than the first delay time m. The first delay time m here can be set according to implementation requirements.

[0077] If it is further determined that the fixed delay time corresponding to the current translation is greater than the first delay time, it is determined whether there is a common prefix between the translation text of the current translation and the historical translation text of the previous translation. When it is determined that there is a common prefix, the common prefix is taken as a fixed result, and the common prefix is speech synthesized according to the acoustic characteristics to obtain the translation speech of the current translation, and the translation speech of the current translation is updated to be the target translation speech of the target speech, and is broadcast in real time. This ensures that when the translation text cannot be fixed and the delay times out, by determining the common prefix, the existing stable translation content can be used for speech synthesis and broadcast, effectively reducing the problem of voice output interruption or incoherence caused by delays and non-fixed translation texts, improving the continuity and comprehensibility of real-time translation, and improving the user experience in instant translation scenarios.

[0078] In some embodiments, step 140 further includes: When it is determined that the common prefix does not exist, speech synthesis is performed on a target number of text translation data blocks ranked top in the translation text of the current translation according to the acoustic features to obtain the target translation speech.

[0079] like Figure 5 As shown, when the translation mode is the voice simultaneous interpretation mode, the following steps can also be performed based on the voice decoder module to perform streaming decoding: When it is further determined that there is no common prefix between the translation text of the current translation and the historical translation text of the previous translation, speech synthesis is performed directly on the target number of text translation data blocks ranked at the top in the translation text of the current translation, that is, 1 / N text translation data blocks, based on the acoustic features to obtain the translation speech of the current translation, and the translation speech of the current translation is updated to be determined as the target translation speech of the target speech, and is broadcast in real time. This ensures that when the translation text cannot be fixed, the delay times out, and there is no common prefix, speech synthesis can be performed on the target number of text translation data blocks in the translation text output in real time by the large language model to control the scope and delay of speech synthesis to a certain extent, ensuring that partial translation speech can still be generated quickly in the absence of a common prefix reference, allowing users to hear partial translation results in time, enhancing the real-time and feedback of the translation, and avoiding users waiting for the complete speech output for a long time.

[0080] In some embodiments, the method further comprises: When the translation mode is the simultaneous voice interpretation mode, for the current translation, when it is determined that the screen-on delay time corresponding to the current translation is greater than the second delay time, the translation text of the current translation is displayed on the display screen.

[0081] like Figure 5 As shown, when the translation mode is the voice interpretation mode, the following steps can be performed based on the voice decoder module to synchronously display the translation content on the screen to further enhance the usability and user experience of the translation: When the translation mode is the voice simultaneous interpretation mode, the screen delay time corresponding to the current translation is monitored in real time. When the screen delay time corresponding to the current translation does not exceed the second delay time k, the screen delay time corresponding to the next translation is continuously monitored until the screen delay time exceeds the second delay time; when the screen delay time corresponding to the current translation exceeds the second delay time, the translation text of the current translation is displayed on the display screen. Specifically, the translation text can be displayed on the display screen in the form of a pop-up window, text box, etc. through a graphical user interface, so that the translation text can be displayed when the screen delay is too long, so that the user can obtain the intermediate translation information in a visual way, enrich the way to obtain translation information, and improve the efficiency and accuracy of the user in obtaining translation information.

[0082] The second delay duration here can be set according to actual needs, wherein the second delay duration can be greater than the first delay duration to ensure audio-visual collaboration.

[0083] In some embodiments, step 130 further includes: When the translation mode is the speech translation mode, based on the adapter corresponding to the speech translation mode in the speech processing model and preset length information, feature fusion is performed on historical speech content features of historical speech, historical translation text of the historical speech, and speech content features of the target speech to obtain a second fused feature; the historical speech is a speech that was translated before the target speech and is associated with the target speech; The second fusion feature and the translation prompt word corresponding to the speech translation mode are input into the large language model to obtain the target translation text.

[0084] like Figure 4 As shown, when performing text translation, if the translation mode is speech translation mode, the target speech can be translated in its entirety. Therefore, there is no need to load the semantic encoder module in the speech processing model. Instead, historical translation records are dynamically loaded directly from the long context module. These historical translation records include the historical speech content features and historical translation texts of historical speech that were translated before the target speech and are related to the target speech. The preset length information, the historical speech content features and historical translation texts of the historical speech, and the speech content features of the target speech are then input into an adapter corresponding to the speech translation mode (also known as an offline adapter). The offline adapter then fuses the historical speech content features and historical translation texts of the historical speech with the speech content features of the target speech based on the preset length information to obtain a second fused feature that is compatible with the speech translation mode and the input dimensions of the large language model. This better constrains the translation length, effectively controls the translation output length, and better meets the real-time and concise requirements of speech translation. Fusion here can include concatenation or feature encoding, etc., which is not specifically limited in this embodiment.

[0085] Then, the second fusion feature and the translation prompt word corresponding to the speech translation mode are input into the large language model, so that the translation prompt word corresponding to the speech translation mode guides the large language model to perform a non-real overall translation based on the second fusion feature to obtain a complete translation text of the target speech output by the large language model, and the complete translation text is directly input into the speech decoder module, so that the speech decoder module performs overall speech synthesis on the complete translation text based on the acoustic features to obtain the target translated speech of the target speech. The translation prompt word corresponding to the speech translation mode here includes a translation instruction corresponding to the speech translation mode. The translation instruction includes at least the translation mode, the source language information of the target speech, and the target language of the translated speech corresponding to the target speech. For example, the translation instruction can be "translate the provided English speech into Chinese" or "translate the provided Chinese speech into English", etc. This embodiment does not specifically limit this.

[0086] The method provided in this embodiment, in the speech translation mode, uses an offline adapter to integrate the historical speech content features and historical translation texts of historical speech and the speech content features of the target speech according to preset length information and input them into a large language model, and combines them with translation prompt words to achieve overall translation of the target speech, and then synthesizes the translated speech through a speech decoder module. This not only improves the effective use of historical information and the long-term consistency of the context of the translation results, but also alleviates the problem of unstable translation result length by explicitly increasing the input speech length information to constrain the length of the translation output by the large language model, thereby significantly improving the accuracy and fluency of the translation and enhancing the user experience.

[0087] In some embodiments, the target speech is obtained based on the following steps: Based on the speech detection module in the speech processing model, removing the silent segments in the original speech of the target speech to obtain multiple target speech segments; According to the silent segments, zero padding is performed between every two adjacent target speech segments to obtain the target speech.

[0088] Figure 6 This is a fourth flow chart of the speech processing method provided by the present invention; Figure 6 As shown, after the original speech is obtained by originally collecting the speech information output by the speaker, in order to improve the efficiency and accuracy of speech translation, the valid speech stream in the original speech can be detected based on the speech detection module in the speech processing model to remove the silent segments in the original speech, so as to remove the invalid silent parts and retain the valid speech segments (that is, the target speech segments), thereby reducing unnecessary processing time.

[0089] During the inference phase, the user's long speech stream needs to be segmented and then fed into the speech encoder module using the segmented, smaller segments. During the training phase, the data used is largely derived from user short sentences for annotation, so the speech encoder module can directly use these short sentences for encoding training. This results in gaps between the data used by the speech encoder module during the inference and training phases, which can easily lead to abnormal encoding results. Therefore, after obtaining multiple valid speech segments during the inference phase, zero padding can be performed between each two adjacent valid speech segments based on their position and length to align them with the original speech. This allows the speech encoder module to view historical speech instead of future speech, resolving the mismatch between the training and decoding phases, preventing speech distortion caused by sudden switching between valid speech segments, and ensuring the continuity and integrity of the speech data. This provides higher-quality and more efficient speech input for subsequent speech translation processing, ultimately obtaining target speech that meets the requirements of speech translation and increasing the robustness of speech processing.

[0090] Accordingly, when performing speech encoding, in order to eliminate the impact of zero-padding segments on translation efficiency and translation quality and improve translation efficiency and quality, the target speech can be first encoded according to the window range corresponding to the translation mode to output a coding feature sequence, and after the encoding is completed, the coding features corresponding to the zero-padding segments in the coding feature sequence are removed to obtain effective speech content features and acoustic features of the target language.

[0091] The method provided in this embodiment is described below with reference to specific examples.

[0092] In speech translation mode, a user-provided speech stream is obtained to obtain the original speech. The speech detection module removes silence segments from the original speech to obtain multiple target speech segments. Zero padding is performed between each adjacent target speech segment to obtain the target speech. The target speech is input into the speech encoder module, which sets a window range for encoding according to the encoding mode corresponding to the speech translation mode. For example, the window range is set to include all data blocks in the target speech. After encoding all data blocks as a whole, the encoding features corresponding to the zero-padding segments are removed to obtain the speech content features and acoustic features of the target speech. After obtaining the speech content features and acoustic features, the historical speech content features of the historical speech, the historical translation text of the historical speech, and the speech content features of the target speech are fused based on the offline adapter and preset length information to obtain a second fused feature. The second fused feature and the translation prompt word corresponding to the speech translation mode are input into the large language model to obtain the target translated text corresponding to the target speech. The target translated text and the acoustic features are then input into the speech decoder module, which synthesizes the target translated speech of the target speech and broadcasts the target translated speech.

[0093] In the simultaneous voice interpretation mode, the voice stream provided by the user is obtained to obtain the original voice, and the silent segments of the original voice are removed through the voice detection module to obtain multiple target voice segments, and zero padding is performed between each two adjacent target voice segments to obtain the target voice. The target voice is input into the voice encoder module so that the voice encoder module sets the encoding window range according to the encoding mode corresponding to the simultaneous voice interpretation mode, such as setting the window range to include three adjacent data blocks, and then performing streaming encoding on the three data blocks within the window range and removing the encoding features corresponding to the zero-padding segments to obtain the voice content features and acoustic features of the target voice. After obtaining the voice content features and acoustic features, the semantic content prediction and semantic end symbol prediction can be further performed on the target voice based on the semantic encoder module, and the voice content features are divided into multiple semantic feature blocks according to the semantic content prediction results and the semantic end symbol prediction results, so as to perform the following streaming translation process on each semantic feature block and output the corresponding translated voice in a streaming manner: In the current translation, based on the online adapter and preset length information, feature fusion is performed on the historical semantic feature blocks corresponding to the previous translations, the historical translation texts of the previous translations, and the current semantic feature blocks to obtain a first fused feature. The first fused feature and the translation prompt words corresponding to the simultaneous voice interpretation mode are input into the large language model to obtain the translated text of the current translation; Based on the speech decoder module, when it is determined that the screen-up delay time corresponding to the current translation is greater than the second delay time, the translated text of the current translation is displayed on the display screen; Based on the speech decoder module, a semantic judgment is performed on the translation text of the current translation. When all text translation data blocks in the translation text of the current translation are fixed according to the semantic judgment result, speech synthesis is performed on the translation text of the current translation according to the acoustic features to obtain the target translation speech; When it is determined based on the semantic judgment result that all text translation data blocks are not fixed, if the fixed delay time corresponding to the current translation is greater than the first delay time, and there is a common prefix between the translation text of the current translation and the historical translation text of the previous translation, speech synthesis is performed on the common prefix based on the acoustic features to obtain the target translation speech; When it is determined not to fix all text translation data blocks based on the semantic judgment result, if the fixed delay time corresponding to the current translation is greater than the first delay time, and there is no common prefix between the translation text of the current translation and the historical translation text of the previous translation, speech synthesis is performed on the target number of text translation data blocks ranked at the top in the translation text of the current translation based on the acoustic features to obtain the target translation speech.

[0094] In summary, the method provided in this embodiment has the following advantages: First, a speech processing model with an integrated framework encompassing both speech translation and simultaneous interpretation adapts the speech encoding, translation, and decoding processes to the appropriate translation mode. This avoids building separate models for different translation modes, effectively reducing deployment costs. Furthermore, the internal knowledge and translation advantages of different translation modes can be shared, thereby better optimizing the real-time performance and quality of speech processing. Second, in simultaneous interpretation mode, the semantic encoder module divides speech content features into semantic blocks for dynamic window translation, effectively achieving an automatic optimal balance between latency and quality, significantly improving the accuracy and real-time performance of simultaneous interpretation. Third, by maintaining historical translation records, we can assist large models in leveraging their long-term memory capabilities for translation, thereby improving the effective use of historical information and enhancing the long-term consistency of the context of the translation results. Fourth, by explicitly increasing the input speech length information to constrain the length of the translation output by the large language model, the problem of unstable translation result length is alleviated. A speaker feature extraction module is introduced at the shallow layer of the speech encoder to perform acoustic features, so that acoustic features can be used for translation speech synthesis during language synthesis, effectively ensuring that the generated translation speech has a natural and fluent speech effect, thereby significantly improving the quality of translation and enhancing the user experience.

[0095] The speech processing device provided by the present invention is described below. The speech processing device described below and the speech processing method described above can be referenced to each other.

[0096] Figure 7 Schematic diagram of the structure of the speech processing device provided by the present invention; Figure 7 As shown, the device includes: The acquisition unit 710 is used to acquire the target speech and the translation mode of the target speech; the translation mode includes a speech translation mode or a speech simultaneous interpretation mode; The encoding unit 720 is used to encode the target speech based on the speech encoder module in the speech processing model and the translation mode to obtain the speech content features and acoustic features of the target speech; The first translation unit 730 is configured to translate the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translation text of the target speech; The second translation unit 740 is configured to perform speech synthesis on the target translation text based on the speech decoder module in the speech processing model, the translation mode, and the acoustic features to obtain a target translation speech of the target speech; The speech processing model is trained based on sample speech, and the sample speech is obtained by training sample translated speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

[0097] The device provided in this embodiment obtains a speech processing model that integrates a speech translation mode or a speech simultaneous interpretation mode through end-to-end training. Within this speech processing model, a speech encoder module dynamically adjusts the encoding strategy according to different translation modes to encode the target speech in different translation modes, thereby obtaining speech content features and acoustic features under different translation modes. Furthermore, input information for a large language model under different translation modes is dynamically determined based on the speech content features, so that the large language model can perform corresponding translations based on the input information under different translation modes, thereby obtaining target translated texts under different translation modes. Furthermore, a speech decoder module dynamically adjusts the decoding strategy according to different translation modes to perform speech synthesis on the target translated texts under different translation modes based on the acoustic features under different translation modes, thereby obtaining target translated speech of the target speech. Thus, through a speech processing model that includes an integrated framework for speech translation and speech simultaneous interpretation, speech encoding, translation, and speech decoding processes are adapted in combination with the translation mode, avoiding the need to build independent models for different translation modes, effectively reducing deployment costs, and simultaneously sharing the internal knowledge and translation advantages of different translation modes, thereby better optimizing the real-time performance and quality of speech processing.

[0098] In some embodiments, the first translation unit is specifically configured to: When the translation mode is the speech simultaneous interpretation mode, based on the semantic encoder module in the speech processing model, semantic content prediction and semantic end symbol prediction are performed on the target speech, and the speech content features are divided into a plurality of semantic feature blocks according to the semantic content prediction results and the semantic end symbol prediction results; Based on the large language model, a plurality of the semantic feature blocks are subjected to streaming translation to obtain the target translation text.

[0099] In some embodiments, the first translation unit is further configured to: For the current translation, obtaining a current semantic feature block corresponding to the current translation from the plurality of semantic feature blocks; Based on the adapter corresponding to the simultaneous speech interpretation mode in the speech processing model and the preset length information, feature fusion is performed on the historical semantic feature block corresponding to the historical translation, the historical translation text of the historical translation, and the current semantic feature block to obtain a first fused feature; Inputting the first fusion feature and the translation prompt word corresponding to the speech simultaneous interpretation mode into the large language model to obtain the translation text of the current translation; The translation text of the current translation is determined as the target translation text.

[0100] In some embodiments, the second translation unit is specifically configured to: When the translation mode is the simultaneous voice interpretation mode, for the current translation, based on the voice decoder module, a semantic judgment is performed on the translation text of the current translation. When all text translation data blocks in the translation text of the current translation are determined to be fixed according to the semantic judgment result, speech synthesis is performed on the translation text of the current translation according to the acoustic features to obtain the target translation speech.

[0101] In some embodiments, the second translation unit is further configured to: When it is determined according to the semantic judgment result that all text translation data blocks are not to be fixed, determining whether the fixed delay duration corresponding to the current translation is greater than the first delay duration; When it is determined that the fixed delay time is greater than the first delay time, determining whether there is a common prefix between the translation text of the current translation and the historical translation text of the previous translation; When it is determined that the common prefix exists, speech synthesis is performed on the common prefix according to the acoustic features to obtain the target translated speech.

[0102] In some embodiments, the second translation unit is further configured to: When it is determined that the common prefix does not exist, speech synthesis is performed on a target number of text translation data blocks ranked top in the translation text of the current translation according to the acoustic features to obtain the target translation speech.

[0103] In some embodiments, the second translation unit is further configured to: When the translation mode is the simultaneous voice interpretation mode, for the current translation, when it is determined that the screen-on delay time corresponding to the current translation is greater than the second delay time, the translation text of the current translation is displayed on the display screen.

[0104] In some embodiments, the first translation unit is further configured to: When the translation mode is the speech translation mode, based on the adapter corresponding to the speech translation mode in the speech processing model and preset length information, feature fusion is performed on historical speech content features of historical speech, historical translation text of the historical speech, and speech content features of the target speech to obtain a second fused feature; the historical speech is a speech that was translated before the target speech and is associated with the target speech; The second fusion feature and the translation prompt word corresponding to the speech translation mode are input into the large language model to obtain the target translation text.

[0105] In some embodiments, the target speech is obtained based on the following steps: Based on the speech detection module in the speech processing model, removing the silent segments in the original speech of the target speech to obtain multiple target speech segments; According to the silent segments, zero padding is performed between every two adjacent target speech segments to obtain the target speech.

[0106] The device provided by the present invention is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific processes and detailed contents, which will not be repeated here.

[0107] Figure 8 An example of a physical structure diagram of an electronic device is shown below. Figure 8 As shown, the electronic device may include: a processor 810 , a communication interface 820 , a memory 830 and a communication bus 840 , wherein the processor 810 , the communication interface 820 and the memory 830 communicate with each other via the communication bus 840 . The processor 810 can call the logic instructions in the memory 830 to execute the speech processing method, which includes: obtaining a target speech and a translation mode of the target speech; the translation mode includes a speech translation mode or a speech simultaneous interpretation mode; encoding the target speech based on the speech encoder module and the translation mode in the speech processing model to obtain speech content features and acoustic features of the target speech; translating the target speech based on the large language model in the speech processing model, the translation mode and the speech content features to obtain a target translated text of the target speech; performing speech synthesis on the target translated text based on the speech decoder module in the speech processing model, the translation mode and the acoustic features to obtain a target translated speech of the target speech; wherein the speech processing model is trained based on sample speech, and sample translated speech of the sample speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

[0108] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech processing method provided by the above methods, which includes: obtaining a target speech and a translation mode of the target speech; the translation mode includes a speech translation mode or a speech simultaneous interpretation mode; based on the speech encoder module and the translation mode in the speech processing model, encoding the target speech to obtain speech content features and acoustic features of the target speech; based on the large language model in the speech processing model, as well as the translation mode and the speech content features, translating the target speech to obtain a target translated text of the target speech; based on the speech decoder module in the speech processing model, as well as the translation mode and the acoustic features, performing speech synthesis on the target translated text to obtain a target translated speech of the target speech; wherein the speech processing model is trained based on sample speech, and sample translated speech of the sample speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

[0110] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech processing method provided by the above-mentioned methods, the method comprising: obtaining a target speech and a translation mode of the target speech; the translation mode comprises a speech translation mode or a speech simultaneous interpretation mode; encoding the target speech based on the speech encoder module and the translation mode in the speech processing model to obtain speech content features and acoustic features of the target speech; translating the target speech based on the large language model in the speech processing model, the translation mode and the speech content features to obtain a target translated text of the target speech; performing speech synthesis on the target translated text based on the speech decoder module in the speech processing model, the translation mode and the acoustic features to obtain a target translated speech of the target speech; wherein the speech processing model is trained based on sample speech, and sample translated speech of the sample speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

[0111] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0112] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A speech processing method, characterized in that: include: Acquiring a target speech and a translation mode of the target speech; The translation mode includes a voice translation mode or a voice simultaneous interpretation mode; Encoding the target speech based on the speech encoder module in the speech processing model and the translation mode to obtain speech content features and acoustic features of the target speech; Translating the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translation text of the target speech; Performing speech synthesis on the target translation text based on the speech decoder module in the speech processing model, the translation mode, and the acoustic features to obtain a target translation speech of the target speech; The speech processing model is trained based on sample speech, and the sample speech is obtained by training sample translated speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

2. The speech processing method according to claim 1, wherein: The method of translating the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translated text of the target speech includes: When the translation mode is the speech simultaneous interpretation mode, based on the semantic encoder module in the speech processing model, semantic content prediction and semantic end symbol prediction are performed on the target speech, and the speech content features are divided into a plurality of semantic feature blocks according to the semantic content prediction results and the semantic end symbol prediction results; Based on the large language model, a plurality of the semantic feature blocks are subjected to streaming translation to obtain the target translation text.

3. The speech processing method according to claim 2, wherein: The step of performing streaming translation on the plurality of semantic feature blocks based on the large language model to obtain the target translation text includes: For the current translation, obtaining a current semantic feature block corresponding to the current translation from the plurality of semantic feature blocks; Based on the adapter corresponding to the simultaneous speech interpretation mode in the speech processing model and the preset length information, feature fusion is performed on the historical semantic feature block corresponding to the historical translation, the historical translation text of the historical translation, and the current semantic feature block to obtain a first fused feature; Inputting the first fusion feature and the translation prompt word corresponding to the speech simultaneous interpretation mode into the large language model to obtain the translation text of the current translation; The translation text of the current translation is determined as the target translation text.

4. The speech processing method according to claim 3, wherein: The method of performing speech synthesis on the target translation text based on the speech decoder module in the speech processing model, the translation mode and the acoustic features to obtain the target translation speech of the target speech includes: When the translation mode is the simultaneous voice interpretation mode, for the current translation, based on the voice decoder module, a semantic judgment is performed on the translation text of the current translation. When all text translation data blocks in the translation text of the current translation are determined to be fixed according to the semantic judgment result, speech synthesis is performed on the translation text of the current translation according to the acoustic features to obtain the target translation speech.

5. The speech processing method according to claim 4, characterized in that: The method further comprises: When it is determined according to the semantic judgment result that all text translation data blocks are not to be fixed, determining whether the fixed delay duration corresponding to the current translation is greater than the first delay duration; When it is determined that the fixed delay time is greater than the first delay time, determining whether there is a common prefix between the translation text of the current translation and the historical translation text of the previous translation; When it is determined that the common prefix exists, speech synthesis is performed on the common prefix according to the acoustic features to obtain the target translated speech.

6. The speech processing method according to claim 5, characterized in that: The method further comprises: When it is determined that the common prefix does not exist, speech synthesis is performed on a target number of text translation data blocks ranked top in the translation text of the current translation according to the acoustic features to obtain the target translation speech.

7. The speech processing method according to claim 3, characterized in that: The method further comprises: When the translation mode is the simultaneous voice interpretation mode, for the current translation, when it is determined that the screen-on delay time corresponding to the current translation is greater than the second delay time, the translation text of the current translation is displayed on the display screen.

8. The speech processing method according to any one of claims 1 to 7, characterized in that: The method of translating the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translated text of the target speech includes: When the translation mode is the speech translation mode, based on the adapter corresponding to the speech translation mode in the speech processing model and preset length information, feature fusion is performed on historical speech content features of historical speech, historical translation text of the historical speech, and speech content features of the target speech to obtain a second fused feature; the historical speech is a speech that was translated before the target speech and is associated with the target speech; The second fusion feature and the translation prompt word corresponding to the speech translation mode are input into the large language model to obtain the target translation text.

9. The speech processing method according to any one of claims 1 to 7, characterized in that: The target speech is obtained based on the following steps: Based on the speech detection module in the speech processing model, removing the silent segments in the original speech of the target speech to obtain multiple target speech segments; According to the silent segments, zero padding is performed between every two adjacent target speech segments to obtain the target speech.

10. A speech processing device, characterized in that: include: an acquiring unit, configured to acquire a target speech and a translation mode of the target speech; The translation mode includes a voice translation mode or a voice simultaneous interpretation mode; An encoding unit, configured to encode the target speech based on the speech encoder module in the speech processing model and the translation mode to obtain speech content features and acoustic features of the target speech; a first translation unit, configured to translate the target speech based on the large language model in the speech processing model, the translation mode, and the speech content features to obtain a target translated text of the target speech; a second translation unit, configured to perform speech synthesis on the target translation text based on the speech decoder module in the speech processing model, the translation mode, and the acoustic features to obtain a target translation speech of the target speech; The speech processing model is trained based on sample speech, and the sample speech is obtained by training sample translated speech in the speech translation mode and sample translated speech in the speech simultaneous interpretation mode.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the speech processing method according to any one of claims 1 to 9 is implemented.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech processing method according to any one of claims 1 to 9 is implemented.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the speech processing method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Voice simultaneous transmission system test method and related device

    CN120913540A

  • Multi-language text adaptive configuration method and electronic equipment

    CN121031624A