Voice simultaneous transmission method and device, related equipment and computer program product
Through the end-to-end speech simultaneous model, combining acoustic features and historical translated text, the translation accuracy and delay problems of the existing speech simultaneous system in long context scenarios are solved, and more efficient speech simultaneous effect is achieved.
Patent Information
- Application Number
- CN202510208516.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
AI Technical Summary
The existing speech simultaneous transmission system has shortcomings in the accuracy and delay of translation, especially in long context scenarios, where the contextual connection is lacking, resulting in poor translation effect.
The end-to-end voice simultaneous model is adopted to obtain the acoustic characteristics of the current voice fragment and the translated text of the historical voice fragment, and combine the set task instructions to translate, to realize the end-to-end voice simultaneous process.
It improves the simultaneous translation effect in long context scenarios, reduces cumulative errors, reduces overall system delays, and enhances the connection between contexts.
Smart Images

Figure CN119993124A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method, apparatus, related equipment and computer program product for simultaneous speech interpretation. Background Art
[0002] Speech interpretation technology, also known as instant speech translation technology, is a technology that converts speech in one language into another language in real time. At present, speech interpretation systems are mainly in cascade form. The cascade method of speech interpretation is a multi-module composition. First, the user's speech is transcribed into text in real time through the streaming speech recognition module, and then the output text is sent to the text translation model to output the target language translation text. On this basis, speech synthesis can be further performed.
[0003] In the cascade framework, each module is trained independently and cannot receive feedback signals from other modules. With the development of large language models and their strong general capabilities, many researchers have fine-tuned large language models through text translation matching with translation instructions, thereby realizing large text translation models. The large text translation model has surpassed the previous small-scale text translation models in terms of effect. Therefore, the text translation module in the cascade method has gradually been replaced by the large text translation model.
[0004] The cascade method based on the large text translation model can fully utilize the translation capabilities of the large model, but the accuracy of the translation semantics is limited by the accuracy of the speech transcription by the speech recognition module. This is also an inherent drawback of the cascade system. The accumulated error has led to semantic deviations in the downstream text translation results to a certain extent. The cascade method also makes the overall delay of the system larger. In addition, the current voice simultaneous interpretation system also lacks contextual connection. The translation effect is poor in multi-round dialogues, speeches, press conferences, and other long context scenarios, and the best translation result cannot be inferred based on context information. Summary of the invention
[0005] In view of the above problems, this application is proposed to provide a method, apparatus, related equipment and computer program product for simultaneous voice interpretation to achieve an end-to-end simultaneous voice interpretation process and improve the simultaneous interpretation effect in long-context scenarios. The specific solution is as follows:
[0006] In a first aspect, a method for simultaneous speech interpretation is provided, comprising:
[0007] Acquiring acoustic features of a current speech segment and a translation text of a historical speech segment before the current speech segment, wherein the current speech segment is a speech segment obtained by segmenting an input speech stream;
[0008] The acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instruction are sent to the configured speech simultaneous interpretation big model to obtain the translation text of the current speech segment output by the speech simultaneous interpretation big model, wherein the set task instruction is used to instruct the big model to perform the translation task from the source language to the target language.
[0009] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the current voice segment is specifically a voice segment obtained after segmenting the input voice stream according to semantic integrity.
[0010] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of segmenting the input speech stream according to semantic integrity to obtain the current speech segment includes:
[0011] The input speech stream is sent to the configured semantic segmentation module, and the semantic segmentation module encodes and decodes the input speech stream to obtain a speech recognition result containing punctuation marks;
[0012] Time-align the speech recognition result with the input speech stream to obtain a timestamp corresponding to each recognition unit token in the speech recognition result;
[0013] The input speech stream is segmented using the timestamps corresponding to the punctuation marks in the speech recognition result as segmentation points to obtain segmented current speech segments.
[0014] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the semantic segmentation module includes an audio encoder, which is used to encode the input speech stream;
[0015] Then, the process of obtaining the acoustic features of the current speech segment includes:
[0016] The acoustic feature sequence of each frame output by the audio encoder is segmented using the timestamp corresponding to the punctuation mark in the speech recognition result as a segmentation point to obtain the acoustic feature of the current speech segment.
[0017] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the acoustic features of the current speech segment, the translated text of the historical speech segment, and the set task instruction are sent to the configured speech simultaneous interpretation large model to obtain the translated text of the current speech segment output by the speech simultaneous interpretation large model, including:
[0018] Sending the acoustic features of the current speech segment to a mapping layer to obtain hidden features after being processed by the mapping layer;
[0019] The translated text of the historical speech segment and the set task instruction are respectively embedded and represented through the embedding layer of the large speech simultaneous interpretation model, and the embedded representation features of the two and the hidden layer features processed by the mapping layer are spliced, and the spliced features are sent to the backbone network of the large speech simultaneous interpretation model for processing to obtain the translated text of the current speech segment output by the large speech simultaneous interpretation model.
[0020] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the semantic segmentation module adopts a neural network structure, and its training process includes:
[0021] Acquire first training data, where the first training data includes a speech sample and a matched recognition result label carrying punctuation marks;
[0022] The semantic segmentation module is trained by combining speech recognition and punctuation prediction tasks with the first training data.
[0023] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the acoustic features of the current speech segment are obtained through an audio encoder;
[0024] The speech simultaneous interpretation large model is a pre-trained large model, and the pre-training process includes:
[0025] Using text translation matching pairs, the general large model is fine-tuned for instruction training to obtain a trained text translation large model, wherein the text translation matching pairs include a source language text and a matching target language translation text;
[0026] Using a text simultaneous interpretation and translation matching pair, the text translation large model is fine-tuned and trained to obtain a trained text simultaneous interpretation and translation large model, wherein the text simultaneous interpretation and translation matching pair includes a source language simultaneous interpretation text and a matched target language simultaneous interpretation and translation text;
[0027] The text simultaneous translation model is used as the initial speech simultaneous translation model, and the speech simultaneous translation training data is used to train the speech simultaneous translation model, the mapping layer and the audio encoder.
[0028] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the text simultaneous translation model is used as the initial speech simultaneous translation model, and the speech simultaneous translation training data is used to train the speech simultaneous translation model, the mapping layer, and the audio encoder, including:
[0029] In the first stage, the parameters of the audio encoder and the large speech simultaneous interpretation model are fixed, and the parameters of the mapping layer are trained using speech simultaneous interpretation training data until a first convergence condition is reached;
[0030] In the second stage, the parameters of the audio encoder and the mapping layer are fixed, and the speech simultaneous interpretation large model is trained using speech simultaneous interpretation training data until a second convergence condition is reached;
[0031] In the third stage, the audio encoder, the mapping layer and the large simultaneous speech interpretation model are jointly trained using simultaneous speech interpretation training data until a third convergence condition is reached.
[0032] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the method further includes:
[0033] Perform speech synthesis on the translated text of the current speech segment and output the synthesized speech.
[0034] In a second aspect, a voice simultaneous interpretation device is provided, comprising:
[0035] An acoustic feature acquisition unit, used to acquire the acoustic features of the current speech segment;
[0036] A historical translation result acquisition unit, used for translating texts of historical voice segments before a current voice segment, wherein the current voice segment is a voice segment obtained by segmenting an input voice stream;
[0037] The simultaneous translation unit is used to input the acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instruction into the configured speech simultaneous interpretation big model, and obtain the translation text of the current speech segment output by the speech simultaneous interpretation big model, wherein the set task instruction is used to instruct the big model to perform the translation task from the source language to the target language.
[0038] In a third aspect, an electronic device is provided, comprising: a memory and a processor;
[0039] The memory is used to store programs;
[0040] The processor is used to execute the program to implement the various steps of the speech simultaneous interpretation method described in any one of the first aspects of the present application.
[0041] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the various steps of the speech simultaneous interpretation method described in any one of the first aspects of the present application are implemented.
[0042] In a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the various steps of the speech simultaneous interpretation method described in any one of the aforementioned first aspects of the present application.
[0043] By means of the above technical scheme, the present application provides an end-to-end large model of simultaneous speech interpretation. On this basis, in order to maintain contextual connectivity and improve the simultaneous translation effect in long-context scenarios, the input speech stream is segmented to obtain the current speech segment and the historical speech segment to be translated, and the acoustic features of the current speech segment are extracted. The acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instructions are organized as the input of the large model of simultaneous speech interpretation, wherein the set task instructions are used to instruct the large model to perform the translation task from the source language to the target language. On this basis, the translation text of the current speech segment can be obtained through the large model of simultaneous speech interpretation. The present application integrates the traditional speech recognition module and the text translation module into an end-to-end large model of simultaneous speech interpretation, realizes end-to-end simultaneous speech interpretation, and solves the drawbacks of the traditional cascaded simultaneous speech interpretation system, such as the semantic deviation of the translation results caused by cumulative errors and the large overall delay of the system. At the same time, by splitting the input voice stream into segments, the large voice simultaneous interpretation model can refer to the translation text of historical voice segments when translating the current voice segment, thereby enhancing the connectivity between contexts and greatly improving the simultaneous interpretation effect in long-context scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0045] Figure 1 A schematic diagram of an implementation system architecture of the speech simultaneous interpretation method provided in an embodiment of the present application;
[0046] Figure 2 A flowchart of a method for simultaneous voice interpretation provided in an embodiment of the present application;
[0047] Figure 3 A schematic diagram of a speech segmentation process based on a semantic segmentation module provided in an embodiment of the present application;
[0048] Figure 4 A schematic diagram of the input and output of a large model of simultaneous speech interpretation provided in an embodiment of the present application;
[0049] Figure 5 A schematic diagram of a process for fine-tuning instruction training of a general large model using text translation matching pairs provided in an embodiment of the present application;
[0050] Figure 6 A schematic diagram of a process for fine-tuning a text translation model using text simultaneous translation matching pairs provided in an embodiment of the present application;
[0051] Figure 7 A schematic diagram of the structure of a voice simultaneous interpretation device provided in an embodiment of the present application;
[0052] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0054] The present application provides a method for simultaneous speech interpretation, which is suitable for simultaneous speech interpretation scenarios and can translate streaming speech in a source language into text in a target language. Furthermore, the text in the target language can also be synthesized into speech to obtain speech in the target language.
[0055] The speech simultaneous interpretation method of the present application can be applied to long-context speech simultaneous interpretation scenarios, that is, speech simultaneous interpretation scenarios with contextual associations, such as dialogues, speeches, press conferences, etc.
[0056] The speech simultaneous interpretation method of the present application can be applied to Figure 1 The system architecture shown in FIG. 1 may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 A server is included as an example for explanation).
[0057] The terminal 100 or the server 200 can be used alone to execute the voice simultaneous interpretation method provided in the embodiment of the present application. In addition, the terminal 100 and the server 200 can also be used in collaboration to execute the voice simultaneous interpretation method provided in the embodiment of the present application.
[0058] In a possible scenario, the terminal 100 may obtain speech in the source language, and perform a speech simultaneous interpretation method on the terminal 100 to obtain a translation text in the target language. Alternatively, the translation text in the target language may be synthesized into speech in the target language.
[0059] In another possible scenario, the terminal 100 can obtain the speech in the source language and send it to the server 200. The server 200 performs a simultaneous voice translation method to translate the speech in the source language into a translation text in the target language. Alternatively, the translation text in the target language can be synthesized into the speech in the target language. The server 200 can return the translation text in the target language or the speech in the target language to the terminal 100.
[0060] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a translation machine, a teaching screen, a wearable device, a vehicle-mounted device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0061] The speech simultaneous interpretation method of the present application adopts an end-to-end speech simultaneous interpretation large model. The speech simultaneous interpretation method may include a model training phase and a model reasoning phase. The model training phase is the process of training the speech simultaneous interpretation large model and other related models; the model reasoning phase is the process of executing the speech simultaneous interpretation method using the trained speech simultaneous interpretation large model and other related models.
[0062] Next, we first introduce a method for simultaneous speech interpretation provided by an embodiment of the present application from the model reasoning stage, and take the method applied to a computer device as an example. The computer device can be specifically Figure 1 The terminal 100 or the system consisting of the terminal 100 and the server 200. Figure 2 The speech simultaneous interpretation method specifically comprises the following steps:
[0063] Step S100: Acquire the acoustic features of the current voice segment and the translated text of the historical voice segment before the current voice segment, wherein the current voice segment is a voice segment obtained by segmenting the input voice stream.
[0064] In the scenario of simultaneous speech interpretation, the streaming input speech can be segmented to obtain the current speech segment to be translated, and the acoustic features of the current speech segment can be further extracted.
[0065] Furthermore, in order to improve the connectivity between contexts and the translation accuracy in long-context scenarios, the translated text of the historical voice segment before the current voice segment is further obtained in this step. The translated text of the historical voice segment can be obtained from the cache, that is, the translated text can be saved after the historical voice segment is translated.
[0066] Step S110: Send the acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instructions to the configured speech simultaneous interpretation large model to obtain the translation text of the current speech segment output by the speech simultaneous interpretation large model.
[0067] Among them, setting the task instruction is used to instruct the large model to perform the translation task from the source language to the target language.
[0068] In this embodiment, an end-to-end speech simultaneous interpretation model is used to perform speech simultaneous interpretation translation tasks. The input of the speech simultaneous interpretation model includes three parts: the acoustic features of the current speech segment, the translation text of the historical speech segment, and the set task instructions.
[0069] The speech simultaneous interpretation big model is a large model that has been trained to perform speech simultaneous interpretation tasks. It inherits the natural language understanding, text generation and other capabilities of the big model, and can further be applied to speech simultaneous interpretation scenarios.
[0070] By adding the translated text of historical speech segments to the input of the speech simultaneous interpretation model, the speech simultaneous interpretation model can combine the translated text of historical speech segments to translate the current speech segment. The speech simultaneous interpretation model can combine the translated text of historical speech segments to achieve coherence and relevance in long-context translation, improve the translation accuracy of the current speech segment, and improve the simultaneous translation effect of long speech.
[0071] In this embodiment, an end-to-end large model of simultaneous speech interpretation is adopted to replace the solution of cascading speech recognition modules and text translation modules in traditional simultaneous speech interpretation systems. This can avoid the defects of the cascaded solution, such as poor translation effect and overall time extension caused by the accumulation of errors.
[0072] In a possible implementation, the method of the present application may further include:
[0073] Perform speech synthesis on the translated text of the current speech segment and output the synthesized speech.
[0074] That is, the output of synthesized speech in the target language can be achieved in the scenario of simultaneous speech interpretation.
[0075] In some embodiments of the present application, a possible implementation process of segmenting an input speech stream is introduced.
[0076] For the input voice stream, a variety of segmentation strategies can be used to segment it into voice segments, thereby obtaining the current voice segment to be translated. For example, the input voice stream is segmented according to a set duration, or silence detection is performed on the input voice stream, and the position where silence exceeds the set duration is used as a cutting point to segment the input voice stream.
[0077] This embodiment further provides a semantic-based speech stream segmentation method, specifically:
[0078] The input speech stream can be segmented according to semantic integrity to obtain segmented speech segments.
[0079] Semantic segmentation is a technology that cuts a long text into multiple semantic sub-segments based on semantic relationships. In order to realize the instant translation function of simultaneous voice interpretation and improve translation accuracy and contextual relevance, the semantic segmentation technology is applied to the voice stream in this embodiment, namely, the voice semantic segmentation technology. It ensures that the content of each voice segment after the input voice stream is segmented corresponds to a semantic segment.
[0080] In a possible implementation, the process of segmenting the input speech stream according to semantic integrity may include:
[0081] S1. Send the input speech stream to the configured semantic segmentation module, and encode and decode the input speech stream through the semantic segmentation module to obtain a speech recognition result containing punctuation marks.
[0082] The semantic segmentation module is configured to process the input speech stream and output a speech recognition result containing punctuation marks.
[0083] Optionally, the semantic segmentation module may adopt a neural network structure. The semantic segmentation module may include an audio encoder and an audio decoder. The audio encoder is used to encode the input voice stream to obtain acoustic features, and the audio decoder is used to decode based on the acoustic features output by the audio encoder to obtain a speech recognition result, which includes punctuation marks and recognized text.
[0084] The audio encoder may adopt a structure with streaming audio encoding capability, such as a streamingconformer. The audio decoder may adopt a ctc decoder or other types of audio decoders.
[0085] For the training process of the semantic segmentation module using a neural network structure, unlike the pure language recognition task, this embodiment can combine the speech recognition and punctuation prediction tasks and use the first training data for training. The first training data includes speech samples and matching recognition result labels carrying punctuation marks.
[0086] Further optionally, before the semantic segmentation module is trained using the first training data, a pre-training process for the audio encoder in the semantic segmentation module may be added.
[0087] Specifically, a large amount of unsupervised data can be used to pre-train the audio encoder through self-supervised learning to improve the robustness and versatility of the audio encoder.
[0088] An optional way to pre-train the audio encoder is as follows:
[0089] Extract the acoustic features of each frame of the audio sample, cluster the acoustic features of each frame of all audio samples, determine the cluster label of each speech frame (i.e. the cluster identifier to which it belongs), and use the cluster label as the sample label corresponding to the speech frame sample. Send the audio sample to the audio encoder to obtain the hidden layer features of each speech frame, predict the category label of each speech frame through a classification layer, and use the sample label of the speech frame as the supervision signal for self-supervised training to obtain the trained audio encoder.
[0090] Of course, the above only illustrates one way to perform unsupervised pre-training on an audio encoder. In addition, other unsupervised training strategies can be set, which are not listed here one by one.
[0091] By pre-training the audio encoder, its robustness and versatility are improved. The pre-trained audio encoder is then combined with other networks in the semantic segmentation module, and the first training data is used to perform fine-tuning training in conjunction with speech recognition and punctuation prediction tasks to obtain a trained semantic segmentation module.
[0092] S2. Time-align the speech recognition result with the input speech stream to obtain the timestamp corresponding to each recognition unit token in the speech recognition result.
[0093] Specifically, the CTC forced alignment algorithm can be used to time-align the speech recognition results (including punctuation marks and recognized text) with the input speech stream, that is, to determine the timestamp corresponding to each recognition unit token (token can be a character or punctuation mark) in the speech recognition result. CTC forced alignment means that in the sequence classification task, the connectionist temporal classification (CTC) technology is used to achieve automatic alignment between the input sequence and the output label sequence.
[0094] S3. Using the timestamps corresponding to the punctuation marks in the speech recognition results as segmentation points, the input speech stream is segmented to obtain the segmented current speech segments.
[0095] After the above steps, the timestamps corresponding to the punctuation marks in the speech recognition result can be obtained, and then the timestamps corresponding to the punctuation marks can be used as segmentation points to segment the input speech stream to obtain the segmented current speech segments.
[0096] Reference Figure 3 , which illustrates a speech segmentation process based on a semantic segmentation module.
[0097] The input speech stream first passes through the two-dimensional convolution module conv2d to extract hidden layer features, is sent to the audio encoder streaming conformer to extract acoustic features, and then is sent to the audio decoder ctc decoder for decoding to obtain the recognition result including punctuation marks: "The morning sun shines through the gaps in the curtains and gently awakens the sleeping town."
[0098] Furthermore, the recognition result is aligned with the input voice stream using the CTC forced alignment algorithm to obtain the timestamp information of each token in the recognition result.
[0099] The segmentation point can be determined according to the timestamp of the punctuation mark, and the input voice stream can be split into two voice segments.
[0100] In a possible implementation, the process of obtaining the acoustic features of the current speech segment in step S100 in the aforementioned embodiment may specifically be to use the timestamps corresponding to the punctuation marks in the speech recognition results output by the semantic segmentation module as the segmentation points. <seg>", segment the acoustic feature sequence of each frame output by the audio encoder to obtain the acoustic features of the current speech segment chunk.
[0101] In some embodiments of the present application, Figure 4 An optional implementation scheme of the aforementioned step S110 is introduced.
[0102] After the acoustic features of the current speech segment chunk are obtained through the audio encoder, the acoustic features of the current speech segment are sent to the mapping layer (project layer) to obtain the hidden layer features after processing by the mapping layer.
[0103] The mapping layer can convert the acoustic features of the current speech segment into hidden features that can be understood by the large model, that is, to achieve the alignment of text features and speech features. The parameters of the mapping layer can be updated during the training phase of the large speech interpretation model, and the parameters of the mapping layer are fixed after the training is completed.
[0104] Furthermore, the translation text of the historical speech segment and the set task instructions are embedded and represented respectively through the embedding layer of the large speech interpretation model, and the embedded representation features of the two and the hidden layer features processed by the mapping layer are concatenated. The concatenated features are sent to the backbone network (LLM) of the large speech interpretation model for processing to obtain the output translation text of the current speech segment.
[0105] Among them, the input of the large model backbone network LLM can be expressed as the following formula:
[0106] x llm_input = concat(x prompt, x chunk_speech_hidden, x context )
[0107] x prompt = embedding(prompt_token)
[0108] x chunk_speech_hidden = project layer(chunk_speech_hidden)
[0109] x context = embedding(...,context chunk t-1 ,context chunk t )
[0110] The current speech segment chunk is defined as the t+1th chunk, and the historical speech segments before the current speech segment are the tth chunk and several chunks before it. The number of historical speech segment chunks can be set as needed. For example, M chunks are set or all speech segments before the current speech segment in the input speech stream are taken as historical speech segments, where the i-th speech segment chunk is represented as chunk i .x prompt represents the embedded representation of the task instruction prompt, x chunk_speech_hidden Represents the acoustic features of the current speech segment chunk after being processed by the project layer, x context Embedded representations of translated text representing historical speech segments.
[0111] In some possible solutions, a speech synthesizer (TTS) can be connected after the large speech interpretation model to perform speech synthesis on the output of the large speech interpretation model to obtain the synthesized speech of the target language.
[0112] The acoustic features of the current speech segment chunk may be acquired through a configured audio encoder.
[0113] In this embodiment, the acoustic features of the current speech segment are sent to the mapping layer for feature conversion, and the acoustic features are aligned to the text feature space, which is convenient for the large model to understand. Furthermore, the embedding layer extracts the embedded representation of the translated text of the historical speech segment and the embedded representation of the set task instruction, and the three are spliced and sent to the large model backbone network for processing, so as to obtain the output translated text of the current speech segment, realizing end-to-end simultaneous speech interpretation.
[0114] In some embodiments of the present application, the training process of the large simultaneous speech interpretation model is further described.
[0115] In one possible implementation, fine-tuning training for the speech simultaneous interpretation task can be performed directly on the general large model to obtain a trained speech simultaneous interpretation large model.
[0116] In another possible implementation, in order to better improve the accuracy of the speech simultaneous interpretation large model translation, this embodiment does not directly perform fine-tuning training for the speech simultaneous interpretation task on the general large model. Due to the high cost of data construction for speech simultaneous interpretation and the small amount of data resources, this embodiment proposes to first train the text simultaneous interpretation translation large model, and further train the speech simultaneous interpretation large model on this basis.
[0117] The specific steps may include:
[0118] S1. First, use text translation matching pairs to perform instruction fine-tuning training on the general large model to obtain the trained text translation large model.
[0119] The text translation matching pair includes the source language text and the matching target language translation text. Text translation matching pairs can be obtained in large quantities, and the acquisition cost is relatively low. Therefore, a large number of text translation matching pairs can be used to fine-tune the instructions of the general large model to obtain a text translation large model. It has more excellent text translation capabilities.
[0120] Combination Figure 5 As shown, it illustrates a process of fine-tuning instructions for a general large model using text translation matching pairs.
[0121] Figure 5 The text translation matching pair shown in is as follows: Chinese sample: "The morning sun filters through the cracks in the curtains, gently walking up the sleeping town."; the corresponding English translation text: "The morning sun filters through the cracks in the curtains, gently walking up the sleeping town.".
[0122] For the Chinese sample, the current segment is taken as "gently awakened the sleeping town." The English translation result corresponding to the historical segment "The morning sun filters through the cracks in the curtains" is "The morning sun filters through the cracks in the curtains," and the task instruction is set to "translate Chinese into English". Figure 5 The method shown is sent to the big model, and the English translation result of the current segment output by the big model is obtained as "gently walking up the sleeping town.".
[0123] S2. Furthermore, a small number of text simultaneous translation matching pairs can be obtained, and the large text translation model can be fine-tuned by instruction training to obtain a trained large text simultaneous translation model.
[0124] The text simultaneous translation matching pair includes the source language simultaneous translation text and the matching target language simultaneous translation text. Compared with the text translation matching pair, the text simultaneous translation matching pair has fewer data resources. Therefore, in this step, a small number of text simultaneous translation matching pairs can be obtained to further fine-tune the instructions of the aforementioned text translation model, thereby obtaining a high-quality text simultaneous translation model with text simultaneous translation capabilities.
[0125] Combination Figure 6 As shown, it illustrates a process of fine-tuning instructions for a large text translation model using text simultaneous translation matching pairs.
[0126] Figure 6 The text simultaneous interpretation and translation matching pair shown in the figure is as follows: Chinese sample: "Gently --- awakened the sleeping town."; the corresponding English translation text: "gently --- walking up the sleeping town.". Among them, "---" represents a short silent segment in the simultaneous interpretation scene speech.
[0127] For the Chinese sample, take the current segment as "awakened the sleeping town.", the historical segment "gently" corresponds to the English translation result "gently", set the task instruction to "translate Chinese into English", and follow Figure 6 The method shown is sent to the big model, and the English translation result of the current segment output by the big model is obtained as "walking up the sleeping town.".
[0128] S3. Use the text simultaneous translation model as the initial speech simultaneous translation model, and use the speech simultaneous translation training data to train the speech simultaneous translation model, mapping layer and audio encoder.
[0129] Compared with directly using a general big model as the initial speech simultaneous interpretation big model, this embodiment uses a text simultaneous interpretation translation big model as the initial speech simultaneous interpretation big model, which can better utilize text information and improve the translation effect of the speech simultaneous interpretation big model.
[0130] Among them, the speech simultaneous interpretation training data includes: speech segment chunks in the source language and the corresponding target language translation text labels. During training, the acoustic features of the current speech segment chunk are extracted through the audio encoder and sent to the mapping layer. The embedding layer of the speech simultaneous interpretation large model extracts the embedded representation of the translation text of the historical speech segment chunk before the current speech segment chunk, as well as the embedded representation of the set task instructions. After splicing, they are sent to the backbone network of the speech simultaneous interpretation large model, and finally the translation text of the current speech segment chunk output by the model is obtained. The value of the loss function is calculated with the model output result and the translation text label corresponding to the current speech segment chunk in the speech simultaneous interpretation training data, and the relevant network parameters are updated according to the value of the loss function.
[0131] When acquiring speech simultaneous interpretation training data, in order to better use a large amount of speech recognition result data without punctuation marks, the present application can use a post-processing tool to predict punctuation marks on the speech recognition results, and then take out each recognition result segment according to the results of the punctuation prediction. The speech segment chunk corresponding to each recognition result segment is further obtained using the CTC-based forced alignment scheme mentioned in the aforementioned embodiment. In this way, the input chunk information for speech simultaneous interpretation task training can be obtained. For the target language translation text label corresponding to each speech segment chunk, the text simultaneous interpretation translation large model trained in the aforementioned embodiment can be used to translate the recognition result segment corresponding to each speech segment chunk to obtain the target language translation text label. At this point, speech simultaneous interpretation training data can be obtained.
[0132] In a possible implementation, multiple training strategies may be used in the process of training the large speech interpretation model, mapping layer and audio encoder using speech interpretation training data.
[0133] Exemplarily, this embodiment provides a three-stage training strategy:
[0134] First, in order to better align the speech and text spaces, in the first stage, the present application fixes the parameters of the audio encoder and the speech simultaneous interpretation model, and uses the speech simultaneous interpretation training data to train the parameters of the mapping layer until the first convergence condition is reached.
[0135] In the second stage, the parameters of the audio encoder and the mapping layer are fixed, and the large simultaneous speech interpretation model is trained using the simultaneous speech interpretation training data until the second convergence condition is reached.
[0136] When training the large model of simultaneous speech interpretation, the LoRA method can be used to update the parameters, or other methods can be used to update the parameters of the large model of simultaneous speech interpretation. Through the second stage of training, the original capabilities of the large model of simultaneous text interpretation can be guaranteed, while also adapting to the scenario of simultaneous speech interpretation.
[0137] In the third stage, all frozen parameters are released, and the audio encoder, mapping layer and speech interpretation large model are jointly trained using speech interpretation training data until the third convergence condition is reached.
[0138] The first, second and third convergence conditions mentioned above can be set by the user.
[0139] Through the above three-stage training strategy, the training of the large speech simultaneous interpretation model, mapping layer and audio encoder can be better completed.
[0140] The following is a description of the voice simultaneous interpretation device provided in an embodiment of the present application. The voice simultaneous interpretation device described below and the voice simultaneous interpretation method described above can be referenced to each other.
[0141] See also Figure 7 , Figure 7 The present invention is a schematic diagram of the structure of a voice simultaneous interpretation device disclosed in an embodiment of the present application.
[0142] like Figure 7 As shown, the device may include:
[0143] An acoustic feature acquisition unit 11, used to acquire the acoustic features of the current speech segment;
[0144] A historical translation result acquisition unit 12 is used for translating texts of historical speech segments before the current speech segment, wherein the current speech segment is a speech segment obtained by segmenting the input speech stream;
[0145] The simultaneous translation unit 13 is used to send the acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instruction to the configured speech simultaneous interpretation big model, and obtain the translation text of the current speech segment output by the speech simultaneous interpretation big model, wherein the set task instruction is used to instruct the big model to perform the translation task from the source language to the target language.
[0146] In a possible implementation, the device of the present application may further include:
[0147] The speech segmentation unit is used to segment the input speech stream according to semantic integrity to obtain the current speech segment.
[0148] In a possible implementation, the process of the speech segmentation unit segmenting the input speech stream according to semantic integrity to obtain the current speech segment includes:
[0149] The input speech stream is sent to the configured semantic segmentation module, and the input speech stream is encoded and decoded by the semantic segmentation module to obtain the speech recognition result containing punctuation marks;
[0150] Time-align the speech recognition result with the input speech stream to obtain a timestamp corresponding to each recognition unit token in the speech recognition result;
[0151] The input voice stream is segmented using the timestamps corresponding to the punctuation marks in the voice recognition result as segmentation points to obtain segmented current voice segments.
[0152] In a possible implementation, the semantic segmentation module includes an audio encoder for encoding the input speech stream; the process of the acoustic feature acquisition unit acquiring the acoustic features of the current speech segment includes:
[0153] The acoustic feature sequence of each frame output by the audio encoder is segmented using the timestamp corresponding to the punctuation mark in the speech recognition result as a segmentation point to obtain the acoustic feature of the current speech segment.
[0154] In a possible implementation, the simultaneous translation unit sends the acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instruction to the configured speech simultaneous interpretation big model to obtain the translation text of the current speech segment output by the speech simultaneous interpretation big model, including:
[0155] Sending the acoustic features of the current speech segment to a mapping layer to obtain hidden features after being processed by the mapping layer;
[0156] The translated text of the historical speech segment and the set task instruction are respectively embedded and represented through the embedding layer of the large speech simultaneous interpretation model, and the embedded representation features of the two and the hidden layer features processed by the mapping layer are spliced, and the spliced features are sent to the backbone network of the large speech simultaneous interpretation model for processing to obtain the translated text of the current speech segment output by the large speech simultaneous interpretation model.
[0157] In a possible implementation, the semantic segmentation module adopts a neural network structure, and the device of the present application may further include:
[0158] The semantic segmentation module training unit is used to train the semantic segmentation module. The training process includes:
[0159] Acquire first training data, where the first training data includes a speech sample and a matched recognition result label carrying punctuation marks;
[0160] The semantic segmentation module is trained by combining speech recognition and punctuation prediction tasks with the first training data.
[0161] In a possible implementation, the acoustic features of the current speech segment are obtained through an audio encoder; the speech simultaneous interpretation large model is a pre-trained large model, and the device of the present application may further include:
[0162] The large model training unit is used to train the large speech simultaneous interpretation model. The training process includes:
[0163] Using text translation matching pairs, the general large model is fine-tuned for instruction training to obtain a trained text translation large model, wherein the text translation matching pairs include a source language text and a matching target language translation text;
[0164] Using text simultaneous interpretation and translation matching pairs, the text translation large model is fine-tuned and trained to obtain a trained text simultaneous interpretation and translation large model, wherein the text simultaneous interpretation and translation matching pairs include a source language simultaneous interpretation text and a matched target language simultaneous interpretation and translation text;
[0165] The text simultaneous translation model is used as the initial speech simultaneous translation model, and the speech simultaneous translation training data is used to train the speech simultaneous translation model, the mapping layer and the audio encoder.
[0166] In a possible implementation, the large model training unit uses the text simultaneous translation large model as the initial speech simultaneous translation large model, adopts speech simultaneous translation training data, and trains the speech simultaneous translation large model, the mapping layer, and the audio encoder, including:
[0167] In the first stage, the parameters of the audio encoder and the large speech simultaneous interpretation model are fixed, and the parameters of the mapping layer are trained using speech simultaneous interpretation training data until a first convergence condition is reached;
[0168] In the second stage, the parameters of the audio encoder and the mapping layer are fixed, and the speech simultaneous interpretation large model is trained using speech simultaneous interpretation training data until a second convergence condition is reached;
[0169] In the third stage, the audio encoder, the mapping layer and the large simultaneous speech interpretation model are jointly trained using simultaneous speech interpretation training data until a third convergence condition is reached.
[0170] In a possible implementation, the device of the present application further includes:
[0171] The speech synthesis unit is used to perform speech synthesis on the translation text of the current speech segment and output the synthesized speech.
[0172] The present application also provides an electronic device in an embodiment. Figure 8 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, tablet computers, translation machines, teaching large screens, wearable devices, etc. Figure 8 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0173] like Figure 8 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603, so as to implement the speech simultaneous interpretation method of the aforementioned embodiment of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0174] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0175] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the speech simultaneous interpretation methods provided in the embodiments of the present application.
[0176] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any one of the simultaneous voice interpretation methods provided in the embodiment of the present application.
[0177] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0178] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0179] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0180] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
[0181] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.< / seg>
Claims
1. A method for simultaneous speech interpretation, characterized in that: include: Acquiring acoustic features of a current speech segment and a translation text of a historical speech segment before the current speech segment, wherein the current speech segment is a speech segment obtained by segmenting an input speech stream; The acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instruction are sent to the configured speech simultaneous interpretation big model to obtain the translation text of the current speech segment output by the speech simultaneous interpretation big model, wherein the set task instruction is used to instruct the big model to perform the translation task from the source language to the target language.
2. The method according to claim 1, characterized in that The current speech segment is specifically a speech segment obtained by segmenting the input speech stream according to semantic integrity.
3. The method according to claim 2, characterized in that The process of segmenting the input speech stream according to semantic integrity to obtain the current speech segment includes: The input speech stream is sent to the configured semantic segmentation module, and the input speech stream is encoded and decoded by the semantic segmentation module to obtain the speech recognition result containing punctuation marks; Time-align the speech recognition result with the input speech stream to obtain a timestamp corresponding to each recognition unit token in the speech recognition result; The input voice stream is segmented using the timestamps corresponding to the punctuation marks in the voice recognition result as segmentation points to obtain segmented current voice segments.
4. The method according to claim 3, characterized in that The semantic segmentation module includes an audio encoder for encoding the input speech stream; Then, the process of obtaining the acoustic features of the current speech segment includes: The acoustic feature sequence of each frame output by the audio encoder is segmented using the timestamp corresponding to the punctuation mark in the speech recognition result as a segmentation point to obtain the acoustic feature of the current speech segment.
5. The method according to claim 1, characterized in that: The process of sending the acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instruction into the configured speech simultaneous interpretation big model to obtain the translation text of the current speech segment output by the speech simultaneous interpretation big model includes: Sending the acoustic features of the current speech segment to a mapping layer to obtain hidden features after being processed by the mapping layer; The translated text of the historical speech segment and the set task instruction are respectively embedded and represented through the embedding layer of the large speech simultaneous interpretation model, and the embedded representation features of the two and the hidden layer features processed by the mapping layer are spliced, and the spliced features are sent to the backbone network of the large speech simultaneous interpretation model for processing to obtain the translated text of the current speech segment output by the large speech simultaneous interpretation model.
6. The method according to claim 3, characterized in that The semantic segmentation module adopts a neural network structure, and its training process includes: Acquire first training data, where the first training data includes a speech sample and a matched recognition result label carrying punctuation marks; The semantic segmentation module is trained by combining speech recognition and punctuation prediction tasks with the first training data.
7. The method according to claim 5, characterized in that The acoustic features of the current speech segment are obtained through an audio encoder; The speech simultaneous interpretation large model is a pre-trained large model, and the pre-training process includes: Using text translation matching pairs, the general large model is fine-tuned for instruction training to obtain a trained text translation large model, wherein the text translation matching pairs include a source language text and a matching target language translation text; Using a text simultaneous interpretation and translation matching pair, the text translation large model is fine-tuned and trained to obtain a trained text simultaneous interpretation and translation large model, wherein the text simultaneous interpretation and translation matching pair includes a source language simultaneous interpretation text and a matched target language simultaneous interpretation and translation text; The text simultaneous translation model is used as the initial speech simultaneous translation model, and the speech simultaneous translation training data is used to train the speech simultaneous translation model, the mapping layer and the audio encoder.
8. The method according to claim 7, characterized in that The process of using the text simultaneous translation model as the initial speech simultaneous translation model, using speech simultaneous translation training data, and training the speech simultaneous translation model, the mapping layer, and the audio encoder includes: In the first stage, the parameters of the audio encoder and the large speech simultaneous interpretation model are fixed, and the parameters of the mapping layer are trained using speech simultaneous interpretation training data until a first convergence condition is reached; In the second stage, the parameters of the audio encoder and the mapping layer are fixed, and the large speech simultaneous interpretation model is trained using speech simultaneous interpretation training data until a second convergence condition is reached; In the third stage, the audio encoder, the mapping layer and the large simultaneous speech interpretation model are jointly trained using simultaneous speech interpretation training data until a third convergence condition is reached.
9. The method according to any one of claims 1 to 8, characterized in that: Also includes: Perform speech synthesis on the translated text of the current speech segment and output the synthesized speech.
10. A voice simultaneous interpretation device, characterized in that: include: An acoustic feature acquisition unit, used to acquire the acoustic features of the current speech segment; A historical translation result acquisition unit, used for translating texts of historical voice segments before a current voice segment, wherein the current voice segment is a voice segment obtained by segmenting an input voice stream; The simultaneous translation unit is used to input the acoustic features of the current speech segment, the translation text of the historical speech segment and the set task instruction into the configured speech simultaneous interpretation big model, and obtain the translation text of the current speech segment output by the speech simultaneous interpretation big model, wherein the set task instruction is used to instruct the big model to perform the translation task from the source language to the target language.
11. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the speech simultaneous interpretation method as described in any one of claims 1 to 9.
12. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method for simultaneous speech interpretation as described in any one of claims 1 to 9 is implemented.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the method for simultaneous speech interpretation as described in any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Audio processing method and device, storage medium and electronic device
CN120932650A
Audio stream processing method and device
CN121122284A