Speech translation method, device, equipment, medium and product
By combining speech recognition and text encoder technology in the speech translation model, the problem of insufficient accuracy and real-time speech translation in the prior art is solved, and more efficient real-time speech translation is achieved.
Patent Information
- Application Number
- CN202510282277.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-01
AI Technical Summary
The translation method of non-real-time cascading architecture is used in the prior art for speech translation, resulting in the defects of translation accuracy and real-time poor translation.
A speech recognizer based on the speech translation model performs acoustic feature extraction and text mapping of the speech sequence to be translated, combined with a text encoder, feature segmentation and feature encoding of the fusion features, and finally stream decoding is performed by the text decoder to achieve real-time translation.
It effectively compensates for the accumulation of errors in the identification and translation stages, improves translation accuracy and real-timeness, and improves user experience.
Smart Images

Figure CN120236586A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a speech translation method, apparatus, device, medium and product. Background Art
[0002] Speech translation technology can convert the speech content of one language into the text or speech of another language, which greatly promotes the barrier-free communication between people with different language backgrounds, and promotes cross-cultural communication and international cooperation. Therefore, how to perform speech translation in real time and accurately is an important issue that the industry urgently needs to study at present.
[0003] In related technologies, a non-real-time cascade architecture translation method is usually adopted for speech translation. This method mainly first uses a speech recognition model based on a non-real-time encoding and decoding architecture to recognize the source speech to obtain the source language text, and then inputs the source language text into a machine translation model based on a non-real-time encoding and decoding architecture for target language text translation. However, the translation accuracy of this translation method will decrease due to error accumulation, and the machine translation can only be performed after waiting for the recognition of a complete clause to be completed during the translation process, resulting in poor translation real-time performance. Summary of the Invention
[0004] The present invention provides a speech translation method, apparatus, device, medium and product, which are used to solve the defects of poor translation accuracy and translation real-time performance caused by using a non-real-time cascade architecture translation method for speech translation in the prior art, and to improve the translation accuracy and translation real-time performance.
[0005] The present invention provides a speech translation method, including: Performing acoustic feature extraction and text mapping on a speech sequence to be translated based on a speech recognizer of a speech translation model, to obtain the acoustic features and text mapping features of the speech sequence to be translated; Performing feature segmentation and feature encoding on the fusion features formed by fusing the acoustic features and the text mapping features based on a text encoder of the speech translation model, to obtain the encoded features of multiple segments of the speech sequence to be translated; Performing streaming decoding on the encoded features of multiple segments based on a text decoder of the speech translation model, to obtain the real-time translation text of the speech sequence to be translated; wherein, the speech translation model is trained based on a sample speech sequence, and the source language text and target language text corresponding to the sample speech sequence.
[0006] According to a speech translation method provided by the present invention, the text encoder based on the speech translation model performs feature segmentation and feature encoding on the fusion feature formed by fusing the acoustic feature and the text mapping feature, to obtain encoded features of multiple segments of the speech sequence to be translated, including: Based on the adapter of the speech translation model, the acoustic feature and the text mapping feature are fused to obtain the fusion feature; Based on the differentiable splitter of the text encoder, the Bernoulli splitting probability of each frame feature in each fusion feature is calculated, and according to the Bernoulli splitting probability, the fusion feature is segmented to obtain segmented features of multiple segments; Based on the segmentation feature encoder of the text encoder, the segmented features of each segment are encoded to obtain the encoded features of each segment.
[0007] According to a speech translation method provided by the present invention, the segmentation feature encoder based on the text encoder encodes the segmented features of each segment to obtain the encoded features of each segment, including: Based on the segmentation feature encoder, bidirectional attention encoding within each segment is performed on the segmented features of each segment to obtain first encoded features of each segment, unidirectional attention encoding between segments is performed on the segmented features of every two segments to obtain second encoded features of each segment, and the first encoded features and the second encoded features of each segment are fused to obtain the encoded features of each segment.
[0008] According to a speech translation method provided by the present invention, the speech translation model is trained based on the following steps: Based on the sample speech sequence and the source language text, the initialized speech recognition model is trained to obtain a pre-trained speech recognition model; Based on the source language text and the target language text, the initialized translator is trained to obtain a pre-trained translator; Based on the pre-trained speech recognition model and the pre-trained translator, the translation model to be trained is initialized to obtain a student model; The pre-trained speech recognition model is used as the teacher model of the initialized speech recognizer in the student model, and the pre-trained translator is used as the teacher model of the initialized text encoder and the initialized text decoder in the student model, and the student model is trained by knowledge distillation to obtain the speech translation model.
[0009] According to a speech translation method provided by the present invention, the student model is trained by knowledge distillation to obtain the speech translation model, including: Input the sample speech sequence into the pre-trained speech recognition model to obtain the first text recognition result of the sample speech sequence, and input the source language text and the target language text into the pre-trained translator to obtain the first text translation result of the sample speech sequence; Input the sample speech sequence and the target language text into the student model to obtain the sample Bernoulli segmentation probability of each frame feature in the fusion feature of the sample speech sequence, the text mapping feature of the sample speech sequence, and the second text translation result of the sample speech sequence; Determine the segmentation constraint function of the student model according to the sample Bernoulli segmentation probability; Determine the distillation loss function of the student model according to the first text recognition result, the first text translation result, the second text recognition result corresponding to the text mapping feature of the sample speech sequence, and the second text translation result; Train the student model according to the segmentation constraint function and the distillation loss function to obtain the speech translation model.
[0010] According to a speech translation method provided by the present invention, the determining the segmentation constraint function of the student model according to the sample Bernoulli segmentation probability includes: Determine the first constraint function according to the difference between the sum of the sample Bernoulli segmentation probabilities of all frame features in the fusion feature of the sample speech sequence and the number of text units of the source language text; Determine the second constraint function according to the number of feature frames in the fusion feature of the sample speech sequence, the sample Bernoulli segmentation probability, and the number of text units; Determine the segmentation constraint function according to the first constraint function and the second constraint function.
[0011] According to a speech translation method provided by the present invention, the determining the second constraint function according to the number of feature frames in the fusion feature of the sample speech sequence, the sample Bernoulli segmentation probability, and the number of text units includes: Perform a max pooling operation on the ratio between the number of feature frames and the number of text units and the sample Bernoulli segmentation probability to obtain the max pooling result corresponding to each frame feature in the fusion feature of the sample speech sequence; Determine the second constraint function according to the difference between the sum of the max pooling results corresponding to all frame features in the fusion feature of the sample speech sequence and the number of text units.
[0012] The present invention also provides a speech translation device, including: A first feature extraction unit, configured to perform acoustic feature extraction and text mapping on a speech sequence to be translated based on a speech recognizer of a speech translation model, so as to obtain acoustic features and text mapping features of the speech sequence to be translated; A second feature extraction unit, configured to perform feature segmentation and feature encoding on a fused feature formed by fusing the acoustic features and the text mapping features based on a text encoder of the speech translation model, so as to obtain encoded features of multiple segments of the speech sequence to be translated; A translation unit, configured to perform streaming decoding on the encoded features of multiple segments based on a text decoder of the speech translation model, so as to obtain a real-time translation text of the speech sequence to be translated; Wherein, the speech translation model is trained based on a sample speech sequence, and a source language text and a target language text corresponding to the sample speech sequence.
[0013] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the speech translation method as described in any one of the above is implemented.
[0014] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the speech translation method as described in any one of the above is implemented.
[0015] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the speech translation method as described in any one of the above is implemented.
[0016] The speech translation method, device, equipment, medium and product provided by the present invention, under an end-to-end real-time translation speech translation model, use a cross-modal fused feature including acoustic features and text mapping features of a speech sequence to be translated recognized by a speech recognizer as an input to a text encoder, so as to effectively compensate for the cumulative errors existing in recognition and translation in the absence of a scenario and context environment, and improve translation accuracy; moreover, the encoded features of multiple segments output by the text encoder after feature segmentation and feature encoding of the fused feature are streamed into a text decoder, so that the text encoder can perform streaming decoding on the encoded features of multiple segments to obtain a real-time translation text of the speech sequence to be translated. While ensuring the semantic integrity of real-time translation, the first response and character brushing rate of real-time speech translation are greatly improved, and the real-time performance and accuracy of the translation text output are further significantly improved, enhancing the user experience. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of the voice translation method provided by the present invention.
[0019] Figure 2 It is one of the schematic structural diagrams of the voice translation model provided by the present invention.
[0020] Figure 3 It is a schematic structural diagram of the adapter provided by the present invention.
[0021] Figure 4 It is the second schematic structural diagram of the voice translation model provided by the present invention.
[0022] Figure 5 It is a schematic distribution diagram of the feature segmentation structure provided by the present invention.
[0023] Figure 6 It is a schematic structural diagram of the voice translation device provided by the present invention.
[0024] Figure 7 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0026] With the development of speech translation (ST) and simultaneous interpretation technologies, the cascaded translation technology based on automatic speech recognition (ASR) and machine translation (MT) has become the mainstream solution and has been widely applied in different scenarios, promoting cross-cultural communication and international cooperation. With the frequent international exchanges, the demand of enterprises and individuals for high-quality real-time language translation is increasing continuously. In the existing cascaded translation technology, the errors in the recognition stage may lead to inaccurate subsequent translation, thus affecting the final output quality, and at the same time, it cannot meet the demand of low latency. Therefore, how to perform high-quality real-time speech translation in low-latency scenarios is an important topic that the industry urgently needs to study at present.
[0027] In related technologies, a non-real-time cascaded architecture translation method is usually adopted for speech translation, and the method specifically includes the following: First, train an automatic speech recognition (ASR) model constructed based on a non-real-time encoder-decoder architecture to recognize the source speech based on the non-real-time encoder-decoder architecture-based speech recognition model in the application to obtain a text recognition result. The speech recognition model includes an encoder, a decoder, a subtask of connectionist temporal classification (CTC) model, and a weighted finite state transducer (WFST) language model. Among them, the encoder is composed of a visual geometry group (VGG) module and a transformer module, and the decoder is composed of an embedding module and a transformer module.
[0028] In training, it can be to extract speech features from the sample speech sequences in the training data set, such as filter bank features (Fbank), and use the extracted speech features as the input of the encoder in the ASR model to encode to obtain corresponding acoustic features, and use the acoustic features as the input of the WTST model to perform hot word normalization on the acoustic features through secondary scoring of the WTST model. And use the normalized acoustic features as the input of the CTC model to perform text mapping to obtain text mapping features, and the text mapping features can be specifically expressed as , where T is the frame length and D is the dimension; and the acoustic features output by the encoder and the source language text are input into the decoder to identify the text recognition result of the sample speech sequence. Finally, based on the error between the text recognition result corresponding to the text mapping feature and the source language text, and the error between the text recognition result output by the decoder and the source language text, the text recognition error is determined, and the ASR model is trained based on the text recognition error. Among them, the text recognition error is calculated as follows: ; ; ; ; Among them, is the acoustic feature output by the encoder; is the model function of the VGG module in the encoder, is the model function of the transformer module in the encoder; is the Fbank feature of the sample speech sequence; is the output of the CTC model; is the linear function of the CTC model; is the text recognition result output by the decoder; is the model function of the decoder; is the source language text; is the cross-entropy (CE) loss function; is the text recognition result corresponding to the mapped text; and are weight parameters.
[0029] Next, a machine translation (MT) model based on a non-real-time encoder-decoder architecture is trained to perform text translation on the text recognition result of the speech recognition model output by the MT model based on the non-real-time encoder-decoder architecture in the application, so as to obtain the translation text of the speech sequence to be translated. This MT model is composed of a text encoder and a text decoder.
[0030] During the training process, the source language text of the sample speech sequence is input into the text encoder for overall encoding to obtain the intermediate encoding feature , and then the intermediate encoding feature , the target language identifier of the sample speech sequence and target language text Input to the text decoder for decoding to obtain the translated text of the sample speech sequence Then, the error between the translated text of the sample speech sequence and the target language text is calculated to train the MT model constructed by the non-real-time codec architecture.
[0031] ; ; in, is the model function of the text encoder; Model function for the text decoder.
[0032] As can be seen from the above, existing translation methods usually use a cascade of speech recognition models and machine translation models to achieve speech translation, resulting in the machine translation model being able to rely only on the text recognition results output by the speech recognition model for translation, and the recognition errors of the speech recognition module will lead to error accumulation in the machine translation module, thereby greatly reducing the quality of speech translation; in addition, during the translation process, the machine translation model needs to wait for the recognition of the complete sentence to be completed before it can perform machine translation as a whole. If the complete sentence is too long, it will greatly increase the first word delay and refresh rate, affecting the real-time performance of the translation.
[0033] In this regard, in order to better realize the real-time translation of multilingual speech and improve the translation accuracy and translation response speed, the present application provides a speech translation method. Unlike the non-real-time cascade speech translation scheme, this method fuses the acoustic features of the speech sequence to be translated and the text mapping features, two cross-modal fusion features, to perform end-to-end streaming translation of speech to text in multiple languages to obtain real-time translated text, effectively avoiding the accumulation of errors in the recognition and translation stages, and ensuring the real-time nature of the translation and the decoding speed, thereby effectively improving the translation accuracy and translation response speed.
[0034] Figure 1 It is a flow chart of the speech translation method provided by the present invention, such as Figure 1 As shown, the method includes step 110, step 120, and step 130.
[0035] Step 110, a speech recognizer based on a speech translation model performs acoustic feature extraction and text mapping on the speech sequence to be translated, to obtain acoustic features and text mapping features of the speech sequence to be translated; Step 120, based on the text encoder of the speech translation model, feature segmentation and feature encoding are performed on the fusion feature formed by fusing the acoustic feature with the text mapping feature to obtain encoding features of multiple segments of the speech sequence to be translated; Step 130: Based on the text decoder of the speech translation model, perform streaming decoding on the multiple segmented encoded features to obtain the real-time translation text of the speech sequence to be translated; Among them, the speech translation model is trained based on sample speech sequences, as well as the source language text and target language text corresponding to the sample speech sequences.
[0036] It can be understood that in the related art, speech translation is performed through a non-real-time cascaded architecture translation method. The translation accuracy will decrease due to error accumulation, and the machine translation cannot be performed until the recognition of the complete clause is completed during the translation process, resulting in slow translation speed. Moreover, multiple modules such as speech recognition and machine translation are connected in series and interact with each other. Due to the independence between modules, the processing steps and computational latency are greatly increased, making the model difficult to maintain and update, and affecting the translation speed and user experience.
[0037] To address this problem, the embodiment of the present application establishes an end-to-end real-time speech translation model that at least includes a speech recognizer, a text encoder, and a text decoder, so as to recognize and output the acoustic features and text mapping features of the speech sequence to be translated through the speech recognizer, and use the cross-modal fusion features that fuse the acoustic features and text mapping features as the text encoder, effectively making up for the error accumulation in the recognition and translation stages. Through the text encoder, perform feature segmentation and feature encoding on the cross-modal fusion features, so as to segment the fusion features according to the acoustic word boundaries, while retaining the integrity of the acoustic information and semantic information, facilitating the text decoder to perform streaming decoding on the segmented encoded features, greatly improving the first response and character refresh rate of real-time speech translation, and thus effectively improving the translation accuracy and translation real-time performance. At the same time, compared with the large and difficult-to-maintain cascaded speech recognition and machine translation solutions, the speech translation model only includes a small number of modules, optimizing the model structure, reducing the number of model parameters, and being more easily implemented on end-to-end devices, effectively improving the usage efficiency of the model, enhancing the model performance, and reducing the maintenance cost.
[0038] Optionally, before performing step 110, the speech translation model can be obtained through end-to-end training. Specifically, an initial translation model can be constructed first. The initial translation model here at least includes an initial speech recognizer, an initial text encoder, and an initial text decoder. For example, it can also include an adapter that can fuse acoustic features and text mapping features, etc. The present embodiment does not make specific limitations on this.
[0039] Among them, the initialized speech recognizer can be a model that is prepared for extracting the acoustic features and text mapping features of a speech sequence after parameter initialization, or a pre-trained model with the functions of extracting the acoustic features and text mapping features of a speech sequence; similarly, the initialized text encoder can be a model that is prepared for feature segmentation and feature encoding after parameter initialization, or a pre-trained model with the functions of feature segmentation and feature encoding; the initialized text decoder can be a model that is prepared for streaming decoding after parameter initialization, or a pre-trained model with the function of streaming decoding. This embodiment does not make specific limitations on this.
[0040] In addition, a large number of sample data in different translation scenarios can be collected, specifically including sample speech sequences, source language texts corresponding to the sample speech sequences, and target language texts. For example, sample speech sequences can be collected, and the source language texts and target language texts corresponding to the sample speech sequences can be obtained through manual annotation; another example is that the source language text or the target language text can be collected first, and then the sample speech sequences corresponding to the source language text or the target language text can be obtained through manual recording or speech synthesis.
[0041] The source language text therein is the language before the translation of the sample speech sequence, which can be Chinese, English, Japanese, etc.; the target language text is the language after the translation of the sample speech sequence, that is, the language to which the translation belongs, which is a language different from the source language. For example, when the source language and the target language are English and Chinese respectively, English-to-Chinese translation can be realized.
[0042] Subsequently, using the sample speech sequence, as well as the source language text and target language text corresponding to the sample speech sequence, the initialized translation model is trained end-to-end to obtain a speech translation model that can perform accurate and efficient end-to-end real-time speech translation. During the training process, it can be achieved through knowledge distillation joint training of a multi-teacher model (such as a multi-teacher model formed by a teacher model constructed by a pre-trained speech recognition model and a teacher model constructed by a pre-trained speech recognition model), or through joint training of multiple tasks (such as a multi-task formed by a speech recognition task and a text translation task). This embodiment does not make specific limitations on this.
[0043] In practical applications, the trained speech translation model can be loaded, and through the speech recognizer in the speech translation model, the acoustic features and text mapping of the speech sequence to be translated are extracted, thereby obtaining the acoustic features and text mapping features of the speech sequence to be translated.
[0044] The speech recognizer here can be constructed based on multiple network models. For example, it can be constructed based on an acoustic encoder and a CTC model, or based on WFST or a neural network language model (NNLM), as well as an acoustic encoder and a CTC model. This embodiment does not make specific limitations in this regard.
[0045] For example, when the speech recognizer is constructed based on an acoustic encoder and a CTC model, the feature extraction steps of the speech recognizer specifically include: first, encoding the speech features of the speech sequence to be translated into hidden layer features through the acoustic encoder to obtain the acoustic features of the speech sequence to be translated, and then applying the acoustic features through the CTC model for source language text mapping to obtain the text mapping features of the speech sequence to be translated.
[0046] Figure 2 is one of the schematic structural diagrams of the speech translation model provided by the present invention; as Figure 2 shown, when the speech recognizer is constructed based on WFST or NNLM, as well as an acoustic encoder and a CTC model, the feature extraction steps of the speech recognizer specifically include: first, encoding the speech features of the speech sequence to be translated into hidden layer features through the acoustic encoder to obtain the acoustic features of the speech sequence to be translated, and then after performing feature transformation on the acoustic features through the time series processing layer and the linear layer in the CTC model, and then performing hot word regularization through WFST or NNLM, and then performing source language text mapping through the classification layer (such as the Softmax layer) in the CTC model to obtain the text mapping features of the speech sequence to be translated.
[0047] The speech features here can be Fbank features, Mel frequency cepstral coefficients, etc. This embodiment does not make specific limitations in this regard. The acoustic features here can be multiple of phoneme features, prosody features, emotion features, context features, etc., so that real-time translation can be performed on the basis of rich acoustic features during the text translation process, thereby improving the translation quality.
[0048] After obtaining the acoustic features and text mapping features of the speech sequence to be translated output by the speech recognizer, multi-modal feature fusion can be performed on the acoustic features and text mapping features to form fusion features. In this process, it can be a fusion module (such as an adapter, etc.) connected between the speech recognizer and the text encoder that performs multi-modal feature fusion on the acoustic features and text mapping features to form fusion features; it can also be a fusion model built in the input end of the text encoder that performs multi-modal feature fusion on the acoustic features and text mapping features to form fusion features, etc. This embodiment does not make specific limitations in this regard.
[0049] After obtaining the fusion features formed by fusing the acoustic features and the text mapping features, the text encoder of the speech translation model can be used to perform feature segmentation on the fusion features, so as to divide the multi-frame features in the fusion features into multi-segmented cutting features according to the corresponding time boundaries, and perform feature encoding on the cutting features of each segment to obtain the encoded features of multiple segments of the speech sequence to be translated. The feature encoding here includes but is not limited to inter-segment encoding and / or intra-segment encoding. The feature segmentation here can be realized by the time boundary predicted by information entropy windowing, or can be realized by the time boundary determined by dynamic time warping, etc. This embodiment does not make specific limitations on this.
[0050] After obtaining the encoded features of multiple segments of the speech sequence to be translated, the text decoder of the speech translation model can be used to perform streaming decoding on the encoded features of multiple segments, and after processing the decoding results through a linear layer and a softmax layer, the real-time translation text of the speech sequence to be translated can be obtained. During the streaming decoding process, the text decoder can decode according to the encoded features of the target number of segments received currently and the historical translation results, so as to gradually generate the corresponding translation text, that is, output the translation text corresponding to the encoded features of the target number of segments received currently step by step, so as to improve the decoding rate and translation real-time performance.
[0051] The text decoder here can be constructed by multiple layers of decoders, such as Nx layers, etc. This embodiment does not make specific limitations on this.
[0052] The historical translation result here is the translation information corresponding to the encoded features of the segments received during the historical decoding process, which at least includes the historical translation text generated during the historical decoding process and the target language identifier corresponding to the speech sequence to be translated.
[0053] The method provided in this embodiment, under the speech translation model of end-to-end real-time translation, takes the cross-modal fusion features including the acoustic features and the text mapping features of the speech sequence to be translated recognized by the speech recognizer as the input of the text encoder, so as to effectively make up for the cumulative errors in recognition and translation in the lack of scenario and context environment, and improve the translation accuracy; and, the multiple segmented encoded features output by the text encoder after feature segmentation and feature encoding of the fusion features are streamed into the text decoder, so that the text encoder can perform streaming decoding on the multiple segmented encoded features to obtain the real-time translation text of the speech sequence to be translated. While ensuring the semantic integrity of real-time translation, it greatly improves the first response and typing rate of real-time speech translation, further significantly improves the real-time performance and accuracy of the translation text output, and enhances the user experience.
[0054] In some embodiments, step 120 specifically includes: An adapter based on the speech translation model performs feature fusion on the acoustic features and the text mapping features to obtain the fused features; Based on the differentiable segmenter of the text encoder, calculate the Bernoulli segmentation probability of each frame feature in each of the fused features, and perform feature segmentation on the fused features according to the Bernoulli segmentation probability to obtain multiple segmented segmentation features; Based on the segmentation feature encoder of the text encoder, perform feature encoding on the segmented segmentation features to obtain the encoded features of each segment.
[0055] The adapter here is specifically set between the speech recognizer and the text encoder. After receiving the acoustic features and text mapping features output by the speech recognizer, it performs feature fusion on the acoustic features and text mapping features, and uses the fused features of the multi-modal features that integrate acoustics and semantics obtained thereby as the input of the text encoder for speech translation, thus compensating for the cumulative errors in recognition and translation in the absence of context and context, and further assisting in improving the translation accuracy.
[0056] Figure 3 It is a schematic structural diagram of the adapter provided by the present invention; as Figure 3 As shown, in the feature fusion process, the adapter can first perform feature transformation on the received acoustic features through a mapping layer to obtain transformed acoustic features, and perform soft information combination on the text mapping features and the transformed acoustic features to obtain fused features. Among them, in the soft information combination process, the text mapping features can be fused with the hot word embedding features to obtain soft embedding information, and the soft embedding information is combined with the transformed acoustic features to obtain fused features.
[0057] Figure 4 It is a second schematic structural diagram of the speech translation model provided by the present invention; as Figure 4 As shown, the text encoder includes a differentiable segmenter and a segmentation feature encoder formed by combining Nx layers of attention networks. The differentiable segmenter specifically includes a feedforward neural network (FFN) layer and a Sigmoid activation function layer, where the FFN layer is constructed by a RELU activation function layer and a linear layer.
[0058] During the feature segmentation process, in order to segment the streaming input features, a differentiable segmenter can use the FFN layer and the Sigmoid activation function layer to calculate the Bernoulli segmentation probability of each frame-level feature in each fused feature as the segmentation probability, and determine the segmentation boundary of the fused feature according to the Bernoulli segmentation probability of each frame feature. Based on this segmentation boundary, the fused feature is segmented to obtain multiple segmented features, so as to predict the time boundary of the whole speech word through the Bernoulli segmentation probability. Thus, while ensuring that the feature information entropy remains unchanged, the fused feature is windowed and segmented to effectively ensure that the segmentation boundary is the time boundary of the whole speech word, thereby ensuring the semantic integrity of real-time translation.
[0059] Figure 5 It is a distribution diagram of the feature segmentation structure provided by the present invention; Figure 5 (a) is a distribution diagram of a hard segmentation structure; Figure 5 (b) is a distribution diagram of a differentiable segmentation structure.
[0060] It should be noted that in the actual application of the model, in order to determine the decomposition boundary, the Bernoulli segmentation probability of each frame-level feature can be compared with a preset threshold to encode the segmentation probability of each frame-level feature as , where is 0 or 1, and the fused feature is hard segmented according to 0 or 1. It can also be to sort the Bernoulli segmentation probabilities of each frame-level feature and use the multiple Bernoulli segmentation probabilities ranked at the front as the segmentation boundary for differentiable segmentation. This implementation does not make specific limitations on this. During the model training process, since hard segmentation will cause the model to be unable to perform backpropagation, that is, unable to perform inference learning, and thus unable to update the segmentation probability , therefore, differentiable segmentation needs to be adopted in the training stage to ensure the effective training of the model.
[0061] Among them, the Bernoulli segmentation probability has the following specific calculation formula: ; ); Among them, is the fused feature output by the adapter; is the model function of the adapter; is the text mapping feature output by the CTC model in the speech recognizer; is the acoustic feature output by the acoustic encoder in the speech recognizer; ) is the model function of the differentiable segmenter; is the Bernoulli segmentation probability output by the differentiable segmenter.
[0062] During the feature encoding process, the Bernoulli segmentation probability can be and the fusion features Input into the segmentation feature encoder, so that the segmentation feature encoder performs feature encoding on the segmentation features of each segment, thereby obtaining the encoded features of each segment. The feature encoding here includes encoding between segments and / or encoding within segments.
[0063] Exemplarily, in some embodiments, the segment encoding in step 120 specifically includes: Based on the segmentation feature encoder, perform bidirectional attention encoding within each segment on the segmentation features of each segment to obtain the first encoded feature of each segment, perform unidirectional attention encoding between segments on the segmentation features of every two segments to obtain the second encoded feature of each segment, and perform feature fusion on the first encoded feature and the second encoded feature of each segment to obtain the encoded feature of each segment.
[0064] Optionally, in order to meet the requirements of the encoding stream input in the real-time translation task and also capture a more comprehensive context representation in the segmentation features of the segments, during the segment encoding process, the Bernoulli segmentation probability and the fusion features Input into the segmentation feature encoder, and use the segmentation feature encoder to perform context information extraction by performing bidirectional attention encoding within each segment on the segmentation features of each segment to obtain the first encoded feature of each segment, perform context information extraction by performing unidirectional attention encoding between segments on the segmentation features of every two segments to obtain the second encoded feature of each segment, and then perform feature fusion on the first encoded feature and the second encoded feature of each segment to obtain the encoded feature containing rich context information of each segment. Among them, the encoded features of multiple segments The specific acquisition steps include: ; Among them, is the model function of the segmentation feature encoder.
[0065] The method provided in this embodiment compensates for the cumulative errors in recognition and translation in the absence of context and context environment by introducing an adapter for feature fusion, and realizes probability-based feature segmentation through a differentiable segmenter, so as to realize the complete segmentation of acoustic features and semantic features according to the acoustic word boundaries through information entropy windowing, so as to greatly improve the first response and typing rate of real-time speech translation and simultaneous interpretation while retaining the acoustic and semantic integrity, thereby improving the decoding speed of the model, enhancing the user experience effect, and performing multi-level attention encoding between segments and within segments through the segmentation feature encoder to capture richer context features and further improve the accuracy of translation.
[0066] In some embodiments, the speech translation model is trained based on the following steps: Based on the sample speech sequence and the source language text, train the initialized speech recognition model to obtain a pre-trained speech recognition model; Based on the source language text and the target language text, train the initialized translator to obtain a pre-trained translator; Based on the pre-trained speech recognition model and the pre-trained translator, initialize the translator model to be trained to obtain a student model; Use the pre-trained speech recognition model as the teacher model for the initialized speech recognizer in the student model, and use the pre-trained translator as the teacher model for the initialized text encoder and initialized text decoder in the student model, and perform knowledge distillation training on the student model to obtain the speech translation model.
[0067] Optionally, during the model training process, the sample speech sequence and the source language text may be input into the initialized speech recognition model, so that the encoder of the initialized speech recognition model encodes the sample speech sequence, outputs the acoustic features of the sample speech sequence, and applies the acoustic features based on the CTC model of the initialized speech recognition model for feature mapping to obtain the text mapping features of the sample speech sequence, and applies the acoustic features and the source language text based on the decoder of the initialized speech recognition model for decoding to obtain the text recognition result of the sample speech sequence. Based on the error between the text recognition result corresponding to the text mapping features and the source language text, and the error between the text recognition result output by the decoder and the source language text, train the initialized speech recognition model to obtain a pre-trained speech recognition model that can accurately extract acoustic features and semantic features and perform text recognition.
[0068] Furthermore, input the source language text and the target language text into the initialized translator, so that the encoder of the initialized translator encodes the source language text, outputs the text encoding features, and applies the text encoding features, the target language text, and the target language identifier based on the decoder of the initialized translator for decoding to obtain the text translation result of the sample speech sequence. Based on the error between the text translation result and the target language text, train the initialized translator to obtain a pre-trained translator that can accurately perform feature encoding and decoding and text translation.
[0069] After obtaining the pre-trained speech recognition model and the pre-trained translator, in order to improve the translation quality and the model convergence efficiency, the encoder in the pre-trained speech recognition model can be used to initialize the to-be-trained acoustic encoder in the to-be-trained translation model, and the encoder and decoder in the pre-trained translator can be used to initialize the to-be-trained text encoder and the to-be-trained text decoder in the to-be-trained translation model to obtain the student model. Moreover, the pre-trained speech recognition model is used as the teacher model for initializing the speech recognizer in the student model, and the pre-trained translator is used as the teacher model for initializing the text encoder and the text decoder in the student model. Through the multi-teacher model constructed by the pre-trained speech recognition model and the pre-trained translator, constraint training of knowledge distillation is performed on the student model, so that the student model can quickly learn the high-performance acoustic feature encoding knowledge in the pre-trained speech recognition model and the high-performance text encoding and decoding knowledge in the pre-trained translator. At the same time, end-to-end real-time translation can be realized, thereby ensuring the overall quality of translation while improving the real-time performance of translation.
[0070] In some embodiments, the knowledge distillation training of the student model to obtain the speech translation model includes: Inputting the sample speech sequence into the pre-trained speech recognition model to obtain the first text recognition result of the sample speech sequence, and inputting the source language text and the target language text into the pre-trained translator to obtain the first text translation result of the sample speech sequence; Inputting the sample speech sequence and the target language text into the student model to obtain the sample Bernoulli segmentation probability of each frame feature in the fusion feature of the sample speech sequence, the text mapping feature of the sample speech sequence, and the second text translation result of the sample speech sequence; Determining the segmentation constraint function of the student model according to the sample Bernoulli segmentation probability; Determining the distillation loss function of the student model according to the first text recognition result, the first text translation result, the second text recognition result corresponding to the text mapping feature of the sample speech sequence, and the second text translation result; Training the student model according to the segmentation constraint function and the distillation loss function to obtain the speech translation model.
[0071] Optionally, during the knowledge distillation training process, the sample speech sequence may be first input into the pre-trained speech recognition model, and after the encoder and decoder in the pre-trained speech recognition model perform encoding and decoding processing on the sample speech sequence, a first text recognition result of the sample speech sequence is output; and, the source language text and the target language text are input into the pre-trained translator, and after the encoder and decoder in the pre-trained translator perform encoding and decoding processing on the source language text, a first text translation result of the sample speech sequence is output.
[0072] In addition, the sample speech sequence and the target language text are input into the student model, and the initialized speech recognizer in the student model outputs the acoustic features and text mapping features of the sample speech sequence based on the input sample speech sequence. The initialized text encoder in the student model performs feature segmentation and feature encoding on the fusion features formed by fusing the acoustic features and text mapping features of the input sample speech sequence, to obtain the sample Bernoulli segmentation probability of each frame feature in the fusion features of the sample speech sequence, and the multi-segment coding features of the sample speech sequence. And the initialized text decoder in the student model performs feature decoding based on the coding features of at least one segment of the currently streaming input sample speech sequence, and the segment translation result and the target language identifier corresponding to the coding features of the historical decoding obtained from the target language text, to obtain a second text translation result of the sample speech sequence.
[0073] In addition, to avoid too many segments from damaging the acoustic integrity or too few segments from degenerating the model into non-real-time speech translation, the number of segments can be constrained according to the sample Bernoulli segmentation probability, thereby obtaining the segment constraint function of the student model, to better constrain the model to perform segment learning from both the acoustic and semantic levels, and improve the rationality and quality of segmentation.
[0074] And, in order to enable the student model to quickly learn the high-performance acoustic feature encoding knowledge in the pre-trained speech recognition model and the high-performance text encoding and decoding knowledge in the pre-trained translator, while realizing end-to-end real-time translation, the errors between the second text recognition result corresponding to the text mapping features of the sample speech sequence and the first text recognition result and the source language text respectively, and the errors between the second text translation result and the first text translation result and the target language text respectively can be calculated to obtain the distillation loss function.
[0075] Subsequently, the segmented constraint function and the distillation loss function are fused to obtain the target loss function. Based on the target loss function, the student model is trained, and a speech translation model that can perform reasonable segmentation, ensure the semantic integrity of real-time translation, learn the high-performance acoustic feature encoding knowledge in the pre-trained speech recognition model, and the high-performance text encoding and decoding knowledge in the pre-trained translator, and can achieve end-to-end real-time translation can be obtained. In this way, real-time and accurate speech translation results can be achieved through this speech translation model in practical applications. The fusion here can be direct addition, weighted addition, etc., and this embodiment does not make specific limitations on this.
[0076] In some embodiments, determining the segmented constraint function of the student model according to the sample Bernoulli segmentation probability includes: Determining a first constraint function according to the difference between the sum of the sample Bernoulli segmentation probabilities of all frame features in the fusion features of the sample speech sequence and the number of text units of the source language text; Determining a second constraint function according to the number of feature frames, the sample Bernoulli segmentation probability, and the number of text units in the fusion features of the sample speech sequence; Determining the segmented constraint function according to the first constraint function and the second constraint function.
[0077] Optionally, during the determination of the segmented constraint function, the sample Bernoulli segmentation probabilities of all frame features in the fusion features of the sample speech sequence can be summed, the sum result is subtracted from the number of text units (tokens) K of the source language text, and the norm of the difference result is calculated to obtain the first constraint function. Thus, through the first constraint function, it is ensured that the sum of the segmentation probabilities of all frames is close to the expected number of segments K, thereby avoiding generating too many or too few segments. In addition, the second constraint function can also be determined based on the number of feature frames in the fusion features of the sample speech sequence, the sample Bernoulli segmentation probability, and the number of text units. Thus, through the second constraint function, it is ensured that segmentation is performed only once in multiple consecutive speech frames, further ensuring the reasonableness of the number of segments and preventing over-segmentation on consecutive silent frames.
[0078] Subsequently, the first constraint function and the second constraint function are fused to obtain the segmented constraint function. Thus, through the segmented constraint function, the number of feature segments is constrained to be approximately the number of text units K, preventing over-segmentation on consecutive silent frames, and ensuring that segmentation is performed only once in consecutive speech frames, so that the segmentation features of each segment can correspond to a complete token in the source language text. Thus, while maintaining the acoustic and semantic integrity of the model trained accordingly, an appropriate number of segments are generated, thereby improving the real-time performance, integrity, and accuracy of translation.
[0079] In some embodiments, determining the second constraint function according to the number of feature frames, the sample Bernoulli segmentation probability, and the number of text units in the fusion features of the sample voice sequence includes: Performing a max pooling operation on the ratio between the number of feature frames and the number of text units and the sample Bernoulli segmentation probability to obtain the max pooling result corresponding to each frame feature in the fusion features of the sample voice sequence; Determining the second constraint function according to the difference between the sum of the max pooling results corresponding to all frame features in the fusion features of the sample voice sequence and the number of text units.
[0080] Optionally, during the calculation of the second constraint function, it may be to calculate the ratio between the number of feature frames and the number of text units, perform a max pooling operation on the ratio result and the sample Bernoulli segmentation probability of each frame feature, thereby obtaining the max pooling result corresponding to each frame feature in the fusion features of the sample voice sequence, that is, the maximum segmentation ratio value; and sum the max pooling results corresponding to all frame features in the fusion features of the sample voice sequence, subtract the number of text units from the sum result, and calculate the norm of the difference result to obtain the second constraint function, thereby ensuring that only one segmentation is performed among multiple consecutive speech frames through the second constraint function, further ensuring the rationality of the segmentation quantity and preventing over-segmentation on consecutive silent frames.
[0081] Among them, the segmentation constraint function The specific calculation formula is as follows: ; Among them, is the number of feature frames in the fusion features of the sample voice sequence; is the sample Bernoulli segmentation probability of the th frame feature in the fusion features of the sample voice sequence; is the number of text units of the source language text; is the max pooling operation function; is the floor calculation; is the two-norm calculation.
[0082] The voice translation device provided by the present invention will be described below. The voice translation device described below can be correspondingly referred to the voice translation method described above.
[0083] Figure 6 is a schematic structural diagram of the voice translation device provided by the present invention; as Figure 6 shown, the device includes: The first feature extraction unit 610 is used to perform acoustic feature extraction and text mapping on the speech sequence to be translated based on the speech recognizer of the speech translation model, so as to obtain the acoustic features and text mapping features of the speech sequence to be translated; The second feature extraction unit 620 is used to perform feature segmentation and feature encoding on the fusion features formed by fusing the acoustic features and the text mapping features based on the text encoder of the speech translation model, so as to obtain the encoded features of multiple segments of the speech sequence to be translated; The translation unit 630 is used to perform streaming decoding on the encoded features of multiple segments based on the text decoder of the speech translation model, so as to obtain the real-time translation text of the speech sequence to be translated; Wherein, the speech translation model is trained based on a sample speech sequence, and the source language text and target language text corresponding to the sample speech sequence.
[0084] The device provided by the present invention, under the speech translation model of end-to-end real-time translation, uses the cross-modal fusion features including the acoustic features and text mapping features of the speech sequence to be translated recognized by the speech recognizer as the input of the text encoder, so as to effectively make up for the cumulative errors existing in recognition and translation in the lack of scenario and context environment, and improve the translation accuracy; moreover, the encoded features of multiple segments output by the text encoder after feature segmentation and feature encoding of the fusion features are streamed into the text decoder, so that the text encoder can perform streaming decoding on the encoded features of multiple segments to obtain the real-time translation text of the speech sequence to be translated. While ensuring the semantic integrity of real-time translation, it greatly improves the first response and typing rate of real-time speech translation, further significantly improves the real-time performance and accuracy of the translation text output, and enhances the user experience.
[0085] In some embodiments, the second feature extraction unit is specifically used for: Based on the adapter of the speech translation model, perform feature fusion on the acoustic features and the text mapping features to obtain the fusion features; Based on the differentiable splitter of the text encoder, calculate the Bernoulli splitting probability of each frame feature in each fusion feature, and perform feature segmentation on the fusion features according to the Bernoulli splitting probability to obtain multiple segmented segmentation features; Based on the segmentation feature encoder of the text encoder, perform feature encoding on the segmentation features of each segment to obtain the encoded features of each segment.
[0086] In some embodiments, the second feature extraction unit is further used for: Based on the segmentation feature encoder, perform bidirectional attention encoding within each segment on the segmentation features of each segment to obtain the first encoded features of each segment, perform unidirectional attention encoding between segments on the segmentation features of every two segments to obtain the second encoded features of each segment, and perform feature fusion on the first encoded features and the second encoded features of each segment to obtain the encoded features of each segment.
[0087] In some embodiments, the apparatus further includes a model training unit, specifically configured to: Train an initialized speech recognition model based on the sample speech sequence and the source language text to obtain a pre-trained speech recognition model; Train an initialized translator based on the source language text and the target language text to obtain a pre-trained translator; Initialize a translation model to be trained based on the pre-trained speech recognition model and the pre-trained translator to obtain a student model; Use the pre-trained speech recognition model as the teacher model for the initialized speech recognizer in the student model, and use the pre-trained translator as the teacher model for the initialized text encoder and initialized text decoder in the student model, and perform knowledge distillation training on the student model to obtain the speech translation model.
[0088] In some embodiments, the model training unit is further configured to: Input the sample speech sequence into the pre-trained speech recognition model to obtain a first text recognition result of the sample speech sequence, and input the source language text and the target language text into the pre-trained translator to obtain a first text translation result of the sample speech sequence; Input the sample speech sequence and the target language text into the student model to obtain the sample Bernoulli segmentation probability of each frame feature in the fusion feature of the sample speech sequence, the text mapping feature of the sample speech sequence, and a second text translation result of the sample speech sequence; Determine the segmentation constraint function of the student model according to the sample Bernoulli segmentation probability; Determine the distillation loss function of the student model according to the first text recognition result, the first text translation result, the second text recognition result corresponding to the text mapping feature of the sample speech sequence, and the second text translation result; Train the student model according to the segmentation constraint function and the distillation loss function to obtain the speech translation model.
[0089] In some embodiments, the model training unit is further configured to: Determine a first constraint function based on the difference between the sum of the sample Bernoulli segmentation probabilities of all frame features in the fusion features of the sample speech sequence and the number of text units of the source language text; Determine a second constraint function based on the number of feature frames, the sample Bernoulli segmentation probability, and the number of text units in the fusion features of the sample speech sequence; Determine the segmentation constraint function according to the first constraint function and the second constraint function.
[0090] In some embodiments, the model training unit is further configured to: Perform a max pooling operation on the ratio between the number of feature frames and the number of text units and the sample Bernoulli segmentation probability to obtain the max pooling result corresponding to each frame feature in the fusion features of the sample speech sequence; Determine the second constraint function based on the difference between the sum of the max pooling results corresponding to all frame features in the fusion features of the sample speech sequence and the number of text units.
[0091] The device provided by the present invention is used to execute the above method embodiments. For the specific process and detailed content, please refer to the above embodiments and will not be elaborated here.
[0092] Figure 7 An example of the physical structure diagram of an electronic device is shown as Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the speech translation method, and the method includes: based on the speech recognizer of the speech translation model, performing acoustic feature extraction and text mapping on the speech sequence to be translated to obtain the acoustic features and text mapping features of the speech sequence to be translated; based on the text encoder of the speech translation model, performing feature segmentation and feature encoding on the fusion features formed by fusing the acoustic features and the text mapping features to obtain the encoded features of multiple segments of the speech sequence to be translated; based on the text decoder of the speech translation model, performing streaming decoding on the encoded features of multiple segments to obtain the real-time translation text of the speech sequence to be translated; wherein, the speech translation model is trained based on a sample speech sequence, and the source language text and target language text corresponding to the sample speech sequence.
[0093] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0094] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech translation method provided by the above-mentioned various methods. The method includes: an automatic speech recognition device based on a speech translation model extracts acoustic features and text mapping of a speech sequence to be translated, and obtains the acoustic features and text mapping features of the speech sequence to be translated; a text encoder based on the speech translation model performs feature segmentation and feature encoding on the fusion features formed by fusing the acoustic features and the text mapping features, and obtains encoded features of multiple segments of the speech sequence to be translated; a text decoder based on the speech translation model performs streaming decoding on the encoded features of multiple segments to obtain a real-time translation text of the speech sequence to be translated; wherein, the speech translation model is trained based on a sample speech sequence, and the source language text and target language text corresponding to the sample speech sequence.
[0095] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the speech translation method provided by the above-mentioned various methods. The method includes: based on a speech recognizer of a speech translation model, performing acoustic feature extraction and text mapping on a speech sequence to be translated, to obtain the acoustic features and text mapping features of the speech sequence to be translated; based on a text encoder of the speech translation model, performing feature segmentation and feature encoding on the fusion features formed by fusing the acoustic features and the text mapping features, to obtain the encoded features of multiple segments of the speech sequence to be translated; based on a text decoder of the speech translation model, performing streaming decoding on the encoded features of multiple segments, to obtain the real-time translation text of the speech sequence to be translated; wherein, the speech translation model is trained based on sample speech sequences, and the source language texts and target language texts corresponding to the sample speech sequences.
[0096] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0098] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech translation method, characterized in that: include: A speech recognizer based on a speech translation model performs acoustic feature extraction and text mapping on a speech sequence to be translated, thereby obtaining acoustic features and text mapping features of the speech sequence to be translated; Based on the text encoder of the speech translation model, feature segmentation and feature encoding are performed on the fusion feature formed by fusing the acoustic feature with the text mapping feature to obtain encoding features of multiple segments of the speech sequence to be translated; Based on the text decoder of the speech translation model, the encoding features of the multiple segments are stream-decoded to obtain a real-time translation text of the speech sequence to be translated; The speech translation model is trained based on a sample speech sequence, and a source language text and a target language text corresponding to the sample speech sequence.
2. The speech translation method according to claim 1, characterized in that: The text encoder based on the speech translation model performs feature segmentation and feature encoding on the fused features formed by fusing the acoustic features with the text mapping features to obtain multiple segmented encoding features of the speech sequence to be translated, including: Based on the adapter of the speech translation model, the acoustic feature and the text mapping feature are subjected to feature fusion to obtain the fused feature; Based on the differentiable segmenter of the text encoder, the Bernoulli segmentation probability of each frame feature in each fused feature is calculated, and according to the Bernoulli segmentation probability, the fused feature is segmented to obtain a plurality of segmented segmentation features; Based on the segmentation feature encoder of the text encoder, feature encoding is performed on the segmentation features of each segment to obtain the encoding features of each segment.
3. The speech translation method according to claim 2, characterized in that: The segmentation feature encoder based on the text encoder performs feature encoding on the segmentation feature of each segment to obtain the encoding feature of each segment, including: Based on the segmentation feature encoder, the segmentation features of each segment are subjected to intra-segment bidirectional attention encoding to obtain the first encoding features of each segment, the segmentation features of every two segments are subjected to inter-segment unidirectional attention encoding to obtain the second encoding features of each segment, and the first encoding features and the second encoding features of each segment are feature fused to obtain the encoding features of each segment.
4. The speech translation method according to any one of claims 1 to 3, characterized in that: The speech translation model is trained based on the following steps: Based on the sample speech sequence and the source language text, an initialization speech recognition model is trained to obtain a pre-trained speech recognition model; Based on the source language text and the target language text, training an initialization translator to obtain a pre-trained translator; Initializing the translation model to be trained based on the pre-trained speech recognition model and the pre-trained translator to obtain a student model; The pre-trained speech recognition model is used as the teacher model for initializing the speech recognizer in the student model, and the pre-trained translator is used as the teacher model for initializing the text encoder and the text decoder in the student model. The student model is trained by knowledge distillation to obtain the speech translation model.
5. The speech translation method according to claim 4, characterized in that: The performing knowledge distillation training on the student model to obtain the speech translation model includes: Inputting the sample speech sequence into the pre-trained speech recognition model to obtain a first text recognition result of the sample speech sequence, and inputting the source language text and the target language text into the pre-trained translator to obtain a first text translation result of the sample speech sequence; Input the sample speech sequence and the target language text into the student model to obtain the sample Bernoulli segmentation probability of each frame feature in the fusion feature of the sample speech sequence, the text mapping feature of the sample speech sequence and the second text translation result of the sample speech sequence; Determining a piecewise constraint function of the student model according to the sample Bernoulli segmentation probability; Determine a distillation loss function of the student model according to the first text recognition result, the first text translation result, a second text recognition result corresponding to the text mapping feature of the sample speech sequence, and the second text translation result; The student model is trained according to the segmentation constraint function and the distillation loss function to obtain the speech translation model.
6. The speech translation method according to claim 5, characterized in that: Determining the segmentation constraint function of the student model according to the sample Bernoulli segmentation probability includes: Determining a first constraint function according to a difference between a sum of sample Bernoulli segmentation probabilities of all frame features in the fusion features of the sample speech sequence and the number of text units of the source language text; Determining a second constraint function according to the number of feature frames in the fusion feature of the sample speech sequence, the sample Bernoulli segmentation probability and the number of text units; The piecewise constraint function is determined according to the first constraint function and the second constraint function.
7. The speech translation method according to claim 6, characterized in that: The determining of the second constraint function according to the number of feature frames in the fusion feature of the sample speech sequence, the sample Bernoulli segmentation probability and the number of text units includes: Performing a maximum pooling operation on the ratio between the number of feature frames and the number of text units and the sample Bernoulli segmentation probability to obtain a maximum pooling result corresponding to each frame feature in the fusion feature of the sample speech sequence; The second constraint function is determined according to the difference between the sum of the maximum pooling results corresponding to all frame features in the fusion features of the sample speech sequence and the number of text units.
8. A speech translation device, characterized in that: include: A first feature extraction unit is used for a speech recognizer based on a speech translation model to extract acoustic features and perform text mapping on a speech sequence to be translated, so as to obtain acoustic features and text mapping features of the speech sequence to be translated; A second feature extraction unit is used to perform feature segmentation and feature encoding on a fusion feature formed by fusing the acoustic feature with the text mapping feature based on the text encoder of the speech translation model, so as to obtain encoding features of multiple segments of the speech sequence to be translated; A translation unit, configured to perform streaming decoding on the encoding features of the plurality of segments based on a text decoder of the speech translation model to obtain a real-time translation text of the speech sequence to be translated; The speech translation model is trained based on a sample speech sequence, and a source language text and a target language text corresponding to the sample speech sequence.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the speech translation method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech translation method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the speech translation method according to any one of claims 1 to 7 is implemented.