A voice translation method, system and related device
By obtaining link information matching the translation task, combining audio features, text conversion features and acoustic features, encoding and decoding target conversion features, the problem of out-of-synchronization of sentences in speech translation is solved, and the accuracy and efficiency of translation are improved.
Patent Information
- Application Number
- CN202510224089.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing pronunciation translation methods are prone to dissynchronization of sentences when expressed in different languages, resulting in poor user experience.
By obtaining link information matching the translation task, combining audio features, text conversion features and acoustic features, the target conversion features are encoded and decoded to generate translated audio.
The translation task is simplified, the accuracy and efficiency of translation are improved, and the expression ability of target conversion characteristics is enhanced.
Smart Images

Figure CN119721071B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech processing, and in particular, to a speech translation method, system, and related devices. Background Art
[0002] With the continuous development of artificial intelligence, speech translation has been applied in more and more scenarios. For example, in scenarios such as large conferences or lectures, there are situations where the language of the presenter and the participants is different, resulting in the participants being unable to accurately understand what the presenter expresses. Currently, the main speech translation method relies on a neural network model to translate word by word according to the input audio stream. Due to the differences in different language expressions, the neural network model is prone to sentence out-of-sync during translation, resulting in a poor user experience.
[0003] In view of this, how to improve the accuracy of audio broadcasting has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a speech translation method, system, and related devices that can improve the accuracy of speech translation.
[0005] To solve the above technical problem, one technical solution adopted by this application is: to provide a speech translation method, including: based on the audio to be translated of the target object, determining the audio features, text conversion features, and acoustic features matching the target object corresponding to the audio to be translated; obtaining link information matching the translation task, and encoding based on the link information, the audio features, the text conversion features, and the acoustic features to obtain target conversion features matching the audio to be translated; decoding the target conversion features to obtain the translated audio corresponding to the audio to be translated.
[0006] To solve the above technical problem, another technical solution adopted by this application is: to provide a speech translation system, including: an acquisition module for determining the audio features, text conversion features, and acoustic features matching the target object corresponding to the audio to be translated based on the audio to be translated of the target object; a processing module for obtaining link information matching the translation task, and encoding based on the link information, the audio features, the text conversion features, and the acoustic features to obtain target conversion features matching the audio to be translated; a translation module for decoding the target conversion features to obtain the translated audio corresponding to the audio to be translated.
[0007] To solve the above technical problem, another technical solution adopted by this application is: to provide an electronic device, including: a memory and a processor coupled to each other, where program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the method mentioned in the above technical solution.
[0008] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the method mentioned in the above technical solution is implemented.
[0009] The beneficial effects of this application are as follows: Different from the prior art, the voice translation method proposed in this application simplifies the translation task by obtaining link information matching the translation task, reducing the difficulty of translating the audio to be translated. Moreover, by combining link information, audio features, text conversion features, and acoustic features, target conversion features are obtained through encoding of various types of features, improving the expression ability of the target conversion features and contributing to improving the accuracy of translation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0011] Figure 1 is a schematic flowchart of an embodiment of the voice translation method of this application;
[0012] Figure 2 is Figure 1 a schematic flowchart of another embodiment corresponding to step S101 in
[0013] Figure 3 is Figure 1 a schematic flowchart of another embodiment corresponding to step S102 in
[0014] Figure 4 is Figure 3 a schematic flowchart of another embodiment corresponding to step S302 in
[0015] Figure 5 is Figure 1 a schematic flowchart of another embodiment corresponding to step S103 in
[0016] Figure 6 is a schematic structural diagram of an embodiment of the voice translation system of this application;
[0017] Figure 7 is a schematic structural diagram of an embodiment of the electronic device of this application;
[0018] Figure 8 is a schematic structural diagram of an embodiment of the computer-readable storage medium of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments, and different embodiments can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0020] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an implementation manner of the voice translation method of the present application. The method includes:
[0021] S101: Based on the audio to be translated of the target object, determine the audio features, text conversion features, and acoustic features matching the target object corresponding to the audio to be translated.
[0022] In one implementation manner, to translate the audio expressed by the target object, the audio expressed by the target object is used as the audio to be translated. According to the audio to be translated, the corresponding audio features, text conversion features, and acoustic features matching the target object are determined.
[0023] Specifically, the above audio features are obtained by extracting audio features from the audio to be translated. The above text conversion features are obtained based on the conversion text converted from the audio to be translated. The above acoustic features are obtained by extracting acoustic features from the audio to be translated, and the acoustic features include at least one of tone color information, speech rate information, and prosody information matching the target object.
[0024] In one implementation scenario, a trained audio recognition model and a trained acoustic feature extraction model are obtained. The audio to be translated is input into the audio recognition model to obtain the audio features corresponding to the audio to be translated, and the corresponding text conversion features are recognized according to the audio features. And, the audio to be translated is input into the acoustic feature extraction model to extract the acoustic features.
[0025] In another implementation manner, to improve the translation accuracy, after the audio to be translated of the target object is obtained, noise reduction processing is performed on the audio to be translated to remove environmental noise. The audio to be translated after removing the noise is processed to obtain the corresponding audio features, text conversion features, and acoustic features.
[0026] S102: Obtain the link information matching the translation task, and based on the link information, audio features, text conversion features, and acoustic features, encode to obtain the target conversion features matching the audio to be translated.
[0027] In one embodiment, a translation task corresponding to the audio to be translated is determined. The translation task is used to represent converting the audio to be translated into a translated audio that matches the target language. The above target language can be determined according to the application scenario or can also be instructed by relevant personnel.
[0028] Further, based on decomposing the translation task, matching link information is obtained. According to the link information, audio features, text conversion features, and acoustic features, a target conversion feature for synthesizing the translated audio is encoded. Among them, the above target conversion feature includes acoustic features that match the target object and translated text information.
[0029] S103: Decode the target conversion feature to obtain the translated audio corresponding to the audio to be translated.
[0030] In one embodiment, the obtained target conversion feature is decoded to generate the translated audio corresponding to the audio to be translated.
[0031] Specifically, in response to the target conversion feature including acoustic features that match the target object, the generated translated audio also has acoustic features that match the target object to achieve simultaneous interpretation.
[0032] The speech translation method proposed in this application simplifies the translation task by obtaining link information that matches the translation task, reducing the difficulty of translating the audio to be translated. Also, by combining link information, audio features, text conversion features, and acoustic features, a target conversion feature is encoded according to various types of features, improving the expression ability of the target conversion feature and helping to improve the accuracy of translation.
[0033] Please refer to Figure 2 , Figure 2 is Figure 1 a schematic flowchart of another embodiment corresponding to step S101 in
[0034] S201: Based on the audio to be translated, obtain the corresponding audio features and text conversion features.
[0035] In one embodiment, the audio features corresponding to the audio to be translated are obtained, and the audio features are recognized to obtain the text conversion features.
[0036] In an implementation scenario, a trained audio recognition model is obtained, and the audio to be translated is input into the feature encoding network in the audio recognition model to obtain the audio features corresponding to the audio to be translated output by the feature encoding network. The audio features are recognized to obtain corresponding text conversion features, which correspond to the conversion text corresponding to the audio to be translated, and the text conversion features are used to generate the above conversion text. Among them, the structure of the above audio recognition model can refer to the existing neural network structure and will not be elaborated in detail here.
[0037] S202: Extract an audio segment from the audio to be translated, and based on the audio segment, obtain acoustic features matching the target object.
[0038] In an implementation manner, the starting acquisition moment corresponding to the audio to be translated is obtained. Based on the starting acquisition moment, an audio segment with an audio duration of the target duration is obtained. By extracting the audio segment from the audio to be translated, directly using the complete audio to be translated to extract acoustic features is avoided, saving computational consumption and helping to improve the translation efficiency. The audio segment is input into the acoustic feature extraction model to extract acoustic features matching the target object.
[0039] Specifically, the moment when the acquisition of the audio to be translated starts is used as the starting acquisition moment. Starting from the starting acquisition moment, an audio segment with a duration of the target duration is intercepted from the collected audio to be translated. Among them, the above target duration can be obtained through estimation or can be inversely deduced by relevant technicians through multiple experiments.
[0040] In another implementation manner, the audio segment can also be randomly intercepted from the audio to be translated, and its corresponding duration is the target duration.
[0041] It should be noted that in actual applications, the execution order of step S201 and step S202 can also be other, for example, step S201 and step S202 are executed simultaneously; or step S202 is executed first, and then step S201 is executed.
[0042] Please refer to Figure 3 , Figure 3 is Figure 1 The flowchart of another implementation manner corresponding to step S102 in
[0043] S301: Use the intelligent analysis model to decompose the translation task and generate corresponding multiple translation sub-tasks.
[0044] In one embodiment, to improve the accuracy of the target conversion features, an intelligent analysis model is used to decompose the translation task, generating multiple translation subtasks, and the multiple translation subtasks constitute the link information mentioned in step S102.
[0045] Specifically, the intelligent analysis model is used to simulate the human thinking and reasoning process, decomposing the translation task into corresponding multiple translation subtasks to reduce the difficulty of translating the audio to be translated.
[0046] In one implementation scenario, the above intelligent analysis model is a large language model with relatively excellent data analysis capabilities. By inputting the translation task into the intelligent analysis model, the intelligent analysis model is enabled to simulate human thinking to generate multiple translation subtasks to construct link information.
[0047] In one specific application scenario, the above large language model may include, but is not limited to, Deep Neural Networks (DNNs), Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM), and Generative Pretrained Transformer models, etc. There is no specific limitation on the specific structure and specific deployment of the large language model here. Additionally, it should be noted that the specific structure and specific deployment of the intelligent analysis model mentioned in other embodiments of this application can all refer to this embodiment.
[0048] S302: Input the translation subtasks, audio features, text conversion features, and acoustic features into the intelligent analysis model, and use the intelligent analysis model to obtain the target conversion features matching the audio to be translated.
[0049] In one embodiment, the translation subtasks, audio features, text conversion features, and acoustic features are input into the intelligent analysis model. According to the association relationships between different translation subtasks, combining the audio features, text conversion features, acoustic features, and the reply text corresponding to the generated translation subtasks, the reply text corresponding to the current translation subtask is generated until the target conversion features matching the audio to be translated are obtained. Among them, the above target conversion features include translation text information and the acoustic information of the target object, and the translation text information matches the translation text corresponding to the conversion text, and the conversion text is determined based on the text conversion features.
[0050] In a specific application scenario, in response to the language of the audio to be translated expressed by the target object being Chinese and the target language of translation being English, the implementation processes of steps S301 and S302 include: using multiple translation subtasks generated by the intelligent analysis model in sequence as "extracting text conversion information", "converting the converted text information into translation text information matching the target language", "converting the translation text information into reference conversion features", and "converting the reference conversion features into target conversion features including acoustic information".
[0051] In another implementation manner, to improve the accuracy of the target conversion features encoded by the intelligent analysis model, the intelligent analysis model is fine-tuned using multiple first training samples to obtain the trained intelligent analysis model. Among them, the first training samples are matched with corresponding target conversion features. During the fine-tuning training process, the intelligent analysis model is pre-tipped to predict the corresponding conversion text according to the input first training samples, and then the intelligent analysis model is tipped to predict the translation text corresponding to the conversion text and output the predicted conversion features. The predicted conversion features include the translation text information and acoustic information corresponding to the first training samples. At least some parameters in the intelligent analysis model are adjusted according to the predicted conversion features and the target conversion features until the preset convergence condition is met, and the fine-tuned intelligent analysis model is obtained.
[0052] In an implementation scenario, in combination with the translation scenario, the first training samples related to the scenario are obtained, and the intelligent analysis model is fine-tuned using the LoRA model (Low-Rank Adaptation of Large Language Models) to obtain the fine-tuned intelligent analysis model. For example, when the translation scenario is related to the earth environment, multiple first training samples related to the earth environment are obtained, and the intelligent analysis model is fine-tuned using the obtained first training samples.
[0053] In the above solution, by using the intelligent analysis model to simulate the human thinking mode and generate link information, the difficulty of obtaining the target conversion features subsequently is reduced. Moreover, using the intelligent analysis model to combine the link information, audio features, text conversion features, and acoustic features to generate the target conversion features helps to enhance the expression ability of the target conversion features, thereby improving the accuracy of subsequent speech translation.
[0054] Please refer to Figure 4 , Figure 4 is Figure 3 the schematic flowchart of another implementation manner corresponding to step S302 in
[0055] S401: Based on the audio features, text conversion features, and acoustic features, use the intelligent analysis model to generate reply information matching each translation subtask.
[0056] In one embodiment, a translation subtask, audio features, text conversion features, and acoustic features are input into an intelligent analysis model to analyze the above-mentioned audio features, text conversion features, and acoustic features by using the intelligent analysis model, and reply information matching each translation subtask is output.
[0057] In a specific application scenario, multiple translation subtasks generated by using the intelligent analysis model include "extracting text conversion information", "obtaining translation text information matching a target language", and "extracting acoustic information of a target object". For the translation subtask of "extracting conversion text information", the intelligent analysis model is used to combine the audio features and text conversion features and further extract to obtain the text conversion information. For the translation subtask of "obtaining translation text information matching a target language", the intelligent analysis model is used to obtain the translation text information according to the text conversion features, and the translation text information is a feature vector representation of the translation text corresponding to the audio to be translated. For the translation subtask of "extracting acoustic information of a target object", the intelligent analysis model is used to further extract the acoustic features to obtain the acoustic information.
[0058] S402: Based on all the reply information, generate target conversion features. Among them, the target conversion features include translation text information and acoustic information of the target object, the translation text information matches the translation text corresponding to the conversion text, and the conversion text is determined based on the text conversion features.
[0059] In one embodiment, the intelligent analysis model is used to perform joint analysis and summary on all the obtained reply information to finally output target conversion features including translation text information and acoustic information.
[0060] Specifically, after obtaining all the reply information generated by the intelligent analysis model, the intelligent analysis model is prompted to generate target conversion features for synthesizing translated audio according to all the reply information, that is, the intelligent analysis model combines the translation text information and the acoustic information to generate the target conversion features.
[0061] Please refer to Figure 5 , Figure 5 is Figure 1 a schematic flowchart of another embodiment corresponding to step S103 in
[0062] S501: Input the target conversion features into the trained codebook compilation model to generate target decoding features.
[0063] In one embodiment, a trained codebook compilation model is obtained, the target conversion features are input into the codebook compilation model to generate corresponding codebook features, and the codebook features are used as the target decoding features.
[0064] In an application scenario, the trained codebook compilation model converts the target conversion features into target decoding features for synthesizing translated audio through the RVQ (Residual Vector Quantization) technology.
[0065] In another implementation, to improve the accuracy of translation, the trained codebook compilation model also has the function of codebook feature prediction. After inputting the target conversion features into the codebook compilation model, the initial codebook features matching the target conversion features are output by using the codebook compilation model. Based on the initial codebook features, multiple predicted codebook features predicted by the codebook compilation model are obtained. Among them, both the initial codebook features and the predicted codebook features contain text information corresponding to the complete translation text and acoustic information corresponding to the target object, and different predicted codebook features have different sampling rates. The initial codebook features and the predicted codebook features are used as the target decoding features.
[0066] Specifically, the trained codebook compilation model is used to generate the initial codebook features in digital representation according to the target conversion features. According to the obtained initial codebook features, the predicted codebook features matching multiple different sampling rates are predicted, and the initial codebook features and all the predicted codebook features are used as the target decoding features. By predicting multiple different predicted codebook features by using the initial codebook features, the expression ability of the target decoding features is improved, which helps to improve the accuracy of subsequent translation.
[0067] In an implementation scenario, the above codebook compilation model is trained by using multiple second training samples, and the second training samples include target training features and corresponding target training audio. The target training features are input into the codebook compilation model to use the codebook compilation model to output training decoding features including multiple codebook features. According to the training decoding features, predicted audio is synthesized. The training loss of the model is calculated according to the predicted audio and the target training audio, and the parameters of the codebook compilation model are adjusted by using the training loss until the preset convergence condition is met. By training the codebook compilation model by using the second training samples, the trained codebook compilation model can output codebook features corresponding to multiple sampling rates as the target decoding features, and synthesize translated audio according to the target decoding features.
[0068] In a specific application scenario, the structure of the above codebook compilation model can refer to the existing non-linear autoregressive model structure.
[0069] S502: Decode the target decoding features to obtain the translated audio corresponding to the audio to be translated.
[0070] In an implementation, the generated target decoding features are decoded to synthesize the corresponding translated audio.
[0071] Specifically, the target decoding feature is input into the decoder to obtain the translated audio output by the decoder.
[0072] In another embodiment, the present application also proposes a speech translation model, which is used to implement the speech translation method mentioned in any of the above embodiments. The speech translation model includes a pre-trained audio recognition model, a pre-trained acoustic feature extraction model, a fine-tuned intelligent analysis model, a pre-trained codebook compilation model, and a pre-trained decoder. The audio recognition model is used to obtain the audio feature and text conversion feature corresponding to the audio to be translated; the acoustic feature extraction model is used to extract the acoustic feature matching the target object; the intelligent analysis model is used to construct link information and encode to obtain the target conversion feature; the codebook compilation model is used to generate the target decoding feature; the decoder is used to decode the target decoding feature.
[0073] In one embodiment, to improve the accuracy of speech translation of the speech translation model, after the speech translation model is constructed, the speech translation model is jointly trained to improve the stability of the speech translation model.
[0074] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of the speech translation system of the present application. The speech translation system includes an acquisition module 10, a processing module 20, and a translation module 30 that are mutually coupled.
[0075] Specifically, the acquisition module 10 is used to determine the audio feature, text conversion feature, and acoustic feature matching the target object corresponding to the audio to be translated based on the audio to be translated of the target object.
[0076] The processing module 20 is used to obtain the link information matching the translation task, and encode to obtain the target conversion feature matching the audio to be translated based on the link information, audio feature, text conversion feature, and acoustic feature.
[0077] The translation module 30 is used to decode the target conversion feature to obtain the translated audio corresponding to the audio to be translated.
[0078] In one embodiment, the link information includes multiple translation subtasks. The processing module 20 obtains the link information matching the translation task, and encodes to obtain the target conversion feature matching the audio to be translated based on the link information, audio feature, text conversion feature, and acoustic feature, including: decomposing the translation task by using the intelligent analysis model to generate corresponding multiple translation subtasks; inputting the translation subtasks, audio feature, text conversion feature, and acoustic feature into the intelligent analysis model, and using the intelligent analysis model to obtain the target conversion feature matching the audio to be translated.
[0079] In one embodiment, the processing module 20 inputs the translation subtask, audio features, text conversion features, and acoustic features into the intelligent analysis model, and uses the intelligent analysis model to obtain the target conversion features matching the audio to be translated, including: based on the audio features, text conversion features, and acoustic features, using the intelligent analysis model to generate response information matching each translation subtask; based on all the response information, generating the target conversion features; wherein, the target conversion features include translation text information and acoustic information of the target object, and the translation text information matches the translation text corresponding to the conversion text, and the conversion text is determined based on the text conversion features.
[0080] In one embodiment, the translation module 30 decodes the target conversion features to obtain the translated audio corresponding to the audio to be translated, including: inputting the target conversion features into the trained codebook compilation model to generate target decoding features; decoding the target decoding features to obtain the translated audio corresponding to the audio to be translated.
[0081] In one embodiment, the translation module 30 inputs the target conversion features into the trained codebook compilation model to generate target decoding features, including: using the codebook compilation model to output initial codebook features matching the target conversion features; based on the initial codebook features, obtaining multiple predicted codebook features predicted by the codebook compilation model; wherein, different predicted codebook features have different sampling rates; taking the initial codebook features and the predicted codebook features as the target decoding features.
[0082] In one embodiment, the acquisition module 10 determines the audio features, text conversion features, and acoustic features matching the target object corresponding to the audio to be translated based on the audio to be translated by the target object, including: based on the audio to be translated, obtaining the corresponding audio features and text conversion features; and extracting an audio segment from the audio to be translated, and based on the audio segment, obtaining the acoustic features matching the target object.
[0083] In one embodiment, the acquisition module 10 obtains the corresponding audio features and text conversion features based on the audio to be translated, including: obtaining the audio features corresponding to the audio to be translated, and identifying the audio features to obtain the text conversion features.
[0084] The acquisition module 10 extracts an audio segment from the audio to be translated, including: obtaining the starting acquisition moment corresponding to the audio to be translated; based on the starting acquisition moment, obtaining an audio segment with an audio duration of the target duration.
[0085] Please refer to Figure 7 , Figure 7It is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device includes: a memory 40 and a processor 50 that are coupled to each other. Program instructions are stored in the memory 40, and the processor 50 is configured to execute the program instructions to implement the methods mentioned in any of the above embodiments. Specifically, the electronic device includes, but is not limited to: desktop computers, laptop computers, tablet computers, servers, etc., which are not limited herein. In addition, the processor 50 may also be referred to as a CPU (Center Processing Unit, central processing unit). The processor 50 may be an integrated circuit chip with signal processing capabilities. The processor 50 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 50 may be implemented jointly by integrated circuit chips.
[0086] Please refer to Figure 8 , Figure 8 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Program instructions 70 that can be run by a processor are stored on the computer-readable storage medium 60. When the program instructions 70 are executed by the processor, the methods mentioned in any of the above embodiments are implemented.
[0087] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0088] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0089] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0090] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0091] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A voice translation method, characterized in that, Including: Based on the audio to be translated of the target object, determine the audio features, text conversion features corresponding to the audio to be translated, and acoustic features matching the target object; wherein, the acoustic features include at least one of timbre information, speech rate information, and prosody information matching the target object; Obtain link information matching the translation task, wherein the link information includes multiple translation subtasks, and input the translation subtasks, the audio features, the text conversion features, and the acoustic features of the link information into an intelligent analysis model; based on the audio features, the text conversion features, and the acoustic features, use the intelligent analysis model to generate reply information matching each translation subtask; use the intelligent analysis model to encode and generate, based on all the reply information, to obtain target conversion features matching the audio to be translated; wherein, the intelligent analysis model is a large language model with data analysis capabilities; Decode the target conversion features to obtain the translated audio corresponding to the audio to be translated.
2. The method according to claim 1, wherein The target conversion features include translation text information and acoustic information of the target object, and the translation text information matches the translation text corresponding to the conversion text, and the conversion text is determined based on the text conversion features.
3. The method according to claim 1, characterized in that, The decoding the target conversion features to obtain the translated audio corresponding to the audio to be translated includes: Input the target conversion features into a trained codebook compilation model to generate target decoding features; Decode the target decoding features to obtain the translated audio corresponding to the audio to be translated.
4. The method according to claim 3, wherein The inputting the target conversion features into a trained codebook compilation model to generate target decoding features includes: Use the codebook compilation model to output initial codebook features matching the target conversion features; Based on the initial codebook features, obtain multiple predicted codebook features predicted by the codebook compilation model; wherein, different predicted codebook features have different sampling rates; Use the initial codebook features and the predicted codebook features as the target decoding features.
5. The method according to claim 1, wherein The determining the audio features, text conversion features corresponding to the audio to be translated, and acoustic features matching the target object based on the audio to be translated of the target object includes: Based on the audio to be translated, obtain the corresponding audio features and text conversion features; and, Extract an audio segment from the audio to be translated, and based on the audio segment, obtain the acoustic features matching the target object.
6. The method according to claim 5, wherein The obtaining the corresponding audio features and text conversion features based on the audio to be translated includes: Obtain the audio features corresponding to the audio to be translated, and identify the audio features to obtain text conversion features; The extracting an audio segment from the audio to be translated includes: Obtain the starting acquisition moment corresponding to the audio to be translated; Based on the starting acquisition moment, obtain the audio segment with an audio duration of the target duration.
7. A voice translation system, characterized in that, Including: An acquisition module, configured to determine an audio feature corresponding to the audio to be translated, a text conversion feature, and an acoustic feature matching the target object based on the audio to be translated of the target object; wherein, the acoustic feature includes at least one of timbre information, speech rate information, and prosody information matching the target object; A processing module, configured to obtain link information matching a translation task, wherein the link information includes a plurality of translation subtasks, input the translation subtasks, the audio feature, the text conversion feature, and the acoustic feature of the link information into an intelligent analysis model; generate reply information matching each translation subtask by using the intelligent analysis model based on the audio feature, the text conversion feature, and the acoustic feature; encode and generate a target conversion feature matching the audio to be translated by using the intelligent analysis model based on all the reply information; wherein, the intelligent analysis model is a large language model with data analysis capabilities; A translation module, configured to decode the target conversion feature to obtain a translated audio corresponding to the audio to be translated.
8. An electronic device, characterized in that, Comprising: A memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Multimodal fusion speech translation method, system and equipment
CN118692446A